A data-driven microgrid voltage control method

CN115986748BActive Publication Date: 2026-08-11NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-29
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

然而,当出现新情况时,传统的模型控制方法必须解决优化问题,这需要大量的计算

Benefits of technology

[0058]本发明基于数据驱动方法和分布式架构,采用集中训练与分散执行结合的方式,通过对历史数据进行离线训练得到有效的运行控制策略,从而能够根据网络的实时运行工况,在线做出实时的联合出力决策,确保了网络运行过程中的可靠性、经济性与安全性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115986748B_ABST
    Figure CN115986748B_ABST
Patent Text Reader

Abstract

This invention discloses a data-driven microgrid voltage control method, relating to the field of microgrid application technology. The method includes: dividing the microgrid into multiple interconnected subnetworks based on a distributed architecture; constructing an intelligent agent for internal control of each subnetwork; building a joint control model with reactive power output as the primary factor and active power output as the secondary factor for the energy storage system; constructing a barrel-shaped voltage barrier function model by combining V-type and U-type voltage barrier function models; using voltage constraints based on the total network power loss and the voltage barrier function model as the target control; using a weighted sum algorithm to balance voltage deviation and network power loss; and finally establishing a distributed VVC model for the microgrid; fitting the distributed VVC model of the microgrid to a partially observable Markov decision-making (POMG) model; and improving the state-action equations from adapting to discrete actions to adapting to continuous actions to meet real-time control requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of microgrid application technology, specifically a data-driven microgrid voltage control method. Background Technology

[0002] A microgrid is a small-scale power generation and distribution system comprised of distributed power sources, energy storage devices, energy conversion devices, related loads, and monitoring and protection devices. It is an autonomous system capable of self-control, protection, and management, and can operate either connected to the external power grid or in isolation. It is an important component of the smart grid. Microgrids have a dual role: For the power grid, as a smart load with variable size, the microgrid provides dispatchable load to the local power system, responding within seconds to meet system needs and providing timely support to the main grid; it can perform system maintenance without affecting customer loads; it can alleviate (extend) the time required for distribution network upgrades; and by adopting the IEEE 1547.4 standard, it guides the islanded operation of distributed power sources, eliminating technical obstacles arising from certain special operational requirements. For users, the microgrid is a customizable power source that can meet diverse user needs, such as enhancing local power supply reliability, reducing feeder losses, supporting local voltage, improving efficiency through waste heat utilization, providing voltage sag correction, or serving as an uninterruptible power source. Microgrids not only solve the problem of large-scale integration of distributed power sources and fully leverage their advantages, but also bring other benefits to users. Microgrids will fundamentally change the traditional approach to load growth, possessing enormous potential in reducing energy consumption and improving the reliability and flexibility of power systems.

[0003] For microgrids, several voltage / reactive power control (VVC) methods have been developed to address power management issues such as voltage regulation and reducing network power losses. Inverters and energy storage devices have been widely studied and used in microgrid regulation and control due to their increasing availability, flexibility, and the fast response speed of advanced power electronics technologies. Currently, the control frameworks for microgrids mainly include three types: centralized control, local control, and distributed control. Distributed methods have received increasing attention because they do not require complex communication networks and can address issues related to privacy, complexity, scalability, and global impact.

[0004] A common research approach in distributed control is to transform non-convex problems into convex optimization problems through reasonable approximation and simplification, and then solve them using distributed algorithms. To address the uncertainties of distributed renewable energy sources and load demand, model-based distributed control methods typically require pre-determining the optimal solution. However, when new situations arise, traditional model control methods must solve optimization problems, which requires substantial computation. Furthermore, model-based control methods usually require accurate and comprehensive grid parameters, which are often difficult to achieve. Summary of the Invention

[0005] To address the shortcomings mentioned in the background section, the present invention aims to provide a data-driven microgrid voltage control method.

[0006] The objective of this invention can be achieved through the following technical solution: a data-driven microgrid voltage control method, the method comprising the following steps:

[0007] The microgrid is divided into multiple sub-networks that are interconnected and coupled through power flow based on a distributed architecture, and an intelligent agent for internal control is built for each sub-network.

[0008] A joint control model is constructed, with the reactive output of the photovoltaic inverter as the main component and the active output of the energy storage system as the auxiliary component. Each sub-network corresponds to a single intelligent agent that controls all photovoltaic inverters and energy storage devices within the sub-network. Effective control of the microgrid voltage is achieved by controlling the reactive output of the photovoltaic inverter and the active output of the energy storage system.

[0009] By combining the V-type and U-type voltage barrier function models, a barrel-type voltage barrier function model is constructed.

[0010] Using voltage constraints based on the total power loss of the network and the voltage barrier function model as the target control, a weighted sum algorithm is used to balance voltage deviation and network power loss. A spatiotemporal uncertainty model based on photovoltaic and load prediction intervals is established, and a network power flow constraint model based on the Newton-Raphson method is established. By combining the spatiotemporal uncertainty model based on photovoltaic and load prediction intervals, a joint control model with reactive power output of photovoltaic inverter as the main component and active power output of energy storage system as the auxiliary component is established. After balancing the control target and network power flow constraint model after voltage deviation and network power loss, a distributed VVC model of microgrid is established.

[0011] The distributed VVC model of the microgrid is fitted to the partially observable Markov decision POMG model, and the state action equation is improved from adapting to discrete action to adapting to continuous action to meet the real-time control requirements.

[0012] The constructed spatiotemporal uncertainty model based on photovoltaic and load prediction intervals is transformed into a stochastic programming model, and a network solution failure penalty is added to the return reward of each training set to improve the network solution success rate.

[0013] The algorithm flow is designed using the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm, and applied to microgrids.

[0014] Preferably, the distributed architecture does not require a central coordinator to collect all information generated in the network, but needs to consider global coordination issues; in distributed optimization, each constructed agent represents a sub-network that only exchanges limited boundary physical information and global reward rewards with its neighboring sub-networks, collectively seeking the globally optimal solution, and during each operation, the trained controller implements the locally measured VVC within the corresponding sub-network.

[0015] Preferably, the joint control model with the photovoltaic inverter primarily providing reactive power output and the energy storage system providing active power output is as follows:

[0016]

[0017]

[0018]

[0019]

[0020]

[0021] This represents the real-time active power output of the photovoltaic system. This represents the real-time reactive power output of the photovoltaic inverter. This represents the complex power of the photovoltaic inverter. δ represents the photovoltaic active power output boundary at time t, and δ is the reactive power capacity factor of the photovoltaic inverter. and This indicates the minimum and maximum active power that the energy storage system can generate and absorb. This represents the active power output value of the stored energy at time t. This represents the real-time capacity of the energy storage at time t. This indicates the maximum energy storage capacity.

[0022] Photovoltaic inverters prioritize providing reactive power when needed. When the reactive power compensation capacity is insufficient, the energy storage system will activate. The reactive power output of each photovoltaic inverter is also limited to a preset proportion of its apparent power capacity. A positive value indicates that reactive power is injected into the grid, and a negative value indicates that reactive power is absorbed from the grid. The energy storage system is set up similarly to the inverter. A positive value indicates that active power is injected into the grid, and a negative value indicates that active power is absorbed from the grid. The remaining power of the energy storage system is always positive.

[0023] Preferably, the barrel-shaped voltage barrier function model is as follows:

[0024]

[0025] In the formula, v a The real-time voltage magnitude of the node; v ref The network voltage reference value is set to 1.00 pu; v (v a This is the real-time reward for the node voltage;

[0026] Furthermore, the voltage barrier function model combines the advantages of V-type and U-type: on the one hand, it has a slow gradient within the safe range, resulting in better voltage conditions; on the other hand, the larger gradient outside the safe range ensures faster policy guidance.

[0027] Preferably, the control objective after balancing voltage deviation and network power loss is:

[0028]

[0029]

[0030]

[0031] In the formula, lv(vi,t) is the real-time voltage barrier function value of node i at time t; N is the number of network nodes; N m A set of network nodes; For the set of network branches; r ij and x ij These represent the resistance and reactance of the branch between nodes i and j, respectively; v i,t and v j,t Let i and j represent the voltage amplitudes at time t, respectively.

[0032] Then, a weighted sum algorithm is used to transform the multi-objective function into an equivalent single-objective function with weighting factors. Utopian points and Nadir points are used to normalize the objective. For any subnetwork m, the weighted sum representation of the normalized objective is obtained as follows:

[0033] In the formula, This represents the voltage deviation of the m-network at time t after normalization. Let α represent the normalized active power loss of network m at time t, and let α and β represent the normalization coefficients.

[0034] Preferably, the process of establishing a spatiotemporal uncertainty model based on photovoltaic and load forecast intervals is as follows:

[0035] Before each operating period, spatial uncertainty scenarios are randomly generated within a given prediction interval using Monte Carlo sampling. Next, within each operating cycle, temporal uncertainty scenarios are generated using Monte Carlo sampling, and time delays are considered using the temporal uncertainty interval. Under each temporal uncertainty scenario, voltage deviation and network loss are calculated, and these two objectives are corrected using the scenario occurrence probability. The corrected normalized objective is expressed as:

[0036]

[0037] This is the sum of normalized values ​​for all cases in stochastic programming (average voltage deviation / network loss of network m at time t); Let ξ be the sum of the normalized values ​​in the case of u; u This represents the probability of scenario u occurring.

[0038] Preferably, the process of improving the state-action equation from adapting to discrete actions to adapting to continuous actions is as follows:

[0039] The state-action function is as follows, used to indicate the current state or the combined return of the state-action pair:

[0040]

[0041] In the formula, τ i This represents the history of agent i; a -i =× j≠i a j To accommodate continuous actions, the state-action function is changed to the following form:

[0042]

[0043] In the formula, based on a′ i Gaussian distribution by π i (a′ i ∣τ i )express, Obtained through Monte Carlo sampling, as follows:

[0044]

[0045] Preferably, the process of adding a penalty for network solution failure to improve the success rate of network solution is as follows:

[0046] Add a network solution failure penalty function F:

[0047] F = -f,t f <T max

[0048] In the formula, f is the penalty occurrence constant, which is a large positive number; t f T represents the duration during which the network successfully solves a problem in the training set, even if the problem is interrupted due to network failure. max The maximum time step for each training set;

[0049] Then, the reward for each training set interrupted due to network solution failure is adjusted to R. mf :

[0050] R mf =R m +F

[0051] In the formula, R m This refers to the normal reward value obtained before the training set was interrupted.

[0052] Preferably, an apparatus includes:

[0053] One or more processors;

[0054] Memory, used to store one or more programs;

[0055] When one or more of the programs are executed by one or more of the processors, the one or more processors implement a data-driven microgrid voltage control method as described above.

[0056] Preferably, a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform a data-driven microgrid voltage control method as described above.

[0057] The beneficial effects of this invention are:

[0058] This invention is based on a data-driven approach and a distributed architecture. It adopts a combination of centralized training and decentralized execution, and obtains effective operation control strategies by offline training on historical data. This enables real-time joint output decisions to be made online based on the real-time operating conditions of the network, ensuring the reliability, economy and security of the network operation. Attached Figure Description

[0059] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0060] Figure 1 This is a schematic diagram of the overall network framework of the present invention;

[0061] Figure 2 This is a schematic diagram of the algorithm flow framework of the present invention. Detailed Implementation

[0062] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0063] like Figure 1 As shown, a data-driven microgrid voltage control method includes the following steps:

[0064] Step (1): The microgrid is divided into multiple sub-networks that are interconnected and coupled through power flow based on a distributed architecture. For each sub-network, a corresponding intelligent agent for internal control is constructed.

[0065] Considering the communication and computational burdens, privacy issues, and the inability of traditional centralized VVC frameworks to achieve global coordination, a regionally coordinated microgrid VVC framework is constructed, as shown in the attached figure. Figure 1 As shown, the microgrid is divided into multiple subnetworks interconnected and coupled by power flow based on a distributed architecture. For each subnetwork, a corresponding intelligent agent is constructed for its internal control. Each agent, representing a subnetwork, exchanges limited boundary physical information and global reward information only with its neighboring subnetworks, collectively seeking the globally optimal solution. During each operation, the trained controller implements locally measured VVC within its corresponding subnetwork.

[0066] Step (2): Construct a joint control model with the reactive output of the photovoltaic inverter as the main component and the active output of the energy storage system as the auxiliary component. Each sub-network corresponds to a single intelligent agent that controls all photovoltaic inverters and energy storage devices within the sub-network. Effective control of the microgrid voltage is achieved by controlling the reactive output of the photovoltaic inverter and the active output of the energy storage system.

[0067] The inverter control model is shown in (1)-(3). The reactive power output of each inverter is also limited to a preset proportion of its apparent power capacity. A positive value indicates that reactive power is injected into the grid, and a negative value indicates that reactive power is absorbed from the grid.

[0068]

[0069]

[0070]

[0071] in, This represents the real-time active power output of the photovoltaic system. This represents the real-time reactive power output of the photovoltaic inverter. This represents the complex power of the photovoltaic inverter. δ represents the photovoltaic active power output boundary at time t, and δ is the photovoltaic inverter reactive power capacity factor.

[0072] The droop control scheduling model of the photovoltaic inverter is shown in (4) and (5). The setting of the energy storage system is similar to that of the inverter. A positive value indicates that active power is injected into the grid, and a negative value indicates that active power is absorbed from the grid. The remaining power of the energy storage system is always positive.

[0073]

[0074]

[0075] in, and This indicates the minimum and maximum active power that the energy storage system can generate and absorb. This represents the active power output value of the stored energy at time t. This represents the real-time capacity of the energy storage at time t. This indicates the maximum energy storage capacity.

[0076] Photovoltaic inverters prioritize providing reactive power when needed. When the reactive power compensation capacity is insufficient, the energy storage system will activate. The reactive power output of each photovoltaic inverter is also limited to a preset proportion of its apparent power capacity. A positive value indicates that reactive power is injected into the grid, and a negative value indicates that reactive power is absorbed from the grid. The energy storage system is set up similarly to the inverter. A positive value indicates that active power is injected into the grid, and a negative value indicates that active power is absorbed from the grid. The remaining power of the energy storage system is always positive.

[0077] Step (3): Combine the V-type and U-type voltage barrier function models to construct the barrel-type voltage barrier function model;

[0078]

[0079]

[0080] In Equation (6), the V-shaped voltage barrier function has a large gradient in the range of 0.95pu-1.05pu, which can achieve better voltage conditions, but it cannot achieve fine adjustment within the range. In Equation (7), the U-shaped voltage barrier function has a small gradient outside the range of 0.95pu-1.05pu, which cannot achieve rapid entry into the safe range.

[0081] Combining the advantages of both, a barrel-shaped voltage barrier function model is proposed:

[0082]

[0083] In equation (8), v a The real-time voltage magnitude of the node; v ref This is the network voltage reference value, typically taken as 1.00 pu; v (v a This is the real-time reward for the voltage of that node.

[0084] This model combines the advantages of V-type and U-type models: on the one hand, it has a slow gradient within the safe range, which can achieve better voltage conditions; on the other hand, the larger gradient outside the safe range can ensure faster policy guidance.

[0085] Step (4): Using the voltage constraint based on the total power loss of the network and the voltage barrier function model as the target control, the weighted sum algorithm is used to balance the voltage deviation and the network power loss. A spatiotemporal uncertainty model based on photovoltaic and load prediction interval is established. A network power flow constraint model based on the Newton-Raphson method is established. The spatiotemporal uncertainty model based on photovoltaic and load prediction interval is combined with the joint control model based on the reactive power output of the photovoltaic inverter and the active power output of the energy storage system. The control target after balancing the voltage deviation and the network power loss and the network power flow constraint model are established to build a microgrid distributed VVC model.

[0086] Each subnet has two VVC objectives: voltage constraint and network active power minimization, as shown in (9). Among them, (10) represents the voltage deviation, and (11) represents the total power loss of the network at time t.

[0087]

[0088]

[0089]

[0090] In equation (10), l v (v i,t) represents the real-time voltage barrier function value of node i at time t; N represents the number of network nodes; N m For the set of network nodes; in equation (11) For the set of network branches, r ij and x ij Let v represent the resistance and reactance of the branch between nodes i and j, respectively. i,t and v j,t Let i and j represent the voltage amplitudes at time t, respectively.

[0091] Using the classic weighted sum algorithm, the multi-objective function is transformed into an equivalent single-objective function with weighting factors. Here, we normalize these objectives using Utopian points and Nadir points. For the subnetwork m, the weighted sum representation of the normalized objective is obtained as follows:

[0092]

[0093] In the formula, This represents the voltage deviation of the m-network at time t after normalization. Let α represent the normalized active power loss of network m at time t, and let α and β represent the normalization coefficients.

[0094] Establish a spatiotemporal uncertainty model based on photovoltaic and load forecast intervals.

[0095] First, a spatial uncertainty model was established based on the prediction intervals of photovoltaic power generation and load, and then optimized at each decision time step. The time interval in the optimized model varies with time and is related to the real-time measurement at point t.

[0096] Given the locational variations of renewable energy generation and load, as well as short-term intermittency and fluctuations, spatial and temporal uncertainties need to be addressed to ensure operational constraints are met.

[0097] First, a spatial uncertainty model was established based on the predicted photovoltaic power generation time interval and load, limiting the uncertainty that may occur at t0 to within the prediction interval:

[0098]

[0099]

[0100] In the formula, This represents the upper and lower limits of the active and reactive power demand of the inverter's maximum power point / bus i, indicating spatial uncertainty.

[0101] In the spatial uncertainty model described above, communication delays cause inverter response delays. To eliminate the impact of these delays, a time uncertainty model is established, optimizing each decision-making time step, which is expressed as:

[0102]

[0103]

[0104] In the formula, This represents the upper and lower limits of the time uncertainty in the inverter's maximum power point / bus active and reactive power demand. The interval of this time uncertainty model varies with time and is related to the real-time measurement at time t.

[0105] The intervals in the time uncertainty model vary with time and are related to real-time measurements at time t.

[0106] The network power flow constraint model based on the Newton-Raphson method is established as shown in equation (18).

[0107]

[0108] In the formula p i,t and q i,t V represents the active power and reactive power injected by node i at time t, respectively. i,t G represents the voltage magnitude at node i at time t. ij and b ij θ represents the conductance and susceptance of the bus between nodes i and j. ij,t This represents the voltage phase angle difference between node i and node j at time t, and N is the index of all nodes in the microgrid.

[0109] Based on the spatiotemporal uncertainty model of photovoltaic and load prediction interval, the joint control model with reactive power output of photovoltaic inverter as the main component and active power output of energy storage system as the auxiliary component, the voltage deviation and network power loss control objectives after trade-off, and the network power flow constraint model, a distributed VVC model of microgrid is established, as shown in Equation (19).

[0110]

[0111] In equation (19), M is the network partition index; T is the total duration; v i,t This is the real-time voltage.

[0112] Step (5): Fit the microgrid distributed VVC model to a partially observable Markov decision POMG model, and improve the state action equation from adapting to discrete action to adapting to continuous action to meet the real-time control requirements.

[0113] First, set the observation space to include v. i,t , and These represent the real-time node voltage, the active power consumption of the load connected to the node, the reactive power consumption of the load connected to the node, the real-time active power injected into the grid by photovoltaic power, and the real-time remaining energy of energy storage, respectively; the continuous action space includes and These represent the inverter's reactive power output and the energy storage's active power output, respectively.

[0114] The objective of each actor is to maximize their expected return within a time frame T.

[0115]

[0116] The state-action equations have been improved from adapting to discrete actions to adapting to continuous actions in order to meet real-time control requirements.

[0117] State-action functions are used to indicate the current state or the combined return of a state-action pair:

[0118]

[0119] In the formula, τ i This represents the history of agent i; a -i =× j≠i a j To accommodate continuous actions, we changed it to the following form:

[0120]

[0121] In the formula, based on a′ i Gaussian distribution by π i (a′ i ∣τ i () indicates. In practical applications, Approximated by Monte Carlo sampling, it can be rewritten as:

[0122]

[0123] Step (6): The constructed spatiotemporal uncertainty model based on photovoltaic and load prediction intervals is transformed into a stochastic programming model. A network solution failure penalty is added to the return reward of each training set to improve the network solution success rate.

[0124] Spatial uncertainty scenarios are generated using Monte Carlo sampling based on the prediction intervals of equations (13) and (14). Time uncertainty scenarios are then generated using Monte Carlo sampling based on equations (15) and (16), and the probability of these scenarios is used to correct for voltage deviation and network loss. ξ u Let's normalize the probability of scenario u∈U.

[0125]

[0126] In equation (24), U considers U time uncertainty scenarios at each decision time step, ξ u It is the probability of scenario u∈U occurring. This is the sum of normalized values ​​for all cases in stochastic programming (average voltage deviation / network loss of network m at time t). The sum of the normalized values ​​for the case of u.

[0127] Finally, the policy gradient function is given:

[0128]

[0129] To reduce network solution failures during training due to limitations in power grid capacity, a network solution failure penalty function F is added:

[0130] F = -f,t f <T max (25)

[0131] In the formula, f is the penalty occurrence constant, which is a large positive number; t f T represents the duration during which the network successfully solves a problem in the training set, even if the problem is interrupted due to network failure. max This represents the maximum time step for each training set.

[0132] Then the reward for each training set interrupted due to network solution failure is adjusted as follows:

[0133] R mf =R m +F(26)

[0134] In equation (26), R m This refers to the normal reward value obtained before the training set was interrupted.

[0135] Step (7): Using the multi-agent deep deterministic policy gradient (MADDPG) algorithm, design a reasonable algorithm flow and apply it to microgrids.

[0136] The detailed algorithm flow is attached. Figure 2 As shown. During initialization, the energy storage system's capacity is set to 50% of the total capacity; if energy storage is activated, the same activation value as photovoltaic (PV) is used. For each training set, a buffer stores PV and load data for 480 time steps (i.e., 1 day). Furthermore, the reward is obtained based on calculations from the buffer data using the PandaPower software package. Before providing feedback to the agents, the received states are assigned to observations according to each agent's region. Each agent receives only local observations and a global reward before making the next decision. This process is repeated until the end of a single training set.

[0137] A reward is awarded upon completion of each action. A training session comprises multiple training sets, the number of which is manually set as needed. Behavior changes and policy model saving occur every 40 training sets, while the target is updated every 120 sets. Ten test cases are executed before each policy model update. Test data is based on the total data sample, and the average of the tests is used to evaluate the policy's effectiveness.

[0138] Based on the same inventive concept, this invention also provides a computer device, comprising: one or more processors, and a memory for storing one or more computer programs; the programs include program instructions, and the processor executes the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, used to implement one or more instructions, specifically for loading and executing one or more instructions stored in a computer storage medium to implement the above-described method.

[0139] It should be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium storing a computer program, which, when executed by a processor, performs the above-described method. This storage medium can be any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the present invention, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0140] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0141] The foregoing has shown and described the basic principles, main features, and advantages of this disclosure. Those skilled in the art should understand that this disclosure is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of this disclosure. Various changes and modifications can be made to this disclosure without departing from its spirit and scope, and all such changes and modifications fall within the scope of this disclosure as claimed.

Claims

1. A data-driven microgrid voltage control method, characterized in that, The method includes the following steps: The microgrid is divided into multiple sub-networks that are interconnected and coupled through power flow based on a distributed architecture, and an intelligent agent for internal control is built for each sub-network. A joint control model is constructed, with the reactive output of the photovoltaic inverter as the main component and the active output of the energy storage system as the auxiliary component. Each sub-network corresponds to a single intelligent agent that controls all photovoltaic inverters and energy storage devices within the sub-network. Effective control of the microgrid voltage is achieved by controlling the reactive output of the photovoltaic inverter and the active output of the energy storage system. By combining the V-type and U-type voltage barrier function models, a barrel-type voltage barrier function model is constructed. Using voltage constraints based on the total power loss of the network and the voltage barrier function model as the target control, a weighted sum algorithm is used to balance voltage deviation and network power loss. A spatiotemporal uncertainty model based on photovoltaic and load prediction intervals is established, and a network power flow constraint model based on the Newton-Raphson method is established. By combining the spatiotemporal uncertainty model based on photovoltaic and load prediction intervals, a joint control model with reactive power output of photovoltaic inverter as the main component and active power output of energy storage system as the auxiliary component is established. After balancing the control target and network power flow constraint model after voltage deviation and network power loss, a distributed VVC model of microgrid is established. The distributed VVC model of the microgrid is fitted to the partially observable Markov decision POMG model, and the state action equation is improved from adapting to discrete action to adapting to continuous action to meet the real-time control requirements. The constructed spatiotemporal uncertainty model based on photovoltaic and load prediction intervals is transformed into a stochastic programming model, and a network solution failure penalty is added to the return reward of each training set to improve the network solution success rate. The process of adding a penalty for network solution failure to improve the success rate of network solution is as follows: Add a network solution failure penalty function : In the formula, To penalize the occurrence of constants; The duration during which the network successfully solves a problem in the training set that was interrupted due to network solution failure; The maximum time step for each training set; Then the reward for each training set interrupted due to network solution failure is adjusted to: : In the formula, The normal reward value obtained before the training set interruption; The algorithm flow is designed using the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm, and applied to microgrids.

2. The data-driven microgrid voltage control method according to claim 1, characterized in that, The distributed architecture does not require a central coordinator to collect all information generated in the network, but needs to consider global coordination issues. In distributed optimization, each constructed agent represents a sub-network that only exchanges limited boundary physical information and global reward rewards with its neighboring sub-networks, collectively seeking the globally optimal solution. During each operation, the trained controller implements a microgrid distributed VVC model with local measurements within the corresponding sub-network.

3. The data-driven microgrid voltage control method according to claim 1, characterized in that, The joint control model, in which the photovoltaic inverter primarily outputs reactive power and the energy storage system secondarily outputs active power, is as follows: This represents the real-time active power output of the photovoltaic system. This represents the real-time reactive power output of the photovoltaic inverter. This represents the complex power of the photovoltaic inverter. This represents the photovoltaic active power output boundary at time t. Reactive power capacity factor of photovoltaic inverter and This indicates the minimum and maximum active power that the energy storage system can generate and absorb. This represents the active power output value of the stored energy at time t. This represents the real-time capacity of the energy storage at time t. Indicates the maximum capacity of energy storage; Photovoltaic inverters prioritize providing reactive power when needed. When the reactive power compensation capacity is insufficient, the energy storage system will activate. The reactive power output of each photovoltaic inverter is also limited to a preset proportion of its apparent power capacity. A positive value indicates that reactive power is injected into the grid, and a negative value indicates that reactive power is absorbed from the grid. The energy storage system is set up similarly to the inverter. A positive value indicates that active power is injected into the grid, and a negative value indicates that active power is absorbed from the grid. The remaining power of the energy storage system is always positive.

4. The data-driven microgrid voltage control method according to claim 1, characterized in that, The barrel-shaped voltage barrier function model is as follows: In the formula, This refers to the real-time voltage level of the node. As the network voltage reference value, take ; The voltage barrier function value of the node.

5. The data-driven microgrid voltage control method according to claim 1, characterized in that, Control objective after balancing voltage deviation and network power loss: In the formula, Let be the real-time voltage barrier function value of node i at time t; This refers to the number of network nodes. A set of network nodes; For a set of network branches; and These represent the resistance and reactance of the branch between nodes i and j, respectively. and Let i and j represent the voltage amplitudes at time t, respectively. Then, a weighted sum algorithm is used to transform the multi-objective function into an equivalent single-objective function with weighting factors. Utopian points and Nadir points are used to normalize the objective. For any subnetwork m, the weighted sum representation of the normalized objective is obtained as follows: In the formula, This represents the voltage deviation of the m-network at time t after normalization. This represents the normalized active power loss of network m at time t. and This represents the normalization coefficient.

6. The data-driven microgrid voltage control method according to claim 1, characterized in that, The process of establishing a spatiotemporal uncertainty model based on photovoltaic and load forecast intervals is as follows: Before each operating period, spatial uncertainty scenarios are randomly generated within a given prediction interval using Monte Carlo sampling. Next, within each operating cycle, temporal uncertainty scenarios are generated using Monte Carlo sampling, and time delays are considered using the temporal uncertainty interval. Under each temporal uncertainty scenario, voltage deviation and network loss are calculated, and these two objectives are corrected using the scenario occurrence probability. The corrected normalized objective is expressed as: This represents the normalized average voltage deviation / network power loss under scenario u. Let represent the normalized average voltage deviation / network power loss of network m at time t, which is the sum of the normalized values ​​in all cases in stochastic programming. This represents the probability of scenario u occurring.

7. The data-driven microgrid voltage control method according to claim 1, characterized in that, The process of improving the state-action equation from adapting to discrete actions to adapting to continuous actions is as follows: The state-action function is as follows, used to indicate the current state or the combined return of the state-action pair: In the formula, Indicates agent Historical records; To accommodate continuous actions, the state-action function is changed to the following form: In the formula, based on Gaussian distribution by express, Obtained through Monte Carlo sampling, as follows: 。 8. A computer device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When one or more of the programs are executed by one or more of the processors, the one or more of the processors implement a data-driven microgrid voltage control method as described in any one of claims 1-7.

9. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform a data-driven microgrid voltage control method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Power distribution network voltage autonomous optimization control method and device

    CN113872213A

  • Voltage control method for cooperative mutual supply of alternating current and direct current micro-grid groups

    CN114421479A