A power distribution network voltage control method based on reactive virtual power plant
By constructing a two-layer collaborative control model for reactive power virtual power plants and combining multi-agent reinforcement learning for discrete and continuous reactive power regulation equipment, the voltage stability problem of reactive power virtual power plants in complex distribution networks is solved, and voltage optimization control across time scales is realized, thereby improving the stability of the distribution network and the lifespan of equipment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
- Filing Date
- 2026-03-05
- Publication Date
- 2026-05-08
AI Technical Summary
In complex distribution networks with a high proportion of distributed power sources, existing reactive power virtual power plants lack cross-device and cross-time scale collaborative response mechanisms, making it difficult to control node voltage stably.
A two-layer collaborative control model based on a reactive power virtual power plant is constructed, including a day-ahead dispatch layer and an intraday correction layer. Combining discrete and continuous reactive power regulation equipment, optimization control is performed through multi-agent reinforcement learning to generate voltage dispatch plans and correction commands, thereby achieving voltage optimization control across time scales.
It effectively reduces node voltage deviation, improves distribution network voltage stability, extends equipment lifespan, enhances resource regulation coordination and system robustness, and ensures the adaptability and flexibility of the power grid in multiple scenarios.
Smart Images

Figure CN121791344B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart grid and distribution network control, and in particular to a distribution network voltage control method based on a reactive power virtual power plant. Background Technology
[0002] With the continued advancement of "dual carbon" (carbon reduction and emission reduction), building a clean, low-carbon, safe, and efficient new power system has become an important direction for energy transformation. Against this backdrop, virtual power plants, as a key technological path for coordinating distributed resources and enhancing system regulation capabilities, are increasingly demonstrating their important role in improving energy system operating efficiency, promoting the consumption of new energy sources, and driving energy conservation and emission reduction. Reactive power virtual power plants, as a special form of virtual power plant, focus on the coordination and management of reactive power. They can effectively integrate multiple types of distributed reactive power resources, construct a coordinated reactive power regulation system of "source-grid-load-storage," thereby improving the voltage stability of the distribution network, enhancing the flexibility of system operation, and promoting the green and low-carbon development of the power system.
[0003] Currently, research on reactive power virtual power plants (VFPs) is still in its early stages both domestically and internationally, especially under conditions of high-proportion renewable energy integration, where their optimization and control mechanisms are still immature. Existing methods mostly employ centralized control architectures, which, while effective in small-scale systems, have significant limitations in scalability, dynamic adaptability, and data privacy protection when facing the actual needs of surging distributed resources and increasingly complex grid structures. Specifically, current reactive power control strategies still lack dynamic matching mechanisms between heterogeneous resources, making it difficult to achieve coordinated responses across devices and time scales under uncertain disturbances. This limits the effectiveness of reactive power virtual power plants in controlling node voltages in distribution networks on a large scale. Summary of the Invention
[0004] Based on the above analysis, the present invention aims to provide a distribution network voltage control method based on a reactive power virtual power plant, in order to solve the technical problem that the existing reactive power virtual power plant control method in complex distribution networks with a high proportion of distributed power sources is difficult to control stably due to the lack of a collaborative response mechanism across devices and time scales.
[0005] This invention provides a distribution network voltage control method based on a reactive power virtual power plant, comprising the following steps:
[0006] S1. Based on discrete and continuous reactive power regulation equipment in the distribution network under reactive power virtual power plant management, construct a two-layer collaborative control model including day-ahead scheduling layer and intraday correction layer, and construct the physical constraints of the distribution network.
[0007] S2. Under the physical constraints of the distribution network, the day-ahead scheduling layer transforms the optimal control of discrete reactive power regulation equipment into a first multi-agent reinforcement learning solution for discrete action space strategy, and generates a day-ahead scheduling plan which is then sent to the intraday correction layer.
[0008] S3. Under the dual constraints of the distribution network physical constraints and the day-ahead scheduling plan as boundary constraints, the intraday correction layer transforms the optimization control of the continuous reactive power regulating equipment into a second multi-agent reinforcement learning solution for a continuous action space strategy to generate intraday correction instructions; based on the intraday correction instructions, the continuous reactive power regulating equipment is adjusted to obtain the node voltage in the distribution network; the node voltage in the distribution network is fed back to the day-ahead scheduling layer for closed-loop coordination to achieve voltage optimization control of the distribution network;
[0009] In this case, a distribution network is managed by at least one reactive virtual power plant.
[0010] Furthermore, the discrete reactive power regulation device includes at least one of an on-load tap-changing transformer (OLTC), a capacitor bank (CB), and a voltage regulator (VR);
[0011] The continuous reactive power regulation equipment includes at least one of a photovoltaic inverter and an energy storage inverter.
[0012] Furthermore, physical constraints for the distribution network are established based on the DistFlow power flow equation; these physical constraints include node power balance constraints and voltage balance constraints.
[0013] Further, step S2 includes:
[0014] Based on the physical constraints of the distribution network, a day-ahead scheduling optimization objective function is constructed with the goal of minimizing the total voltage deviation of nodes in the distribution network within the scheduling cycle; and a constraint on the maximum allowable operating frequency of each discrete reactive power regulating device within the scheduling cycle is constructed.
[0015] The day-ahead scheduling optimization is modeled as a Markov game process, and a corresponding discrete agent is defined for each discrete reactive power regulation device to obtain the first multi-agent; wherein, the action space of each agent is a discrete gear or switching operation, the state space includes the current state of the device and the cumulative number of actions, and the reward function includes a comprehensive evaluation of the node voltage deviation in the distribution network, the cost of device actions and the voltage over-limit situation.
[0016] The MASAC-Discrete algorithm is used to perform reinforcement learning on the first multi-agent system to obtain the agent cooperative strategy for the discrete reactive power regulation device. ;
[0017] Based on the agent-based collaborative strategy of the discrete reactive power regulation equipment, a day-ahead scheduling plan is generated, which includes the gear position or switching status of each discrete reactive power regulation equipment in the future scheduling cycle.
[0018] Furthermore, the day-ahead scheduling optimization objective function is as follows:
[0019]
[0020] in, For time period index, Number of time periods; For reactive power virtual power plant index, This represents the total number of virtual reactive power plants. For node indexes in the distribution network, For reactive virtual power plants The set of nodes in the distribution network under its jurisdiction; In order to be in The time period belongs to the node in the distribution network of the reactive virtual power plant. The per-unit voltage value.
[0021] Furthermore, under the physical constraints of the distribution network, the maximum permissible operating frequency constraints for each discrete reactive power regulating device at the day-ahead dispatching level within the dispatching cycle are as follows:
[0022]
[0023] in, , and Variables are 0 or 1; δ OLTC ( t (for the time period) Indicator variable for whether an on-load tap changer (OLTC) has operated. M OLTC In order to be within the planning period OLTC Maximum number of actions allowed; δ CB ( t (for the time period) The indicator variable indicating whether a capacitor bank (CB) switching operation has occurred. M CB This represents the maximum number of actions allowed for CB within the planning period. For the time period The indicator variable indicating whether a voltage regulator VR level adjustment action has occurred. M VR This represents the maximum number of actions allowed in VR within the planning period.
[0024] Furthermore, the agent-based cooperative strategy of the discrete reactive power regulation device Including OLTC gear adjustment amount and CB switching status change amount VR level adjustment ,as follows:
[0025]
[0026] The local observation state space of a discrete agent includes: during the current scheduling period During which time period, the tap position of the on-load tap-changing transformer (OLTC) Switching status of capacitor bank CB Voltage regulator VR gear position Up to Cumulative number of actions in OLTC during the time period CB's cumulative number of actions Cumulative number of actions in VR ,as follows:
[0027]
[0028] reward function ,as follows:
[0029]
[0030] in, It represents the sum of squares of voltage deviations across the entire distribution network and is used to measure voltage quality; express The motion vector for the time period includes OLTC gear shift adjustment, CB casting / switching status, and VR gear shift adjustment, with a L2 norm. Characterizes the operating cost of the equipment; This is a voltage over-limit indication function, which takes a value of 1 when the voltage exceeds the allowable range, and 0 otherwise. , They are respectively and Weighting coefficients;
[0031] The day-ahead scheduling plan is expressed as follows:
[0032]
[0033] in, T OLTC ( t ), S CB ( t ) and L VR ( t) respectively in the current The time period is determined by the day-ahead discrete layer. OLTC tap position, CB Throwing and Cutting Status and VR The actual value of the gear adjustment.
[0034] Further, step S3 includes:
[0035] Based on the physical constraints of the distribution network, the day-ahead scheduling plan as boundary conditions, and the set of uncertainty scenarios, an intraday correction optimization objective function is constructed with the goal of minimizing the expected voltage deviation.
[0036] The intraday correction optimization is modeled as a Markov game process, and a corresponding continuous agent is defined for each continuous reactive power regulation device to obtain a second agent; wherein, the action space of each agent is continuous reactive power output, the state space includes local active or reactive load, and photovoltaic active power output, and the reward function includes node voltage deviation in the distribution network and penalty for the continuous reactive power regulation device's action to violate the corresponding safe operation constraints.
[0037] The MASAC algorithm is used to perform reinforcement learning on the second agent to obtain the cooperative control strategy of the agent in the continuous reactive power regulation equipment. ;
[0038] Based on the collaborative control strategy of the intelligent agent of the continuous reactive power regulation equipment, intraday correction instructions for each continuous reactive power regulation equipment are generated.
[0039] The continuous reactive power regulation equipment is adjusted based on the intraday adjustment command to obtain the node voltage of the distribution network under the reactive power virtual power plant management.
[0040] The node voltage of the distribution network is compared with the node voltage of the distribution network corresponding to the day-ahead scheduling plan to obtain node voltage deviation information; the day-ahead scheduling layer re-performs the first multi-agent reinforcement learning solution based on the node voltage deviation information to regenerate the day-ahead scheduling plan.
[0041] Furthermore, based on the scenario set Ω including uncertainties in source, grid, load, and storage, the optimization objective function based on the intraday correction layer of Ω is as follows:
[0042]
[0043] in, ω An index for scenarios with uncertainties; Let ω represent the expected value of all possible scenarios in the scenario set Ω for uncertain factors;
[0044] Uncertainties in each scenario include fluctuations in the output of distributed power sources on the source side, disturbances in line parameters and topology on the grid side, random changes in load on the load side, and deviations in SOC and efficiency on the energy storage side.
[0045] Power constraints of continuous reactive power regulation equipment in the intraday correction layer, including power constraints of photovoltaic inverters and power constraints of energy storage inverters;
[0046] The power constraints of the photovoltaic inverter are as follows:
[0047]
[0048] in, and For photovoltaic inverters t Active power output and reactive power output during a given time period; For photovoltaic systems in t Maximum active power output limit in maximum power point tracking mode; This refers to the rated apparent power of the photovoltaic inverter. The minimum power factor allowed at the photovoltaic grid connection point;
[0049] The power constraints of the energy storage inverter are as follows:
[0050]
[0051] in, and For energy storage inverters t Active and reactive power output during a given time period; and These are the maximum discharge power and maximum charging power of the energy storage inverter, respectively. This is the rated apparent power of the energy storage inverter.
[0052] Furthermore, the collaborative control strategy of intelligent agents in continuous reactive power regulation equipment. Including the reactive power output of photovoltaic inverters and reactive power of energy storage inverter ,as follows:
[0053]
[0054] in, For a continuous intelligent agent, it represents the local observation state space.
[0055] Including active load during time period t reactive load and photovoltaic active power output ,as follows:
[0056]
[0057] Each continuous agent reward function ,as follows:
[0058]
[0059] in, For the number of continuous agents, For continuous intelligent agents The set of nodes in the distribution network. For nodes exist Per-unit value of voltage amplitude during the time period For continuous intelligent agents The penalty coefficient for violations of the action.
[0060] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:
[0061] 1. This invention achieves multi-timescale voltage support of "day-ahead scheduling planning - intraday correction" by introducing OLTC, CB, and VR discrete reactive power regulation devices in the discrete layer for optimized control, combined with the rapid response of photovoltaic inverters and energy storage inverters in the continuous layer. This effectively reduces node voltage deviation, improves power quality, and enhances the stability of distribution network voltage.
[0062] 2. The present invention introduces equipment action costs and constraints into the current scheduling layer optimization objective, which avoids frequent switching of discrete reactive power regulation equipment such as OLTC, CB, and VR, reduces the number of equipment actions, extends the service life of key voltage regulation equipment, and reduces operation and maintenance costs.
[0063] 3. This invention introduces uncertainty scenario modeling into the intraday correction layer, fully considering photovoltaic distributed power output fluctuations, load forecasting errors and network disturbances, so that the learned agent cooperative strategy has strong adaptability and robustness in multiple scenarios.
[0064] 4. This invention constructs a hierarchical optimization mechanism that complements discrete reactive power regulation equipment and continuous reactive power regulation equipment. Discrete reactive power regulation equipment provides day-ahead scheduling planning, while continuous reactive power regulation equipment provides short-term real-time flexible correction, thereby realizing efficient coordination and comprehensive utilization of reactive power resources.
[0065] 5. This invention combines a centralized training and distributed execution CTDE framework, and adopts a combination of offline training and online rolling optimization. The distribution network can continuously update and correct the control strategy in actual operation, ensuring long-term stability and flexibility, and achieving self-adaptation and continuous voltage optimization.
[0066] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description
[0067] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.
[0068] Figure 1 This is a flowchart of a distribution network voltage control method based on a reactive power virtual power plant in an embodiment of the present invention;
[0069] Figure 2 This is a schematic diagram illustrating how the reward function converges continuously as the algorithm trains in the test scenario of this embodiment of the invention.
[0070] Figure 3 The reactive power output curves of the energy storage inverter, photovoltaic inverter, and capacitor bank under the test scenario in this embodiment of the invention are shown.
[0071] Figure 4 This is a schematic diagram illustrating the gear changes of the voltage regulating device in a test scenario according to an embodiment of the present invention;
[0072] Figure 5 This is a schematic diagram of the functional modules of a distribution network voltage control system based on a reactive power virtual power plant in an embodiment of the present invention. Detailed Implementation
[0073] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.
[0074] Example 1:
[0075] To address the aforementioned technical challenges, there is an urgent need to establish a reactive power optimization and control framework for large-scale distributed environments, possessing hierarchical collaborative capabilities and adaptive learning capabilities, to achieve coordinated regulation of different types of resources at different time scales. This invention addresses this critical need by disclosing a distribution network voltage control method based on a reactive power virtual power plant. It employs a dual-time-scale optimization control mechanism of "day-ahead" and "intra-day," utilizing traditional reactive power regulation equipment with discrete control characteristics for strategy formulation during the day-ahead phase, and inverter-type equipment with continuous control capabilities for real-time intra-day correction. Furthermore, it leverages deep reinforcement learning to achieve regulation strategy learning and rolling updates under complex operating environments. This method ensures stable voltage at distribution network nodes while enhancing the synergy of resource regulation, the robustness of distribution network operation, and the scalability of the algorithm, demonstrating broad engineering application prospects and significant theoretical research value.
[0076] The present invention aims to optimize the coordinated regulation of reactive power equipment through intelligent decision-making, so as to minimize the overall voltage deviation of the distribution network and achieve optimized control of the distribution network voltage.
[0077] A specific embodiment of the present invention discloses a distribution network voltage control method based on a reactive power virtual power plant, such as... Figure 1 As shown, it includes the following steps:
[0078] Step S1: Based on the discrete and continuous reactive power regulation equipment in the distribution network under the reactive power virtual power plant management, construct a two-layer collaborative control model including the day-ahead scheduling layer and the intraday correction layer, and construct the physical constraints of the distribution network.
[0079] Step S2: Under the physical constraints of the distribution network, the day-ahead scheduling layer transforms the optimal control of discrete reactive power regulation equipment into a first multi-agent reinforcement learning solution for discrete action space strategy, and generates a day-ahead scheduling plan which is then sent to the intraday correction layer.
[0080] Step S3: Under the dual constraints of the distribution network physical constraints and the day-ahead scheduling plan as boundary constraints, the intraday correction layer transforms the optimized control of the continuous reactive power regulating equipment into a second multi-agent reinforcement learning solution for a continuous action space strategy, generating intraday correction instructions; based on the intraday correction instructions, the continuous reactive power regulating equipment is adjusted to obtain the node voltage in the distribution network; the node voltage in the distribution network is fed back to the day-ahead scheduling layer for closed-loop coordination, realizing the voltage optimization control of the distribution network;
[0081] In this case, a distribution network is managed by at least one reactive virtual power plant.
[0082] Step S1, specifically.
[0083] This step is the foundational modeling stage for constructing the entire reactive power virtual power plant voltage control method, aiming to provide a complete mathematical model and physical boundaries for subsequent agent learning and collaborative optimization. The specific implementation includes the following three key parts:
[0084] (a) Identify the control targets: discrete and continuous reactive power regulation equipment
[0085] The discrete reactive power regulation equipment includes at least one of an on-load tap-changing transformer (OLTC), a capacitor bank (CB), and a voltage regulator (VR).
[0086] The continuous reactive power regulation equipment includes at least one of a photovoltaic inverter and an energy storage inverter.
[0087] First, the control resources aggregated by the reactive power virtual power plant are defined and divided into two categories based on their control characteristics:
[0088] (1) Discrete reactive power regulation equipment: refers to equipment whose control output is a step change or a step-by-step adjustment, mainly including:
[0089] On-load tap changer (OLTC): By adjusting the tap position, the transformer ratio is changed, thereby continuously or in stages adjusting the voltage of the connected bus.
[0090] Capacitor Bank (CB): A device that injects or absorbs discrete reactive power into the power grid by switching between groups of capacitors using mechanical or power electronic switches; Voltage Regulator (VR): A series or parallel voltage regulating device whose regulating output (such as compensation voltage or injection current) has discrete levels.
[0091] The common characteristics of discrete reactive power regulation equipment are that the control actions (such as "upgrade one level" or "put a group into operation") are discrete events, and the number of actions per day / per scheduling cycle is strictly limited by the mechanical life and operating procedures.
[0092] (2) Continuous reactive power regulation equipment: refers to equipment whose control output can be continuously and smoothly adjusted within a certain range, mainly including:
[0093] Photovoltaic inverters: Under the premise of prioritizing the output of maximum active power, they utilize their remaining capacity to generate or absorb reactive power in real time and continuously through power electronic converters.
[0094] Energy storage inverter: While meeting the active power charging and discharging plan, it uses its inverter capacity to continuously regulate reactive power.
[0095] The advantages of continuous reactive power regulation equipment are fast response speed and high regulation accuracy, but its reactive power regulation capability is constrained by its current active power output, rated apparent power and operating boundary.
[0096] (II) Constructing a two-layer collaborative regulation model architecture
[0097] To address the limitations of single-timescale control, this invention constructs a two-layer collaborative model consisting of a "day-ahead scheduling layer and an intraday correction layer":
[0098] The day-ahead dispatch layer (upper layer) operates on a 24-hour dispatch cycle, with a time resolution typically of 15 minutes or 1 hour. This layer focuses on optimizing discrete reactive power regulation equipment. Its core task is to develop a global, coarse-grained voltage control "skeleton" plan based on the load and renewable energy forecasts for the next day, while meeting long-term constraints such as the number of equipment operations. This plan provides a stable voltage reference and reactive power support framework for intraday operation.
[0099] Intraday Correction Layer (Lower Layer): This layer performs real-time or near-real-time control with a rolling optimization cycle of 5-15 minutes. It focuses on optimizing continuous reactive power regulation equipment. Its core task is to rapidly respond to real-time voltage fluctuations caused by prediction errors, load surges, etc., based on ultra-short-term, more accurate measurement and forecasting data. Within the framework set by the day-ahead layer, it performs meticulous, fine-tuning of the voltage to ensure that it always meets operational standards.
[0100] The two layers collaborate through constraint transmission and information feedback, forming an organic whole of "daily scheduling and planning framework with intraday rolling correction".
[0101] (III) Constructing the physical constraints of the distribution network
[0102] To ensure that all control decisions are physically feasible, a mathematical model of the distribution network's operating behavior, i.e., power flow constraints, is constructed. This invention uses the DistFlow power flow equation, or its linearized / simplified form, applicable to radial distribution networks, as the core physical constraint.
[0103] OLTC, CB, and VR are discrete reactive power regulation devices. Their common characteristic is that their output or state is segmented or switched, making continuous and smooth regulation impossible. This contrasts with continuous reactive power regulation devices such as photovoltaic inverters and energy storage inverters, which can be continuously adjusted. This invention aims to coordinate the control of these devices with varying discrete and continuous characteristics.
[0104] The physical constraints of the distribution network are established based on the DistFlow power flow equation; the physical constraints of the distribution network include node power balance constraints and voltage balance constraints.
[0105] Before performing optimization modeling for the discrete day-ahead scheduling layer and the continuous intraday correction layer, it is first necessary to establish a basic physical constraint model of the distribution network. To ensure the physical feasibility of subsequent control strategies, the operating characteristics of the distribution network are characterized based on the DistFlow power flow equations, including node power balance constraints and voltage balance constraints, as shown below:
[0106] Formula (1)
[0107] in, 、 They represent the times respectively. t nodes in the distribution network i to downstream nodes j The active and reactive power of the branch; , They represent the times respectively. t node j Net productive effort and net unproductive effort at each location; Represents a node j The set of child nodes; , These represent all child nodes from node j to its downstream end in the distribution network. Active power and reactive power; , They are nodes i and j At any moment t The per-unit voltage value; , They are nodes i To the node j The branch resistance and reactance.
[0108] The first and second formulas of formula (1) represent the node power balance constraint; for any node j, the difference between the power flowing into the node and the power flowing out of the node is equal to the net load of that node.
[0109] The third formula in formula (1) represents the voltage balance constraint, which describes the voltage drop between node i and node j. This voltage drop is related to the active and reactive power flow of the branch.
[0110] Formula (1) power balance equation ensures that the power difference between the inflow and outflow at any node is equal to the net load of that node, while voltage balance equation characterizes the branch. i arrive j The relationship between voltage drop and branch power is investigated. Based on the distribution network topology and line parameters, physical constraint equations for node power and voltage are established to address the infeasibility of voltage exceeding limits and power imbalance in subsequent optimization, thus providing a computational basis to ensure the physical feasibility of the control strategy.
[0111] Step S1 establishes a basic model for the voltage control of the reactive virtual power plant, which is constrained by physical laws, clearly distinguishes between discrete and continuous equipment, and is based on the "day-to-day" dual time scale for coordinated control.
[0112] Based on the physically feasible region defined by the DistFlow equation, it is further necessary to model the discrete reactive power regulation equipment in the distribution network under the management of the reactive power virtual power plant, i.e., step S2. This step is based on the voltage demand of distribution network nodes and the operating characteristics of discrete reactive power regulation equipment, aiming to solve the problems of excessive node voltage deviation and over-operation of reactive power regulation equipment in the distribution network, thereby obtaining an optimized model that can maintain power quality in the day-ahead stage, i.e., step S3.
[0113] To optimize reactive power and maintain power quality, reactive power regulation equipment within the distribution network should be controlled in a coordinated manner.
[0114] This invention constructs a reactive power virtual power plant framework based on a "source-grid-load-storage" collaborative mechanism, aggregates various distributed reactive power regulation devices, and introduces a "dual time scale + dual-layer learning control" mechanism to design different control strategies for traditional discrete reactive power regulation devices with discrete action characteristics and inverter reactive power regulation devices with continuous regulation capabilities.
[0115] This invention uses a reactive power virtual power plant as the control core, integrating distributed photovoltaic, energy storage systems, and distribution network voltage regulation equipment. It employs a two-layer, multi-time-scale collaborative decision-making architecture, divided into a discrete equipment layer and a continuous equipment layer, and uses the MASAC algorithm and the MASAC-Discrete algorithm for optimization decisions respectively.
[0116] The discrete equipment layer, also known as the day-ahead dispatch layer, is used for global voltage dispatch on medium- and long-term time scales for discrete reactive power regulation equipment such as on-load tap changer (OLTC), capacitor bank (CB), and voltage regulator (VR).
[0117] The continuous equipment layer, also known as the intraday correction layer, is designed for continuous reactive power regulation equipment in photovoltaic inverters and energy storage inverters. It has high-frequency dynamic regulation capabilities and is used to quickly respond to voltage fluctuations and achieve fine voltage control during intraday operation.
[0118] The two-layer equipment achieves multi-timescale coordination through the central coordinator of the reactive virtual power plant, forming a closed-loop control strategy of day-ahead decision-making and intraday correction.
[0119] Step S2, specifically.
[0120] Step S2 includes:
[0121] Based on the physical constraints of the distribution network, an objective function for day-ahead scheduling optimization is constructed with the goal of minimizing the total voltage deviation of nodes in the distribution network within the scheduling cycle; and a constraint on the maximum allowable operating frequency of each discrete reactive power regulating device within the scheduling cycle is constructed.
[0122] The day-ahead scheduling optimization is modeled as a Markov game process, and a corresponding discrete agent is defined for each discrete reactive power regulation device to obtain the first multi-agent; wherein, the action space of each agent is a discrete gear or switching operation, the state space includes the current state of the device and the cumulative number of actions, and the reward function includes a comprehensive evaluation of the node voltage deviation in the distribution network, the cost of device actions and the voltage over-limit situation.
[0123] The MASAC-Discrete algorithm is used to perform reinforcement learning on the first multi-agent system to obtain the agent cooperative strategy for the discrete reactive power regulation device. ;
[0124] Based on the agent-based collaborative strategy of the discrete reactive power regulation equipment, a day-ahead scheduling plan is generated, which includes the gear position or switching status of each discrete reactive power regulation equipment in the future scheduling cycle.
[0125] During the daytime scheduling phase, an optimization model based on minimizing voltage deviation is constructed to schedule the action plans of discrete devices such as OLTC, capacitor banks, and voltage regulators.
[0126] During the daytime operation phase, voltage data is collected in real time, and a reinforcement learning method is used to generate control strategies for continuous equipment.
[0127] A two-layer policy learning mechanism is implemented using a deep reinforcement learning framework. A centralized training and distributed execution architecture is adopted to improve the adaptability and generalization of the control policy. An uncertainty modeling mechanism for source-network-load-storage collaboration is constructed to improve system robustness. A device frequency protection mechanism and a rolling learning update mechanism are introduced to ensure policy deployability and device reliability.
[0128] The objective function for the discrete equipment layer in the day-ahead optimization phase is set to minimize the total voltage deviation of the distribution network nodes.
[0129] The objective function for day-ahead scheduling optimization is as follows:
[0130] Formula (2)
[0131] in, For time period index, Number of time periods; For reactive power virtual power plant index, This represents the total number of virtual reactive power plants. For node indexes in the distribution network, For reactive virtual power plants The set of nodes in the distribution network under its jurisdiction; In order to be in The time period belongs to the node in the distribution network of the reactive virtual power plant. The per-unit voltage value.
[0132] For example, a scheduling cycle is one day, and each sampling period The duration is 15 minutes; T represents the total number of control periods within a scheduling cycle, which consists of 96 sampling periods. The 1 in the figure represents the per-unit value of the rated voltage.
[0133] Per-unit value is a normalized representation method for power systems. A per-unit value of 1.0 represents the rated voltage; a value higher than 1.0 indicates overvoltage; and a value lower than 1.0 indicates undervoltage.
[0134] Under the physical constraints of the distribution network, the maximum allowable operating frequency constraints for each discrete reactive power regulating device at the day-ahead dispatching level within the dispatching cycle are as follows:
[0135] Formula (3)
[0136] in, , and Variables are 0 or 1; δ OLTC ( t (for the time period) Indicator variable for whether an on-load tap changer (OLTC) has operated. M OLTC In order to be within the planning period OLTC Maximum number of actions allowed; δ CB ( t (for the time period) The indicator variable indicating whether a capacitor bank (CB) switching operation has occurred. M CB This represents the maximum number of actions allowed for CB within the planning period. For the time period Whether an indicator variable has been generated indicating a voltage regulator VR level adjustment action. M VR This represents the maximum number of actions allowed in VR within the planning period.
[0137] δ OLTC ( t ) indicates at time This is an indicator variable indicating whether an on-load tap changer (OLTC) has operated. It is set to 1 if an operation has occurred, and 0 otherwise. M OLTC Indicates within the planning period OLTC Maximum number of actions allowed; The number of times the OLTC can be operated per day is limited to avoid frequent switching affecting equipment lifespan and grid stability.
[0138] δ CB ( t ) indicates at time This is an indicator variable indicating whether a capacitor bank CB switching operation has occurred. It is set to 1 if an operation has occurred, and 0 otherwise. M CB This indicates the maximum number of actions allowed for the CB within the planning period; This limits the operation of the CB to this mutual interaction within a day in order to protect equipment and maintain grid stability;
[0139] Indicates at time This is an indicator variable indicating whether a voltage regulator VR level adjustment action has occurred. It is set to 1 if an action has occurred, and 0 otherwise. M VR This indicates the maximum number of actions allowed for VR within the planning period; The number of times VR can be used in a day is limited to reduce wear and tear on the device and power grid disturbances.
[0140] Formula (3) protects discrete voltage regulating equipment from damage caused by overuse by limiting the frequency of operation of the equipment within the planned scheduling cycle, while ensuring the smooth voltage regulation process of the distribution network and avoiding voltage fluctuations caused by frequent operation, thereby improving the stability and reliability of the power grid. This limitation helps to extend the service life of the equipment, reduce maintenance costs, and ensure that the power grid can operate safely and efficiently.
[0141] Building upon the previously established objective function and constraints at the scheduling layer, this step further transforms them into a reinforcement learning-solvable agent model. Since the presence of other agents renders the environment non-static from any agent's perspective, Markov game theory is introduced for modeling to ensure the learning and convergence of each agent within the interactive environment.
[0142] The agent-based cooperative strategy for discrete reactive power regulation equipment Including OLTC gear adjustment amount and CB switching status change amount VR level adjustment ,as follows:
[0143] Formula (4)
[0144] The local observation state space of a discrete agent includes: during the current scheduling period During which time period, the tap position of the on-load tap-changing transformer (OLTC) Switching status of capacitor bank CB Voltage regulator VR gear position Up to Cumulative number of actions in OLTC during the time period CB's cumulative number of actions Cumulative number of actions in VR ,as follows:
[0145] Formula (5)
[0146] reward function ,as follows:
[0147] Formula (6)
[0148] in, It represents the sum of squares of voltage deviations across the entire distribution network and is used to measure voltage quality. express The motion vector for the time period includes OLTC gear shift amount, CB casting / switching status, and VR gear shift amount, with L2 norm. Characterizes the operating cost of the equipment; This is a voltage over-limit indication function, which takes a value of 1 when the voltage exceeds the allowable range, and 0 otherwise. , They are respectively and Weighting coefficients;
[0149] The day-ahead scheduling plan is expressed as follows:
[0150] Formula (7)
[0151] in, T OLTC ( t ), S CB ( t ) and L VR ( t ) respectively in the current The time period is determined by the day-ahead discrete layer. OLTC tap position, CB Throwing and Cutting Status and VR The actual value of the gear adjustment.
[0152] To overcome interference and ensure good algorithm convergence, the reward function of the day-ahead scheduling layer is mainly used to characterize the regulation effect of discrete devices such as OLTC, CB, and VR during the day-ahead phase. This function, based on node voltage deviation, device action costs, and voltage limit exceedance scenarios, aims to address how to reduce device operating frequency and avoid voltage limit exceedances while ensuring voltage quality.
[0153] Weighting coefficient λ 1. λ 2. The importance of adjusting equipment action costs and over-limit penalties in the reward function; Formula (6) guides discrete reactive power regulation equipment to form a balanced control strategy through triple constraints of voltage quality, equipment action costs, and safe operation. The goal is to generate a day-ahead global control skeleton, that is, to complete the voltage support of the distribution network with as few equipment actions as possible while ensuring voltage stability and safety.
[0154] Through the above processing, a complete discrete-layer intelligent agent modeling framework is obtained. Combined with action definition, state observation, and a unified reward function, this enables OLTC, CB, and VR discrete reactive power regulation devices to generate optimal control strategies within a reinforcement learning framework. The technological advantage lies in transforming the complex voltage control problem into a learnable intelligent agent interaction process, thereby improving the adaptability and robustness of the optimization.
[0155] After the day-ahead scheduling layer generates the day-ahead scheduling plan, this plan is passed to the intraday correction layer as the boundary condition for its operation. This step, based on the voltage regulation results and no-function constraints of the discrete layer, aims to solve the problems of unclear boundaries and potential deviations from the global objective in the intraday optimization process, thereby obtaining an intraday operation framework that conforms to the day-ahead scheduling plan while also possessing flexibility.
[0156] Equation (7) ensures that intraday correction is based on the day-ahead dispatch plan set by the discrete reactive power regulation equipment issued a day before. The optimized control of continuous reactive power regulation equipment must be executed within the day-ahead dispatch plan value obtained by the day-ahead dispatch, and must not violate the voltage reference range and reactive power regulation capacity boundary set by the higher-level dispatch, thereby ensuring the hierarchy and consistency of control decisions. Through this mechanism, a top-down multi-time-scale constraint transmission is formed: the day-ahead dispatch provides the boundary for intraday correction optimization, and the intraday correction optimization achieves rapid dynamic response within this boundary. Its technical role is to avoid conflicts between control strategies of different time scales and ensure the coordination and stability of the distribution network system operation.
[0157] Step S2 is to model the day-ahead scheduling optimization of discrete reactive power regulation equipment as a multi-agent reinforcement learning problem, solve it to obtain the voltage-optimal cooperative strategy under the constraint of the number of equipment actions, and generate the day-ahead scheduling plan to guide the intraday correction layer.
[0158] Step S3, specifically.
[0159] Step S3 includes:
[0160] Based on the physical constraints of the distribution network, the day-ahead scheduling plan as boundary conditions, and the set of uncertainty scenarios, an intraday correction optimization objective function is constructed with the goal of minimizing the expected voltage deviation.
[0161] The intraday correction optimization is modeled as a Markov game process, and a corresponding continuous agent is defined for each continuous reactive power regulation device to obtain a second agent; wherein, the action space of each agent is continuous reactive power output, the state space includes local active or reactive load, and photovoltaic active power output, and the reward function includes node voltage deviation in the distribution network and penalty for the continuous reactive power regulation device's action to violate the corresponding safe operation constraints.
[0162] The MASAC algorithm is used to perform reinforcement learning on the second agent to obtain the cooperative control strategy of the agent in the continuous reactive power regulation equipment. ;
[0163] Based on the collaborative control strategy of the intelligent agent of the continuous reactive power regulation equipment, intraday correction instructions for each continuous reactive power regulation equipment are generated.
[0164] The continuous reactive power regulation equipment is adjusted based on the intraday adjustment command to obtain the node voltage of the distribution network under the reactive power virtual power plant management.
[0165] The node voltage of the distribution network is compared with the node voltage of the distribution network corresponding to the day-ahead scheduling plan to obtain node voltage deviation information; the day-ahead scheduling layer re-performs the first multi-agent reinforcement learning solution based on the node voltage deviation information to regenerate the day-ahead scheduling plan.
[0166] After determining the day-ahead scheduling plan between the day-ahead scheduling layer and the intraday correction layer, it is necessary to further consider the impact of uncertainties in various aspects such as power sources, grid, load, and storage on continuous-layer optimization. This step is based on the fluctuations in distributed photovoltaic output, changes in available energy storage capacity, randomness of load power, and uncertainties in grid operating conditions in the distribution network operation environment. It aims to solve the technical problems of traditional optimization models lacking robustness and being unable to cope with disturbances in multiple scenarios, thereby obtaining a continuous-layer optimization framework that can maintain stability under complex operating conditions.
[0167] The uncertainty scenario set Ω characterizes various possible operating states of the distribution network. By introducing the scenario set Ω into the optimization objective function of the intraday correction layer, the optimization of the intraday correction layer is transformed into a robust multi-scenario optimization problem under the desired meaning, ensuring that the voltage support strategy of the distribution network remains feasible and effective under various disturbances.
[0168] Based on the scenario set Ω including uncertainties in source, network, load, and storage, the optimization objective function based on the intraday correction layer of Ω is as follows:
[0169] Formula (8)
[0170] in, ω An index for scenarios with uncertainties; Let ω represent the expected value of all possible scenarios in the scenario set Ω for uncertain factors;
[0171] Uncertainties in each scenario include fluctuations in the output of distributed power sources on the source side, disturbances in line parameters and topology on the grid side, random changes in load on the load side, and deviations in SOC and efficiency on the energy storage side.
[0172] Power constraints of continuous reactive power regulation equipment in the intraday correction layer, including power constraints of photovoltaic inverters and power constraints of energy storage inverters;
[0173] The power constraints of the photovoltaic inverter are as follows:
[0174] Formula (9)
[0175] in, and For photovoltaic inverters t Active power output and reactive power output during a given time period; For photovoltaic systems in t Maximum active power output limit in maximum power point tracking mode; This refers to the rated apparent power of the photovoltaic inverter; The minimum power factor allowed at the photovoltaic grid connection point;
[0176] The power constraints of the energy storage inverter are as follows:
[0177] Formula (10)
[0178] in, and For energy storage inverters t Active and reactive power output during a given time period; and These are the maximum discharge power and maximum charging power of the energy storage inverter, respectively. This is the rated apparent power of the energy storage inverter.
[0179] Uncertainty in scenarios involving uncertainties includes:
[0180] (1) Output fluctuations of distributed power sources on the source side: such as the volatility of photovoltaic and wind power generation;
[0181] (2) Random changes in load-side conditions: uncertainty in user electricity demand;
[0182] (3) SOC deviation and efficiency changes on the energy storage side: changes in the state of charge (SOC) and charge / discharge efficiency of the energy storage system;
[0183] (4) Network-side line parameters and topology disturbances: changes in power grid line parameters and changes in network topology.
[0184] The result of this step is to obtain an optimized objective function for the intraday correction layer containing the scenario set Ω. Its technical role is to enhance the system's adaptability to random disturbances and ensure that the voltage support strategy is not only effective under ideal prediction conditions, but also remains stable and feasible under actual operational fluctuations.
[0185] After introducing the set of uncertain scenarios Ω and establishing the continuous-layer optimization objective function in the desired sense, it is necessary to further model the power regulation process by combining the operating characteristics of photovoltaic inverters and energy storage inverters. Based on the flexibility requirements of the intraday correction layer for voltage support, the aim is to solve the power regulation range and operating constraints of continuous reactive power regulation equipment under uncertain environments, thereby obtaining an intraday optimization model that can be executed under multiple scenarios.
[0186] This represents the active power output of the photovoltaic inverter at any given time interval t. It must be equal to its maximum active power output. This indicates that the photovoltaic inverter should always operate at maximum active power to fully utilize solar energy resources.
[0187] This indicates ensuring the reactive power output of the photovoltaic inverter. Not exceeding its rated apparent power Limitations. It allows for reactive power to be adjusted to support grid voltage, provided that the reactive power output does not exceed the capacity limit of the photovoltaic inverter;
[0188] The power factor is the ratio of active power to apparent power, representing the efficiency of power utilization in a power distribution network system. The power factor must not be lower than a specified minimum value. This ensures that the photovoltaic inverter will not have an adverse impact on the power grid when connected to the grid.
[0189] Formula (9) stipulates that while ensuring the photovoltaic inverter operates at maximum active power, it supports the voltage only by adjusting reactive power: the active power output shall not exceed the current maximum value, the reactive power output shall be constrained by the capacity limit of the photovoltaic inverter, and the operating power factor of the photovoltaic grid connection point shall not be lower than the specified value.
[0190] For formula (10) P ESS ( t )>0 indicates that the energy storage is discharging into the distribution network. P ESS ( t )<0 indicates that the energy storage is charged from the distribution network.
[0191] This means ensuring that the charging and discharging power of the energy storage inverter does not exceed its rated charging and discharging power limits, avoiding equipment overload, extending equipment life, and ensuring safe operation.
[0192] Ensure that the reactive power output of the energy storage inverter does not exceed the limit of its rated apparent power.
[0193] Inequality (10) limits the charging and discharging power of the energy storage device at any time to not exceed its rated power upper and lower limits, and its reactive power output is limited by the capacity curve of the energy storage inverter, so as to ensure that the energy storage inverter operates in the safe operating range.
[0194] The result of this step is a continuous layer optimization model that includes constraints on photovoltaic and energy storage inverters. Its technical role is to provide a safe and reasonable adjustment space for continuous equipment under uncertain scenarios, and to ensure the feasibility and flexibility of voltage support during the intraday phase.
[0195] After establishing the objective function and constraints of the intraday correction layer, it needs to be further transformed into an agent model that can be solved by reinforcement learning. Since the intraday correction layer involves multiple regulation units such as photovoltaic inverters and energy storage inverters, the presence of other agents will cause the environment to exhibit non-static characteristics when any agent makes a decision. Therefore, a Markov game model is established to characterize the dynamic features of the system under multi-agent interaction.
[0196] Cooperative control strategy of intelligent agents for continuous reactive power regulation equipment Including the reactive power output of photovoltaic inverters and reactive power of energy storage inverter ,as follows:
[0197] Formula (11)
[0198] in, For a continuous intelligent agent, it represents the local observation state space.
[0199] Including active load during time period t reactive load and photovoltaic active power output ,as follows:
[0200] Formula (12)
[0201] Each continuous agent reward function ,as follows:
[0202] Formula (13)
[0203] in, For the number of continuous agents, For continuous intelligent agents The set of nodes in the distribution network. For nodes exist Per-unit value of voltage amplitude during the time period For continuous intelligent agents The penalty coefficient for violations of the action.
[0204] Through formula (13), the intraday correction layer can not only correct voltage deviations in real time within a local range, but also avoid over-operation of inverters and energy storage devices, thus extending equipment lifespan. Its physical significance lies in guiding photovoltaic and energy storage inverters to achieve rapid dynamic adjustment and local voltage support during the intraday phase, while taking into account both equipment operation economy and lifespan management.
[0205] After establishing the discrete and continuous agent models, a centralized training-decentralized execution (CTDE) framework needs to be introduced to achieve collaborative optimization among multiple agents. This step is based on the interactive characteristics of multi-agent systems and aims to resolve the contradiction between centralized global information and decentralized local execution, thereby obtaining a joint optimization mechanism that can utilize global information for training while ensuring real-time independent execution.
[0206] During the training phase, the Critic network receives global system state information and joint actions from all agents to evaluate the value of the overall policy, thereby guiding policy updates for each agent. The Actor network, on the other hand, models individual agents and selects actions based solely on local observations. In this way, agents can fully utilize global information during the centralized training phase, accelerating convergence; and during actual operation, they can execute independently under local observation conditions, ensuring real-time performance in a distributed environment.
[0207] This framework is applicable not only to action planning at the day-ahead scheduling layer but also to power regulation at the intraday correction layer. Through the unified CTDE framework, agents at both layers can achieve collaborative optimization across multiple time scales.
[0208] The result of this step is the formation of a joint optimization mechanism with centralized training and decentralized execution, enabling collaborative problem-solving between the day-ahead scheduling layer and the intraday correction layer within a unified framework. Its technical advantage lies in ensuring global optimality while also considering local real-time performance, allowing for efficient integration of day-ahead scheduling planning and intraday control.
[0209] Within the CTDE framework, the system further proposes a multi-timescale hierarchical optimization and inter-layer coordination mechanism. This mechanism, based on node voltage prediction results, equipment operating characteristics, and control requirements at different time scales, solves the problem of how to achieve coordinated voltage support at the two levels of "day-ahead scheduling - intraday correction".
[0210] During the day-ahead phase, the dispatching layer utilizes the tap positions and switching control of OLTC, CB, and VR equipment to generate a 24-hour rolling voltage control plan in advance, based on voltage forecasts and equipment operation limits. This proactive planning allows for the early detection of potential voltage exceedance risks and preventative dispatching, thereby ensuring the long-term stability of the power grid.
[0211] The intraday correction layer plays a flexible adjustment role during the intraday period. Photovoltaic inverters and energy storage inverters adjust power within a short timescale of 5–15 minutes based on real-time voltage deviations, distributed power output fluctuations, and load surges to quickly correct disturbances caused by prediction errors. Through the dynamic response capability of continuous resources, the system voltage is ensured to remain within a safe operating range.
[0212] To ensure coordination across different time scales, a two-way information channel of "status-command-feedback" was designed. Regarding inter-layer information transmission, the direction from the day-ahead scheduling layer to the intraday correction layer primarily involves the issuance of day-ahead scheduling planning commands. This includes the OLTC tap position, CB switching status, voltage reference range, and no-function boundary plan information generated the previous day, which are transmitted to the intraday correction layer as fixed boundary conditions. This ensures that intraday correction layer control must be locally optimized within this framework, thus avoiding conflicts between short-term control and long-term planning. Simultaneously, the direction from the intraday correction layer to the day-ahead scheduling layer involves feedback on distribution network node voltage and deviations. During operation, the intraday correction layer monitors the voltage status in real time and transmits deviation correction results, strategy execution quality, and disturbance trend data back to the day-ahead scheduling layer. This information serves as crucial input for the next day-ahead cycle prediction and optimization, thereby achieving dynamic coupling and continuous improvement of layered control.
[0213] This step yields a multi-timescale hierarchical collaborative mechanism encompassing "day-ahead scheduling planning—intraday correction—information exchange." Its technical role lies in combining forward-looking planning with flexible correction of power grid operation, ensuring both global consistency in voltage control and enhancing the system's dynamic adaptability to uncertain disturbances.
[0214] Based on the aforementioned hierarchical and coordination mechanisms, an optimization scheme combining offline training and online rolling control is implemented. This approach addresses the issue of control strategies gradually becoming ineffective as the operating environment changes, by utilizing historical operating data, real-time feedback information, and the physical constraints of the distribution network.
[0215] During the offline phase, a complete training environment is constructed by utilizing historical operating data, load characteristics, and distributed power generation output patterns, combined with the physical constraints of OLTC, CB, VR, and photovoltaic and energy storage inverters. Large-scale training is performed using deep reinforcement learning algorithms, and comprehensive indicators such as voltage deviation, equipment operation costs, and energy storage aging costs are incorporated into the reward function to achieve a multi-objective trade-off between voltage stability, equipment lifespan, and operational economy. After training, the model is deployed to the reactive power virtual power plant central coordinator as a foundational decision-making strategy.
[0216] During the online rolling control phase, real-time data on distribution network node voltage, load fluctuations, distributed output changes, and network topology adjustments are collected and input into the rolling optimization module. This module rapidly generates control commands based on an offline-trained model and, combined with a rolling update mechanism, feeds back voltage deviations and equipment response effects generated during operation to the learning module, thereby continuously correcting and evolving the strategy. In this way, the control system can be continuously optimized in actual operation, maintaining long-term effectiveness.
[0217] This establishes a closed-loop control mechanism of "offline training - online optimization - rolling updates." Its technological advantage lies in overcoming the limitations of traditional static control models, enabling the power distribution network system to continuously learn and adaptively evolve, effectively coping with the uncertainties and dynamic changes of complex power grid environments during long-term operation.
[0218] Figure 2 This diagram illustrates the convergence of the reward function as the algorithm trains in the test scenario. The vertical axis represents the agent's average reward, and the horizontal axis represents the number of training rounds. Initially, the curve fluctuates significantly, reflecting that the agent's trade-offs regarding voltage deviation, device action costs, and limit violation penalties are not yet stable during the exploration phase. As training progresses, the curve gradually rises and stabilizes, indicating that the policy has learned to control the intensity of device actions and avoid limit violations while reducing voltage deviation. This demonstrates the effectiveness of the reward design and the CTDE training paradigm, and provides a basis for model convergence and generalization in subsequent deployments in real-world operations.
[0219] Figure 3 The graph shows the reactive power output curves of the energy storage inverter, photovoltaic inverter, and capacitor bank under the test scenario. As can be seen from the graph, the reactive power output of the capacitor bank exhibits a stepped change, reflecting the characteristic of discrete equipment providing baseline reactive power support during the day-ahead phase. The reactive power output of the photovoltaic and energy storage inverters, on the other hand, shows a continuously adjustable curve, enabling rapid response to load fluctuations and uncertainties in distributed power output during the day-ahead phase, achieving dynamic voltage correction.
[0220] Figure 4 This test scenario shows the voltage regulation equipment's gear changes. The OLTC and VR gear changes are represented by discrete step curves, showing the minimum necessary actions under the upper limit of the number of actions and the gear boundary constraints. The trajectory reflects the day-ahead scheduling layer's day-ahead optimization to stabilize the voltage with the fewest actions, significantly reducing unnecessary frequent operations, extending equipment life, and reducing operational disturbances while meeting voltage qualification and power flow constraints.
[0221] Through the above method, this invention breaks through the limitations of traditional centralized reactive voltage support methods based on static models in terms of response speed and scalability. It realizes the voltage support target of multi-device collaboration, multi-timescale linkage and long-term self-evolution optimization in complex power grid environments, and ensures the safety and flexibility of power quality and system operation.
[0222] The purpose of step S3 is to introduce uncertain scenarios and continuous action space reinforcement learning to achieve intraday real-time voltage correction of photovoltaic inverters and energy storage inverters within the day-ahead scheduling plan boundary, and to feed back the operating deviation to the day-ahead scheduling layer to form a closed loop, so as to improve the adaptability of the power distribution network to random disturbances and the accuracy of overall voltage control.
[0223] Example 2:
[0224] A specific embodiment of the present invention discloses a distribution network voltage control system based on a reactive power virtual power plant, thereby implementing the distribution network voltage control method based on a reactive power virtual power plant in Embodiment 1. The specific implementation methods of each module are as described in the corresponding descriptions in Embodiment 1.
[0225] like Figure 5 As shown, the system includes a distribution network physical and equipment modeling module M1, a day-ahead scheduling layer planning module M2, and an intraday correction layer execution and feedback module M3;
[0226] The distribution network physical and equipment modeling module M1 is used to construct a two-layer collaborative control model containing a day-ahead scheduling layer and an intraday correction layer based on discrete and continuous reactive power regulation equipment in the distribution network under reactive power virtual power plant management, and to construct distribution network physical constraints.
[0227] The day-ahead scheduling layer planning module M2 is used to solve the first multi-agent reinforcement learning of the discrete action space strategy of the day-ahead scheduling layer by converting the optimal control of discrete reactive power regulation equipment into discrete action space strategy under the physical constraints of the distribution network, and generate the day-ahead scheduling plan and send it to the intraday correction layer.
[0228] The intraday correction layer execution and feedback module M3 is used to, under the dual constraints of the distribution network physical constraints and the day-ahead scheduling plan as boundary constraints, transform the optimal control of continuous reactive power regulating equipment into a second multi-agent reinforcement learning solution for a continuous action space strategy, generate intraday correction instructions; adjust the continuous reactive power regulating equipment based on the intraday correction instructions to obtain the node voltage in the distribution network; and feed back the node voltage in the distribution network to the day-ahead scheduling layer for closed-loop coordination to achieve voltage optimization control of the distribution network.
[0229] In this case, a distribution network is managed by at least one reactive virtual power plant.
[0230] Since the system in this embodiment and the method in Embodiment 1 are related and can be referenced from each other, this description is redundant and will not be repeated here. Because this system embodiment shares the same principle as the above method embodiment, it also possesses the corresponding technical effects of the above method embodiment.
[0231] Example 3:
[0232] Another specific embodiment of the present invention discloses an electronic device for voltage control of a distribution network based on a reactive power virtual power plant, the electronic device comprising:
[0233] Memory, used to store computer programs;
[0234] A processor is used to execute the computer program to implement a distribution network voltage control method based on a reactive power virtual power plant.
[0235] Example 4:
[0236] Another specific embodiment of the present invention discloses a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of a distribution network voltage control method based on a reactive power virtual power plant.
[0237] In summary, the distribution network voltage control method based on a reactive power virtual power plant according to an embodiment of the present invention has the following beneficial effects:
[0238] 1. This invention achieves multi-timescale voltage support of "day-ahead scheduling planning - intraday correction" by introducing OLTC, CB, and VR discrete reactive power regulation devices in the discrete layer for optimized control, combined with the rapid response of photovoltaic inverters and energy storage inverters in the continuous layer. This effectively reduces node voltage deviation, improves power quality, and enhances the stability of distribution network voltage.
[0239] 2. The present invention introduces equipment action costs and constraints into the current scheduling layer optimization objective, which avoids frequent switching of discrete reactive power regulation equipment such as OLTC, CB, and VR, reduces the number of equipment actions, extends the service life of key voltage regulation equipment, and reduces operation and maintenance costs.
[0240] 3. This invention introduces uncertainty scenario modeling into the intraday correction layer, fully considering photovoltaic distributed power output fluctuations, load forecasting errors and network disturbances, so that the learned agent cooperative strategy has strong adaptability and robustness in multiple scenarios.
[0241] 4. This invention constructs a hierarchical optimization mechanism that complements discrete reactive power regulation equipment and continuous reactive power regulation equipment. Discrete reactive power regulation equipment provides day-ahead scheduling planning, while continuous reactive power regulation equipment provides short-term real-time flexible correction, thereby realizing efficient coordination and comprehensive utilization of reactive power resources.
[0242] 5. This invention combines a centralized training and distributed execution CTDE framework, and adopts a combination of offline training and online rolling optimization. The distribution network can continuously update and correct the control strategy in actual operation, ensuring long-term stability and flexibility, and achieving self-adaptation and continuous voltage optimization.
[0243] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.
[0244] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A distribution network voltage control method based on a reactive power virtual power plant, characterized in that, Includes the following steps: S1. Based on discrete and continuous reactive power regulation equipment in the distribution network under reactive power virtual power plant management, construct a two-layer collaborative control model including day-ahead scheduling layer and intraday correction layer, and construct the physical constraints of the distribution network. S2. Under the physical constraints of the distribution network, the day-ahead scheduling layer transforms the optimal control of discrete reactive power regulation equipment into a first multi-agent reinforcement learning solution for discrete action space strategy, and generates a day-ahead scheduling plan which is then sent to the intraday correction layer. Step S2 includes: Based on the physical constraints of the distribution network, an objective function for day-ahead scheduling optimization is constructed with the goal of minimizing the total voltage deviation of nodes in the distribution network within the scheduling cycle; and a constraint on the maximum allowable operating frequency of each discrete reactive power regulating device within the scheduling cycle is constructed. The day-ahead scheduling optimization is modeled as a Markov game process, and a corresponding discrete agent is defined for each discrete reactive power regulation device to obtain the first multi-agent; wherein, the action space of each agent is a discrete gear or switching operation, the state space includes the current state of the device and the cumulative number of actions, and the reward function includes a comprehensive evaluation of the node voltage deviation in the distribution network, the cost of device actions and the voltage over-limit situation. The MASAC-Discrete algorithm is used to perform reinforcement learning on the first multi-agent system to obtain the agent cooperative strategy for the discrete reactive power regulation device. ; Based on the intelligent agent coordination strategy of the discrete reactive power regulation equipment, a day-ahead scheduling plan is generated, which includes the gear or switching status of each discrete reactive power regulation equipment in the future scheduling cycle. The objective function for day-ahead scheduling optimization is as follows: in, For time period index, Number of time periods; For reactive power virtual power plant index, This represents the total number of virtual reactive power plants. For node indexes in the distribution network, For reactive virtual power plants The set of nodes in the distribution network under its jurisdiction; In order to be in The time period belongs to the node in the distribution network of the reactive virtual power plant. The per-unit voltage value; S3. Under the dual constraints of the distribution network physical constraints and the day-ahead scheduling plan as boundary constraints, the intraday correction layer transforms the optimization control of the continuous reactive power regulating equipment into a second multi-agent reinforcement learning solution for a continuous action space strategy to generate intraday correction instructions; based on the intraday correction instructions, the continuous reactive power regulating equipment is adjusted to obtain the node voltage in the distribution network; the node voltage in the distribution network is fed back to the day-ahead scheduling layer for closed-loop coordination to realize the voltage optimization control of the distribution network; Step S3 includes: Based on the physical constraints of the distribution network, the day-ahead scheduling plan as boundary conditions, and the set of uncertainty scenarios, an intraday correction optimization objective function is constructed with the goal of minimizing the expected voltage deviation. The intraday correction optimization is modeled as a Markov game process, and a corresponding continuous agent is defined for each continuous reactive power regulation device to obtain a second agent. The action space of each agent is continuous reactive power output, the state space includes local active or reactive load, and photovoltaic active power output, and the reward function includes the node voltage deviation in the distribution network and the penalty for the continuous reactive power regulation device to violate the corresponding safe operation constraints. The MASAC algorithm is used to perform reinforcement learning on the second agent to obtain the cooperative control strategy of the agent in the continuous reactive power regulation equipment. ; Based on the collaborative control strategy of the intelligent agent of the continuous reactive power regulation equipment, intraday correction instructions for each continuous reactive power regulation equipment are generated. The continuous reactive power regulation equipment is adjusted based on intraday adjustment instructions to obtain the node voltage of the distribution network under reactive power virtual power plant management; The node voltage of the distribution network is compared with the node voltage of the distribution network corresponding to the day-ahead scheduling plan to obtain node voltage deviation information; the day-ahead scheduling layer re-performs the first multi-agent reinforcement learning solution based on the node voltage deviation information to regenerate the day-ahead scheduling plan. In this case, a distribution network is managed by at least one reactive virtual power plant.
2. The distribution network voltage control method based on a reactive power virtual power plant according to claim 1, characterized in that, The discrete reactive power regulation equipment includes at least one of an on-load tap-changing transformer (OLTC), a capacitor bank (CB), and a voltage regulator (VR). The continuous reactive power regulation equipment includes at least one of a photovoltaic inverter and an energy storage inverter.
3. The distribution network voltage control method based on a reactive power virtual power plant according to claim 1, characterized in that, The physical constraints of the distribution network are established based on the DistFlow power flow equation; the physical constraints of the distribution network include node power balance constraints and voltage balance constraints.
4. The distribution network voltage control method based on a reactive power virtual power plant according to claim 1, characterized in that, Under the physical constraints of the distribution network, the maximum allowable operating frequency constraints for each discrete reactive power regulating device at the day-ahead dispatching level within the dispatching cycle are as follows: in, , and Variables are 0 or 1; δ OLTC ( t (for the time period) Indicator variable for whether an on-load tap changer (OLTC) has operated. M OLTC In order to be within the planning period OLTC Maximum number of actions allowed; δ CB ( t (for the time period) The indicator variable indicating whether a capacitor bank (CB) switching operation has occurred. M CB This represents the maximum number of actions allowed for CB within the planning period. For the time period Whether an indicator variable has been generated indicating a voltage regulator VR level adjustment action. M VR This represents the maximum number of actions allowed in VR within the planning period.
5. The distribution network voltage control method based on a reactive power virtual power plant according to claim 4, characterized in that, The agent-based cooperative strategy for discrete reactive power regulation equipment Including OLTC gear adjustment amount and CB switching status change amount VR level adjustment ,as follows: The local observation state space of a discrete agent includes: during the current scheduling period During which time period, the tap position of the on-load tap-changing transformer (OLTC) Switching status of capacitor bank CB Voltage regulator VR gear position Up to Cumulative number of actions in OLTC during the time period CB's cumulative number of actions Cumulative number of actions in VR ,as follows: reward function ,as follows: in, It represents the sum of squares of voltage deviations across the entire distribution network and is used to measure voltage quality; express The motion vector for the time period includes OLTC gear shift adjustment, CB casting / switching status, and VR gear shift adjustment, with a L2 norm. Characterizes the operating cost of the equipment; This is a voltage over-limit indication function, which takes a value of 1 when the voltage exceeds the allowable range, and 0 otherwise. , They are respectively and Weighting coefficients; The day-ahead scheduling plan is expressed as follows: in, T OLTC ( t ), S CB ( t ) and L VR ( t ) respectively in the current The time period is determined by the day-ahead discrete layer. OLTC tap position, CB Throwing and Cutting Status and VR The actual value of the gear adjustment.
6. The distribution network voltage control method based on a reactive power virtual power plant according to claim 1, characterized in that, Based on the scenario set Ω including uncertainties in source, network, load, and storage, the optimization objective function based on the intraday correction layer of Ω is as follows: in, ω An index for scenarios with uncertainties; Let ω represent the expected value of all possible scenarios in the scenario set Ω for uncertain factors; Uncertainties in each scenario include fluctuations in the output of distributed power sources on the source side, disturbances in line parameters and topology on the grid side, random changes in load on the load side, and deviations in SOC and efficiency on the energy storage side. Power constraints of continuous reactive power regulation equipment in the intraday correction layer, including power constraints of photovoltaic inverters and power constraints of energy storage inverters; The power constraints of the photovoltaic inverter are as follows: in, and For photovoltaic inverters t Active power output and reactive power output during a given time period; For photovoltaic systems in t Maximum active power output limit in maximum power point tracking mode; This refers to the rated apparent power of the photovoltaic inverter. The minimum power factor allowed at the photovoltaic grid connection point; The power constraints of the energy storage inverter are as follows: in, and For energy storage inverters t Active and reactive power output during a given time period; and These are the maximum discharge power and maximum charging power of the energy storage inverter, respectively. This is the rated apparent power of the energy storage inverter.
7. The distribution network voltage control method based on a reactive power virtual power plant according to claim 6, characterized in that, Cooperative control strategy of intelligent agents for continuous reactive power regulation equipment Including the reactive power output of photovoltaic inverters and reactive power of energy storage inverter ,as follows: in, For a continuous agent, the local observation state space is used. Including active load during time period t reactive load and photovoltaic active power output ,as follows: Each continuous agent reward function ,as follows: in, For the number of continuous agents, For continuous intelligent agents The set of nodes in the distribution network. For nodes exist Per-unit value of voltage amplitude during the time period For continuous intelligent agents The penalty coefficient for violations of the action.
Citation Information
Patent Citations
Distribution network multi-time scale reactive power control method for inhibition of photovoltaic grid connection point voltage fluctuation
CN108321810A
Intelligent power dispatching method and system for virtual power plant
CN120728758A