A power distribution network active voltage control method and system
Patent Information
- Application Number
- CN202610919703.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-25
AI Technical Summary
该类控制方式存在明显缺陷:一是响应滞后,无法提前预判电压变化趋势,难以应对新能源出力突变带来的电压骤变;二是多设备协同能力差,各调压设备独立运行、缺乏统一调度,容易出现频繁反复动作、调压震荡的现象,大幅缩短机械类调压设备的检修周期与服役年限;三是控制策略高度依赖精准的电网拓扑参数、线路阻抗参数以及负荷模型,一旦电网拓扑发生变更、线路参数出现漂移,原有控制策略便会失效,系统鲁棒性较差
本发明模型依赖性低,鲁棒性强:本发明采用数据驱动的深度强化学习架构,无需依赖精确的配电网线路参数、拓扑模型与负荷模型,仅依靠运行数据即可完成策略学习与实时决策,面对电网拓扑变更、参数漂移、新能源出力随机波动等场景仍可稳定运行,适配性与抗干扰能力显著优于传统优化算法。
Smart Images

Figure CN122823508A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power distribution network control technology, specifically to a method and system for active voltage control of power distribution networks. Background Technology
[0002] Against the backdrop of the ongoing construction of new power systems, distributed power sources such as wind power, photovoltaic power, and energy storage devices are being connected to medium and low voltage distribution networks on a large scale and in a high proportion, completely changing the traditional unidirectional power flow and load-driven operation characteristics of distribution networks. Traditional distribution networks are generally supplied by a unified upstream substation, with power flowing step-by-step from the high-voltage side to low-voltage load nodes. The power flow direction is fixed, and the voltage distribution pattern is relatively stable. Maintenance personnel can manage voltage using fixed voltage regulation strategies, manual on-site operation, or timed switching of reactive power compensation equipment. However, with the integration of distributed power sources, the distribution network evolves into an active distribution network, exhibiting bidirectional power flow characteristics. The output of new energy sources such as photovoltaic and wind power is highly subject to random fluctuations and intermittentity due to factors such as sunlight, wind speed, and weather conditions. The load side also experiences significant fluctuations at different times of day, night, and holidays. The superposition of multiple uncertainties leads to frequent problems of voltage exceeding limits at various nodes in the distribution network.
[0003] Voltage, as one of the core indicators of power quality in power distribution networks, directly affects the safety of electrical equipment, the user experience, and the economic efficiency of power grid operation. When node voltage consistently exceeds the safe range stipulated by national standards, it can lead to reduced efficiency and shortened lifespan of household appliances, industrial motors, and other electrical equipment, or even equipment burnout and insulation breakdown. Furthermore, unreasonable voltage distribution significantly increases active power losses in distribution network lines and transformers, raising power grid operating costs.
[0004] Currently, the mainstream voltage regulation methods in power distribution networks mainly rely on two types of traditional voltage regulating equipment: on-load tap-changing transformers and parallel capacitor banks. The control logic mostly adopts threshold-triggered passive control, meaning that the equipment only initiates voltage regulation when the monitored voltage reaches preset upper or lower limits. This type of control has significant drawbacks: First, it has a slow response time, making it impossible to predict voltage change trends in advance and difficult to cope with sudden voltage fluctuations caused by changes in renewable energy output. Second, it has poor multi-equipment coordination capabilities; each voltage regulating device operates independently without unified scheduling, easily leading to frequent repeated actions and voltage regulation oscillations, significantly shortening the maintenance cycle and service life of mechanical voltage regulating equipment. Third, the control strategy is highly dependent on accurate grid topology parameters, line impedance parameters, and load models; once the grid topology changes or line parameters drift, the original control strategy becomes ineffective, resulting in poor system robustness. Summary of the Invention
[0005] To address the aforementioned technical problems, the purpose of this application is to provide an active voltage control method and system for distribution networks, the specific technical solution of which is as follows: This application provides an active voltage control method for a distribution network, including the following steps: The system acquires real-time operating status data of the power distribution network; inputs the real-time operating status data into a pre-trained deep reinforcement learning agent; uses the deep reinforcement learning agent to output an optimal control strategy based on the real-time operating status data; generates control commands based on the optimal control strategy and sends them to voltage regulating devices in the power distribution network to regulate the node voltage of the power distribution network.
[0006] Furthermore, acquiring real-time operating status data of the distribution network specifically includes: collecting voltage amplitude, voltage phase angle, branch current, and active and reactive power of each distributed power source at each node in the distribution network; constructing the collected data into a state vector as the real-time operating status data.
[0007] Furthermore, the voltage regulating equipment includes an on-load tap-changing transformer, a parallel capacitor bank, a photovoltaic inverter, and an energy storage device; the step of generating control commands according to the optimal control strategy and sending them to the voltage regulating equipment in the distribution network specifically includes: generating tap adjustment commands for the on-load tap-changing transformer, switching commands for the parallel capacitor bank, and reactive power adjustment commands for the photovoltaic inverter and the energy storage device.
[0008] Furthermore, before inputting the real-time operating status data into the pre-trained deep reinforcement learning agent, the method further includes a training step for the deep reinforcement learning agent: constructing an active voltage control model for the distribution network, the model including a state space, an action space, and a reward function; acquiring historical operating data of the distribution network; and using a preset intelligent optimization algorithm to train the agent offline using the historical operating data until the agent converges, thereby obtaining the pre-trained deep reinforcement learning agent.
[0009] Furthermore, the construction logic of the reward function is as follows: setting a voltage deviation penalty term, a network loss penalty term, and a control action penalty term; when the voltage of a distribution network node exceeds a preset safety range, increasing the value of the voltage deviation penalty term; when the network loss of the distribution network exceeds a preset threshold or the number of actions of the voltage regulating equipment exceeds a preset limit, correspondingly increasing the value of the network loss penalty term or the control action penalty term, so as to guide the agent to output an optimal control strategy that reduces network loss and ensures smooth action.
[0010] Furthermore, the step of using a preset intelligent optimization algorithm to train the agent offline using the historical running data specifically includes: using a deep deterministic policy gradient algorithm or a proximal policy optimization algorithm as the intelligent optimization algorithm; the agent interacts with the environment, selects an action based on the current state, the environment provides feedback on the reward value and the state at the next moment, and the network parameters of the agent are updated by minimizing the loss function of the value function.
[0011] Furthermore, the step of using the deep reinforcement learning agent to output the optimal control strategy based on the real-time operating state data specifically includes: mapping the real-time operating state data to the current state of the deep reinforcement learning agent; the agent directly calculating and outputting the action value corresponding to the current state based on the current policy network, as the optimal control strategy, wherein the action value includes the continuous or discrete control quantities of each voltage regulating device.
[0012] Furthermore, the method also includes an online update step: after adjusting the node voltage of the distribution network according to the optimal control strategy, the actual operating effect of the distribution network is monitored in real time; if the actual operating effect meets the preset triggering conditions, the deep reinforcement learning agent is fine-tuned and updated online using the current real-time operating status data.
[0013] Furthermore, the method is applied to centralized controllers or distributed local controllers in power distribution networks; when applied to distributed local controllers, each controller only obtains the operating status data of local and adjacent nodes, and outputs a locally optimal control strategy through a multi-agent cooperative mechanism.
[0014] This application also provides an active voltage control system for a distribution network, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any of the above-described active voltage control methods for a distribution network.
[0015] As can be seen from the above, the active voltage control method and system for distribution networks provided in this application have at least the following beneficial effects: This invention has low model dependency and strong robustness: It adopts a data-driven deep reinforcement learning architecture, which does not rely on precise distribution network line parameters, topology models and load models. It can complete strategy learning and real-time decision-making based solely on operating data. It can still operate stably in the face of scenarios such as grid topology changes, parameter drift and random fluctuations in new energy output. Its adaptability and anti-interference ability are significantly better than traditional optimization algorithms.
[0016] This invention features multi-objective collaborative optimization with high overall benefits: The design incorporates a multi-dimensional reward function that includes voltage deviation, network loss, and equipment operation frequency. Under the premise of ensuring that the node voltage is strictly within the safe range, it simultaneously achieves three major objectives: reducing grid losses, extending equipment life, and balancing power supply safety, economical operation, and equipment reliability.
[0017] This invention features fast control response speed, meeting real-time requirements: the intelligent agent adopts an "offline training + online inference" mode, completing massive operating condition learning in the offline stage, and only needing to complete network forward calculation during online operation, with extremely short decision latency, enabling rapid response to voltage changes and solving the problem of lag in traditional passive voltage regulation.
[0018] This invention features coordinated scheduling of equipment, resulting in stable voltage regulation: It unifies the scheduling of four main types of voltage regulation equipment—on-load tap-changing transformers, parallel capacitor banks, photovoltaic inverters, and energy storage devices—to achieve mixed control of discrete and continuous operations, avoiding the problems of frequent operation of individual equipment and voltage regulation oscillation.
[0019] The invention features a flexible and highly scalable architecture: it supports both centralized control and distributed multi-agent collaborative control deployment modes, making it suitable for small urban distribution networks and park distribution networks, as well as wide-area rural distribution networks and multi-zone interconnected distribution networks. Communication pressure can be flexibly adjusted.
[0020] This invention has self-learning capabilities and excellent long-term performance: it adds an online fine-tuning and update mechanism, which allows the intelligent agent to continuously iterate and optimize strategies based on changes in the long-term operating conditions of the power grid, thereby improving voltage control accuracy. It also eliminates the need for manual offline retraining and reduces subsequent operation and maintenance costs. Attached Figure Description
[0021] Figure 1 A flowchart illustrating the steps of an active voltage control method for a power distribution network provided in this application. Detailed Implementation
[0022] The following description, in conjunction with the accompanying drawings, details a specific scheme for an active voltage control method and system for a power distribution network provided in this application.
[0023] Please see Figure 1 The diagram illustrates a flowchart of an active voltage control method for a distribution network according to an embodiment of this application, including the following steps: Step 1: Obtain real-time operating status data of the power distribution network.
[0024] Specifically, it includes one on-load tap-changing transformer (OLTC), three sets of parallel capacitors, eight photovoltaic inverters, and two energy storage devices, totaling 26 load and power nodes. It adopts a centralized controller + a single deep reinforcement learning agent to achieve active voltage control.
[0025] Through data acquisition terminals in the distribution network, the voltage amplitude, voltage phase angle, branch current, and active and reactive power of each distributed power source are collected in real time. This collected data is then used to construct a state vector as real-time operating status data. To ensure the real-time performance of voltage control, this embodiment employs synchronous data acquisition at 200-millisecond intervals. Field data acquisition is performed by the distribution terminal unit (FTU / DTU) and uploaded to the centralized controller via a dedicated power communication channel. The acquisition method is periodic and cyclical, ensuring continuous and complete data collection. Even during brief communication interruptions, the system automatically retains the most recent valid data, maintaining uninterrupted control flow.
[0026] Because the raw collected data contains issues such as different units, large differences in numerical range, and random noise, data preprocessing is necessary before inputting it into the agent. This preprocessing includes: (1) Outlier removal: Threshold judgment and sliding window filtering are used to remove abrupt data, null values and erroneous data; (2) Data normalization: Map all electrical quantities to the [0,1] interval to eliminate the influence of dimensions; (3) Vector construction: All preprocessed data are spliced together in a fixed order to form a high-dimensional real-time state vector.
[0027] Through this step, the overall operating status of the distribution network is completely transformed into a standardized vector, enabling the acquisition of real-time operating status data.
[0028] Step 2: Input the real-time running status data into the pre-trained deep reinforcement learning agent.
[0029] Before input, the agent has completed offline training. During offline training, the system constructs an active voltage control model for the distribution network, including a state space, action space, and reward function. Using historical operating data of the distribution network, the agent is trained using either the Deep Deterministic Policy Gradient (DDPG) algorithm or the Proximal Policy Optimization (PPO) algorithm. During interaction with the environment, the agent selects actions based on its current state, and the environment provides feedback on the reward value and the next state. The reward function comprehensively considers voltage deviation penalties, network loss penalties, and control action penalties, continuously updating network parameters by minimizing the loss function of the value function until the agent converges.
[0030] This step serves as a bridge and entry point connecting the perception and decision-making stages. Its core task is to securely, stably, and accurately transmit the real-time operating status data generated in step 1 to the deep reinforcement learning agent that has completed offline training, enabling the agent to gain a complete perception of the current power grid operating conditions and prepare for subsequent policy reasoning.
[0031] Prior to this step, the deep reinforcement learning agent has undergone offline training using massive amounts of historical operational data. The training process includes constructing the state space, action space, and multi-objective reward function, and employing advanced algorithms such as Deep Deterministic Policy Gradient (DDPG) or Proximal Policy Optimization (PPO) to continuously interact and iterate with the power grid simulation environment until the model converges. The trained agent possesses powerful generalization and decision-making capabilities, enabling it to handle various complex scenarios such as sudden changes in illumination, load fluctuations, power grid topology switching, and random output from distributed power sources. Furthermore, it does not rely on precise power grid mathematical models, exhibiting significantly better robustness than traditional model-driven methods.
[0032] The specific implementation of this step includes data transmission, format verification, dimension verification, legality check, and status input.
[0033] First, the central controller transmits the real-time running state vector generated in step 1 to the agent inference module through its internal high-speed data interface. The data transmission process adopts a zero-copy mechanism to ensure low latency and high reliability.
[0034] Secondly, a strict input validation mechanism is implemented to prevent erroneous data from causing policy anomalies. (1) Dimension verification: Check whether the length of the input vector is consistent with the dimension of the agent's input layer to prevent data loss or redundancy; (2) Numerical range verification: Confirm that all data are within the normalized reasonable range and exclude abnormal values; (3) Equipment status consistency verification: compare the real-time status of the collected voltage regulating equipment with the historical status to avoid abnormal equipment location feedback.
[0035] If the verification fails, the system automatically retains the valid state data from the previous moment and issues an alarm, without performing subsequent reasoning, thus ensuring control security.
[0036] After successful verification, the real-time operating status data is formally fed into the input layer of the deep reinforcement learning agent to complete the state mapping. The agent recognizes this vector as a complete observation of the current environment, including all implicit feature information such as voltage level, power flow distribution, power output, and equipment adjustability. At this point, the agent has all the input conditions required to make the optimal decision and enters the core inference stage.
[0037] Although this step does not directly generate the control strategy, its reliability, real-time performance, and safety directly determine the stability of the entire control system, making it an indispensable key link in achieving closed-loop control.
[0038] Step 3: Utilize the deep reinforcement learning agent to output the optimal control strategy based on the real-time operating status data.
[0039] The agent maps the received real-time operating status data to the current state, and directly calculates and outputs the action value corresponding to the current state based on the trained policy network. This action value is the optimal control policy, which includes the continuous or discrete control quantities of each voltage regulating device.
[0040] The state data is used by a deep reinforcement learning agent that has completed offline training to make rapid inferences and autonomous decisions, and finally output the optimal control strategy that meets the power grid safety constraints, economic operation requirements and equipment action limitations.
[0041] The agent maps the received real-time operating status data to the current state, and directly calculates and outputs the action value corresponding to the current state based on the trained policy network. This action value is the optimal control policy, which includes the continuous or discrete control quantities of each voltage regulating device.
[0042] In practice, the intelligent agent does not simply perform numerical mapping. Instead, it learns control experience from massive amounts of power grid operation data during the offline phase and combines this with the current real-time operating conditions of the power grid to complete a full decision-making process from state perception to strategy output. This process does not rely on power grid topology parameters, line impedance parameters, load forecasting models, or distributed power generation output models. It is a purely data-driven end-to-end decision-making method, thus possessing characteristics such as fast response speed, strong anti-interference ability, and outstanding adaptability.
[0043] First, the agent performs state mapping and feature parsing on the received real-time operating state vector. Since the input state vector is standardized data that has been normalized, dimensionally unified, and time-aligned, it comprehensively reflects information such as the current node voltage, branch power flow, distributed generation output, and voltage regulation equipment status of the distribution network. The agent directly identifies it as the current environmental state in the reinforcement learning environment. This mapping process forms the basis for the forward inference of the policy network, ensuring that the agent can accurately perceive the current operating status of the power grid, including whether the voltage is too high or too low, whether the distributed generation output is fluctuating, whether the load is changing abruptly, and whether the voltage regulation equipment has adjustment capacity.
[0044] Subsequently, the agent invokes its internally trained and converged policy network, using the current state as input, to perform forward computation of the neural network. During the offline training phase, the policy network has learned optimal control principles under different power grid conditions through a large number of samples, including complex decision-making logic such as which type of equipment should be prioritized for adjustment when voltage exceeds limits, how reactive power should be allocated when network losses are significant, and how to smooth output when equipment operates frequently. In the online operation phase, the policy network no longer engages in exploration, trial and error, or iterative optimization; instead, it directly and quickly outputs the action combination with the highest matching degree to the current state based on the learned stable policy.
[0045] The output action values collectively constitute the optimal control strategy, which is a multi-dimensional hybrid action vector containing both discrete and continuous control quantities, corresponding to different types of voltage regulating equipment in the distribution network. The discrete control quantity primarily targets mechanical voltage regulating equipment such as the tap positions of on-load tap-changing transformers and the number of parallel capacitor banks switched on and off; its output is a discrete value with a finite number of taps and banks. The continuous control quantity primarily targets power electronic voltage regulating equipment such as photovoltaic inverters and energy storage devices; its output is a continuously changing value within the rated regulation range, enabling precise and oscillating reactive power regulation.
[0046] In the process of outputting the optimal control strategy, the agent strictly follows the optimization objective established by the reward function during the offline training phase. That is, while ensuring that the voltage of all nodes is maintained within a safe operating range, the agent minimizes the active power loss of the entire distribution network and limits the frequent operation of voltage regulating equipment in a short period of time. Therefore, the action value output by the strategy network does not simply pursue rapid voltage regulation, but comprehensively considers grid security, operational economy, and equipment reliability to form a globally optimal solution after balancing multiple objectives.
[0047] To further ensure the reliability and security of the strategy output, the agent performs legality verification and constraint judgment before outputting action values. This includes verifying whether transformer tap changes exceed limits, whether capacitor switching is too frequent, whether photovoltaic and energy storage reactive power regulation exceeds the rated range, and whether voltage regulation amplitude is too large. If a potential unsafe action trend is detected, the agent will automatically correct the action output based on the constraint rules learned during the training phase, ensuring that the final optimal control strategy fully complies with the actual operation specifications of the distribution network.
[0048] In summary, this step achieves a direct mapping from the real-time state of the power grid to the optimal control strategy through real-time reasoning by a deep reinforcement learning agent. The entire decision-making process is short, fast, and requires no human intervention. It can quickly respond to complex operating conditions such as distributed power source fluctuations and load changes, providing accurate, reliable, and optimal decision-making basis for subsequent control command generation and equipment execution.
[0049] Step 4: Generate control commands based on the optimal control strategy and send them to the voltage regulating equipment in the distribution network.
[0050] The voltage regulation equipment in this embodiment includes an on-load tap-changing transformer, a parallel capacitor bank, a photovoltaic inverter, and an energy storage device. Based on the action values, the system generates tap adjustment commands for the on-load tap-changing transformer, switching commands for the parallel capacitor bank, and reactive power adjustment commands for the photovoltaic inverter and energy storage device, thereby achieving active regulation of the distribution network node voltage.
[0051] This step is the execution output and closed-loop terminal of the entire active voltage control method. Its core task is to transform the optimal control strategy output by the intelligent agent in step 3 into executable control commands that conform to the power equipment communication specifications, and then send them to each voltage regulating device through reliable communication, so as to achieve rapid, stable and precise regulation of the voltage at the distribution network nodes.
[0052] The strategy itself is an abstract combination of action values, which cannot be directly recognized by physical devices such as transformers, capacitors, and inverters. Therefore, it must go through a series of processes such as instruction parsing, instruction encapsulation, protocol conversion, security verification, instruction issuance, and execution feedback to complete the transformation from "digital strategy" to "physical action".
[0053] Specifically, firstly, based on the action value type in the optimal control strategy, corresponding control commands are generated according to the device type: (1) For on-load tap-changing transformers: Based on the tap change value output by the strategy, generate tap change, tap up, down or hold instructions, clarify the target tap, and ensure accurate adjustment; (2) For parallel capacitor banks: Based on the number of switching banks output by the strategy, generate corresponding capacitor switching or holding commands, and strictly follow the switching interval constraints. (3) For photovoltaic inverters: Based on the continuous reactive power regulation coefficient output by the strategy, generate reactive power setpoint instructions to achieve smooth and continuous regulation; (4) For energy storage devices: Based on the continuous reactive power output value output by the strategy, generate reactive power target instructions to give full play to the rapid response advantage of energy storage.
[0054] Secondly, the generated instructions are encapsulated and formatted for communication to meet the execution requirements of industrial equipment. This embodiment adopts the mainstream power system standard protocol IEC60870-5-104 to encapsulate the control instructions into standard messages, including information such as device address, instruction type, set value, execution priority, and checksum, to ensure that the instructions can be correctly parsed and executed without errors by the field terminal.
[0055] Subsequently, a safety interlock check is performed before issuing the command: this includes checking whether the equipment is in remote control mode, whether there are fault alarms, whether it is under maintenance, whether the command is duplicated with the current position, and whether there are any protective action interlocks. Only after all checks are passed will the command be allowed to be issued, thus preventing misoperation from the source.
[0056] After the inspection is passed, the central controller sends the instructions to the corresponding voltage regulating equipment through the power dedicated communication network. The instruction transmission delay is less than 150 milliseconds, and it supports automatic retransmission and response mechanisms to ensure that the instructions are not lost, delayed, or confused.
[0057] Upon receiving the instruction, the voltage regulating equipment immediately executes the corresponding action: • Adjusting the tap changer of an on-load tap-changing transformer changes the turns ratio, thereby raising or lowering the overall voltage of the entire network; • Switching parallel capacitor banks provides large-capacity reactive power support and quickly smooths out voltage fluctuations; • Photovoltaic inverters and energy storage systems can quickly respond to continuous reactive power commands, enabling fine-tuning of voltage and smooth convergence.
[0058] The coordinated operation of multiple devices creates a complementary voltage regulation effect of "coarse adjustment + fine adjustment" and "discrete + continuous adjustment", enabling the voltage of all network nodes to quickly return to a safe and qualified range within hundreds of milliseconds, while reducing network losses and minimizing device operation.
[0059] After the instruction is executed, the system automatically returns to the success / failure status and immediately enters the next control cycle, re-executes step 1 to obtain the latest operating data, forming a complete closed-loop control of "data acquisition → status input → strategy reasoning → instruction execution", realizing uninterrupted active optimization of the distribution network voltage 24 / 7.
[0060] Furthermore, in order to adapt to changes in power grid topology or operating conditions, this method also includes an online update step: after adjusting the voltage of the distribution network nodes, the actual operating effect is monitored in real time; if the actual operating effect meets the preset triggering conditions (in this embodiment, the triggering condition is multiple consecutive voltage overruns), the deep reinforcement learning agent is fine-tuned and updated online using the current real-time operating status data to ensure the long-term effectiveness of the control strategy.
[0061] Furthermore, this method can be applied to centralized controllers or distributed local controllers in power distribution networks. When applied to distributed local controllers, each controller only acquires the operating status data of its local and neighboring nodes, and outputs a locally optimal control strategy through a multi-agent cooperative mechanism to reduce communication pressure and improve response speed.
[0062] Based on the same inventive concept as the above method, this application also provides an active voltage control system for a distribution network, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any of the above-described active voltage control methods for a distribution network.
[0063] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Any equivalent structural or procedural transformations made based on the description and drawings of this application, or direct or indirect applications in other related technical fields, are similarly included within the protection scope of this application.
Claims
1. A method for active voltage control in a distribution network, characterized in that, Includes the following steps: Obtain real-time operating status data of each node in the power distribution network; The real-time operating status data is input into a pre-trained deep reinforcement learning agent; The deep reinforcement learning agent outputs the optimal control strategy based on the real-time operating status data. Control commands are generated according to the optimal control strategy and sent to the voltage regulating equipment in the distribution network to regulate the node voltage of the distribution network.
2. The active voltage control method for a distribution network as described in claim 1, characterized in that, Acquiring real-time operating status data of the distribution network specifically includes: collecting voltage amplitude, voltage phase angle, branch current, and active and reactive power of each distributed power source at each node in the distribution network; constructing a state vector from the collected data as the real-time operating status data.
3. The active voltage control method for a distribution network as described in claim 1, characterized in that, The voltage regulating equipment includes an on-load tap-changing transformer, a parallel capacitor bank, a photovoltaic inverter, and an energy storage device; the step of generating control commands according to the optimal control strategy and sending them to the voltage regulating equipment in the distribution network specifically includes: generating tap adjustment commands for the on-load tap-changing transformer, switching commands for the parallel capacitor bank, and reactive power adjustment commands for the photovoltaic inverter and the energy storage device.
4. The active voltage control method for a distribution network as described in claim 1, characterized in that, Before inputting the real-time running status data into the pre-trained deep reinforcement learning agent, the process also includes a training step for the deep reinforcement learning agent: An active voltage control model for the distribution network is constructed, which includes a state space, an action space, and a reward function. Obtain historical operating data of the power distribution network; A preset intelligent optimization algorithm is used to train the agent offline using the historical running data until the agent converges, thus obtaining the pre-trained deep reinforcement learning agent.
5. The active voltage control method for a distribution network as described in claim 4, characterized in that, The specific logic for constructing the reward function is as follows: Set voltage deviation penalty, network loss penalty, and control action penalty; When the voltage at a distribution network node exceeds the preset safety range, the value of the voltage deviation penalty term is increased; When the network loss exceeds a preset threshold or the number of voltage regulating device actions exceeds a preset limit, the value of the network loss penalty term or control action penalty term is increased accordingly to guide the agent to output the optimal control strategy that reduces network loss and ensures smooth operation.
6. The active voltage control method for a distribution network as described in claim 4, characterized in that, The process of employing a preset intelligent optimization algorithm and using the historical operational data to train the agent offline specifically includes: The intelligent optimization algorithm is described using either a deep deterministic policy gradient algorithm or a near-end policy optimization algorithm. The agent interacts with the environment, selects actions based on the current state, and receives a reward value and the next state from the environment. The network parameters of the agent are updated by minimizing the loss function of the value function.
7. The active voltage control method for a distribution network as described in claim 1, characterized in that, The step of using the deep reinforcement learning agent to output the optimal control strategy based on the real-time operating status data specifically includes: mapping the real-time operating status data to the current state of the deep reinforcement learning agent; the agent directly calculating and outputting the action value corresponding to the current state based on the current policy network, as the optimal control strategy, wherein the action value includes the continuous or discrete control quantities of each voltage regulating device.
8. The active voltage control method for a distribution network as described in claim 1, characterized in that, The method also includes an online update step: After adjusting the node voltages of the distribution network according to the optimal control strategy, the actual operating effect of the distribution network is monitored in real time. If the actual operating effect meets the preset triggering conditions, the deep reinforcement learning agent is fine-tuned and updated online using the current real-time operating status data.
9. The active voltage control method for a distribution network as described in claim 1, characterized in that, The method is applied to centralized controllers or distributed local controllers in power distribution networks; When applied to distributed local controllers, each controller only obtains the operating status data of its local and neighboring nodes, and outputs the locally optimal control strategy through a multi-agent cooperative mechanism.
10. An active voltage control system for a distribution network, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the active voltage control method for a distribution network as described in any one of claims 1-9.