Power distribution network reactive power optimization method and device based on deep reinforcement learning

By constructing a two-layer scheduling framework based on deep reinforcement learning, the actions of slow and fast adjustment devices are coordinated, solving the problem that traditional methods struggle to balance economy and security in distribution networks with a high proportion of renewable energy access, and achieving efficient and safe operation of the distribution network.

CN121395408APending Publication Date: 2026-01-23STATE GRID BEIJING ELECTRIC POWER CO
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511585108.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Traditional reactive power optimization methods struggle to simultaneously consider the economic efficiency, safety, and equipment lifespan of the system when faced with a high proportion of renewable energy integration into the distribution network, and they also lack the ability to respond in real time to load and renewable energy output fluctuations.

Method used

A two-layer scheduling framework based on deep reinforcement learning is adopted. By using day-ahead mixed integer second-order cone programming (MISOCP) and intraday multi-agent deep deterministic policy gradient (MADDPG) algorithms, the actions of slow and fast adjustment equipment are coordinated to achieve reactive power optimization at multiple time scales.

Benefits of technology

While ensuring grid voltage quality and operational safety, the total system operating cost is significantly reduced, frequent equipment operation is decreased, and operational economy and flexibility are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121395408A_ABST
    Figure CN121395408A_ABST
Patent Text Reader

Abstract

The invention provides a power distribution network reactive power optimization method and device based on deep reinforcement learning, and the method comprises the steps: building a mixed integer second-order cone programming model according to the state parameters of low-speed reactive power regulation equipment in a power distribution network, solving the mixed integer second-order cone programming model, and obtaining the day-ahead optimal scheduling result of each piece of low-speed reactive power regulation equipment; based on the day-ahead optimal scheduling result and the operation characteristics of the current power distribution network, establishing a multi-agent reinforcement learning model with the goal of minimizing the current operation cost of the power distribution network; a multi-agent depth deterministic strategy gradient algorithm is adopted to train the multi-agent reinforcement learning model, and an intra-day multi-agent real-time scheduling strategy is generated; and fusing the day-ahead optimal scheduling result with the intra-day multi-agent real-time scheduling strategy, and outputting a multi-time-scale coordinated power distribution network reactive power optimization control instruction. According to the technical scheme, a multi-time-scale coordinated reactive power optimization mechanism considering day-ahead economical efficiency and intra-day real-time performance is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of railway transportation and power system dispatch automation technology, specifically relating to a method and device for reactive power optimization of distribution networks based on deep reinforcement learning. Background Technology

[0002] With the deepening implementation of the "dual carbon" target, the penetration rate of distributed photovoltaic, wind power, and other renewable energy sources in the distribution network has significantly increased. The inherent volatility and intermittency of distributed energy sources pose serious challenges to the reactive power balance and voltage stability of the distribution network. Inappropriate reactive power dispatch strategies can easily lead to frequent operation of compensation devices, insufficient or excessive reactive power output, thereby increasing system network losses, widening voltage deviations, and even threatening power supply security. Traditional reactive power optimization methods heavily rely on precise mathematical models and manually preset control rules. When faced with the strong uncertainty and multi-timescale fluctuations brought about by the high proportion of new energy integration, they often lack sufficient real-time performance and adaptability, making it difficult to simultaneously consider the economic efficiency, safety, and equipment lifespan of system operation.

[0003] Therefore, there is an urgent need for a reactive power optimization method that coordinates multiple time scales and multiple equipment types to improve the operational reliability, economy, and control flexibility of high-proportion renewable energy distribution networks. Summary of the Invention

[0004] This invention provides a reactive power optimization method for distribution networks based on deep reinforcement learning. The scheme constructs a two-layer scheduling framework of "day-ahead mixed integer second-order cone programming (MISOCP) + intraday multi-agent deep deterministic policy gradient (MADDPG)". First, in the day-ahead phase, for slow-regulating devices such as on-load tap changers (OLTCs) and capacitor banks (CBs), a mixed integer second-order cone programming (MISOCP) model is established to minimize the system's total daily network loss and voltage deviation, generating a globally optimal 24-hour scheduling plan. Second, in the intraday phase, for fast-regulating devices such as voltage source converters (VSCs) with rapid response, the MADDPG algorithm is used. Each VSC is treated as an independent agent, learning and making decisions online through real-time interaction with the grid environment to quickly respond to real-time fluctuations in load and renewable energy output, forming a complete multi-timescale coordinated reactive power optimization mechanism that balances day-ahead economy and intraday real-time performance.

[0005] In a first aspect, the present invention provides a reactive power optimization method for distribution networks based on deep reinforcement learning, the method comprising: Based on the state parameters of the slow reactive power regulating equipment in the distribution network, a mixed integer second-order cone programming model is established with the objective of minimizing the sum of the total daily network loss and the total node voltage deviation. The mixed integer second-order cone programming model is solved based on the tap change constraint, voltage safety constraint and power flow constraint of the slow reactive power regulating equipment to obtain the day-ahead optimal scheduling result of each of the slow reactive power regulating equipment. The current operating characteristics of the distribution network are obtained, and the current operating characteristics of the distribution network and the day-ahead optimal scheduling result are input into a pre-trained multi-agent reinforcement learning model with the objective of minimizing the current operating cost of the distribution network. The model outputs a day-ahead real-time multi-agent scheduling strategy. The multi-agent reinforcement learning model includes an independent agent corresponding to each fast adjustment device, and the multi-agent deep deterministic policy gradient algorithm is used to train the multi-agent reinforcement learning model. The day-ahead optimal scheduling result is fused with the intraday multi-agent real-time scheduling strategy to output a multi-timescale coordinated reactive power optimization control command for the distribution network.

[0006] By employing the above scheme, this invention presents a reactive power optimization method for distribution networks based on deep reinforcement learning. It constructs a two-layer scheduling framework of "day-ahead mixed integer second-order cone programming (MISOCP) + intraday multi-agent deep deterministic policy gradient (MADDPG)". First, in the day-ahead phase, a MISOCP model is established for slow-regulating devices such as on-load tap changers (OLTCs) and capacitor banks (CBs) to minimize the total daily network loss and voltage deviation, generating a globally optimal 24-hour scheduling plan. Second, in the intraday phase, for fast-regulating devices such as voltage source converters (VSCs) with rapid response, the MADDPG algorithm is used. Each VSC is treated as an independent agent, learning and making decisions online through real-time interaction with the grid environment to quickly respond to real-time fluctuations in load and renewable energy output. This invention achieves coordinated actions of devices with different response speeds, achieving optimal economic efficiency while ensuring system safety constraints, and improving the adaptive capability of the scheduling strategy to uncertain disturbances.

[0007] In some embodiments of the present invention, the objective function for minimizing the sum of the total daily network loss and the total node voltage deviation of the distribution network is: , In the formula, For time period t network loss, branch road l The resistance, branch road l The square of the current, For nodes i During the period tVoltage deviation, Represents a node i At any moment t The square of the voltage, As a penalty weight for voltage deviation, m For the number of branch roads, n This represents the number of nodes.

[0008] In some embodiments of the present invention, the reactive power optimization method for distribution networks based on deep reinforcement learning according to claim 1 is characterized in that the day-ahead optimal scheduling result includes OLTC tap curves, capacitor bank switching plans, and node voltage reference trajectories.

[0009] In some embodiments of the present invention, the distribution network minimizes its current operating cost using the following reward function formula:

[0010]

[0011] In the formula, This is a penalty item for voltage exceeding the limit. C Loss This is the marginal network loss coefficient. C β An additional cost factor is added for voltage exceeding limits. , These are the upper and lower limits for safe voltage operation in the system, respectively. n This represents the number of nodes.

[0012] In some embodiments of the present invention, training the multi-agent reinforcement learning model using a multi-agent deep deterministic policy gradient algorithm includes: the multi-agent reinforcement learning model employs a combined structure of an Actor network and a Critic network; the Actor network generates reactive power control action instructions for each agent based on the state information of each agent; the Critic network receives the state information and reactive power control action instructions of all agents and generates value assessment information for the reactive power control action instructions of all agents; wherein the action instructions generated by the Actor network are updated based on the value assessment information output by the Critic network.

[0013] In some embodiments of the present invention, the multi-agent deep deterministic policy gradient algorithm includes: setting up a circular buffer to store interaction samples based on a continuous action space; introducing OU noise into the interaction samples to generate an experience replay sample set.

[0014] In some embodiments of the present invention, the multi-agent deep deterministic policy gradient algorithm further includes: A preset batch of training sample datasets is extracted from the experience replay sample set, and the training sample set is input into the Actor network to generate a set of reactive power control action instructions for all agents based on the training sample dataset. The Critic network evaluates the value of the reactive power control action instruction set of all agents based on the training sample dataset, and generates the current predictive value evaluation of the Critic network. Calculate the target value assessment corresponding to the target Critic network based on the predictive value assessment of the current Critic network; The goal is to minimize the mean square error between the current Critic network's predicted value assessment and the target value assessment, and the parameters of the current Critic network are updated using the gradient descent method.

[0015] In some embodiments of the present invention, the method further includes updating the current Actor network parameters by calculating the deterministic policy gradient and using the gradient ascent method after updating the Critic network parameters, with the goal of maximizing the predictive value assessment of the current Critic network.

[0016] In some embodiments of the present invention, the method further includes, after updating the current Actor network parameters and the current Critic network parameters, making a small adjustment to the target Actor network parameters and the target Critic network parameters in a soft update manner towards the updated current Actor network parameters and the current Critic network parameters; wherein, the soft update is implemented by the following formula:

[0017] , In the formula, To update the coefficients, ≤1 indicates that a small step size update is used. θ and φ These represent the parameters of the current Actor network and the current Critic network, respectively. θ' and φ' These represent the parameters of the target Actor network and the target Critic network, respectively.

[0018] Compared with existing technologies, the beneficial effects of this invention lie in its construction of a two-layer scheduling framework of "Day-ahead Mixed Integer Second-Order Cone Programming (MISOCP) + Intra-day Multi-Agent Deep Deterministic Policy Gradient (MADDPG)". Through time-scale decomposition, this framework effectively coordinates the planned actions of slow-speed equipment with the real-time responses of fast-speed equipment. The technical solution provided by this invention first performs day-ahead global optimization by solving the MISOCP model, providing an economically optimal benchmark for intra-day operation while ensuring compliance with physical constraints such as the number of equipment actions. Subsequently, during the intra-day phase, the deep online self-learning capability of the MADDPG algorithm enables rapid and accurate tracking of load and renewable energy output fluctuations, further reducing real-time network losses and voltage deviations. This invention can significantly reduce the total operating cost of the system and decrease unnecessary frequent equipment actions while ensuring grid voltage quality and operational safety, thereby achieving a dual improvement in operational economy and safety. It provides an efficient and feasible technical solution for the intelligent and refined scheduling of modern distribution networks.

[0019] A second aspect of the present invention provides a reactive power optimization system for a distribution network based on deep reinforcement learning, comprising: The day-ahead optimal scheduling module is used to establish a mixed integer second-order cone programming model with the objective of minimizing the sum of the total daily network loss and the total node voltage deviation, based on the state parameters of the slow reactive power regulating equipment in the distribution network. The mixed integer second-order cone programming model is solved based on the tap change constraint, voltage safety constraint and power flow constraint of the slow reactive power regulating equipment to obtain the day-ahead optimal scheduling result of each of the slow reactive power regulating equipment. Multi-agent reinforcement learning modeling module: Based on the day-ahead optimal scheduling results and the current operating characteristics of the distribution network, it models each fast-regulating device in the distribution network as an independent agent, establishes a multi-agent reinforcement learning model with the goal of minimizing the current operating cost of the distribution network, and trains the multi-agent reinforcement learning model using a multi-agent deep deterministic policy gradient algorithm. Real-time strategy generation module: used to acquire the current operating characteristics of the distribution network, input the current operating characteristics of the distribution network and the day-ahead optimal scheduling result into a pre-trained multi-agent reinforcement learning model with the objective of minimizing the current operating cost of the distribution network, and output the intraday multi-agent real-time scheduling strategy; wherein, the multi-agent reinforcement learning model includes an independent agent corresponding to each fast adjustment device, and the multi-agent deep deterministic policy gradient algorithm is used to train the multi-agent reinforcement learning model; Multi-timescale coordinated control module: used to integrate the day-ahead optimal scheduling result with the intraday multi-agent real-time scheduling strategy, and output multi-timescale coordinated reactive power optimization control instructions for the distribution network.

[0020] A third aspect of the present invention provides a power distribution network reactive power optimization device based on deep reinforcement learning, characterized in that the device includes a computer device, the computer device includes a processor and a memory, the processor stores computer instructions, and when the computer instructions are executed, the device implements the power distribution network reactive power optimization method based on deep reinforcement learning.

[0021] Additional advantages, objects, and features of the invention will be set forth in part in the description which follows, and will also become apparent in part to those skilled in the art upon studying the text, or may be learned by practice of the invention. The objects and other advantages of the invention will become apparent from the description and the accompanying drawings.

[0022] Those skilled in the art will understand that the objectives and advantages achievable with the present invention are not limited to those specifically described above, and that the above and other objectives achievable with the present invention will become clearer from the following detailed description. Attached Figure Description

[0023] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0024] In the attached diagram: Figure 1 This is a flowchart illustrating a reactive power optimization method for power distribution networks based on deep reinforcement learning, provided as an embodiment of the present invention.

[0025] Figure 2 This is a schematic diagram of a power distribution network reactive power optimization system based on deep reinforcement learning, provided as an embodiment of the present invention.

[0026] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0027] The present invention will now be described in detail with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.

[0028] The following detailed description is exemplary and intended to provide further detailed explanation of the invention. Unless otherwise specified, all technical terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terminology used in this invention is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention.

[0029] Figure 1This is a flowchart illustrating a reactive power optimization method for power distribution networks based on deep reinforcement learning, provided in an embodiment of the present invention.

[0030] Example 1, such as Figure 1 As shown, this invention provides a reactive power optimization method for distribution networks based on deep reinforcement learning, the method comprising the following steps: S1: Based on the state parameters of the slow reactive power regulating equipment in the distribution network, establish a mixed integer second-order cone programming model with the objective of minimizing the sum of the total daily network loss and the total node voltage deviation. Solve the mixed integer second-order cone programming model based on the tap change constraint, voltage safety constraint and power flow constraint of the slow reactive power regulating equipment to obtain the day-ahead optimal scheduling result of each of the slow reactive power regulating equipment.

[0031] S2: Obtain the current operating characteristics of the distribution network, input the current operating characteristics of the distribution network and the day-ahead optimal scheduling result into a pre-trained multi-agent reinforcement learning model with the goal of minimizing the current operating cost of the distribution network, and output the intraday multi-agent real-time scheduling strategy; wherein, the multi-agent reinforcement learning model includes an independent agent corresponding to each fast adjustment device, and the multi-agent deep deterministic policy gradient algorithm is used to train the multi-agent reinforcement learning model; S4: Integrate the day-ahead optimal scheduling result with the intraday multi-agent real-time scheduling strategy to output a multi-timescale coordinated reactive power optimization control command for the distribution network.

[0032] By employing the above scheme, this invention presents a reactive power optimization method for distribution networks based on deep reinforcement learning. It constructs a two-layer scheduling framework of "day-ahead mixed integer second-order cone programming (MISOCP) + intraday multi-agent deep deterministic policy gradient (MADDPG)". First, in the day-ahead phase, a MISOCP model is established for slow-regulating devices such as on-load tap changers (OLTCs) and capacitor banks (CBs) to minimize the total daily network loss and voltage deviation, generating a globally optimal 24-hour scheduling plan. Second, in the intraday phase, for fast-regulating devices such as voltage source converters (VSCs) with rapid response, the MADDPG algorithm is used. Each VSC is treated as an independent agent, learning and making decisions online through real-time interaction with the grid environment to quickly respond to real-time fluctuations in load and renewable energy output. This invention achieves coordinated actions of devices with different response speeds, achieving optimal economic efficiency while ensuring system safety constraints, and improving the adaptive capability of the scheduling strategy to uncertain disturbances.

[0033] Specifically, step S1 includes the following sub-steps: Sub-step A1: Establish the objective function. The objective is to minimize the sum of the total network loss and the total node voltage deviation penalty over 24 hours.

[0034] In some implementations, the objective function for minimizing the sum of the total daily network loss and the total node voltage deviation of the distribution network is:

[0035] In the formula, For time period t network loss, branch road l The resistance, branch road l The square of the current, For nodes i During the period t Voltage deviation, Represents a node i At any moment t The square of the voltage, As a penalty weight for voltage deviation, m For the number of branch roads, n This represents the number of nodes.

[0036] Sub-step A2: Establish constraints.

[0037] It mainly includes gear shift constraints, voltage safety constraints, and power flow constraints.

[0038] To ensure that the tap changers of the on-load tap changer (OLTC) move within the allowable range of the equipment and to limit the number of tap change steps between adjacent times, thereby ensuring the flexibility of tap changing and the lifespan of the equipment, the following constraints are introduced:

[0039]

[0040] In the formula, For OLTC tap positions, The maximum number of bits allowed for the upper and lower gears in OLTC.

[0041] For the group switching capacitors (CB) at the access node, it is necessary to limit the number of units and calculate their reactive power injection to reflect the equipment capacity and combination characteristics, and limit the number of capacitors in the capacitor bank at the access node:

[0042]

[0043] In the formula, For the first k The maximum number of capacitor banks that can be switched on. For the first k Capacity of a single capacitor in a group.

[0044] To ensure that the node voltage remains within the safe operating range and thus guarantee the safe and stable operation of the system, the upper and lower limits of the voltage square of the second-order cone model are constrained as follows:

[0045] In the formula, V max , V min These are the upper and lower limits of the square of the node voltage, respectively.

[0046] Considering the voltage drop of each branch in the distribution network and the impact of branch current on network transmission capacity and losses, a second-order cone constraint is introduced. Each branch... l (from node) i To the node j Voltage drop equation:

[0047] In the formula, x l branch road l Reactance, u i,t For the first node of the branch road i At any moment t The square of the voltage, u j,t For the end node of the branch j At any moment t The square of the voltage, P l,t , Q l,t Branch roads l At any moment t The active and reactive power.

[0048] Each node needs to satisfy both active power balance and reactive power balance, with the following constraints:

[0049]

[0050] In the formula, p(i) For nodes i Parent branch index, For nodes i The set of all sub-branches, P p(i),t For nodes i All parent branches at time t The active power input to it, Qp(i),t For nodes i All parent branches at time t The reactive power input to it, For nodes i At any moment t The active power delivered to all its sub-branches, For nodes i At any moment t The reactive power delivered to all its sub-branches, Q C,i,t For nodes i At any moment t The capacitor bank injected power, P D,i,t For nodes i At any moment t Active load, Q D,i,t For nodes i At any moment t The reactive load.

[0051] To ensure the physical relationship between power flow and current and voltage, a second-order cone constraint is established for the model to approximate the actual power flow under relaxation conditions: .

[0052] In some embodiments of the present invention, the day-ahead optimal scheduling result includes the OLTC tap curve, capacitor bank switching plan, and node voltage reference trajectory.

[0053] The above method uses the MISOCP model to perform day-ahead global optimization, and provides an economically optimal benchmark for subsequent intraday operations while ensuring that physical constraints such as the number of equipment actions are met.

[0054] Based on the current optimal scheduling result, step S2 includes constructing a multi-agent reinforcement learning model and training the multi-agent reinforcement learning model.

[0055] Furthermore, step S2 includes the following sub-steps: Sub-step B1: Define the multi-agent environment.

[0056] The N VSCs in the distribution network are considered as N cooperating intelligent agents. The time scale is refined to 15 minutes, and each intelligent agent interacts with the power grid environment.

[0057] Sub-step B2: Construct the state vector.

[0058] State vector s tThe system aims to comprehensively reflect the current operating characteristics of the system, including but not limited to information such as the voltage of each node, voltage change rate, branch loss, load and photovoltaic output, current time, day-ahead optimal scheduling result, and historical output of each VSC.

[0059] For example:

[0060] In the formula, V t Let be the voltage at each node at time t. ΔV t The voltage change rate of each node (limited by the node voltage reference trajectory). L t For each branch road at all times t The loss, The load variation coefficient, PV t For photovoltaic access nodes at any time t Photovoltaic power output, VSC t For VSC access nodes at time t The amount of reactive power compensation, tp t This refers to the OLTC gear in the current optimal scheduling results. CB t For capacitor bank switching plans.

[0061] Sub-step B3: Define the action space.

[0062] Each agent i The action is the adjustment of the reactive power output of the VSC it controls in the current time period. a i,t This action is a continuous value and is subject to the upper and lower limits of the VSC capacity.

[0063]

[0064] In the formula, ΔQ max This is the maximum single-step adjustable reactive power capacity. The raw action value directly output by the agent, with a value range of [ 1,1].

[0065] Sub-step B4: Design the reward function. Reward function r t Used to guide intelligent agents in learning, its goal is to minimize the system's current operating cost, defined as the sum of negative network loss cost and voltage over-limit penalty cost.

[0066] In some embodiments of the present invention, the distribution network minimizes its current operating cost using the following reward function formula:

[0067]

[0068] In the formula, This is a penalty item for voltage exceeding the limit. C Loss This is the marginal network loss coefficient. C β An additional cost factor is added for voltage exceeding limits. , These are the upper and lower limits for safe voltage operation in the system, respectively. n This represents the number of nodes.

[0069] In some embodiments of the present invention, training the multi-agent reinforcement learning model using a multi-agent deep deterministic policy gradient algorithm includes: the multi-agent reinforcement learning model employs a combined structure of an Actor network and a Critic network; the Actor network generates reactive power control action instructions for each agent based on the state information of each agent; the Critic network receives the state information and reactive power control action instructions of all agents and generates value assessment information for the reactive power control action instructions of all agents; wherein the action instructions generated by the Actor network are updated based on the value assessment information output by the Critic network.

[0070] Specifically, the Actor network is constructed (each agent has the same structure):

[0071]

[0072]

[0073]

[0074] In the formula, s This is the local state vector of the current agent; W k , b k The first k Layer weight matrix and bias terms; ReLU ( x )=max(0, x ) is the element-wise rectified activation function; tanh(x) is the hyperbolic tangent function, used to constrain the output to [ 1,1]; h k For the first kHidden layer state; This is the action vector output by the Actor network.

[0075] Critic network construction (same structure for each agent):

[0076]

[0077]

[0078]

[0079]

[0080] In the formula, x This is the concatenated global state and action vectors of all agents; , These are the Critic network numbers. k Layer weight matrix and bias; h k For the first k Layered hidden representation; Q ( x ) is the input for Critic x Value estimate.

[0081] In some embodiments of the present invention, the multi-agent deep deterministic policy gradient algorithm includes: setting a circular buffer (Replay Buffer) to store interaction samples based on a continuous action space; introducing OU noise into the interaction samples to generate an experience replay sample set.

[0082] Specifically, the Replay Buffer uses a circular buffer to store interactive samples, in the following format:

[0083] In the formula, s t For a moment t state, a t For a moment t joint actions, r t The immediate reward obtained after performing an action. s t+1 For the state at the next moment, d t This is the end marker.

[0084] In the continuous action space, to generate smooth exploratory noise with temporal correlation, the Ornstein–Uhlenbeck(OU) process is introduced, whose continuous form is:

[0085] Discretized as:

[0086] In the formula, x t In a noisy state, θ The mean regression rate, μ This represents the long-term mean of the noise. σ The noise intensity coefficient, dW t For Wiener process increment, t For time step.

[0087] To train the Actor and Critic networks in the aforementioned multi-agent reinforcement learning model, the Actor and Critic networks need to undergo repeated trial and error and experience learning. This allows the agents (i.e., the rapid adjustment devices) to gradually evolve from an initial random and inefficient control strategy to an efficient and collaborative real-time control capability. This compensates for the real-time fluctuations that day-ahead scheduling (step S1) cannot handle, and ultimately achieves the "multi-timescale coordinated optimization" desired in step S4.

[0088] Furthermore, in some embodiments of the present invention, the multi-agent deep deterministic policy gradient algorithm further includes: A preset batch of training sample datasets is extracted from the experience replay sample set, and the training sample set is input into the Actor network to generate a set of reactive power control action instructions for all agents based on the training sample dataset. The Critic network evaluates the value of the reactive power control action instruction set of all agents based on the training sample dataset, and generates the current predictive value evaluation of the Critic network. Calculate the target value assessment corresponding to the target Critic network based on the predictive value assessment of the current Critic network; The goal is to minimize the mean square error between the current Critic network's predicted value assessment and the target value assessment, and the parameters of the current Critic network are updated using the gradient descent method.

[0089] Specifically, to avoid time correlation and provide diverse training samples, a sample batch of size B is uniformly drawn from the experience replay pool in each training iteration:

[0090] In the formula, B For batch size, d k This is a termination marker.

[0091] Generate the next joint action using the target Actor network And calculate the first i The target of the sample Q value:

[0092] In the formula, For the target Actor network in The resulting joint action For the first i The goal of an intelligent agent Q value, r k For the first k Instant rewards for each experience point.

[0093] The above scheme randomly extracts a batch of past scenario data fragments from the algorithm's "memory" (experience replay pool), such as: the power grid state s at a certain moment, the reactive power adjustment actions a taken by each agent, the power grid reward r after adjustment (such as reduced grid loss and more stable voltage), and the new power grid state s' after adjustment. Random sampling avoids the "mental fixation" caused by continuously learning from the most recent similar scenarios, making the learning process more stable.

[0094] The current Critic network parameters are updated using gradient descent by calculating the mean square error that minimizes the difference between the current Critic and the target Q-value, i.e.:

[0095] In the formula, The learning rate for the Critic network. a k For the first k Action vectors in a set of experiences.

[0096] Update the parameters of the current Critic network so that the Critic network's prediction of the action value (Q value) under the current policy is closer to the calculated "benchmark" (target Q value).

[0097] The Critic network implements a set of reactive power regulation actions taken jointly by all fast devices, and evaluates the contribution of a set of reactive power regulation actions to "minimizing the current operating cost" (the objective in step S2).

[0098] In some embodiments of the present invention, the method further includes updating the current Actor network parameters by calculating the deterministic policy gradient and using the gradient ascent method after updating the Critic network parameters, with the goal of maximizing the predictive value assessment of the current Critic network.

[0099] Specifically, updating the current Actor network parameters using the gradient ascent method can adjust the Actor parameters in a direction that increases the Critic value, i.e.:

[0100] In the formula, is the learning rate of the Actor network.

[0101] The Actor network is the "brain" of each fast reactive power regulation device. This step tells these "brains": "Based on the judgment of the 'Critic network' just now, you should now adjust your strategy to make reactive power regulation decisions that will achieve higher evaluations in the future (i.e., more effective reduction of network losses and voltage stabilization)." For example, learning "When the node voltage is too high, I should absorb reactive power instead of emitting reactive power."

[0102] In some embodiments of the present invention, the method further includes, after updating the current Actor network parameters and the current Critic network parameters, making a small adjustment to the target Actor network parameters and the target Critic network parameters in a soft update manner towards the updated current Actor network parameters and the current Critic network parameters; wherein, the soft update is implemented by the following formula:

[0103] , In the formula, To update the coefficients, ≤1 indicates that a small step size update is used. θ and φ These represent the parameters of the current Actor network and the current Critic network, respectively. θ' and φ' These represent the parameters of the target Actor network and the target Critic network, respectively.

[0104] In reinforcement learning, the "benchmark (target Q-value)" itself must be relatively stable. This step is like adding inertia to the "scoring standard," preventing it from changing drastically with each improvement of the current network. If the benchmark changes too quickly (i.e., it is directly overwritten without soft updates), the learning objective will become erratic, leading to policy oscillations or even failure, making it impossible to generate the reliable "intraday multi-agent real-time scheduling policy" required by S2.

[0105] Through the above training and optimization, the intelligent agent in the distribution network can learn from historical experience, continuously correct its judgment of the value of actions, optimize its decision-making strategy based on more accurate judgment, and steadily improve its decision-making strategy within a stable framework.

[0106] Compared with the prior art, the beneficial effects of the present invention are as follows: The technical solution of this invention constructs a two-layer scheduling framework of "Day-ahead Mixed Integer Second-Order Cone Programming (MISOCP) + Intra-day Multi-Agent Deep Deterministic Policy Gradient (MADDPG)". Through time-scale decomposition, it effectively coordinates the planned actions of slow devices with the real-time responses of fast devices. The technical solution provided by this invention first performs day-ahead global optimization by solving the MISOCP model. While ensuring physical constraints such as the number of device actions are met, it provides an economically optimal benchmark for intra-day operation. Subsequently, during the intra-day phase, the deep online self-learning capability of the MADDPG algorithm is utilized to achieve rapid and accurate tracking of load and renewable energy output fluctuations, further reducing real-time network losses and voltage deviations. This invention can significantly reduce the total operating cost of the system and reduce unnecessary frequent device actions while ensuring grid voltage quality and operational safety, thereby achieving safe, economical, and efficient operation of the distribution network. It provides an efficient and feasible technical solution for the intelligent and refined scheduling of modern distribution networks.

[0107] Figure 2 This is a flowchart illustrating a power distribution network reactive power optimization system based on deep reinforcement learning, provided by an embodiment of the present invention.

[0108] Example 2, as Figure 2 As shown, the present invention also provides a reactive power optimization system for distribution networks based on deep reinforcement learning, including: a day-ahead optimization scheduling module S11, a multi-agent reinforcement learning modeling module S12, a real-time policy generation module S13, and a multi-timescale coordination control module S14.

[0109] The day-ahead optimal scheduling module S11 is used to establish a mixed integer second-order cone programming model with the goal of minimizing the sum of the total daily network loss and the total node voltage deviation of the distribution network based on the state parameters of the slow reactive power regulating equipment in the distribution network. Based on the tap change constraint, voltage safety constraint and power flow constraint of the slow reactive power regulating equipment, the mixed integer second-order cone programming model is solved to obtain the day-ahead optimal scheduling result of each of the slow reactive power regulating equipment. Multi-agent reinforcement learning modeling module S12: Based on the day-ahead optimal scheduling results and the current operating characteristics of the distribution network, it models each fast-regulating device in the distribution network as an independent agent, establishes a multi-agent reinforcement learning model with the goal of minimizing the current operating cost of the distribution network, and trains the multi-agent reinforcement learning model using a multi-agent deep deterministic policy gradient algorithm. Real-time strategy generation module S13: used to acquire the current operating characteristics of the distribution network, input the current operating characteristics of the distribution network and the day-ahead optimal scheduling result into a pre-trained multi-agent reinforcement learning model with the goal of minimizing the current operating cost of the distribution network, and output the intraday multi-agent real-time scheduling strategy; wherein, the multi-agent reinforcement learning model includes an independent agent corresponding to each fast adjustment device, and the multi-agent deep deterministic policy gradient algorithm is used to train the multi-agent reinforcement learning model; Multi-timescale coordinated control module S14: used to integrate the day-ahead optimal scheduling result with the intraday multi-agent real-time scheduling strategy, output multi-timescale coordinated power distribution network reactive power optimization control instructions, and send the reactive power optimization control instructions to the corresponding reactive power regulation equipment for execution.

[0110] Example 3: The present invention also provides a power distribution network reactive power optimization device based on deep reinforcement learning. The device includes a computer device, which includes a processor and a memory. The processor stores computer instructions. When the computer instructions are executed, the device implements the power distribution network reactive power optimization method based on deep reinforcement learning.

[0111] Example 4, as Figure 3 As shown, the present invention also provides an electronic device 100 for implementing a power distribution network reactive power optimization method based on deep reinforcement learning.

[0112] The electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on at least one processor 102, and at least one communication bus 104.

[0113] The memory 101 can be used to store the computer program 103. The processor 102 implements the steps of the power distribution network reactive power optimization method based on deep reinforcement learning described in the first aspect of the present invention by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101.

[0114] The memory 101 may primarily include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created based on the use of the electronic device 100 (such as audio data), etc. In addition, the memory 101 may include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other non-volatile solid-state storage device.

[0115] At least one processor 102 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 102 may be a microprocessor or any conventional processor. Processor 102 is the control center of electronic device 100, connecting various parts of electronic device 100 via various interfaces and lines.

[0116] The memory 101 in the electronic device 100 stores multiple instructions to implement a reactive power optimization method for a power distribution network based on deep reinforcement learning, and the processor 102 can execute multiple instructions to achieve the following: Based on the state parameters of the slow reactive power regulating equipment in the distribution network, a mixed integer second-order cone programming model is established with the objective of minimizing the sum of the total daily network loss and the total node voltage deviation. The mixed integer second-order cone programming model is solved based on the tap change constraint, voltage safety constraint and power flow constraint of the slow reactive power regulating equipment to obtain the day-ahead optimal scheduling result of each of the slow reactive power regulating equipment. The current operating characteristics of the distribution network are obtained, and the current operating characteristics of the distribution network and the day-ahead optimal scheduling result are input into a pre-trained multi-agent reinforcement learning model with the objective of minimizing the current operating cost of the distribution network. The model outputs a day-ahead real-time multi-agent scheduling strategy. The multi-agent reinforcement learning model includes an independent agent corresponding to each fast adjustment device, and the multi-agent deep deterministic policy gradient algorithm is used to train the multi-agent reinforcement learning model. The day-ahead optimal scheduling result is fused with the intraday multi-agent real-time scheduling strategy to output a multi-timescale coordinated reactive power optimization control command for the distribution network.

[0117] Example 5: If the modules / units integrated in the electronic device 100 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, and read-only memory (ROM).

[0118] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0119] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0120] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0121] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0122] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.

Claims

1. A reactive power optimization method for distribution networks based on deep reinforcement learning, characterized in that, Includes the following steps: Based on the state parameters of the slow reactive power regulating equipment in the distribution network, a mixed integer second-order cone programming model is established with the objective of minimizing the sum of the total daily network loss and the total node voltage deviation. The mixed integer second-order cone programming model is solved based on the tap change constraint, voltage safety constraint and power flow constraint of the slow reactive power regulating equipment to obtain the day-ahead optimal scheduling result of each of the slow reactive power regulating equipment. The current operating characteristics of the distribution network are obtained, and the current operating characteristics of the distribution network and the day-ahead optimal scheduling result are input into a pre-trained multi-agent reinforcement learning model with the objective of minimizing the current operating cost of the distribution network. The model outputs a day-ahead real-time multi-agent scheduling strategy. The multi-agent reinforcement learning model includes an independent agent corresponding to each fast adjustment device, and the multi-agent deep deterministic policy gradient algorithm is used to train the multi-agent reinforcement learning model. The day-ahead optimal scheduling result is fused with the intraday multi-agent real-time scheduling strategy to output a multi-timescale coordinated reactive power optimization control command for the distribution network.

2. The reactive power optimization method for distribution networks based on deep reinforcement learning according to claim 1, characterized in that, The sum of the total daily network loss and the total node voltage deviation of the distribution network J The objective function to be minimized is: , In the formula, For time period t Distribution network losses, branch road l The resistance, branch road l The square of the current, For nodes i During the period t Voltage deviation, Represents a node i At any moment t The square of the voltage, As a penalty weight for voltage deviation, m For the number of branch roads, n This represents the number of nodes.

3. The reactive power optimization method for distribution networks based on deep reinforcement learning according to claim 1, characterized in that, The day-ahead optimal scheduling results include the on-load tap-changing transformer tap position curve, capacitor bank switching plan, and node voltage reference trajectory.

4. The reactive power optimization method for distribution networks based on deep reinforcement learning according to claim 1, characterized in that, The distribution network minimizes its current operating cost using the following reward function formula: In the formula, r t For time period t The reward value, For time period t Voltage over-limit penalty item, C Loss This is the marginal network loss coefficient. C β An additional cost factor is added for voltage exceeding limits. , These are the upper and lower limits for safe voltage operation in the system, respectively. For time period t Distribution network losses, V i,t For nodes i During the period t voltage, n This represents the number of nodes.

5. The reactive power optimization method for distribution networks based on deep reinforcement learning according to any one of claims 1 to 4, characterized in that, The step of training the multi-agent reinforcement learning model using a multi-agent deep deterministic policy gradient algorithm includes: The multi-agent reinforcement learning model adopts a combination structure of Actor network and Critic network. The Actor network generates reactive power control action instructions for each agent based on the state information of each agent. The Critic network receives the state information of all agents and the reactive power control action instructions of all agents, and generates value evaluation information for the reactive power control action instructions of all agents. The action instructions generated by the Actor network are updated based on the value assessment information output by the Critic network.

6. The reactive power optimization method for distribution networks based on deep reinforcement learning according to claim 5, characterized in that, The multi-agent deep deterministic policy gradient algorithm includes: A circular buffer is used to store interaction samples based on the continuous action space; OU noise is introduced into the interaction samples to generate an experience replay sample set.

7. The reactive power optimization method for distribution networks based on deep reinforcement learning according to claim 6, characterized in that, The multi-agent deep deterministic policy gradient algorithm also includes: A preset batch of training sample datasets is extracted from the experience replay sample set, and the training sample set is input into the Actor network to generate a set of reactive power control action instructions for all agents based on the training sample dataset. The Critic network evaluates the value of the reactive power control action instruction set of all agents based on the training sample dataset, and generates the current predictive value evaluation of the Critic network. Calculate the target value assessment corresponding to the target Critic network based on the predictive value assessment of the current Critic network; The goal is to minimize the mean square error between the current Critic network's predicted value assessment and the target value assessment, and the parameters of the current Critic network are updated using the gradient descent method.

8. The reactive power optimization method for distribution networks based on deep reinforcement learning according to claim 7, characterized in that, It also includes updating the current Actor network parameters by calculating the deterministic policy gradient and using the gradient ascent method after updating the Critic network parameters, with the goal of maximizing the predictive value assessment of the current Critic network.

9. The reactive power optimization method for distribution networks based on deep reinforcement learning according to claim 8, characterized in that, This also includes, after updating the current Actor network parameters and the current Critic network parameters, using a soft update method to slightly adjust the target Actor network parameters and the target Critic network parameters to align with the updated current Actor network parameters and the current Critic network parameters; wherein, the soft update is implemented using the following formula: , In the formula, To update the coefficients, ≤1 indicates that a small step size update is used. θ and φ These represent the parameters of the current Actor network and the current Critic network, respectively. θ' and φ' These represent the parameters of the target Actor network and the target Critic network, respectively.

10. A reactive power optimization device for power distribution networks based on deep reinforcement learning, characterized in that, The device includes a computer device, which includes a processor and a memory. The processor stores computer instructions. When the computer instructions are executed, the device implements the power distribution network reactive power optimization method based on deep reinforcement learning as described in any one of claims 1 to 9.

Citation Information

Cited By

  • Power distribution network voltage control method based on reactive virtual power plant

    CN121791344A