A voltage and reactive power control method based on constraint heterogeneous ordered multi-agent reinforcement learning
By employing a constraint-based heterogeneous ordered multi-agent reinforcement learning method, a voltage regulator, switchable capacitor, and battery are coordinated for control. This solves the problem of inconsistent control of heterogeneous devices in existing technologies, thereby improving the performance of voltage reactive power control and enhancing system stability.
Patent Information
- Application Number
- CN202411439608.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-15
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-10-15
AI Technical Summary
Existing reinforcement learning methods assume that all agents are homogeneous, which cannot effectively coordinate heterogeneous devices in the power grid. This leads to inconsistent control strategies, cannot guarantee monotonic performance improvement, and results in system uncertainty and security risks.
A constraint-based heterogeneous ordered multi-agent reinforcement learning approach is adopted to control heterogeneous devices, including voltage regulators, switchable capacitors, and batteries, through multi-agent reinforcement learning algorithms. Reward functions and voltage-reactive power sensitivity analysis are designed to optimize control strategies and achieve coordinated control of devices.
It improves the performance of voltage and reactive power control, enhances the adaptability of the distribution network to heterogeneous equipment, reduces power loss, and ensures the monotonic improvement of the control strategy and system stability.
Smart Images

Figure CN119324469B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of power system control, and in particular to a voltage and reactive power control method based on constrained heterogeneous ordered multi-agent reinforcement learning. BACKGROUND
[0002] Modern power networks face challenges due to the wide range of power infrastructure components and various device types. These challenges are exacerbated by the increasing integration of renewable energy sources, which can introduce unstable voltage fluctuations. In addition, as the power grid expands, it becomes increasingly important to maintain more stable voltage levels and minimize power loss. To address these issues, researchers have proposed voltage and reactive power control methods based on reinforcement learning, which utilize the adaptive learning capabilities of agents to coordinate the operation of heterogeneous devices.
[0003] However, existing reinforcement learning methods typically assume that all agents are homogeneous, which severely limits their applicability in real power grids. There are many different types of devices in the power grid, such as voltage regulators, capacitor banks, and batteries, which have different action spaces and physical constraints, requiring a coordinated and consistent control strategy. In addition, existing algorithms often cannot guarantee monotonic performance improvement during training, which can lead to uncertainty in system performance and potential safety hazards in real-time operation. SUMMARY
[0004] To address the above problems, the present application aims to provide a voltage and reactive power control method based on constrained heterogeneous ordered multi-agent reinforcement learning, which can effectively improve the performance of voltage and reactive power control and enhance the adaptability of distribution networks to heterogeneous devices, providing an effective solution for optimal operation of distribution networks.
[0005] To achieve the above-mentioned purposes, the present application adopts the following technical solutions:
[0006] A voltage and reactive power control method based on constrained heterogeneous ordered multi-agent reinforcement learning, which controls heterogeneous devices in the distribution network through a multi-agent reinforcement learning algorithm to achieve dynamic optimization control of voltage and reactive power. The heterogeneous devices include voltage regulators, switchable capacitors, and batteries. The specific steps are as follows:
[0007] Heterogeneous multi-agent state space and action space modeling: one heterogeneous device is modeled as an agent, and a collection of multiple different heterogeneous devices is modeled as a heterogeneous multi-agent. The voltage and reactive power control problem of the distribution network is modeled as a heterogeneous multi-agent reinforcement learning problem, and the state space of each node and the action space of the heterogeneous devices are defined. The state space of each node includes node voltage, heterogeneous device operating state, and load demand. The action space of the heterogeneous devices specifically includes the gear state of the voltage regulator, the switch state of the switchable capacitor, and the charge and discharge state of the battery.
[0008] Designing reward function according to control target: the overall goal of voltage reactive power control is to maintain the voltage stability of the power distribution network by adjusting the voltage regulator, switchable capacitor and battery, reduce the reactive power loss, and meet the heterogeneous device operation constraints; design the reward function in combination with the overall goal of voltage reactive power control, comprehensively consider the punishment of voltage violation, reactive power loss and control error, and introduce Lagrange multiplier to process the constraint conditions, to ensure that the agent can effectively optimize the voltage reactive power control strategy in the training process, and ensure that the power distribution network meets the physical and operating limits;
[0009] Voltage-reactive power sensitivity analysis: according to the voltage-reactive power sensitivity matrix, the influence of each node on the voltage and reactive power of the power distribution network is calculated, and the control priority of each node is dynamically adjusted;
[0010] Determine the update order of multi-agent: based on the voltage-reactive power sensitivity, the nodes with greater influence on the power distribution network are selected to update their control strategies first, to improve the response speed and control accuracy of the power distribution network;
[0011] Device cooperative control: during the operation of the power distribution network, the voltage regulator adjusts the gear to control the node voltage; the switchable capacitor compensates the reactive power according to the demand; the battery adjusts the reactive power through charging and discharging to realize the reactive power balance and voltage stability of the power distribution network;
[0012] Strategy updating and optimization: the update order of multi-agent is calculated continuously, and the strategy network of multi-agent is updated according to the update order through the iterative training of multi-agent reinforcement learning algorithm, to realize the optimal control of each heterogeneous device under different operating conditions and improve the voltage reactive power control effect of the power distribution network;
[0013] Real-time control and feedback adjustment: in actual operation, each agent makes autonomous decisions according to the real-time observed state of the power distribution network, realizes dynamic control of voltage and reactive power, and adjusts the future control strategy according to the actual control effect.
[0014] Preferably, the multi-agent reinforcement learning algorithm adopts a proximal policy optimization algorithm with monotonic improvement characteristics, ensuring that the voltage reactive power control effect of the power distribution network is continuously optimized after each strategy update, avoiding instability of the power distribution network caused by immature strategies. The cooperative control strategy of heterogeneous devices is optimized through centralized training and distributed execution.
[0015] Preferably, the voltage-reactive power sensitivity matrix is used to evaluate the sensitivity of each node to voltage and reactive power, and the voltage-reactive power sensitivity matrix is dynamically updated according to the voltage fluctuation and load demand of the node, to ensure the flexibility and accuracy of the control strategy.
[0016] Preferably, the voltage regulator controls the node voltage by adjusting the gear, so that the distribution network voltage is maintained within a predetermined range, and the voltage fluctuation is reduced.
[0017] Preferably, the switchable capacitor is switched according to the reactive power demand of the distribution network, providing reactive power compensation to maintain the reactive power balance of the distribution network.
[0018] Preferably, the battery adjusts the reactive power output by controlling the charging and discharging state, to assist the voltage regulator and the switchable capacitor to achieve the reactive power balance of the distribution network.
[0019] The present application provides a voltage and reactive power control method based on constrained heterogeneous ordered multi-agent reinforcement learning, which has the following beneficial effects:
[0020] Modeling the voltage and reactive power regulation problem as a heterogeneous multi-agent reinforcement learning problem can effectively describe the dynamic characteristics of various heterogeneous devices in the distribution network, improving the control performance.
[0021] The voltage and reactive power sensitivity matrix is introduced to preferentially update the agent with high sensitivity at the node, enhancing the learning ability of the dynamic characteristics of the distribution network.
[0022] The multi-agent reinforcement learning algorithm based on strategy is used to train the agent, so that it can control multiple types of heterogeneous devices at the same time, improving the adaptability and robustness of the distribution network.
[0023] During the training process, the performance index of the multi-agent reinforcement learning algorithm can be monotonically improved, and finally converges to the optimal value, realizing autonomous optimization.
[0024] This method can effectively improve the performance of voltage and reactive power control, reduce power loss, and provide an effective solution for optimal operation of the distribution network. BRIEF DESCRIPTION OF DRAWINGS
[0025] Figure 1 The workflow diagram of the voltage and reactive power control method based on constrained heterogeneous ordered multi-agent reinforcement learning. DETAILED DESCRIPTION
[0026] The present application will be further described in detail below in combination with the drawings and specific embodiments.
[0027] As Figure 1 shown, the present application provides a voltage and reactive power control method based on constrained heterogeneous ordered multi-agent reinforcement learning, mainly used in the distribution network, which realizes the optimization of voltage and reactive power through the collaborative control of heterogeneous devices such as voltage regulators, switchable capacitors and batteries. The following are the specific implementation steps:
[0028] Step S1: Modeling of heterogeneous multi-agent state space and action space
[0029] When modeling the voltage reactive power regulation problem as a heterogeneous multi-agent reinforcement learning problem, a state space containing various observations needs to be defined. These observations include node bus voltage, voltage regulator tap state, switchable capacitor state, battery bank load state, and battery bank discharge power. Each state can be represented as a binary vector or a continuous state normalized to a bounded range depending on the device type. The node bus voltage is represented using the base voltage unit on the node bus. Therefore, the state space can be represented as a set containing the above-mentioned various observations as follows:
[0030]
[0031] s = (s1, s2, s3, s4, s5)
[0032] where, V i represents the bus voltage of the i-th node, and N represents the number of nodes; a R i represents the i-th voltage regulator tap state, and N reg represents the number of voltage regulator taps; a C i represents the i-th switchable capacitor state, and N cap represents the number of switchable capacitors; a B i represents the i-th battery bank load state, a B i represents the i-th battery bank discharge power, and N bat represents the number of battery banks; and Concats represents concatenating state vectors.
[0033] The action space of heterogeneous devices is defined as a set of control variables that can influence the state of the distribution network within certain limits. In the distribution network, the action space is the state of each node and the control variables of heterogeneous devices. Specifically, the control of each switchable capacitor is represented by a switch variable a C ∈{0,1}. For voltage regulators, each tap has a controllable on / off state, so the action space of the voltage regulator is a binary vector a R with dimensions from 0 to N tap , where N tap is the number of taps, and the default value is 19. The action of the battery is discretized as a binary vector a bat with N dimensions. Similar to the voltage regulator, the default value of N bat is also set to 19, where the first 10 dimensions represent charging, and the rest represent discharging.
[0034] Step S2: Designing a reward function according to the control objective
[0035] The control target of the method is to maintain the voltage within the specified range and avoid the occurrence of voltage violations. And optimize the reactive power compensation, reduce the power loss of the power distribution system.
[0036] The present application adopts a multi-agent reinforcement learning algorithm, each heterogeneous device, including voltage regulators, switchable capacitors, batteries, is controlled by an independent agent. The goal of each agent is to optimize its control strategy to maximize the overall effect of voltage and reactive power control while cooperating. Use the proximal policy optimization algorithm for centralized training, and make decisions under local observation through a distributed execution framework.
[0037] The reward function is set in combination with the overall goal of voltage and reactive power control, aiming to reduce the comprehensive impact of voltage violations, control errors and power loss. The reward function is as follows:
[0038]
[0039] Where, r(s t ,a t ) represents the reward of taking action a t in state s t , V viol is the voltage violation penalty term, P loss is the power loss penalty term, C error is the control error penalty term, α, β and γ are weight parameters corresponding to each penalty term, used to balance the influence of each part in the total reward, constraint k represents the violation degree of the kth constraint.
[0040] The voltage violation penalty term V viol is used to measure whether the node voltage exceeds the allowed range. Specifically, the degree to which each node voltage v i exceeds the maximum voltage V max or is lower than the minimum voltage V min is calculated to determine:
[0041]
[0042] Where N is the set of all nodes in the network.
[0043] The power loss penalty term P loss is used to evaluate the power loss in the power distribution network, and the calculation method is the ratio of the power loss on all edges (i,j) to the total injected power P ij :
[0044]
[0045] Where G ij represents the conductance of edge (i,j), vi and v j are the voltages of node i and node j, respectively, θ ij is the phase angle between node i and node j, P ij is the total power injected into the distribution network, E is the set of all edges in the network.
[0046] Control error penalty term C error is used to measure the error between the control device action and the desired state, which is determined by calculating the absolute difference between the states of each heterogeneous device controllable node f at time t and t+1, and weighted sum:
[0047]
[0048] where N a represents the set of nodes with action space, s f (t) and s f (t+1) are the states of the heterogeneous device controllable node f at time t and t+1, respectively, η f is the weight parameter of the control error.
[0049] The constraint penalty part is used to ensure that the distribution network meets various physical and operational constraints in actual operation. By introducing Lagrange multiplier λ k , the degree of violation of each constraint constraint k is incorporated into the reward function to increase the punishment for violating the constraints. The mathematical expression of the constraint penalty is as follows:
[0050]
[0051] where λ k represents the weight parameter related to the kth constraint condition, which is used to adjust the punishment degree of different constraint violations.
[0052] Step S3: Voltage-reactive power sensitivity analysis
[0053] After the initialization of the distribution network, the influence of each node and heterogeneous device on the voltage and reactive power of the distribution network is calculated. The specific steps are as follows:
[0054] 1. Calculate the voltage-reactive power sensitivity of the node
[0055] First, the relationship between the active P and reactive Q power of the node and the voltage V and phase angle θ is calculated by the following equation:
[0056]
[0057] where, is the Jacobian matrix calculated by the impedance between nodes, where J Qθdenotes the partial derivative of reactive power Q with respect to phase angle θ, i.e. and so on. VQ is the voltage-reactive sensitivity matrix, which is used to assess the voltage stability of the nodes.
[0058] Next, the voltage-reactive sensitivity of each node during the iteration is calculated:
[0059]
[0060] where S VQ,t (i) denotes the voltage-reactive sensitivity of node i at time t, and N is the total number of nodes in the network.
[0061] 2. Calculate the weight w VQ,t
[0062] According to the proportion of each step reward to the total reward, the weight w VQ,t is calculated:
[0063]
[0064] where w VQ,t denotes the weight at time t, r t is the reward of the t-th step, r H is the total reward, and H is the number of iterations.
[0065] 3. Calculate the priority ranking of the agent
[0066]
[0067] where O(i) denotes the priority ranking of node i, for masking non-agent nodes, Sort denotes the sorting operation, w VQ,t is the weight, and S VQ,t (i) is the voltage-reactive sensitivity of node i at time t.
[0068] Step S4: Multi-agent strategy training and updating
[0069] 1. Initialization of the agent's policy network: Each agent is assigned an initial policy network, which is constructed based on the predefined state space and action space. The initial policy may be based on historical data or a random policy generation, providing a starting point for subsequent training.
[0070] 2. Environment interaction and data collection: Each heterogeneous device executes its strategy and collects state transition, reward, and action data.
[0071] 3. Calculate the update order: Determine the update priority of the agent based on the voltage-reactive sensitivity analysis results.
[0072] 4. Policy network iterative update: Using the reinforcement learning algorithm of proximal policy optimization, the policy network of multi-agent is iteratively updated according to the collected data and the calculated update order. In each iteration, the multi-agent updates its policy network to optimize future control decisions according to the feedback obtained from the environment.
[0073] Step S5: Real-time control and feedback adjustment
[0074] In the actual operation of the distribution network, the trained strategy is executed by the distributed heterogeneous devices. The agent monitors the voltage, reactive power demand and operating state of the heterogeneous devices in the distribution network in real time through the sensor. According to the observed current environment state, the agent makes autonomous decisions based on the trained strategy, and the voltage regulator adjusts the gear to control the voltage level. The switchable capacitor performs switching operation according to the reactive power demand, and the battery performs charging and discharging operation according to the change of the system reactive power demand, so as to realize the dynamic balance of the reactive power. The agent controls the three kinds of devices cooperatively to ensure the optimal adjustment of voltage and reactive power.
[0075] After completing the operation, the agent continues to monitor the feedback information in the environment and evaluates the voltage and reactive state changes of the distribution network in real time. If the state of the distribution network does not achieve the expected effect, the agent will adjust the control strategy according to the feedback and continue to perform further decision operations until the optimization target is reached. With the change of the power system load and the fluctuation of the distribution network state, the agent can continuously adapt to the environment, and continuously optimize the control strategy through the reinforcement learning algorithm to ensure the long-term stability and reliability of the voltage and reactive power control.
[0076] The method of the present application is suitable for voltage stability and reactive power optimization control in distribution network. The method realizes the dynamic adjustment of voltage and the cooperative control of reactive power in distribution network through the multi-agent reinforcement learning algorithm combined with voltage regulator, switchable capacitor and battery and other heterogeneous devices. The multi-agent interacts with the dynamic environment and continuously optimizes its control strategy to maximize the voltage stability of the distribution network, minimize the reactive power loss and reduce the operation cost of the device. The algorithm ensures the optimization of the global system in the framework of centralized training and distributed execution. In actual operation, the multi-agent makes autonomous decisions according to the real-time monitoring of the system state, realizes the coordinated control of the voltage regulator, switchable capacitor and battery, and ensures the stability of the voltage and reactive power of the distribution network in the complex operating environment. The method of the present application has good adaptability and scalability, and can effectively cope with the challenges of load change, device state fluctuation and other challenges in the power grid, and improve the operation efficiency and safety of the distribution network.
Claims
1. A voltage and reactive power control method based on constrained heterogeneous ordered multi-agent reinforcement learning, characterized in that, The method controls heterogeneous devices in a power distribution network through a multi-agent reinforcement learning algorithm, realizes dynamic optimization control of voltage and reactive power, and the heterogeneous devices include a voltage regulator, a switchable capacitor and a battery, and the specific steps are as follows: Heterogeneous multi-agent state space and action space modeling: one heterogeneous device is modeled as one agent, a plurality of different heterogeneous devices are modeled as a heterogeneous multi-agent, a voltage and reactive power control problem of the power distribution network is modeled as a heterogeneous multi-agent reinforcement learning problem, a state space of each node and an action space of the heterogeneous devices are defined, the state space of each node includes a node voltage, a running state of the heterogeneous devices and a load demand, and the action space of the heterogeneous devices specifically includes a gear state of the voltage regulator, a switch state of the switchable capacitor and a charging and discharging state of the battery; Designing a reward function according to a control target: the overall target of voltage and reactive power control is to maintain voltage stability of the power distribution network and reduce reactive power loss by adjusting the voltage regulator, the switchable capacitor and the battery, while meeting the running constraint conditions of the heterogeneous devices; the reward function is designed in combination with the overall target of voltage and reactive power control, the punishment of voltage violation, reactive power loss and control error is comprehensively considered, and a Lagrange multiplier is introduced to process the constraint conditions, so that the agent can effectively optimize the voltage and reactive power control strategy in the training process, and ensure that the power distribution network meets the physical and operating limits; Voltage-reactive power sensitivity analysis: according to a voltage-reactive power sensitivity matrix, the influence of each node on the voltage and reactive power of the power distribution network is calculated, and the control priority of each node is dynamically adjusted; Determining a multi-agent update sequence: based on the voltage and reactive power sensitivity, the nodes with a large influence on the power distribution network are selected to update the control strategy first, so as to improve the response speed and control accuracy of the power distribution network; Device cooperative control: in the running process of the power distribution network, the voltage regulator controls the node voltage by adjusting the gear; the switchable capacitor compensates reactive power according to the demand; and the battery adjusts the reactive power by charging and discharging, so as to realize reactive power balance and voltage stability of the power distribution network; Strategy updating and optimization: the update sequence of the multi-agent is continuously calculated, the strategy network of the multi-agent is updated according to the update sequence through iterative training of the multi-agent reinforcement learning algorithm, the optimal control of each heterogeneous device under different running states is realized, and the voltage and reactive power control effect of the power distribution network is improved; Real-time control and feedback adjustment: in actual operation, each agent makes autonomous decisions according to the real-time observed state of the power distribution network, realizes dynamic control of voltage and reactive power, and adjusts the future control strategy according to the actual control effect.
2. The voltage and reactive power control method based on constrained heterogeneous ordered multi-agent reinforcement learning according to claim 1, characterized in that, The multi-agent reinforcement learning algorithm adopts a proximal policy optimization algorithm and has a monotonic improvement feature, so that the voltage and reactive power control effect of the power distribution network is continuously optimized after each strategy update, and the instability of the power distribution network caused by immature strategies is avoided; the cooperative control strategy of the heterogeneous devices is optimized through centralized training and distributed execution.
3. The voltage and reactive power control method based on constrained heterogeneous ordered multi-agent reinforcement learning according to claim 1, characterized in that, The voltage-reactive power sensitivity matrix is used to evaluate the sensitivity of each node to voltage and reactive power, the voltage-reactive power sensitivity matrix is dynamically updated according to the voltage fluctuation and load demand of the node, and the flexibility and accuracy of the control strategy are ensured.
4. The voltage and reactive power control method based on constrained heterogeneous ordered multi-agent reinforcement learning according to claim 1, characterized in that, The voltage regulator controls the node voltage by adjusting the gear ratio, so that the distribution network voltage is maintained within a predetermined range, and the voltage fluctuation is reduced.
5. The voltage and reactive power control method based on constrained heterogeneous ordered multi-agent reinforcement learning according to claim 1, characterized in that, The switchable capacitor is switched according to the reactive demand of the distribution network, and provides reactive compensation to maintain the reactive power balance of the distribution network.
6. The voltage and reactive power control method based on constrained heterogeneous ordered multi-agent reinforcement learning according to claim 1, characterized in that, The battery adjusts the reactive power output by controlling the charging and discharging state, to assist the voltage regulator and the switchable capacitor to achieve the reactive balance of the distribution network.
Citation Information
Patent Citations
Reactive voltage control method based on multi-time-scale multi-agent deep reinforcement learning
CN113363997A
Active power distribution network cooperative voltage regulation method and system based on multi-agent deep reinforcement learning
CN114362187A