Industrial park virtual power plant reactive voltage control method of multi-target multi-agent deep reinforcement learning considering energy storage life

By employing a multi-objective, multi-agent deep reinforcement learning method, the problem of traditional methods struggling to balance network losses and energy storage lifespan in industrial parks was solved. This enabled efficient voltage control of virtual power plants in industrial parks and provided synergistic optimization of energy storage device reliability and park energy efficiency.

CN121012045APending Publication Date: 2025-11-25HARBIN ELECTRIC SCI & TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511195669.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Traditional reactive power and voltage control methods for virtual power plants in industrial parks struggle to balance network loss optimization and energy storage lifetime protection. Existing model-driven methods simplify energy storage lifetime to capacity constraints, neglecting the nonlinear coupling relationship between temperature and lifetime. Single-objective data-driven methods cannot capture multi-objective conflicts.

Method used

A multi-objective, multi-agent deep reinforcement learning approach is adopted. By dividing the power grid area, an agent structure containing multiple Actor networks is designed. A parallel training scheme is used to construct an Actor-Critic network structure to achieve reactive power voltage control of energy storage. Multi-objective optimization is performed by combining local observation and normalized reward function to generate Pareto front.

Benefits of technology

It achieves coordinated control in a virtual power plant in an industrial park, taking into account real-time performance, park energy efficiency, and the reliability of energy storage equipment. It optimizes park network losses, node voltage deviations, and energy storage life through a distributed control scheme, and provides quantitative scheduling basis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121012045A_ABST
    Figure CN121012045A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial park virtual power plant reactive voltage control method of multi-target multi-agent deep reinforcement learning considering energy storage life, and belongs to the technical field of virtual power plant reactive voltage control methods based on industrial data acquisition. The problem that in the prior art, a traditional industrial park virtual power plant reactive voltage control method is difficult to consider park network loss optimization and energy storage life protection at the same time is solved. According to the method, a multi-objective optimization model including network loss minimization, node voltage deviation minimization and energy storage battery service life maximization is constructed, a virtual power plant is divided through a distributed multi-agent framework, each agent trains multiple strategies in parallel through a multi-objective multi-agent deep reinforcement learning algorithm based on local observation, the Pareto front is approached, and the energy storage battery service life is maximized. High-efficiency control is realized, and each intelligent agent independently generates an energy storage reactive power regulation and active power charging and discharging instruction. The reliability of the energy storage equipment is improved, and the method can be applied to an industrial park active power distribution network scene with high renewable energy source permeation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a reactive power and voltage control method for a virtual power plant in an industrial park, and more particularly to a multi-objective, multi-agent deep reinforcement learning method for reactive power and voltage control of a virtual power plant in an industrial park that considers energy storage lifetime. It belongs to the technical field of reactive power and voltage control methods for virtual power plants based on industrial data acquisition. Background Technology

[0002] With the increasing global demand for clean energy and the growing awareness of sustainable development, the penetration rate of distributed energy storage systems in industrial park virtual power plants (VPPs) is showing a rapid growth trend. Such large-scale energy storage access has powerfully promoted the transformation of traditional industrial park distribution networks into active virtual power plants with flexible adjustment capabilities. However, the transformation process is also accompanied by many severe challenges.

[0003] Energy storage systems offer bidirectional flexibility in both active charge / discharge and reactive power regulation, but their operation is accompanied by significant lifespan degradation. The cycle life of energy storage batteries is highly dependent on the depth of charge / discharge (DOD), operating temperature, and operating frequency. For example, frequent high-power charge / discharge can lead to increased battery cell temperature and accelerated electrolyte decomposition, thereby shortening the lifespan. At the same time, while the rapid power regulation of energy storage can smooth out intermittent fluctuations in distributed energy sources (such as photovoltaic and wind power) within industrial parks (such as sudden power changes caused by weather variations), it may also cause new problems. If the pursuit of virtual power plant voltage control precision or park network loss optimization is excessive, it may force energy storage to operate under high-loss conditions for a long time, accelerating equipment aging. Conversely, if energy storage operation is simply restricted to protect lifespan, it may lead to excessive voltage fluctuations in the park's power grid or a deterioration in power flow distribution.

[0004] Traditional voltage / reactive power control (VVC) methods mainly rely on on-load tap changers (OLTC) and capacitor banks (CB). However, these mechanical devices have slow response times (minute-level adjustments) and coarse adjustment granularity, making them unsuitable for the rapid dynamic characteristics of energy storage. For example, OLTC tap changer switching takes tens of seconds, making it difficult to track distributed energy power fluctuations on a second-level scale; discrete switching of CB can also cause voltage oscillations. Although energy storage systems can provide flexible reactive power support through converters, their control strategies still face the challenge of achieving a quantitative balance between virtual power plant voltage control, campus network loss optimization, and energy storage lifetime protection. Existing model-driven methods (such as linear programming-based optimization) often simplify energy storage lifetime to capacity constraints, ignoring the nonlinear coupling relationship between temperature and lifetime; while single-objective data-driven methods (such as the DRL algorithm that only optimizes voltage deviation) cannot capture multi-objective conflicts, leading to one-sided optimization results.

[0005] In summary, a reactive power and voltage control method for a virtual power plant in an industrial park that considers energy storage lifetime and employs multi-objective, multi-agent deep reinforcement learning is needed. Summary of the Invention

[0006] A brief overview of the invention is given below to provide a basic understanding of certain aspects of it. It should be understood that this overview is not an exhaustive summary of the invention. It is not intended to identify key or essential parts of the invention, nor is it intended to limit the scope of the invention. Its purpose is merely to present certain concepts in a simplified form as a prelude to the more detailed description that follows.

[0007] In view of this, in order to solve the problem that the traditional reactive power and voltage control methods of virtual power plants in industrial parks in the prior art are difficult to take into account both the optimization of network losses in the park and the protection of energy storage lifetime, the present invention provides a reactive power and voltage control method for virtual power plants in industrial parks that takes into account energy storage lifetime through multi-objective and multi-agent deep reinforcement learning.

[0008] The technical solution is as follows: A reactive power and voltage control method for a virtual power plant in an industrial park, considering energy storage lifetime and employing multi-objective, multi-agent deep reinforcement learning, includes the following steps:

[0009] S1. To address the conflict between voltage fluctuations, network losses, and energy storage lifespan losses caused by the high proportion of distributed energy access in virtual power plants in industrial parks, three optimization objectives are defined. At the same time, constraints such as AC power flow balance, energy storage power limit, voltage safety range, and branch current are set to construct a nonlinear multi-objective optimization problem suitable for virtual power plants.

[0010] S2. Based on the nonlinear multi-objective optimization problem applicable to virtual power plants, the power grid area is divided and modeled by a distributed partially observable Markov decision process to obtain the design parameters of the multi-objective multi-agent deep reinforcement learning algorithm framework.

[0011] S3. Based on the design parameters of the multi-objective multi-agent deep reinforcement learning algorithm framework, design an agent structure containing multiple Actor networks, adopt a parallel training scheme to simultaneously optimize multiple strategies, construct an Actor-Critic network structure, and use soft updates to complete the design of a reactive voltage control algorithm for multi-objective multi-agent deep reinforcement learning.

[0012] S4. The reactive voltage control algorithm using multi-objective multi-agent deep reinforcement learning is trained offline. Based on the training results of the multi-objective strategy, the Pareto front is generated and the sensitivity is analyzed to obtain the trained reactive voltage control model.

[0013] S5. The Actor network in the trained reactive voltage control model is retained for online control. The agent in the trained reactive voltage control model calls the pre-trained strategy based on real-time local observation data to generate energy storage reactive power regulation and active power charging and discharging commands, thereby realizing real-time response control.

[0014] Furthermore, in S1, the three optimization objectives include minimizing energy loss in the campus network, minimizing node voltage deviation, and maximizing the lifespan of energy storage devices.

[0015] By calculating the total energy loss of the entire virtual power plant during its operating cycle, the energy loss of the campus network can be minimized.

[0016] Total energy loss Represented as:

[0017]

[0018]

[0019] in, For a moment Branch power loss is the active power of energy storage. For a moment Branch losses, For a moment Total active power generation of the system For a moment Active load, The active power used for interaction with the main network. branch road The resistance, For a moment branch road active power, For a moment branch road reactive power, For nodes The active power of the energy storage system For nodes The reactive power of the energy storage system For nodes Voltage amplitude control at the location, To control the total number of optimization steps;

[0020] With maximum voltage deviation Evaluate voltage control performance to minimize node voltage deviation;

[0021] Maximum voltage deviation Represented as:

[0022]

[0023] in, This indicates the entire optimization time period. Take the maximum value and find the deviation value at the moment with the most severe voltage deviation among all moments. Indicates at a single moment For all nodes in the park Take the maximum voltage deviation to determine the node deviation with the worst voltage quality at that moment. The reference voltage;

[0024] Based on the battery cycle life model, energy storage lifetime loss is defined. To maximize energy storage lifespan;

[0025] Energy storage lifespan loss Represented as:

[0026]

[0027] in, The number of energy storage units. For energy storage units At any moment The discharge depth was obtained by modeling based on the Arrhenius equation. For at any time Corresponding energy storage unit The remaining cycle life, in hours. , As the baseline cycle life, For temperature sensitivity coefficient, As the reference depth of discharge;

[0028] The constraints are as follows: AC power flow balance is constrained by establishing active / reactive power balance equations; power output from energy storage is limited by... Constraints and node voltage safety ranges are implemented. Apply constraints, including branch current constraints;

[0029] Branch current constraints are expressed as follows:

[0030]

[0031] in, For a moment branch road The branch current, For a moment branch road active power, For a moment branch road reactive power, For nodes The active power of the energy storage system For nodes The reactive power of the energy storage system For nodes Voltage amplitude control at the location, branch road Maximum permissible current.

[0032] Furthermore, step S2 includes the following steps:

[0033] S21. Divide the virtual power plant area and obtain local information;

[0034] In step S21, the virtual power plant is divided into Each region is composed of an intelligent agent. Management allows for the acquisition of local node voltage, photovoltaic output, and load information without the need for inter-regional communication;

[0035] S22. Combining local information, a distributed partially observable Markov decision process is used to model the multi-objective, multi-agent deep reinforcement learning algorithm framework, and the design parameters of the multi-objective, multi-agent deep reinforcement learning algorithm framework are obtained.

[0036] In step S22, the modeling process includes setting up the state space, observation space, action space, and reward function, as detailed below:

[0037] State space: Sets the global runtime state It contains information on the voltage and power of all nodes in the system.

[0038] Observation space: Intelligent agent Local observation ,in, Node voltage amplitude within the region Photovoltaic power generation meritorious / Reactive load, Energy storage state of charge and Depth of discharge of stored energy;

[0039] Action Space: Intelligent Agent The output is the proportional factor for energy storage scheduling. ,in, This is the reactive power regulation proportional factor. This is the active charge / discharge ratio factor;

[0040] Reward function: The global reward is a linear weighted sum of multiple objectives, which, after normalization, yields the reward function. ;

[0041] reward function Represented as:

[0042]

[0043] in, For a moment Normalized network function, For the first The first weighting factor of each Actor For the first The second weighting factor for each Actor For a moment The normalized objective function for node voltage deviation, For a moment The normalized objective function for energy storage lifetime loss, and These are the target's Utopia point and Nadir point, respectively. The target point.

[0044] Furthermore, step S3 includes the following steps:

[0045] S31. Based on the design parameters of the multi-objective multi-agent deep reinforcement learning algorithm framework, design a parallel training architecture;

[0046] In S31, each intelligent agent Include There are 1 Actor networks, each corresponding to a set of weight factors. ;

[0047] A parallel training scheme based on Actor networks is adopted, which trains multiple Actor networks simultaneously and updates parameters through stochastic gradient ascent.

[0048] S32. Based on the parallel training architecture, construct an Actor-Critic network structure that includes an Actor network and a Critic network;

[0049] In S32, the Actor network receives local observations as input. The output is a scaling factor. Maximize expected reward through policy gradient ,in, For intelligent agents strategy Below, based on current local observations and the actions taken The average expected value of the Q-value evaluated by the Critic network. The action output by the Actor network. These are the parameters of the Actor network;

[0050] Critic network: Input is The output is Value, assessing the value of an action; the loss function is the mean squared error. ,in, target value ;

[0051] S33. The parameters of the Actor-Critic network structure are updated using soft updates to complete the design of a reactive voltage control algorithm based on multi-objective multi-agent deep reinforcement learning.

[0052] In S33, the update process is represented as follows: ,in, For the parameters of the target network, For coefficients;

[0053] Furthermore, step S4 includes the following steps:

[0054] S41. Data preparation and environmental simulation for reactive voltage control algorithm based on multi-objective multi-agent deep reinforcement learning;

[0055] In S41, historical photovoltaic output and load data of the park are used to simulate uncertainty through normal distribution random disturbance, run virtual power plant power flow in simulation environment, calculate each target value and constraint condition, and generate training samples.

[0056] S42. A multi-objective, multi-agent deep reinforcement learning algorithm for reactive voltage control is used for multi-objective policy training;

[0057] In step S42, initialization is performed. Each agent, Each Actor network corresponds to a different weight factor, which serves as the input training sample for the algorithm, i.e., the multi-objective multi-agent deep reinforcement learning model. The model is trained in parallel for hundreds of cycles, with photovoltaic / load scenarios randomly assigned in each cycle. The network parameters are optimized through empirical replay and batch gradient descent.

[0058] S43. Based on the training results of the multi-objective strategy, generate the Pareto front and analyze the sensitivity to obtain the trained reactive voltage control model;

[0059] In step S43, after training is completed, all non-dominated solutions are extracted, and a three-dimensional Pareto front containing network loss, voltage deviation, and energy storage lifetime is constructed. By adjusting the weighting factors, the impact on energy storage lifetime is analyzed.

[0060] The beneficial effects of this invention are as follows: The multi-objective multi-agent deep reinforcement learning (MOMADRL) method proposed in this invention constructs a multi-objective optimization model for the virtual power plant scenario in industrial parks, which includes minimizing network loss in the park, minimizing node voltage deviation, and maximizing energy storage lifetime. By dividing the virtual power plant into multiple regions, each region is configured with an agent to achieve decentralized control: the agent optimizes multiple strategies simultaneously based on local observations through an Actor-Critic architecture and an Actor-based parallel training scheme. That is, each Actor corresponds to a different combination of weight factors (such as focusing on energy storage lifetime protection or park voltage control). The multi-objective is transformed into a unified optimization objective through a normalized reward function, approximating the Pareto front. This invention provides a collaborative control scheme for virtual power plants in industrial parks with high energy storage penetration, which takes into account real-time performance, park energy efficiency, and the reliability of energy storage equipment. Attached Figure Description

[0061] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:

[0062] Figure 1 A flowchart illustrating a multi-objective, multi-agent deep reinforcement learning method for reactive power and voltage control in a virtual power plant in an industrial park, considering energy storage lifetime.

[0063] Figure 2 This is a schematic flowchart of an embodiment of a reactive power and voltage control method for a virtual power plant in an industrial park that considers energy storage lifetime and employs multi-objective, multi-agent deep reinforcement learning. Detailed Implementation

[0064] To make the technical solutions and advantages of the embodiments of the present invention clearer, the exemplary embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not an exhaustive list of all embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0065] refer to Figure 1 and Figure 2 This embodiment details a reactive power and voltage control method for a virtual power plant in an industrial park that considers energy storage lifetime using multi-objective, multi-agent deep reinforcement learning. The method includes the following steps:

[0066] S1. To address the conflict between voltage fluctuations, network losses, and energy storage lifespan losses caused by the high proportion of distributed energy sources (photovoltaics, wind power, etc.) connected to virtual power plants in industrial parks, three major optimization objectives are defined. At the same time, constraints such as AC power flow balance, energy storage power limits (including apparent power constraints and upper and lower limits of state of charge (SOC)), voltage safety range, and branch current are set to construct a nonlinear multi-objective optimization problem suitable for virtual power plants.

[0067] S2. Based on the nonlinear multi-objective optimization problem applicable to virtual power plants, the power grid area is divided and modeled through a distributed partially observable Markov decision process to obtain the design parameters of the multi-objective multi-agent deep reinforcement learning algorithm framework. That is, the global optimization problem of the virtual power plant is decomposed into the local decision problem of each region's agents to achieve distributed cooperative control.

[0068] S3. Based on the design parameters of the multi-objective multi-agent deep reinforcement learning algorithm framework, design an agent structure containing multiple Actor networks, adopt the parallel training scheme (ABPTS) to simultaneously optimize multiple strategies, construct the Actor-Critic network structure, and use soft update to complete the design of the reactive voltage control algorithm for multi-objective multi-agent deep reinforcement learning, transforming the multi-objective problem into single-objective optimization and approaching the Pareto front.

[0069] S4. The reactive voltage control algorithm using multi-objective multi-agent deep reinforcement learning is trained offline. Based on the training results of the multi-objective strategy, the Pareto front is generated and the sensitivity is analyzed to obtain the trained reactive voltage control model.

[0070] S5. The Actor network in the trained reactive voltage control model is retained for online control. The agent in the trained reactive voltage control model calls the pre-trained strategy based on real-time local observation data (such as regional voltage, energy storage SOC, and current load) to generate energy storage reactive power regulation and active power charging and discharging commands, thereby realizing real-time response control.

[0071] Specifically, in step S3, the Actor network generates the reactive power regulation ratio factor and the active power charge-discharge ratio factor for energy storage based on local observations, ensuring that the output does not exceed the energy storage capacity limit (ensuring that the capacity limit is not exceeded). The Critic network is used to evaluate the value of actions and optimize network parameters through a soft update mechanism.

[0072] In step S4, the operation scenario of a virtual power plant is simulated using historical distributed energy output, load, and energy storage charging and discharging data (3-minute resolution) from the industrial park. The reactive voltage control algorithm of the multi-objective multi-agent deep reinforcement learning model is trained offline through hundreds of training cycles. Random disturbances are introduced during training to simulate distributed energy fluctuations, load changes, and energy storage state uncertainties. The network parameters are updated through experience replay and stochastic gradient ascent. After training, all non-dominated solutions are extracted, and a three-dimensional Pareto front is constructed, which includes park network loss, voltage deviation, and energy storage lifetime (cycle lifetime model based on depth of discharge DOD). The impact of adjusting the weight factors on the energy storage battery cell temperature and remaining cycle lifetime is analyzed, providing a multi-objective trade-off basis for scheduling between network operation performance and energy storage equipment reliability.

[0073] In step S5, regarding the issue of computational efficiency: the time for a single decision is only a few minutes, which meets the second-level / minute-level control requirements of the distribution network.

[0074] Furthermore, in S1, the three major optimization objectives include minimizing the energy loss of the campus network to reduce the overall operating cost, minimizing the node voltage deviation to ensure the power supply quality of the campus, and maximizing the lifespan of energy storage devices (based on the cycle life model of depth of discharge (DOD)).

[0075] By calculating the total energy loss of the entire virtual power plant during its operating cycle, the energy loss of the campus network can be minimized.

[0076] Total energy loss Represented as:

[0077]

[0078]

[0079] in, For a moment Branch power loss is the active power of energy storage (negative for charging, positive for discharging). For a moment Branch losses, For a moment Total active power generation of the system For a moment Active load, The active power used for interaction with the main network. branch road The resistance, For a moment branch road active power, For a moment branch road reactive power, For nodes The active power of the energy storage system For nodes The reactive power of the energy storage system For nodes Voltage amplitude control at the location, To control the total number of optimization steps;

[0080] With maximum voltage deviation Evaluate voltage control performance to minimize node voltage deviation;

[0081] Maximum voltage deviation Represented as:

[0082]

[0083] in, This indicates the entire optimization time period. Take the maximum value and find the deviation value at the moment with the most severe voltage deviation among all moments. Indicates at a single moment For all nodes in the park Take the maximum voltage deviation to determine the node deviation with the worst voltage quality at that moment. The reference voltage is 1.0 pu, where pu represents the ratio relative to the reference value.

[0084] Based on the battery cycle life model, energy storage lifetime loss is defined. To maximize energy storage lifespan;

[0085] Energy storage lifespan loss Represented as:

[0086]

[0087] in, The number of energy storage units. For energy storage units At any moment The discharge depth was obtained by modeling based on the Arrhenius equation. For at any time Corresponding energy storage unit The remaining cycle life (in hours) is as follows. , As the baseline cycle life, This is the temperature sensitivity coefficient (empirical parameter). As the reference depth of discharge;

[0088] The constraints are as follows: AC power flow balance is constrained by establishing active / reactive power balance equations; power output from energy storage is limited by... Constraints and node voltage safety ranges are implemented. Apply constraints, including branch current constraints;

[0089] Branch current constraints are expressed as follows:

[0090]

[0091] in, For a moment branch road The branch current, For a moment branch road active power, For a moment branch road reactive power, For nodes The active power of the energy storage system For nodes The reactive power of the energy storage system For nodes Voltage amplitude control at the location, branch road Maximum permissible current.

[0092] Furthermore, step S2 includes the following steps:

[0093] S21. Divide the virtual power plant area and obtain local information;

[0094] In step S21, the virtual power plant is divided into Each region (e.g., a 141-node system is divided into 9 regions) is controlled by an agent. (Regional controller) management, obtaining local information such as local node voltage, photovoltaic output, and load, without the need for inter-regional communication;

[0095] S22. Combining local information, a decentralized partially observable Markov decision process (Dec-POMDP) ​​is used for modeling to obtain the design parameters of the multi-objective multi-agent deep reinforcement learning algorithm framework;

[0096] In step S22, the modeling process includes setting up the state space, observation space, action space, and reward function, as detailed below:

[0097] State space: Sets the global runtime state It contains information such as voltage and power of all nodes in the system.

[0098] Observation space: Intelligent agent Local observation ,in, Node voltage amplitude within the region Photovoltaic power generation Meritorious / Reactive load, Energy storage state of charge and Depth of discharge of stored energy;

[0099] Action Space: Intelligent Agent The output is the proportional factor for energy storage scheduling. ,in, This is the reactive power regulation proportional factor, which is used to control the reactive power output of energy storage. It is the active charge / discharge ratio factor, which is used to control the active power of energy storage;

[0100] Reward function: The global reward is a linear weighted sum of multiple objectives, which, after normalization, yields the reward function. ;

[0101] reward function Represented as:

[0102]

[0103] in, For a moment Normalized network function, For the first The first weighting factor of each Actor For the first The second weighting factor for each actor, For a moment The normalized objective function for node voltage deviation, For a moment The normalized objective function for energy storage lifetime loss, and These are the target's Utopia point (minimum) and Nadir point (maximum), respectively. The target point.

[0104] Furthermore, step S3 includes the following steps:

[0105] S31. Based on the design parameters of the multi-objective multi-agent deep reinforcement learning algorithm framework, design a parallel training architecture;

[0106] In S31, each intelligent agent Include There are 1 Actor networks, each corresponding to a set of weight factors. , used to approximate different solutions to the Pareto front;

[0107] A parallel training scheme based on Actor networks (ABPTS) is adopted to train multiple Actor networks simultaneously and update parameters through stochastic gradient ascent, which significantly reduces training time and obtains multiple strategies.

[0108] S32. Based on the parallel training architecture, construct an Actor-Critic network structure that includes an Actor network and a Critic network;

[0109] In S32, the Actor network receives local observations as input. The output is a scaling factor. Maximize expected reward through policy gradient ,in, For intelligent agents strategy Below, based on current local observations and the actions taken The average expected value of the Q-score (action value) evaluated by the Critic network. The action output by the Actor network. These are the parameters of the Actor network;

[0110] Critic network: Input is The output is Value, assessing the value of an action; the loss function is the mean squared error. ,in, target value ;

[0111] S33. The parameters of the Actor-Critic network structure are updated using soft updating (Polyakaveraging) to complete the design of a reactive voltage control algorithm based on multi-objective multi-agent deep reinforcement learning.

[0112] In S33, the update process is represented as follows: ,in, For the parameters of the target network, For coefficients;

[0113] Furthermore, step S4 includes the following steps:

[0114] S41. Data preparation and environmental simulation for reactive voltage control algorithm based on multi-objective multi-agent deep reinforcement learning;

[0115] In S41, historical photovoltaic power output and load data of the park (such as 3-year, 3-minute resolution data) are used to simulate uncertainty through normal distribution random disturbance, run virtual power plant power flow in simulation environment, calculate each target value and constraint condition, and generate training samples.

[0116] S42. A multi-objective, multi-agent deep reinforcement learning algorithm for reactive voltage control is used for multi-objective policy training;

[0117] In step S42, initialization is performed. Each agent, Each Actor network corresponds to different weight factors (e.g., ... With a step size of 0.01, the algorithm, i.e., a multi-objective multi-agent deep reinforcement learning model, inputs training samples and trains in parallel for hundreds of epochs. In each epoch, photovoltaic / load scenarios are randomly assigned, and the network parameters are optimized through experience replay and batch gradient descent.

[0118] S43. Based on the training results of the multi-objective strategy, generate the Pareto front and analyze the sensitivity to obtain the trained reactive voltage control model;

[0119] In step S43, after training is complete, all non-dominated solutions are extracted, and a three-dimensional Pareto front including network loss, voltage deviation, and energy storage lifetime is constructed. This is achieved by adjusting weighting factors (such as...). From 0 to 1), analyze its impact on energy storage lifespan, low When focusing on network loss or voltage deviation optimization, energy storage needs to frequently adjust active / reactive power to quickly respond to grid fluctuations, leading to an increase in depth of discharge (DOD) or the number of charge-discharge cycles. According to energy storage lifetime models (such as cycle lifetime formulas based on the Arrhenius equation), the core temperature of core components (such as batteries) rises, accelerating aging and shortening lifespan; high When focusing on energy storage lifetime protection, the algorithm limits the adjustment range and frequency of energy storage, reduces DOD and core temperature, but may lead to increased voltage deviation due to insufficient reactive power support, or a slight increase in network losses due to reduced active power regulation. This process provides a quantitative basis for scheduling strategy through multi-objective trade-offs at the Pareto front, balancing grid operating efficiency and energy storage device reliability.

[0120] Although the invention has been described with reference to a limited number of embodiments, those skilled in the art will understand from the foregoing description that other embodiments are conceivable within the scope of the invention described herein. Furthermore, it should be noted that the language used in this specification has been chosen primarily for readability and instructional purposes, and not for the purpose of interpreting or limiting the subject matter of the invention. Therefore, many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the appended claims. The disclosure of the invention is illustrative and not restrictive, and the scope of the invention is defined by the appended claims.

Claims

1. A reactive power and voltage control method for a virtual power plant in an industrial park, considering energy storage lifetime and employing multi-objective, multi-agent deep reinforcement learning, characterized in that... Includes the following steps: S1. To address the conflict between voltage fluctuations, network losses, and energy storage lifespan losses caused by the high proportion of distributed energy access in virtual power plants in industrial parks, three optimization objectives are defined. At the same time, constraints such as AC power flow balance, energy storage power limit, voltage safety range, and branch current are set to construct a nonlinear multi-objective optimization problem suitable for virtual power plants. S2. Based on the nonlinear multi-objective optimization problem applicable to virtual power plants, the power grid area is divided and modeled by a distributed partially observable Markov decision process to obtain the design parameters of the multi-objective multi-agent deep reinforcement learning algorithm framework. S3. Based on the design parameters of the multi-objective multi-agent deep reinforcement learning algorithm framework, design an agent structure containing multiple Actor networks, adopt a parallel training scheme to simultaneously optimize multiple strategies, construct an Actor-Critic network structure, and use soft updates to complete the design of a reactive voltage control algorithm for multi-objective multi-agent deep reinforcement learning. S4. The reactive voltage control algorithm using multi-objective multi-agent deep reinforcement learning is trained offline. Based on the training results of the multi-objective strategy, the Pareto front is generated and the sensitivity is analyzed to obtain the trained reactive voltage control model. S5. The Actor network in the trained reactive voltage control model is retained for online control. The agent in the trained reactive voltage control model calls the pre-trained strategy based on real-time local observation data to generate energy storage reactive power regulation and active power charging and discharging commands, thereby realizing real-time response control.

2. The reactive power and voltage control method for a virtual power plant in an industrial park, considering energy storage lifetime and employing multi-objective, multi-agent deep reinforcement learning, as described in claim 1, is characterized in that... In S1, the three optimization objectives include minimizing energy loss in the campus network, minimizing node voltage deviation, and maximizing the lifespan of energy storage devices. By calculating the total energy loss of the entire virtual power plant during its operating cycle, the energy loss of the campus network can be minimized. Total energy loss Represented as: in, For a moment Branch power loss is the active power of energy storage. For a moment Branch losses, For a moment Total active power generation of the system For a moment Active load, The active power used for interaction with the main network. branch road The resistance, For a moment branch road active power, For a moment branch road reactive power, For nodes The active power of the energy storage system For nodes The reactive power of the energy storage system For nodes Voltage amplitude control at the location, To control the total number of optimization steps; With maximum voltage deviation Evaluate voltage control performance to minimize node voltage deviation; Maximum voltage deviation Represented as: in, This indicates the entire optimization time period. Take the maximum value and find the deviation value at the moment with the most severe voltage deviation among all moments. Indicates at a single moment For all nodes in the park Take the maximum voltage deviation to determine the node deviation with the worst voltage quality at that moment. The reference voltage; Based on the battery cycle life model, energy storage lifetime loss is defined. To maximize energy storage lifespan; Energy storage lifespan loss Represented as: in, The number of energy storage units. For energy storage units At any moment The discharge depth was obtained by modeling based on the Arrhenius equation. For at any time Corresponding energy storage unit The remaining cycle life, in hours. , As the baseline cycle life, For temperature sensitivity coefficient, As the reference depth of discharge; The constraints are as follows: AC power flow balance is constrained by establishing active / reactive power balance equations; power output from energy storage is limited by... Constraints and node voltage safety ranges are implemented. Apply constraints, including branch current constraints; Branch current constraints are expressed as follows: in, For a moment branch road The branch current, For a moment branch road active power, For a moment branch road reactive power, For nodes The active power of the energy storage system For nodes The reactive power of the energy storage system For nodes Voltage amplitude control at the location, branch road Maximum permissible current.

3. The reactive power and voltage control method for a virtual power plant in an industrial park, considering energy storage lifetime and employing multi-objective, multi-agent deep reinforcement learning, as described in claim 2, is characterized in that... S2 includes the following steps: S21. Divide the virtual power plant area and obtain local information; In step S21, the virtual power plant is divided into Each region is composed of an intelligent agent. Management allows for the acquisition of local node voltage, photovoltaic output, and load information without the need for inter-regional communication; S22. Combining local information, a distributed partially observable Markov decision process is used to model the multi-objective, multi-agent deep reinforcement learning algorithm framework, and the design parameters of the multi-objective, multi-agent deep reinforcement learning algorithm framework are obtained. In step S22, the modeling process includes setting up the state space, observation space, action space, and reward function, as detailed below: State space: Sets the global runtime state It contains information on the voltage and power of all nodes in the system; Observation space: Intelligent agent Local observation ,in, Node voltage amplitude within the region Photovoltaic power generation meritorious / Reactive load, Energy storage state of charge and Depth of discharge of stored energy; Action Space: Intelligent Agent The output is the proportional factor for energy storage scheduling. ,in, This is the reactive power regulation proportional factor. This is the active charge / discharge ratio factor; Reward function: The global reward is a linear weighted sum of multiple objectives, which, after normalization, yields the reward function. ; reward function Represented as: in, For a moment Normalized network function, For the first The first weighting factor of each Actor For the first The second weighting factor for each Actor For a moment The normalized objective function for node voltage deviation, For a moment The normalized objective function for energy storage lifetime loss, and These are the target's Utopia point and Nadir point, respectively. The target point.

4. The reactive power and voltage control method for a virtual power plant in an industrial park, considering energy storage lifetime and employing multi-objective, multi-agent deep reinforcement learning, as described in claim 3, is characterized in that... S3 includes the following steps: S31. Based on the design parameters of the multi-objective multi-agent deep reinforcement learning algorithm framework, design a parallel training architecture; In S31, each intelligent agent Include There are 1 Actor networks, each corresponding to a set of weight factors. ; A parallel training scheme based on Actor networks is adopted, which trains multiple Actor networks simultaneously and updates parameters through stochastic gradient ascent. S32. Based on the parallel training architecture, construct an Actor-Critic network structure that includes an Actor network and a Critic network; In S32, the Actor network receives local observations as input. The output is a scaling factor. Maximize expected reward through policy gradient ,in, For intelligent agents strategy Below, based on current local observations and the actions taken The average expected value of the Q-value evaluated by the Critic network. The action output by the Actor network. These are the parameters of the Actor network; Critic network: Input is The output is Value, assessing the value of an action; the loss function is the mean squared error. ,in, Target value ; S33. The parameters of the Actor-Critic network structure are updated using soft updates to complete the design of a reactive voltage control algorithm based on multi-objective multi-agent deep reinforcement learning. In S33, the update process is represented as follows: ,in, For the parameters of the target network, is a coefficient.

5. The reactive power and voltage control method for a virtual power plant in an industrial park, considering energy storage lifetime and employing multi-objective, multi-agent deep reinforcement learning, as described in claim 4, is characterized in that... S4 includes the following steps: S41. Data preparation and environmental simulation for reactive voltage control algorithm based on multi-objective multi-agent deep reinforcement learning; In S41, historical photovoltaic output and load data of the park are used to simulate uncertainty through normal distribution random disturbance, run virtual power plant power flow in simulation environment, calculate each target value and constraint condition, and generate training samples. S42. A multi-objective, multi-agent deep reinforcement learning algorithm for reactive voltage control is used for multi-objective policy training; In step S42, initialization is performed. Each agent, Each Actor network corresponds to a different weight factor, which serves as the input training sample for the algorithm, i.e., the multi-objective multi-agent deep reinforcement learning model. The model is trained in parallel for hundreds of cycles, with photovoltaic / load scenarios randomly assigned in each cycle. The network parameters are optimized through empirical replay and batch gradient descent. S43. Based on the training results of the multi-objective strategy, generate the Pareto front and analyze the sensitivity to obtain the trained reactive voltage control model; In step S43, after training is completed, all non-dominated solutions are extracted, and a three-dimensional Pareto front containing network loss, voltage deviation, and energy storage lifetime is constructed. By adjusting the weighting factors, the impact on energy storage lifetime is analyzed.

Citation Information

Cited By

  • Multi-energy storage collaborative virtual power plant economic optimization system and method

    CN121395467A

  • Distributed photovoltaic scheduling method based on multi-agent consensus optimization

    CN121436615A