Heterogeneous multi-agent reinforcement learning method and equipment for intelligent regulation and control of small hydropower station cluster, and medium
By constructing a heterogeneous multi-agent reinforcement learning method, adopting the HMATD3 algorithm and dual-Q network structure, and combining voltage-power sensitivity index, the problem of coordinated control between small hydropower clusters and energy storage equipment was solved, realizing efficient, safe and economical active-reactive power regulation of the distribution network.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUIZHOU POWER GRID CO LTD
- Filing Date
- 2025-12-18
- Publication Date
- 2026-04-24
AI Technical Summary
Traditional voltage-reactive power control methods for distribution networks are ill-suited to address the nonlinear fluctuations and localization of power output from small hydropower clusters. They lack flexible and proactive adjustment mechanisms, leading to frequent occurrences of voltage overruns and localized congestion. Furthermore, existing multi-agent reinforcement learning methods fail to effectively coordinate the differences between small hydropower and energy storage resources, resulting in low control efficiency.
A heterogeneous multi-agent reinforcement learning method for small hydropower clusters is constructed. The delayed deep deterministic policy gradient algorithm (HMATD3) is adopted, and a local TD error decomposition mechanism and a dual-Q network structure are introduced. Combined with the dynamic priority update mechanism of voltage-power sensitivity index, the collaborative control of small hydropower units and energy storage equipment is realized.
It improves the operating efficiency and intelligent control level of the distribution network in multi-source heterogeneous scenarios, ensures system voltage safety and economical operation, adapts to dynamic scenario changes, realizes deep coordinated regulation of active and reactive power, and improves the system's resilience and energy utilization efficiency.
Smart Images

Figure CN121923086A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent reinforcement learning technology, and particularly to a heterogeneous multi-agent reinforcement learning method, device, and medium for intelligent regulation of small hydropower clusters. Background Technology
[0002] Small hydropower, as a green, renewable, and relatively regulated distributed power source, is widely deployed in mountainous areas of central and western my country and regions rich in water resources, forming large-scale small hydropower clusters. This type of resource is characterized by its output being influenced by natural hydrological conditions and its adjustable boundaries being constrained by physical structures; therefore, its regulation potential has not yet been fully explored.
[0003] Traditional voltage-reactive power control in distribution networks mainly relies on centralized optimization strategies and linear control models, which are insufficient to address the nonlinear fluctuations and localization issues of small hydropower clusters. Furthermore, most small hydropower plants are currently used for grid-connected operation, lacking flexible and proactive regulation mechanisms and the ability to participate in real-time grid voltage regulation and power balance. This results in frequent voltage exceedances and localized congestion, impacting system operational safety and economic efficiency.
[0004] In recent years, with the decline in the cost of energy storage technology, grid-side energy storage systems (BSS) have become an effective means to improve control flexibility and cope with source-load uncertainties. However, small hydropower and energy storage equipment have significant heterogeneity in terms of operating characteristics, response speed, and control variables. Traditional centralized optimization methods are difficult to coordinate their actions simultaneously, which can easily lead to low control efficiency and decreased system stability.
[0005] Against this backdrop, reinforcement learning, with its ability to learn control strategies from data and adapt to complex dynamic environments, is increasingly being applied to power system operation optimization. Multi-agent reinforcement learning (MARL) is particularly suitable for complex control scenarios involving multiple devices and distributed decision-making. However, existing MARL methods are mostly based on homogeneous control unit designs, lacking differentiated modeling and parallel coordination mechanisms for heterogeneous resources such as hydropower and energy storage, making it difficult to achieve both refined control and overall system performance optimization.
[0006] Therefore, there is an urgent need for a new control method that can take into account the different characteristics of small hydropower and energy storage resources, and has online learning and distributed regulation capabilities, in order to improve the operating efficiency and regulation intelligence of the distribution network in multi-source heterogeneous scenarios, and ensure system voltage safety, economical operation and strategy adaptability. Summary of the Invention
[0007] In view of the aforementioned existing problems, the present invention is proposed.
[0008] Therefore, this invention provides a heterogeneous multi-agent reinforcement learning method, device, and medium for intelligent regulation of small hydropower clusters to solve the problems of coarse control granularity, lagging regulation response, strong resource heterogeneity, difficulty in coordination, poor strategy generalization, and inability to adapt to dynamic scenarios in the current operation of power distribution networks.
[0009] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, this invention provides a heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters, including: Obtain the operating parameters of small hydropower clusters, construct a power flow calculation model for distribution networks in the scenario of small hydropower cluster access, and determine the control boundary conditions for small hydropower units and energy storage devices. A collaborative control framework based on heterogeneous multi-agent reinforcement learning is constructed. Small hydropower units and energy storage devices are modeled as heterogeneous agents with independent strategies, and the linkage control task between multiple devices is formalized as a Markov game process. For heterogeneous agents, a delayed deep deterministic policy gradient algorithm is set up, and a local TD error decomposition mechanism and a dual-Q network structure are introduced to optimize the parallel training process of heterogeneous control policies of agents. A dynamic priority update mechanism based on voltage-power sensitivity index is constructed to guide key agents to iterate preferentially under a centralized training-distributed execution architecture; The heterogeneous intelligent agents that have completed training iterations are deployed to the field environment to achieve adaptive, distributed active and reactive power coordinated control for typical small hydropower regulation scenarios.
[0010] As a preferred embodiment of the heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters described in this invention, the method includes: constructing a power flow calculation model for the distribution network in the scenario of small hydropower cluster access and control boundary conditions for small hydropower units and energy storage devices, including: In a distribution network with a known topology, the DistFlow power flow model is introduced to model the active / reactive power balance, voltage-current relationship, and line power loss at each node, constructing a set of power flow equations, expressed as: In the formula, 、 These represent the node numbers in the distribution network; This indicates that the node is flowing in from other nodes. The set of upstream nodes, Indicates from node The set of downstream nodes that output power. Represents a node Reactive power output of small hydropower units; Represents a node Active power output of the energy storage system; , Representing nodes respectively Active and reactive power of electrical load; , They represent the nodes respectively To the node The active and reactive power transmitted on the branch line; Indicates a branch The square of the current; , Representing branch roads Resistance and reactance; This represents the set of all nodes in a power distribution network. Indicates from node downstream nodes The active power transmitted on the branch line; Indicates from node downstream nodes The reactive power transmitted on the branch line; Among them, nodes Reactive power output of small hydropower unit nodes The control boundary constraints are expressed as follows, which are adjusted by the excitation current: In the formula, Represents a node Small hydropower units at time The lower limit of reactive power output. Represents a node Small hydropower units at time The actual reactive power output, Represents a node Small hydropower units at time The upper limit of reactive power output; Among them, nodes Active power output of energy storage devices Simultaneously constrained by the dynamic boundary of SOC and the charging / discharging power limit, the dynamic expression of the control boundary conditions of charging / discharging SOC is as follows: In the formula, For energy storage systems at all times The state of charge; For energy storage systems at all times The state of charge; For energy storage systems at all times Active power output; To improve the charging and discharging efficiency of energy storage; This refers to the rated capacity of the energy storage system. , These are the minimum and maximum boundaries allowed by SOC, respectively; The time step is in hours.
[0011] As a preferred embodiment of the heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters described in this invention, the method includes: constructing a collaborative control framework based on heterogeneous multi-agent reinforcement learning, modeling small hydropower units and energy storage devices as heterogeneous agents with independent strategies, and formalizing the linkage regulation task between multiple devices as a Markov game process, including: Define the state space of the control framework Each agent's decision in the game depends on the state of the energy storage system, represented as: Among them, including the first At the [time]th moment The nodes accessed by each intelligent agent have active power. Node reactive power Node voltage Energy storage state of charge ; express One intelligent device; Setting the motion space Action space express The joint action of all intelligent agents at any given moment consists of the reactive power of the small hydropower unit and the active power of the energy storage, and is expressed as: In the formula, For at any time The joint action vector of all agents involved in regulation; This refers to the set of nodes currently involved in regulation. To indicate at time node Decision variables for reactive power output of small hydropower units; Indicates at time node The active power output decision variable of the energy storage system.
[0012] As a preferred embodiment of the heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters described in this invention, it further includes: setting a joint reward function, expressed as follows, to guide each agent to learn an economical operation strategy that satisfies the operating constraints of the energy storage system: In the formula, Represents the global reward function; The penalty coefficient representing voltage exceeding the limit; The penalty coefficient representing network loss; represent Voltage out-of-bounds reward function at any given time; represent Time-bound network loss reward function; , These represent the minimum and maximum values of the node voltage change, respectively. Indicates distribution network node The voltage amplitude; Indicates at time node The voltage state components are used for calculating voltage constraints and over-limit penalty functions; Indicates at time By node Power state and The determined branch active power loss function value, Indicates at time node The active power state component is used to calculate branch power loss; Indicates the energy storage system at time Total active power loss; and Indicates the lower and upper limits of voltage operation; express One intelligent agent; Define a state transition function. The state of the energy storage system changes due to the current actions of all agents. The state transition function is expressed as: in, This represents the state transition function of the energy storage system, i.e., at time t. state and joint actions Given the conditions, the energy storage system transitions to the next state. The probability distribution, through the power flow equations Mapping implementation; This represents the power flow model function for the distribution network.
[0013] As a preferred embodiment of the heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters described in this invention, wherein: a delayed deep deterministic policy gradient algorithm is set for the heterogeneous agents, and each agent in the algorithm... It has six neural networks, including: Actor network Target Actor Network Dual-action value Critic network and Target Critic Network and ; During the interaction phase with the distribution network environment, the Actor network relies on local observations. Output the action and add exploration noise to perturb the policy, represented as: in, It is Gaussian noise; Indicates at time intelligent agent Action decision-making; To indicate at time intelligent agent Local observation information; Represents intelligent agents The strategy function.
[0014] As a preferred embodiment of the heterogeneous multi-agent reinforcement learning method for intelligent control of small hydropower clusters described in this invention, the method introduces a local TD error decomposition mechanism and a dual-Q network structure to optimize the parallel training process of the heterogeneous control strategy of the agents, including: The Critic network constructs a loss function by minimizing the TD error, corresponding to the... The loss of a Q-network is expressed as: In the formula, For the first The loss function of the Q-network is used to minimize the TD error and evaluate the current parameters. The difference between the Q value and the target value; Batch size; To indicate the state and actions Next, the A Q network for intelligent agents Action value estimation; Indicates the state index; Among them, the TD target value Constructed using the minimum double-Q estimation structure, it is expressed as: In the formula, Indicates the state Lower intelligent agent The immediate reward obtained after performing an action; Indicates the discount factor; and These are the two output values of the target Q-network, used in the double-Q learning structure to avoid overestimation problems; Indicates the state After the action is performed, the energy storage system transitions to the next state; This indicates the local observation of the target policy network in the next time step. Intelligent agents generated below The action; During the heterogeneous control update phase, the Actor network aims to maximize the action value, and the gradient direction is guided by the Q network, as follows: In the formula, For intelligent agents Policy network parameters The gradient of the objective function; Indicates the intelligent agent action The partial derivative operator; To indicate the state and actions Below, the first Critic network evaluates the agent. Estimation of the value of an action; Indicates at time intelligent agent The action output; Represents intelligent agents The policy function maps local observations to actions; Represents intelligent agents Local observation information; To represent the local observations of the Actor network at the input At that time, the output action is relative to the parameters The gradient is used to update the policy parameters.
[0015] As a preferred embodiment of the heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters described in this invention, the parallel training process further includes: A strategy-delayed update mechanism is adopted, where the Actor network is updated only once after multiple rounds of updates to the Critic network. Meanwhile, the target network uses a soft update method to achieve asymptotic approximation: In the formula, Represents intelligent agents The target network parameters, Represents intelligent agents The current network parameters, This is a soft update coefficient; While asymptotically approximating the target, the policy network is updated only once to suppress training oscillations.
[0016] As a preferred embodiment of the heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters described in this invention, the method includes: constructing a dynamic priority update mechanism based on voltage-power sensitivity indicators to guide key agents in priority iteration under a centralized training-distributed execution architecture, comprising: By setting the update order of agents during the training phase, each agent is guided to update the policy in the optimal order during intensive training, as shown below: In the formula, , , , These are the active and reactive power partial derivative matrices for the node voltage and the active and reactive power partial derivative matrices for the phase angle pair, respectively, used to characterize the sensitivity of voltage changes to active / reactive power output. A sensitivity matrix representing the relationship between node voltage and active / reactive power; After each round of training, the total VQ sensitivity of each node is calculated to measure its impact on the voltage stability of the energy storage system. Sensitivity score of each agent corresponding to a node Represented as: In the formula, This represents the total number of nodes in the distribution network. Indicates at time ,node For nodes Voltage-power sensitivity index value for power disturbance; Considering the time-dynamic characteristics of the power grid operating environment, a reward-weighted coefficient is introduced. The adjustment weight, which is the sensitivity index of the agent at the current moment, is defined as the reward at the current moment. Total round rewards The ratio is expressed as: In the formula, Indicates the state Execute joint actions The instant rewards received Indicates the total duration of a training round or evaluation period; Based on the weighted sensitivity index Sort all agents and generate an update priority list. During the intensive training phase, agents are scheduled one by one in sequence to update the policy.
[0017] In a second aspect, the present invention provides an electronic device, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters are implemented.
[0018] Thirdly, the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters.
[0019] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention constructs a multi-agent cooperative optimization model based on Markov games, modeling small hydropower units and energy storage devices as heterogeneous agents, and employs the Delayed Deep Deterministic Policy Gradient Algorithm (HMATD3) for joint training. This enables each control unit to achieve deep cooperative regulation of active and reactive power while ensuring system operational safety, significantly improving the coordination of regulation and the overall system operating efficiency. By introducing a Local Error Decomposition (TD) mechanism and a Dual-Q Network (DQN) structure, the policy stability of heterogeneous multi-agents during parallel training can be enhanced. Combined with a dynamic priority update mechanism, key control nodes learn policies first, effectively mitigating policy conflicts and oscillations, and accelerating the convergence speed and control accuracy of the overall system policy. Simultaneously, by adopting a centralized training-distributed execution architecture, each agent can independently execute control actions, possessing online deployment and local perception capabilities. This allows it to adapt to dynamic environmental changes such as small hydropower output fluctuations and load disturbances, achieving all-weather, full-process distributed active and reactive power regulation, significantly improving the voltage stability, energy utilization efficiency, and system resilience of the distribution network. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1This is a schematic diagram of the overall process of a heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters according to an embodiment of the present invention.
[0022] Figure 2 This is a schematic diagram of a simulation test system in a heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters, as described in an embodiment of the present invention. Detailed Implementation
[0023] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0024] Example 1, referring to Figure 1 As one embodiment of the present invention, a heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters is provided, comprising: S100: Obtain the operating parameters of small hydropower clusters, construct a power flow calculation model for distribution networks in the scenario of small hydropower cluster access, and determine the control boundary conditions for small hydropower units and energy storage devices. S200: Construct a collaborative control framework based on heterogeneous multi-agent reinforcement learning, model small hydropower units and energy storage devices as heterogeneous agents with independent strategies, and formalize the linkage control task between multiple devices into a Markov game process. S300: A delayed, deep deterministic policy gradient algorithm is set for heterogeneous agents. A local TD error decomposition mechanism and a dual-Q network structure are introduced to optimize the parallel training process of heterogeneous control policies for agents. S400: Construct a dynamic priority update mechanism based on voltage-power sensitivity index to guide key agents to iterate preferentially under a centralized training-distributed execution architecture; S500: Deploy the heterogeneous intelligent agents that have completed training iterations to the field environment to achieve adaptive, distributed active and reactive power coordinated control for typical small hydropower regulation scenarios.
[0025] It should be noted that current traditional centralized voltage-reactive power control strategies are difficult to achieve refined modeling and rapid response for large-scale heterogeneous resources, such as small hydropower and energy storage devices. Small hydropower and energy storage devices differ significantly in terms of operating boundaries, response time, and controllable variables, lacking a unified and efficient collaborative control framework. Furthermore, fixed rules and model-driven methods are difficult to adapt to the natural fluctuations in small hydropower output and dynamic load changes, making it difficult to achieve long-term robust control. Without an effective coordination mechanism, it can easily lead to problems such as node voltage exceeding limits, increased energy loss, and even grid instability.
[0026] To address the aforementioned main issues, steps S100-S500 mainly propose an intelligent control method that integrates Markov game modeling and the Heterogeneous Multi-Agent Delayed Deep Deterministic Policy Gradient Algorithm (HMATD3). Small hydropower units and energy storage devices are modeled as agents with heterogeneous policies. Through a centralized training-distributed execution architecture, combined with voltage-power sensitivity indicators and priority update mechanisms, the active and reactive power coordinated control of the distribution network under the access of small hydropower clusters is realized.
[0027] Example 2, refer to Figure 1 As an embodiment of the present invention, based on the above embodiment, a heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters is provided.
[0028] In this embodiment of the invention, step S100 involves constructing a power flow calculation model for a distribution network oriented towards small hydropower cluster access scenarios, as well as control boundary conditions for small hydropower units and energy storage devices, including: In a distribution network with a known topology, a power flow model (DistFlow) is introduced to model the active / reactive power balance, voltage-current relationship, and line power loss at each node. The following set of power flow equations is constructed, expressed as follows: In the formula, 、 These represent the node numbers in the distribution network; This indicates that the node is flowing in from other nodes. The set of upstream nodes, Indicates from node The set of downstream nodes that output power. Represents a node Reactive power output of small hydropower units; Represents a node Active power output of the energy storage system; , Representing nodes respectively Active and reactive power of electrical load; , They represent the nodes respectively To the node The active and reactive power transmitted on the branch line; Indicates a branch The square of the current; , Representing branch roads Resistance and reactance; This represents the set of all nodes in a power distribution network. Indicates from node downstream nodes The active power transmitted on the branch line; Indicates from node downstream nodes The reactive power transmitted on the branch line; Among them, nodes Reactive power output of small hydropower unit nodes The control boundary constraints are expressed as follows, which are adjusted by the excitation current: In the formula, Represents a node Small hydropower units at time The lower limit of reactive power output; Represents a node Small hydropower units at time The actual reactive power output; Represents a node The upper limit of reactive power output of a small hydropower unit at time t; Among them, nodes Active power output of energy storage devices Simultaneously constrained by the dynamic boundary of the state of charge (SOC) and the charging / discharging power limit, the dynamic control boundary conditions of the charging / discharging SOC are expressed as follows: In the formula, For energy storage systems at all times The state of charge; For energy storage systems at all times The state of charge; For energy storage systems at all times Active power output; To improve the charging and discharging efficiency of energy storage; This refers to the rated capacity of the energy storage system. , These are the minimum and maximum boundaries allowed by SOC, respectively; The time step is in hours.
[0029] In addition, the voltage at grid nodes must meet operational safety constraints: In the formula, For nodes At any moment The voltage amplitude; , These are the lower and upper limits for voltage operation, respectively.
[0030] It should be noted that this step is used to clarify the active and reactive power regulation capabilities, response timeliness, and operational constraints of controllable resources; the aim is to establish a foundation for power flow analysis and control modeling of distribution networks suitable for the joint operation of small hydropower clusters and energy storage systems, and to comprehensively characterize key operational variables such as active / reactive power transmission, voltage levels, line currents, and equipment constraints in the system.
[0031] Furthermore, the aforementioned distribution network operation model is primarily used to characterize the dynamic evolution of small hydropower units and energy storage devices under voltage, current, and power constraints. However, this invention does not directly solve the model established in S100 numerically, but instead transforms it into a Markov Decision Process (MDP) in the next step S200. Specifically, the state space and action space of the MDP are constructed based on the physical constraints of the aforementioned model: state variables reflect key information such as grid voltage, power flow, and equipment energy status; action variables correspond to the reactive power regulation of small hydropower units and the active power output regulation of energy storage devices. Thus, the aforementioned model mainly serves to provide the basis for defining states and actions in this scheme, and the parameters involved will be specifically reflected in the subsequent definitions of the state space and action space.
[0032] In this embodiment of the invention, step S200 involves constructing a cooperative control framework based on heterogeneous multi-agent reinforcement learning. Small hydropower units and energy storage devices are modeled as heterogeneous agents with independent strategies, and the task of coordinated control between multiple devices is formalized as a Markov game process, including: Define the state space of the control framework Each agent's decision in the game depends on the state of the energy storage system, represented as: Among them, including the first At the [time]th moment The nodes accessed by each intelligent agent have active power. Node reactive power Node voltage Energy storage state of charge ; express One intelligent device; Setting the motion space Action space express The joint action of all intelligent agents at any given moment consists of the reactive power of the small hydropower unit and the active power of the energy storage, and is expressed as: In the formula, For at any time The joint action vector of all agents involved in regulation; This refers to the set of nodes currently involved in regulation. To indicate at time node Decision variables for reactive power output of small hydropower units; Indicates at time node The active power output decision variable of the energy storage system.
[0033] In this embodiment of the invention, step S200 further includes setting a joint reward function, expressed as follows, to guide each agent to learn an economical operation strategy that satisfies the operating constraints of the energy storage system: In the formula, Represents the global reward function; The penalty coefficient representing voltage exceeding the limit; The penalty coefficient representing network loss; represent The voltage out-of-bounds reward function at any given time. represent Time-bound network loss reward function; , These represent the minimum and maximum values of the node voltage change, respectively. Indicates distribution network node voltage amplitude, Indicates at time node The voltage state components are used for calculating voltage constraints and over-limit penalty functions; Indicates at time By node Power state and The determined branch active power loss function value, Indicates at time node The active power state component is used to calculate branch power loss. Indicates the energy storage system at time Total active power loss; and Indicates the lower and upper limits of voltage operation; express One intelligent agent; Define a state transition function. The state of the energy storage system changes due to the current actions of all agents. The state transition function is expressed as: in, This represents the state transition function of the energy storage system, i.e., at time t. state and joint actions Given the conditions, the energy storage system transitions to the next state. The probability distribution, through the power flow equations Mapping implementation; This represents the power flow model function for the distribution network.
[0034] It should be noted that in step S200, within the context of distribution network operation with the integration of small hydropower clusters and energy storage systems, each small hydropower unit and energy storage device in the system is considered as an intelligent agent with perception and decision-making capabilities. A multi-agent cooperative control framework based on a Markov game model is constructed. Under this framework, each agent independently learns policies in a distributed environment based on local state perception and system incentives, and achieves active and reactive power cooperative control optimization of the distribution network through game interaction with other agents.
[0035] Furthermore, addressing the significant heterogeneity, large differences in control boundaries, and varying dynamic response rates between small hydropower units and energy storage devices, a depth-deterministic policy gradient algorithm with a set delay is designed to improve the stability and convergence efficiency of multi-agent systems during parallel training. This algorithm supports the independent evolution of strategies for each agent under joint game theory and possesses the following core mechanisms: In this embodiment of the invention, step S300 sets a delay-deterministic policy gradient algorithm for heterogeneous agents, where each agent in the algorithm... It has six neural networks, including: reinforcement learning (Actor) networks. Target reinforcement learning (Actor) network Dual-action value (Critic) network and Target Dual Action Value (Critic) Network and ; During the interaction phase with the distribution network environment, the Actor network relies on local observations. Output the action and add exploration noise to perturb the policy, represented as: in, It is Gaussian noise; Indicates at time intelligent agent Action decision-making; To indicate at time intelligent agent Local observation information; Represents intelligent agents The strategy function.
[0036] In this embodiment of the invention, step S300, to evaluate the prediction accuracy of each region policy for the action value in the current state, introduces a local error (TD) decomposition mechanism and a dual-Q network structure to optimize the parallel training process of heterogeneous control policies of the agent, including: In the formula, For the first The loss function of the Q-network is used to minimize the TD error and evaluate the current parameters. The difference between the Q value and the target value; Batch size; To indicate the state and actions Next, the A Q network for intelligent agents Action value estimation; Indicates the state index; Among them, the TD target value Constructed using the minimum double-Q estimation structure, it is expressed as: In the formula, Indicates the state Lower intelligent agent The immediate reward obtained after performing an action; Indicates the discount factor; and These are the two output values of the target Q-network, used in the double-Q learning structure to avoid overestimation problems; Indicates the state After the action is performed, the energy storage system transitions to the next state; This indicates the local observation of the target policy network in the next time step. Intelligent agents generated below The action; During the heterogeneous control update phase, the Actor network aims to maximize the action value, and the gradient direction is guided by the Q network, as follows: In the formula, For intelligent agents Policy network parameters The gradient of the objective function; Indicates the intelligent agent action The partial derivative operator; To indicate the state and actions Below, the first Critic network evaluates the agent. Estimation of the value of an action; Indicates at time intelligent agent The action output; Represents intelligent agents The policy function maps local observations to actions; Represents intelligent agents Local observation information; To represent the local observations of the Actor network at the input At that time, the output action is relative to the parameters The gradient is used to update the policy parameters.
[0037] In this embodiment of the invention, the parallel training process in step S300 further includes: To further improve the stability of the algorithm, HMATD3 adopts a policy delayed update mechanism, which updates the Actor network only once after updating the Critic network for multiple rounds. Meanwhile, the target network uses a soft update method to achieve asymptotic approximation: In the formula, Represents intelligent agents The target network parameters, Represents intelligent agents The current network parameters, This is a soft update coefficient; While asymptotically approximating the target, the policy network is updated only once to suppress training oscillations.
[0038] In this embodiment of the invention, step S400 involves constructing a dynamic priority update mechanism based on the voltage-power sensitivity index to guide key agents to iterate preferentially under a centralized training-distributed execution architecture, including the following steps A1-A3: A1: By setting the update order of agents during the training phase, each agent is guided to update the policy in the optimal order during intensive training, as shown below: In the formula, , , , These are the active and reactive power partial derivative matrices for the node voltage and the active and reactive power partial derivative matrices for the phase angle pair, respectively, used to characterize the sensitivity of voltage changes to active / reactive power output. A sensitivity matrix representing the relationship between node voltage and active / reactive power; It should be noted that setting the update order of agents during the training phase can accelerate convergence and enhance the system's response to voltage disturbances.
[0039] A2: After each round of training, calculate the total voltage-reactive power (VQ) sensitivity of each node to measure its impact on system voltage stability. Sensitivity score of each agent corresponding to a node Represented as: In the formula, This represents the total number of nodes in the distribution network. Indicates at time ,node For nodes Voltage-power sensitivity index for power disturbances.
[0040] It should be noted that the sensitivity index value Indicators reflect nodes The ability to respond to voltage disturbances at other nodes; higher sensitivity indicates that the corresponding agent at that node has greater potential to improve the stability of the energy storage system through policy updates.
[0041] A3: Considering the time-dynamic characteristics of the power grid operating environment, a reward-weighted coefficient is introduced. The adjustment weight, which is the sensitivity index of the agent at the current moment, is defined as the reward at the current moment. Total round rewards The ratio is expressed as: In the formula, Indicates the state Execute joint actions The instant rewards received This indicates the total duration of a training round or evaluation period.
[0042] It should be noted that adjusting the weights can enhance the system's ability to respond to phased voltage disturbances.
[0043] Based on the weighted sensitivity index Sort all agents and generate an update priority list. During the intensive training phase, agents are scheduled one by one in sequence to update the policy.
[0044] In this embodiment of the invention, step S500 involves deploying the heterogeneous intelligent agents that have completed training iterations to the field environment to achieve adaptive, distributed active-reactive power coordinated control for typical small hydropower regulation scenarios. Specifically, the steps include: (1) State perception: In each control cycle, the small hydropower station and the energy storage agent collect key operating information such as local node voltage, active load, reactive demand, and SOC status (energy storage) to construct their own state vector. (2) Strategy invocation: Each agent inputs its current state into the locally deployed reinforcement learning (Actor) neural network and outputs continuous control actions, including the excitation regulation of small hydropower (which indirectly determines reactive power output) and adjustable output, the charging and discharging power of energy storage equipment and reactive power support instructions. (3) Edge execution: The generated control actions are sent to the corresponding execution devices, such as the excitation system and the physical coding sublayer (PCS) interface, through control commands to complete the power regulation operation of the physical layer; (4) Control closed loop: The control process is executed iteratively in each scheduling cycle to form a real-time closed loop of state-action-feedback, realizing online active-reactive distributed collaborative control.
[0045] Overall, without relying on centralized coordination, each agent can operate autonomously locally and supports offline retraining or parameter fine-tuning according to the policy update cycle to adapt to changes in the external environment.
[0046] Example 3, referring to Figure 2 This is an embodiment of the present invention, designed to verify the actual control performance and engineering feasibility of the method. A simulation experiment was conducted based on the IEEE Power System Test Model (IEEE) 118-node distribution network system, with the system topology shown in the figure. The system includes several typical distributed energy units and control devices, including three small hydropower units (located at nodes 17, 62, and 85, respectively) and four energy storage devices (located at nodes 25, 48, 95, and 105, respectively). Node 1 is the main power supply node, and the simulation period is set to a typical daily 96-period period (each period lasting 15 minutes). Load and hydropower output were constructed using actual historical fluctuation data.
[0047] This experiment uses the HMATD3 intelligent agent built on the platform as the control core, combining the excitation regulation of small hydropower and the charging and discharging control capabilities of energy storage. The reactive power regulation boundary of the hydropower unit is defined (set to 0.8–0.8 pu), the active power regulation boundary of the energy storage device is set to ±0.5 MW, the SOC range is limited to [0.2, 0.8], and the charging and discharging efficiency is 95%. The hyperparameters of the HMATD3 algorithm of this invention are shown in Table 1, where MLP is a multilayer perceptron, ReLU is a modified linear unit function, and Tanh is a hyperbolic tangent function.
[0048] Table 1 Hyperparameter settings of the algorithm
[0049] After training, the proposed heterogeneous multi-agent reinforcement learning control strategy was deployed to the IEEE 118-node distribution network simulation environment. Typical daily full-cycle (96 time periods) online control tests were conducted, and its performance was compared with the following two methods: Method A: Centralized Deep Deterministic Policy Gradient (DDPG) control strategy (no multi-agent modeling, unified policy); Method B: Traditional multi-agent deep reinforcement learning (MATD3) strategy (without heterogeneous modeling and priority update mechanism).
[0050] The experiment focuses on examining the following three key indicators: Voltage Violation Rate: Measures the percentage of nodes and time periods where the voltage exceeds [0.95, 1.05] pu over the entire cycle. Total Active Power Loss: The sum of energy losses in all branches throughout the entire lifecycle; Policy convergence speed and stability (Training Stability): Examines the trend and fluctuation of the average reward of the agent during the training phase.
[0051] The simulation results are shown below:
[0052] The results show that: The method of this invention controls the voltage over-limit rate to below 1%, which is significantly better than centralized control and traditional multi-agent methods, ensuring node voltage compliance. The system's average daily active power loss was reduced to 0.991MW, a decrease of 29.8% and 11.4% compared to DDPG and MATD3, respectively, achieving higher energy efficiency; During the training phase, the proposed strategy reduced the number of convergence rounds by more than 25% and significantly reduced training fluctuations, indicating that the heterogeneous modeling and priority update mechanism effectively improved learning stability and efficiency.
[0053] Example 4: The above is an illustrative scheme of a heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters.
[0054] This embodiment also provides an electronic device suitable for a heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to realize the heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters as proposed in the above embodiment.
[0055] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, it implements the heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters as proposed in the above embodiment.
[0056] The electronic device and medium proposed in this embodiment belong to the same inventive concept as the heterogeneous multi-agent reinforcement learning method for realizing intelligent regulation of small hydropower clusters proposed in the above embodiments. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0057] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.
[0058] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A heterogeneous multi-agent reinforcement learning method for intelligent control of small hydropower clusters, characterized in that, include: Obtain the operating parameters of small hydropower clusters, construct a power flow calculation model for distribution networks in the scenario of small hydropower cluster access, and determine the control boundary conditions for small hydropower units and energy storage devices. A collaborative control framework based on heterogeneous multi-agent reinforcement learning is constructed. Small hydropower units and energy storage devices are modeled as heterogeneous agents with independent strategies, and the linkage control task between multiple devices is formalized as a Markov game process. For heterogeneous agents, a delayed deep deterministic policy gradient algorithm is set up, and a local TD error decomposition mechanism and a dual-Q network structure are introduced to optimize the parallel training process of heterogeneous control policies of agents. A dynamic priority update mechanism based on voltage-power sensitivity index is constructed to guide key agents to iterate preferentially under a centralized training-distributed execution architecture; The heterogeneous intelligent agents that have completed training iterations are deployed to the field environment to achieve adaptive, distributed active and reactive power coordinated control for typical small hydropower regulation scenarios.
2. The heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters as described in claim 1, characterized in that, Construct a power flow calculation model for distribution networks in the scenario of small hydropower cluster access, as well as control boundary conditions for small hydropower units and energy storage devices, including: In a distribution network with a known topology, the DistFlow power flow model is introduced to model the active / reactive power balance, voltage-current relationship, and line power loss at each node, constructing a set of power flow equations, expressed as: In the formula, 、 These represent the node numbers in the distribution network; This indicates that the node is flowing in from other nodes. The set of upstream nodes, Indicates from node The set of downstream nodes that output power. Represents a node Reactive power output of small hydropower units; Represents a node Active power output of the energy storage system; , Representing nodes respectively Active and reactive power of electrical load; , They represent the nodes respectively To the node The active and reactive power transmitted on the branch line; Indicates a branch The square of the current; , Representing branches Resistance and reactance; This represents the set of all nodes in a power distribution network. Indicates from node downstream nodes The active power transmitted on the branch line; Indicates from node downstream nodes The reactive power transmitted on the branch line; Among them, nodes Reactive power output of small hydropower unit nodes The control boundary constraints are expressed as follows, which are adjusted by the excitation current: In the formula, Represents a node Small hydropower units at time The lower limit of reactive power output. Represents a node Small hydropower units at time The actual reactive power output, Represents a node Small hydropower units at time The upper limit of reactive power output; Among them, nodes Active power output of energy storage devices Simultaneously constrained by the dynamic boundary of SOC and the charging / discharging power limit, the dynamic expression of the control boundary conditions of charging / discharging SOC is as follows: In the formula, For energy storage systems at all times The state of charge; For energy storage systems at all times The state of charge; For energy storage systems at all times Active power output; To improve the charging and discharging efficiency of energy storage; This refers to the rated capacity of the energy storage system. , These are the minimum and maximum boundaries allowed by SOC, respectively; The time step is in hours.
3. The heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters as described in claim 2, characterized in that, A collaborative control framework based on heterogeneous multi-agent reinforcement learning is constructed. Small hydropower units and energy storage devices are modeled as heterogeneous agents with independent policies, and the task of coordinated control between multiple devices is formalized as a Markov game process, including: Define the state space of the control framework Each agent's decision in the game depends on the state of the energy storage system, represented as: Among them, including the first At the [time]th moment The nodes accessed by each intelligent agent have active power. Node reactive power Node voltage Energy storage state of charge ; express One intelligent device; Setting the motion space Action space express The joint action of all intelligent agents at any given moment consists of the reactive power of the small hydropower unit and the active power of the energy storage, and is expressed as: In the formula, For at any time The joint action vector of all agents involved in regulation; This refers to the set of nodes currently involved in regulation. To indicate at time node Decision variables for reactive power output of small hydropower units; Indicates at time node The active power output decision variable of the energy storage system.
4. The heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters as described in claim 3, characterized in that, It also includes setting a joint reward function, expressed as follows, to guide each agent in learning an economical operating strategy that satisfies the constraints of the energy storage system: In the formula, Represents the global reward function; The penalty coefficient representing voltage exceeding the limit; The penalty coefficient representing network loss; represent Voltage out-of-bounds reward function at any given time; represent Time-bound network loss reward function; , These represent the minimum and maximum values of the node voltage change, respectively. Indicates distribution network node The voltage amplitude; Indicates at time node The voltage state components are used for calculating voltage constraints and over-limit penalty functions; Indicates at time By node Power state and The determined branch active power loss function value, Indicates at time node The active power state component is used to calculate branch power loss; Indicates the energy storage system at time Total active power loss; and Indicates the lower and upper limits of voltage operation; express One intelligent agent; Define a state transition function. The state of the energy storage system changes due to the current actions of all agents. The state transition function is expressed as: in, This represents the state transition function of the energy storage system, i.e., at time t. state and joint actions Given the conditions, the energy storage system transitions to the next state. The probability distribution, through the power flow equations Mapping implementation; This represents the power flow model function for the distribution network.
5. The heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters as described in claim 4, characterized in that, A delayed deep deterministic policy gradient algorithm is designed for heterogeneous agents. Each agent in the algorithm... It has six neural networks, including: Actor network Target Actor Network Dual-action value Critic network and Target Critic Network and ; During the interaction phase with the distribution network environment, the Actor network relies on local observations. Output the action and add exploration noise to perturb the policy, represented as: in, It is Gaussian noise; Indicates at time intelligent agent Action decision-making; To indicate at time intelligent agent Local observation information; Represents intelligent agents The strategy function.
6. The heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters as described in claim 5, characterized in that, A local TD error decomposition mechanism and a dual-Q network structure are introduced to optimize the parallel training process of heterogeneous control strategies for intelligent agents, including: The Critic network constructs a loss function by minimizing the TD error, corresponding to the... The loss of a Q-network is expressed as: In the formula, For the first The loss function of the Q-network is used to minimize the TD error and evaluate the current parameters. The difference between the Q value and the target value; Batch size; To indicate the state and actions Next, the A Q network for intelligent agents Action value estimation; Indicates the state index; Among them, the TD target value Constructed using the minimum double-Q estimation structure, it is expressed as: In the formula, Indicates the state Lower intelligent agent The immediate reward obtained after performing an action; Indicates the discount factor; and These are the two output values of the target Q-network, used in the double-Q learning structure to avoid overestimation problems; Indicates the state After the action is performed, the energy storage system transitions to the next state; This indicates the local observation of the target policy network in the next time step. Intelligent agents generated below The action; During the heterogeneous control update phase, the Actor network aims to maximize the action value, and the gradient direction is guided by the Q network, as follows: In the formula, For intelligent agents Policy network parameters The gradient of the objective function; Indicates the intelligent agent action The partial derivative operator; To indicate the state and actions Below, the first Critic network evaluates the agent. Estimation of the value of an action; Indicates at time intelligent agent The action output; Represents intelligent agents The policy function maps local observations to actions; Represents intelligent agents Local observation information; To represent the local observations of the Actor network at the input At that time, the output action is relative to the parameters The gradient is used to update the policy parameters.
7. The heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters as described in claim 6, characterized in that, The parallel training process also includes: A strategy-delayed update mechanism is adopted, where the Actor network is updated only once after multiple rounds of updates to the Critic network. Meanwhile, the target network uses a soft update method to achieve asymptotic approximation: In the formula, Represents intelligent agents The target network parameters, Represents intelligent agents The current network parameters, This is a soft update coefficient; While asymptotically approximating the target, the policy network is updated only once to suppress training oscillations.
8. The heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters as described in claim 7, characterized in that, A dynamic priority update mechanism based on voltage-power sensitivity metrics is constructed to guide key agents in priority iteration within a centralized training-distributed execution architecture, including: By setting the update order of agents during the training phase, each agent is guided to update the policy in the optimal order during intensive training, as shown below: In the formula, , , , These are the active and reactive power partial derivative matrices for the node voltage and the active and reactive power partial derivative matrices for the phase angle pair, respectively, used to characterize the sensitivity of voltage changes to active / reactive power output. A sensitivity matrix representing the relationship between node voltage and active / reactive power; After each round of training, the total VQ sensitivity of each node is calculated to measure its impact on the voltage stability of the energy storage system. Sensitivity score of each agent corresponding to a node Represented as: In the formula, This represents the total number of nodes in the distribution network. Indicates at time ,node For nodes Voltage-power sensitivity index value for power disturbance; Considering the time-dynamic characteristics of the power grid operating environment, a reward-weighted coefficient is introduced. The adjustment weight, which is the sensitivity index of the agent at the current moment, is defined as the reward at the current moment. Total round rewards The ratio is expressed as: In the formula, Indicates the state Execute joint actions The instant rewards received Indicates the total duration of a training round or evaluation period; Based on the weighted sensitivity index Sort all agents and generate an update priority list. During the intensive training phase, agents are scheduled one by one in sequence to update the policy.
9. An electronic device, characterized in that, include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, they implement the steps of the heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The device stores computer-executable instructions, which, when executed by a processor, implement the steps of the heterogeneous multi-agent reinforcement learning method for intelligent regulation of small hydropower clusters as described in any one of claims 1 to 8.