A micro-grid energy scheduling method and system based on multi-agent reinforcement learning

CN122823503APending Publication Date: 2026-09-25ANHUI ZHONGXIN JIYUAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611272757.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-21
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

但孤岛微电网的频率是一个全局统一标量,任何节点处的有功不平衡均会瞬间反映为全网频率的同步变化,各分布式电源的调节动作通过电网的功率平衡方程发生即时耦合,导致不同智能体各自动作产生的边际频率贡献在瞬态响应轨迹中相互混合,难以被有效分解,使得信用分配机制无法准确评估单个智能体的真实调节效果,各智能体因此倾向于学习到彼此近似且被平均化的模糊策略,无法形成针对实时扰动的精准协同响应,最终引发频率调节过程中的功率过调与持续振荡,严重威胁孤岛微电网的安全稳定运行

Benefits of technology

[0044]1.针对孤岛微电网频率统一标量特性导致的协同困境,通过识别电压相角差变化率与本地的频率变化率的极性冲突,将暂态协同失配的瞬时工况从正常运行中分离出来,直接捕捉了本节点调节动作未能改善相邻节点频率响应甚至产生反向效果的时刻,为后续计算边际贡献提供了明确的物理切入点,利用极性相反时刻前后的暂态有功增量与本节点和相邻节点间线路电抗的比值,从有功功率变化的维度估计本智能体动作对频率恢复的独立作用分量,使得在全局耦合的频率响应中能够提取出可归因的个体调节效果,解决了信用分配机制在全局标量环境下贡献信号被湮灭的难题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122823503A_ABST
    Figure CN122823503A_ABST
Patent Text Reader

Abstract

The application discloses a kind of microgrid energy scheduling method and system based on multi-agent reinforcement learning, is specifically related to power system dispatching control technical field, for solving in island microgrid global frequency scalar environment, credit distribution mechanism cannot accurately distinguish the real contribution of each agent to frequency regulation, leading to the problem that each agent strategy tends to be fuzzy, system produces power over-regulation and sustained oscillation;By detecting the frequency deviation exceeds the dead zone and the voltage phase angle difference rate and the local frequency change rate polarity is opposite, determine the transient state of collaborative adjustment, use the ratio of the transient active increment injected by the adjacent nodes before and after the time of opposite polarity in this state to the line reactance to construct marginal contribution estimate, the estimate is used as the advantage function correction term and the value gradient of the centralized Critic network to obtain the credit distribution signal, each agent updates the Actor network strategy accordingly and generates energy scheduling action.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system dispatch and control technology, and more specifically, to a microgrid energy dispatch method and system based on multi-agent reinforcement learning. Background Technology

[0002] Microgrids are regional power systems composed of distributed generation sources, energy storage units, and local loads, and can operate in both grid-connected and islanded modes. With the increasing penetration of renewable energy, the proportion of intermittent power sources such as photovoltaics and wind power in microgrids has significantly increased. The uncertainty on both the source and load sides poses a significant challenge to energy dispatch. Traditional energy dispatch methods based on mixed-integer linear programming or deterministic dynamic programming rely on accurate system models and predictive data, making it difficult to adapt to rapid fluctuations in generation output and load demand in real-time operation. To address these issues, each generation unit, energy storage device, and load can be treated as an independent decision-making agent. Through centralized training and a distributed execution framework, each agent can autonomously learn dispatch strategies under local observation conditions, thereby achieving adaptive energy management without relying on a globally accurate model.

[0003] However, when a microgrid disconnects from the main grid and enters islanded operation, the system inertia decreases significantly. Frequency stabilization becomes a globally rigid task that all power generation and energy storage units must collaboratively complete. In existing scheduling methods based on multi-agent reinforcement learning, the system relies on a credit allocation mechanism to distinguish the contributions of each agent to frequency regulation, thereby guiding strategy updates. However, the frequency of an islanded microgrid is a globally unified scalar. Active power imbalance at any node will instantly reflect a synchronous change in the frequency of the entire grid. The regulation actions of each distributed power source are coupled instantly through the power balance equation of the grid, causing the marginal frequency contributions generated by the actions of different agents to mix in the transient response trajectory, making it difficult to decompose effectively. This makes it impossible for the credit allocation mechanism to accurately evaluate the true regulation effect of a single agent. As a result, each agent tends to learn fuzzy strategies that are similar to each other and averaged out, failing to form a precise collaborative response to real-time disturbances. Ultimately, this leads to power overshoot and continuous oscillations during frequency regulation, seriously threatening the safe and stable operation of the islanded microgrid. Summary of the Invention

[0004] To overcome the aforementioned deficiencies of the prior art, this invention provides a microgrid energy scheduling method and system based on multi-agent reinforcement learning to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] A microgrid energy dispatching method based on multi-agent reinforcement learning includes the following steps:

[0007] S1: Each agent acquires its local frequency value, active power, and frequency change rate, as well as the voltage phase angle difference change rate with neighboring nodes;

[0008] S2: When the frequency deviation determined based on the frequency value exceeds the dead zone, and the voltage phase angle difference change rate exceeds the preset change rate threshold and is opposite in polarity to the local frequency change rate, it is determined that the transient coordinated adjustment state is entered.

[0009] S3: Under transient coordinated regulation, the transient active power increment injected by this node into adjacent nodes before and after the moment of opposite polarity is calculated based on the active power. The ratio of the transient active power increment to the line reactance between this node and adjacent nodes is used as the estimated marginal contribution of this agent's action to frequency recovery, including:

[0010] The active power of the measurement cycle before the moment of opposite polarity is taken as the first active power, and the active power of the measurement cycle after the moment of opposite polarity is taken as the second active power.

[0011] Subtracting the first active power from the second active power yields the transient active power increment;

[0012] Read the line reactance between this node and adjacent nodes;

[0013] Dividing the transient active power increment by the line reactance yields the equivalent electromagnetic torque component, which is then used as an estimate of the marginal contribution of the agent's actions to frequency recovery.

[0014] S4: The marginal contribution estimate is used as the advantage function correction term and fused with the value gradient output by the global Critic network in centralized training to obtain the credit allocation signal for this agent.

[0015] S5: Each agent updates the policy gradient of its local Actor network based on the credit allocation signal to generate the energy scheduling action for the next time step.

[0016] Furthermore, each agent acquires its local frequency value, active power, and frequency change rate, as well as the voltage phase angle difference change rate with neighboring nodes, including:

[0017] The local three-phase voltage is collected by a voltage transformer, and the local frequency value is obtained by performing a phase-locked loop operation on the local three-phase voltage.

[0018] The rate of change of frequency is obtained by differentiating the local frequency value.

[0019] The active power is calculated by collecting the local three-phase current through a current transformer and combining it with the local three-phase voltage.

[0020] The three-phase voltages of adjacent nodes are collected by voltage transformers. The phase angles of the adjacent nodes are obtained by performing phase-locked loop operations on the three-phase voltages of the adjacent nodes. The voltage phase angle difference is obtained by subtracting the voltage phase angle of the adjacent nodes from the local voltage phase angle. The voltage phase angle difference is obtained by performing differential operations on the voltage phase angle difference to obtain the rate of change of the voltage phase angle difference.

[0021] Furthermore, a phase-locked loop operation is performed on the local three-phase voltage to obtain the local frequency value, including: performing a Clarke transform on the local three-phase voltage to obtain the alpha-axis voltage component and beta-axis voltage component in a two-phase stationary coordinate system; performing a Park transform on the alpha-axis voltage component and beta-axis voltage component to obtain the direct-axis voltage component and quadrature-axis voltage component in a two-phase rotating coordinate system; inputting the quadrature-axis voltage component into a proportional-integral controller, using the output of the proportional-integral controller as a correction amount for the local angular frequency; adding the correction amount to the rated angular frequency to obtain the local angular frequency; and dividing the local angular frequency by twice pi to obtain the local frequency value.

[0022] Furthermore, when the frequency deviation determined based on the frequency value exceeds the dead zone, and the rate of change of the voltage phase angle exceeds a preset rate of change threshold and is opposite in polarity to the local rate of change of frequency, it is determined that a transient coordinated adjustment state has been entered, including:

[0023] The frequency deviation is obtained by subtracting the local frequency value from the rated frequency.

[0024] Determine whether the absolute value of the frequency deviation exceeds the dead zone. If it does, record the start time when the frequency deviation exceeds the dead zone.

[0025] The absolute value of the rate of change of voltage phase angle difference is compared with a preset rate of change threshold.

[0026] Perform a sign bit XOR operation on the rate of change of voltage phase angle difference and the local rate of change of frequency. If the result is true, it is determined that the polarity of the rate of change of voltage phase angle difference and the local rate of change of frequency are opposite.

[0027] When the frequency deviation exceeds the dead zone and the rate of change of voltage phase angle exceeds the preset rate of change threshold and is opposite in polarity to the local rate of change of frequency, it is determined that the transient coordinated adjustment state has been entered.

[0028] Furthermore, the absolute value of the voltage phase angle difference change rate is compared with a preset change rate threshold, which is obtained by reading the line reactance between the current node and the adjacent node, as well as the active power of the local load of the adjacent node; dividing the active power of the local load of the adjacent node by the line reactance between the current node and the adjacent node to obtain the sensitivity coefficient of the power angle of the adjacent node to frequency change; and using the product of the sensitivity coefficient and the allowable range of frequency deviation as the preset change rate threshold.

[0029] Furthermore, the marginal contribution estimate is used as a correction term for the advantage function and fused with the value gradient output by the global Critic network during centralized training to obtain a credit allocation signal specific to the agent, including:

[0030] The value gradient is obtained from the global Critic network trained in a centralized manner, with local frequency values ​​and frequency change rate as input conditions for the output.

[0031] The marginal contribution estimate is multiplied by the attenuation factor to obtain the dominance function correction term. The attenuation factor is determined based on the ratio of the line reactance between the current node and adjacent nodes to the average line reactance of the entire network.

[0032] The credit allocation signal for this agent is obtained by adding the advantage function correction term to the value gradient.

[0033] Furthermore, the attenuation factor is determined based on the ratio of the line reactance between the current node and its adjacent nodes to the average line reactance of the entire network. This includes: obtaining the line reactance between each pair of adjacent nodes in the entire network; summing the line reactance between each pair of adjacent nodes in the entire network and dividing the sum by the total number of pairs of adjacent nodes in the entire network to obtain the average line reactance of the entire network; dividing the line reactance between the current node and its adjacent nodes by the average line reactance of the entire network to obtain the reactance ratio; and using the reactance ratio as the input of an exponential function, with the output of the exponential function serving as the attenuation factor.

[0034] Furthermore, each agent updates its local Actor network's policy gradient based on the credit allocation signal to generate the energy scheduling action for the next time step, including:

[0035] Obtain the credit allocation signal; replace the value gradient of the global Critic network output in centralized training with the credit allocation signal, and calculate the policy gradient of the local Actor network;

[0036] The policy gradient is used to update the parameters along the backpropagation path of the local Actor network parameters to obtain the updated local Actor network; the local frequency value and voltage phase angle difference change rate are input into the updated local Actor network to output the energy scheduling action for the next time step.

[0037] On the other hand, the present invention provides a microgrid energy dispatching system based on multi-agent reinforcement learning, comprising the following modules:

[0038] The electrical quantity acquisition module is used by each intelligent agent to acquire local frequency values, active power and frequency change rate, as well as the voltage phase angle difference change rate with adjacent nodes;

[0039] The state determination module is used to determine that the transient coordinated adjustment state is entered when the frequency deviation determined based on the frequency value exceeds the dead zone and the voltage phase angle difference change rate exceeds the preset change rate threshold and is opposite to the polarity of the local frequency change rate.

[0040] The contribution estimation module is used to calculate the transient active power increment injected by this node to the adjacent node before and after the moment of opposite polarity under transient coordinated adjustment state. The ratio of the transient active power increment to the line reactance between this node and the adjacent node is used as the marginal contribution estimate of this agent's action to frequency recovery.

[0041] The allocation and fusion module is used to use the marginal contribution estimate as the advantage function correction term and fuse it with the value gradient output by the global Critic network in centralized training to obtain the credit allocation signal for this agent.

[0042] The action generation module is used by each agent to update the policy gradient of the local Actor network based on the credit allocation signal, so as to generate the energy scheduling action for the next moment.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] 1. To address the coordination dilemma caused by the unified scalar characteristics of frequency in isolated microgrids, this paper identifies the polarity conflict between the rate of change of voltage phase angle difference and the local rate of change of frequency, thus separating the transient coordination mismatch from normal operation. This directly captures the moment when the local node's adjustment action fails to improve the frequency response of adjacent nodes or even produces a reverse effect, providing a clear physical entry point for subsequent calculation of marginal contributions. By using the ratio of the transient active power increment before and after the moment of opposite polarity to the line reactance between the local node and adjacent nodes, the independent action component of the agent's action on frequency recovery is estimated from the dimension of active power change. This allows attributable individual adjustment effects to be extracted from the globally coupled frequency response, solving the problem of the credit allocation mechanism's contribution signal being annihilated in a global scalar environment.

[0045] 2. The marginal contribution estimate is used as a correction term for the advantage function. After being weighted by an attenuation factor determined by the ratio of line reactance to the average line reactance of the entire network, it is superimposed with the value gradient of the global Critic network in centralized training to form a corrected credit allocation signal. This not only retains the optimization direction of reinforcement learning for long-term returns, but also specifically makes up for the inherent defects of pure data-driven credit allocation in strongly physically coupled scenarios. This allows each agent to obtain more distinctive individualized learning signals. The local Actor network updates its policies and generates energy scheduling actions accordingly. Each agent can learn differentiated cooperative strategies in distributed execution, avoiding policy convergence and power overshoot and oscillation in frequency regulation, and improving the stability of the frequency recovery process of the islanded microgrid. Attached Figure Description

[0046] Figure 1 This is a flowchart of a microgrid energy scheduling method based on multi-agent reinforcement learning according to the present invention;

[0047] Figure 2 This is a schematic diagram of the structure of a microgrid energy dispatching system based on multi-agent reinforcement learning according to the present invention. Detailed Implementation

[0048] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0049] Example 1: Figure 1 This invention presents a microgrid energy scheduling method based on multi-agent reinforcement learning, which includes the following steps:

[0050] S1: Each agent acquires its local frequency value, active power, and frequency change rate, as well as the voltage phase angle difference change rate with neighboring nodes;

[0051] S2: When the frequency deviation determined based on the frequency value exceeds the dead zone, and the voltage phase angle difference change rate exceeds the preset change rate threshold and is opposite in polarity to the local frequency change rate, it is determined that the transient coordinated adjustment state is entered.

[0052] S3: Under transient coordinated regulation, the transient active power increment injected by this node into the adjacent node before and after the moment of opposite polarity is calculated based on the active power. The ratio of the transient active power increment to the line reactance between this node and the adjacent node is used as the marginal contribution estimate of this agent's action to frequency recovery.

[0053] S4: The marginal contribution estimate is used as the advantage function correction term and fused with the value gradient output by the global Critic network in centralized training to obtain the credit allocation signal for this agent.

[0054] S5: Each agent updates the policy gradient of its local Actor network based on the credit allocation signal to generate the energy scheduling action for the next time step.

[0055] In the specific implementation of S1, each agent collects the local three-phase voltage through a voltage transformer installed at its node. The local three-phase voltage is the instantaneous three-phase voltage value of the AC bus at this node in the microgrid. The voltage transformer samples the local three-phase voltage according to a fixed sampling period, for example, set to 1 / 20 of the microgrid's power frequency cycle. The collected local three-phase voltage is input into a phase-locked loop (PLL) for calculation. The purpose of the PLL calculation is to extract the local frequency value and the local voltage phase angle from the local three-phase voltage. The specific process of the PLL calculation is as follows: the local three-phase voltage is subjected to a Clarke transform to obtain the alpha-axis voltage component and the beta-axis voltage component in a two-phase stationary coordinate system. The Clarke transform is an equal-amplitude transformation from a three-phase stationary coordinate system to a two-phase stationary coordinate system. The alpha-axis voltage component is in phase with the A-phase voltage of the local three-phase voltage, and the beta-axis voltage component leads the alpha-axis voltage component by 90 degrees. The alpha-axis and beta-axis voltage components are subjected to a Park transform to obtain the direct-axis and quadrature-axis voltage components in a two-phase rotating coordinate system. The rotation angle used in the Park transform is an estimate of the local voltage phase angle fed back from the phase-locked loop (PLL). The direct-axis voltage component reflects the voltage amplitude, while the quadrature-axis voltage component reflects the voltage phase tracking error. The quadrature-axis voltage component is input to a proportional-integral (PI) controller. The proportional and integral coefficients of the PI controller are tuned according to the bandwidth requirements of the PLL. The output of the PI controller serves as a correction for the local angular frequency. This correction is added to the rated angular frequency (the angular frequency corresponding to the microgrid's nominal frequency). Dividing the local angular frequency by twice pi yields the local frequency value. This local frequency value is the current real-time frequency of this node, measured in Hertz (Hz). Simultaneously, the local voltage phase angle is obtained by integrating the rotation angle used in the Park transform. This local voltage phase angle is the real-time phase angle of the voltage at this node, measured in radians.

[0056] Each agent differentiates its local frequency value to obtain the frequency change rate. The differentiation operation is implemented using first-order backward difference. The frequency change rate is calculated as: Δf = (fk - f{k-1}) / Ts; where Δf represents the frequency change rate, fk represents the local frequency value at the current sampling time, f{k-1} represents the local frequency value at the previous sampling time, and Ts represents the sampling period. The frequency change rate reflects the speed and direction of the local frequency value's change over time, and its unit is Hertz per second. If the current sampling time is the first time the local frequency value is acquired, the frequency change rate is initialized to zero.

[0057] Each agent collects local three-phase current through current transformers installed at its node. The local three-phase current is the instantaneous three-phase current value of the AC bus at that node in the microgrid. The current transformer samples the local three-phase current according to the same sampling period as the voltage transformer. The collected local three-phase current is subjected to a Clarke transform to obtain the alpha-axis current component and beta-axis current component in a two-phase stationary coordinate system. Combined with the previously obtained alpha-axis voltage component and beta-axis voltage component, the active power is calculated. The active power is calculated as: P = (3 / 2) × (vα × iα + vβ × iβ); where P represents active power, vα represents the alpha-axis voltage component, vβ represents the beta-axis voltage component, iα represents the alpha-axis current component, and iβ represents the beta-axis current component. The coefficient multiplied by 3 / 2 is the inherent conversion factor for the three-phase power calculation under the equal-amplitude form of the Clarke transform, used to compensate for the amplitude scaling caused by the equal-amplitude transformation. The calculated active power is the current instantaneous active power of this node, in watts. A positive value indicates that this node absorbs active power from the microgrid, while a negative value indicates that this node injects active power into the microgrid.

[0058] Each agent acquires the rate of change of voltage phase angle difference with neighboring nodes. Neighboring nodes are adjacent nodes in the microgrid that have a direct electrical connection to the current agent. Each agent knows the node number of its neighboring nodes through a communication network or locally stored topology information. For each neighboring node, each agent acquires the three-phase voltage of the neighboring node from its voltage transformer via a communication interface connected to that node. The acquisition method for the three-phase voltage of the neighboring node is the same as that for the local three-phase voltage. A phase-locked loop (PLL) operation is performed on the three-phase voltage of the neighboring node to obtain its voltage phase angle. The PLL operation process is completely consistent with the PLL operation on the local three-phase voltage. The voltage phase angle difference is obtained by subtracting the voltage phase angle of the neighboring node from the local voltage phase angle. The voltage phase angle difference is calculated as: Δθ = θnei - θloc; where Δθ represents the voltage phase angle difference, θnei represents the voltage phase angle of the neighboring node, and θloc represents the local voltage phase angle. The unit of the voltage phase angle difference is radians. A positive value indicates that the voltage phase of the neighboring node leads the local voltage phase, and a negative value indicates that the voltage phase of the neighboring node lags behind the local voltage phase. The voltage phase angle difference is differentiated to obtain the rate of change of the voltage phase angle difference. This differentiation is implemented using a first-order backward difference. The rate of change of the voltage phase angle difference is calculated as: Δfθ = (Δθk - Δθ{k-1}) / Ts; where Δfθ represents the rate of change of the voltage phase angle difference, Δθk represents the voltage phase angle difference at the current sampling moment, Δθ{k-1} represents the voltage phase angle difference at the previous sampling moment, and Ts represents the sampling period. The rate of change of the voltage phase angle difference reflects the speed and direction of the change in the voltage phase difference between the current node and its neighboring nodes, expressed in radians per second. If the current sampling moment is the first time the voltage phase angle difference is acquired, the rate of change of the voltage phase angle difference is initialized to zero. If the current node has multiple neighboring nodes, the above operation is performed on each neighboring node, and each agent ultimately obtains the rate of change of the voltage phase angle difference corresponding to each neighboring node.

[0059] The entire acquisition process described above is based on standard measurement equipment already installed in the microgrid. Voltage transformers and current transformers are standard measurement devices configured at each node in the microgrid. Phase-locked loop (PLL) arithmetic is a mature and widely used method for extracting frequency and phase angles. Clarke transform and Park transform are standard coordinate transformation tools in power system analysis. Differential arithmetic is implemented through difference approximation, which is computationally simple and has good real-time performance. The entire acquisition process does not rely on any external predictive data or precise system models; it can be completed using only real-time electrical measurements from local and adjacent nodes. This provides a complete and reliable data foundation for subsequent steps to determine the transient coordinated regulation state and calculate marginal contribution estimates.

[0060] In the specific implementation of S2, each agent has already acquired its local frequency value, frequency change rate, and voltage phase angle difference rate with adjacent nodes in step S1. The frequency deviation is obtained by subtracting the local frequency value from the rated frequency. The frequency deviation is calculated as: Δfdev = floc - fnom; where Δfdev represents the frequency deviation, floc represents the local frequency value, and fnom represents the rated frequency, which is the nominal operating frequency of the microgrid, such as 50 Hz or 60 Hz, in Hertz. The absolute value of the frequency deviation reflects the degree to which the local frequency value deviates from the rated frequency.

[0061] The system determines whether the absolute value of the frequency deviation exceeds the dead zone. The dead zone is a pre-defined allowable range for frequency deviation, determined based on the tolerance of the devices connected to the microgrid to frequency deviation. For example, the dead zone is set to 0.05% to 0.1% of the rated frequency. This dead zone setting ensures that if the frequency deviation is within the allowable small fluctuation range of normal microgrid operation, it will not trigger subsequent judgments, avoiding misjudgments due to measurement noise or normal load fluctuations. If the absolute value of the frequency deviation does not exceed the dead zone, it indicates that the local frequency value is within the normal operating range, and there is no need to trigger transient coordinated regulation. If the absolute value of the frequency deviation exceeds the dead zone, the starting time of the frequency deviation exceeding the dead zone is recorded. The starting time of the frequency deviation exceeding the dead zone is the sampling time corresponding to the instant when the absolute value of the frequency deviation changes from within the dead zone to outside the dead zone.

[0062] The absolute value of the voltage phase angle difference change rate is compared with a preset change rate threshold. The preset change rate threshold is obtained by reading the line reactance between the current node and adjacent nodes, as well as the active power of the local load at adjacent nodes. The line reactance between the current node and adjacent nodes is a microgrid topology parameter, obtainable from a grid topology database or a pre-stored network parameter table. The active power of the local load at adjacent nodes is the real-time active power value of the load connected to the adjacent node, measured by the agent at the adjacent node and transmitted to the current agent via the communication network. The preset change rate threshold is calculated as follows: the active power of the local load at adjacent nodes is divided by the line reactance between the current node and adjacent nodes to obtain the sensitivity coefficient of the adjacent node's power angle to frequency changes. The sensitivity coefficient reflects the sensitivity of the voltage phase angle relative to frequency changes caused by changes in the active load at adjacent nodes. The sensitivity coefficient is calculated as: Snei = Ploadnei / Xline; where Snei represents the sensitivity coefficient of the adjacent node's power angle to frequency changes, Ploadnei represents the active power of the local load at adjacent nodes, and Xline represents the line reactance between the current node and adjacent nodes. The product of the sensitivity coefficient and the allowable frequency deviation range is used as the preset rate of change threshold. The preset rate of change threshold is calculated as: ThΔfθ = Snei × Δfallow; where ThΔfθ represents the preset rate of change threshold, Snei represents the sensitivity coefficient of the power angle of adjacent nodes to frequency changes, and Δfallow represents the allowable frequency deviation range. The allowable frequency deviation range is the maximum frequency deviation change allowed for normal operation of the microgrid, measured in Hertz. For example, the allowable frequency deviation range is set to 0.1 Hertz to 0.5 Hertz. The setting of the preset rate of change threshold allows the judgment threshold of the voltage phase angle difference change rate to adaptively match the electrical distance between the node and adjacent nodes, as well as the load level of adjacent nodes. When the load of adjacent nodes is heavy or the line reactance is small, the sensitivity coefficient increases, and the preset rate of change threshold is increased accordingly, avoiding false triggering of transient coordinated adjustment state due to load fluctuations under normal operating conditions. When the load of adjacent nodes is light or the line reactance is large, the preset rate of change threshold is decreased accordingly, ensuring reliable detection of weak but sensitive coordinated mismatch phenomena between weakly connected nodes.

[0063] A sign-bit XOR operation is performed on the voltage phase angle difference rate of change and the local frequency change rate. This operation extracts the sign bit of the voltage phase angle difference rate of change and the sign bit of the frequency change rate, and then performs an XOR operation on these two sign bits. If the sign-bit XOR result is true, it indicates that the voltage phase angle difference rate of change and the local frequency change rate are opposite in polarity, one being positive and the other negative. The physical meaning of this opposite polarity is that the trend of local frequency change at this node conflicts directionally with the dynamic trend of voltage phase difference change between this node and neighboring nodes. This reflects that the adjustment action of this node injecting active power into neighboring nodes has failed to effectively improve the frequency response of neighboring nodes, and may even have a reverse effect. This indicates that the adjustment actions of each distributed power source are in a state of coordinated mismatch coupling, and the credit allocation mechanism cannot accurately distinguish the individual contributions of each agent under this condition.

[0064] When three conditions are simultaneously met—the absolute value of the frequency deviation exceeding the dead zone, the absolute value of the voltage phase angle difference change rate exceeding the preset change rate threshold, and the sign bit XOR operation result indicating that the voltage phase angle difference change rate has opposite polarity to the local frequency change rate—the system is determined to enter a transient coordinated regulation state. This determination logic accurately distinguishes the abnormally coupled operating condition—where the contributions of various agents' actions are mixed and difficult to be effectively decomposed by the credit allocation mechanism—from the normal operating condition and the transient process of a single disturbance during frequency regulation in an islanded microgrid. This provides a clear state identifier and triggering condition for subsequent steps to calculate the marginal contribution estimate using the transient active power increment at the moment of opposite polarity.

[0065] After determining the transition to the transient coordinated regulation state in step S2, each agent has identified the moment of opposite polarity. The moment of opposite polarity is the sampling moment when the rate of change of voltage phase angle and the local rate of change of frequency are detected simultaneously with opposite polarities and the absolute value of the frequency deviation exceeds the dead zone. Step S3 uses the change in active power before and after the moment of opposite polarity to calculate the estimated marginal contribution of the agent's action to frequency recovery. The specific implementation process is described below.

[0066] The active power in the measurement cycle preceding the moment of opposite polarity is taken as the first active power, and the active power in the measurement cycle following the moment of opposite polarity is taken as the second active power. The measurement cycle is the same as the sampling cycle of the voltage transformer and current transformer in step S1, and the sampling cycle is set to, for example, 1 / 20 of the microgrid power frequency cycle. Both the first and second active powers are active powers already calculated in step S1, and can be read from the locally stored active power sequence according to the time index. The first active power corresponds to the active power value at the most recent sampling moment before the moment of opposite polarity, and the second active power corresponds to the active power value at the most recent sampling moment after the moment of opposite polarity.

[0067] The transient active power increment is obtained by subtracting the first active power from the second active power. The transient active power increment reflects the instantaneous change in active power injected by the current node into adjacent nodes before and after the polarity reversal. The transient active power increment is calculated as: ΔPtrans = P2 - P1; where ΔPtrans represents the transient active power increment, P2 represents the second active power, and P1 represents the first active power. The unit of the transient active power increment is watts. A positive value indicates that the active power injected by the current node into adjacent nodes increases after the polarity reversal, while a negative value indicates that the injected active power decreases.

[0068] Read the line reactance between this node and its adjacent nodes. The line reactance between this node and its adjacent nodes has already been read when the preset rate of change threshold was obtained in step S2. Here, the same parameter value is used directly and there is no need to obtain it again.

[0069] The equivalent electromagnetic torque component is obtained by dividing the transient active power increment by the line reactance between the local node and the adjacent node. The equivalent electromagnetic torque component is calculated as: Teq = ΔPtrans / Xline; where Teq represents the equivalent electromagnetic torque component, ΔPtrans represents the transient active power increment, and Xline represents the line reactance between the local node and the adjacent node. The equivalent electromagnetic torque component has the same dimensions as torque, and its physical meaning is that the change in active power injected from the local node into the adjacent node before and after the moment of opposite polarity is equivalent to the electromagnetic torque component applied to the rotor of the adjacent node. The electromagnetic torque component directly affects the frequency response of the adjacent node. The physical basis of this calculation method is that the change in active power transmitted between nodes in a power system is proportional to the change in the voltage phase angle difference between nodes, and the proportionality coefficient is the reciprocal of the line reactance. The rate of change of the voltage phase angle difference is directly related to frequency dynamics. Therefore, the ratio of the transient active power increment to the line reactance between the local node and the adjacent node can capture the degree of independent influence of the agent's adjustment actions on the frequency recovery of the adjacent node from the perspective of active power change.

[0070] The equivalent electromagnetic torque component is used as the marginal contribution estimate of the agent's actions to frequency recovery. This marginal contribution estimate quantitatively characterizes the magnitude and direction of the independent effect of the agent's active power regulation actions near the moment of opposite polarity on the system's frequency recovery process. Positive values ​​indicate that the agent's regulation actions are beneficial to frequency recovery, while negative values ​​indicate that they are detrimental. This marginal contribution estimate differs from traditional credit allocation mechanisms that rely solely on time difference error or generalized advantage estimation. Instead, it is based on the physical laws of electromechanical transients in power systems, utilizing the instantaneous characteristic of cooperative mismatch at the moment of opposite polarity to separate the agent's independent contribution component from the globally coupled frequency response. This effectively solves the problem of the difficulty in decomposing the contributions of each agent's actions in a global frequency scalar environment. The marginal contribution estimate provides a clear and interpretable physical quantity input for correcting the credit allocation signal in subsequent steps.

[0071] After obtaining the marginal contribution estimate of the agent's action to frequency recovery in step S3, step S4 integrates the marginal contribution estimate into the credit allocation mechanism of multi-agent reinforcement learning to obtain the credit allocation signal for the agent. The specific implementation process is described below.

[0072] The value gradient is obtained from the globally Critic network trained in a centralized manner, using local frequency values ​​and rate of change of frequency as input conditions. The globally Critic network trained in a centralized manner is a neural network responsible for evaluating the global state value in multi-agent reinforcement learning. It takes the observation information of all agents as input and outputs an estimate of the global state value. The local frequency values ​​and rate of change of frequency, already obtained in step S1, are used as input conditions for the globally Critic network trained in a centralized manner. The network uses its internally trained network parameters to perform forward propagation calculations on the input conditions and outputs the value gradient. The value gradient is the partial derivative of the global state value with respect to the agent's action, reflecting the direction and magnitude of the marginal influence of the agent's action on the global state value from the perspective of the value evaluation of the globally Critic network trained in a centralized manner.

[0073] The attenuation factor is determined based on the ratio of the line reactance between the current node and its adjacent nodes to the average line reactance of the entire network. The determination process is as follows: First, obtain the line reactance between all pairs of adjacent nodes in the entire network. Each pair of adjacent nodes refers to the set of all directly electrically connected node pairs in the microgrid. The line reactance between each pair of adjacent nodes is a microgrid topology parameter, obtained from a grid topology database or a pre-stored network parameter table. Then, sum the line reactances between all pairs of adjacent nodes and divide by the total number of adjacent node pairs in the entire network to obtain the average line reactance of the entire network. The average line reactance of the entire network is calculated as: Xavg = ΣXlineij / Npair; where Xavg represents the average line reactance of the entire network, Xlineij represents the line reactance between the i-th node and its j-th adjacent node pair, and Npair represents the total number of adjacent node pairs in the entire network. The reactance ratio is obtained by dividing the line reactance between this node and its adjacent nodes by the average line reactance of the entire network. The reactance ratio is calculated as: RX = Xline / Xavg; where RX represents the reactance ratio, Xline represents the line reactance between this node and its adjacent nodes, and Xavg represents the average line reactance of the entire network. The reactance ratio is used as the input to an exponential function, and the output of the exponential function is used as the attenuation factor. The attenuation factor is calculated as: α = e (-RX) Where α represents the attenuation factor and RX represents the reactance ratio. The exponential function e (-RX)The value range of α is between 0 and 1. The larger the reactance ratio RX, the closer the attenuation factor α is to 0; the smaller the reactance ratio RX, the closer the attenuation factor α is to 1. The setting of the attenuation factor is such that: when the line reactance between the local node and its neighboring nodes is greater than the average line reactance of the entire network, the electrical distance between the local node and its neighboring nodes is relatively large, the physical coupling is weak, and the reliability of the marginal contribution estimate calculated in step S3 in the global value assessment is relatively low. The attenuation factor correspondingly reduces the weight of the marginal contribution estimate in the credit allocation signal. When the line reactance between the local node and its neighboring nodes is less than the average line reactance of the entire network, the electrical distance between the local node and its neighboring nodes is relatively small, the physical coupling is tight, and the reliability of the marginal contribution estimate reflecting the independent adjustment effect of the agent is relatively high. The attenuation factor correspondingly maintains the weight of the marginal contribution estimate close to the original level.

[0074] The marginal contribution estimate is multiplied by the attenuation factor to obtain the dominance function correction term, which is calculated as: Acorr = α × Teq; where Acorr represents the dominance function correction term, α represents the attenuation factor, and Teq represents the equivalent electromagnetic torque component calculated in step S3. The dominance function correction term is the marginal contribution value estimated based on the physical model and weighted by electrical distance confidence.

[0075] The credit allocation signal for this agent is obtained by adding the advantage function correction term to the value gradient. The credit allocation signal is calculated as follows: Where Csignal represents the credit allocation signal, Let represent the value gradient and Acorr represent the advantage function correction term. The credit allocation signal integrates the value gradient evaluated from a data-driven perspective by the global Critic network during centralized training and the advantage function correction term estimated from a physical model perspective. This ensures that the credit allocation signal obtained by the agent retains the optimization direction of reinforcement learning for global long-term rewards while specifically correcting the inherent defect of the credit allocation mechanism in the environment of global frequency scalar coupling in isolated microgrids, which cannot accurately evaluate the actual regulation effect of a single agent. The credit allocation signal provides an accurate and reliable individualized learning signal for the policy gradient update of the local Actor network in subsequent steps.

[0076] After obtaining the credit allocation signal for this agent in step S4, step S5 uses the credit allocation signal to guide the policy update of the local Actor network and generate the energy scheduling action for the next moment. The specific implementation process is described as follows.

[0077] The credit assignment signal is obtained. The credit assignment signal has been calculated in step S4. The credit assignment signal combines the value gradient of the global Critic network in centralized training and the advantage function correction term based on the physical model estimation.

[0078] The policy gradient of the local Actor network is calculated by replacing the value gradient output by the global Critic network in centralized training with the credit allocation signal. The local Actor network is a neural network in multi-agent reinforcement learning responsible for generating energy scheduling actions. Each agent maintains its own local Actor network, which takes local observations as input and outputs active power scheduling commands for each generation and storage unit in the microgrid. In conventional multi-agent reinforcement learning frameworks, the calculation of the policy gradient depends on the value gradient output by the global Critic network in centralized training. The value gradient serves as a directional guide for policy optimization, updating the parameters of the local Actor network. In this step, the credit allocation signal obtained in step S4 is used to replace the value gradient. The credit allocation signal not only includes the global long-term reward gradient information evaluated from a data-driven perspective by the global Critic network in centralized training, but also incorporates the advantage function correction term based on the estimation of power system physical characteristics in step S3. This advantage function correction term provides targeted correction for the credit allocation failure problem caused by the globally unified frequency scalar in the islanded microgrid. The policy gradient is obtained by taking the partial derivative of the product of the credit assignment signal and the log probability of the output action of the local Actor network with respect to the parameters of the local Actor network. The calculation of the policy gradient is as follows: ;in, Let represent the policy gradient, Csignal represent the credit assignment signal, πθ(a|s) represent the probability that the local Actor network will output action a given observation s, and θ represent the parameters of the local Actor network. This represents the partial derivative of the local Actor network parameters. This partial derivative calculation is automatically performed using the backpropagation algorithm of the local Actor network. The backpropagation algorithm propagates the credit assignment signal and the gradient of the logarithmic probability of the local Actor network output layer to the network input layer by layer, calculating the partial derivative of the network parameters at each layer.

[0079] The policy gradient is used to update the parameters along the backpropagation path of the local Actor network parameters, resulting in the updated local Actor network. The parameter update uses stochastic gradient descent, and the parameter update calculation for the local Actor network is as follows: Where θnew represents the updated local Actor network parameters, θold represents the original local Actor network parameters, and lr represents the learning rate. This represents the policy gradient. The learning rate is a hyperparameter that controls the step size of parameter updates. The learning rate can be set, for example, between 0.0001 and 0.001. The specific value of the learning rate is adjusted according to the number of layers and neurons in the local Actor network and the stability of training convergence. When the number of layers is deep or the number of neurons is large, a smaller learning rate is selected to ensure training stability. When oscillations occur during training convergence, the learning rate is appropriately reduced.

[0080] The local frequency value and voltage phase angle difference change rate are input into the updated local Actor network, which outputs the energy dispatch action for the next time step. The local frequency value and voltage phase angle difference change rate have already been obtained in step S1. The updated local Actor network receives the local frequency value and voltage phase angle difference change rate as input observations. After forward propagation calculation by the updated local Actor network, it outputs the energy dispatch action for the next time step. The energy dispatch action is the active power dispatch command for each generation unit and energy storage unit in the microgrid, such as the active power output reference value of the photovoltaic inverter or the charging and discharging power command of the energy storage converter. Each agent independently executes steps S1 to S5, autonomously generating its own node's energy dispatch action based on local observations and information interacting with neighboring nodes, thus realizing distributed real-time energy dispatch of the microgrid.

[0081] Example 2: Figure 2 A schematic diagram of a microgrid energy dispatching system based on multi-agent reinforcement learning is provided. This system includes the following modules:

[0082] The electrical quantity acquisition module is used by each intelligent agent to acquire local frequency values, active power and frequency change rate, as well as the voltage phase angle difference change rate with adjacent nodes;

[0083] The state determination module is used to determine that the transient coordinated adjustment state is entered when the frequency deviation determined based on the frequency value exceeds the dead zone and the voltage phase angle difference change rate exceeds the preset change rate threshold and is opposite to the polarity of the local frequency change rate.

[0084] The contribution estimation module is used to calculate the transient active power increment injected by this node to the adjacent node before and after the moment of opposite polarity under transient coordinated adjustment state. The ratio of the transient active power increment to the line reactance between this node and the adjacent node is used as the marginal contribution estimate of this agent's action to frequency recovery.

[0085] The allocation and fusion module is used to use the marginal contribution estimate as the advantage function correction term and fuse it with the value gradient output by the global Critic network in centralized training to obtain the credit allocation signal for this agent.

[0086] The action generation module is used by each agent to update the policy gradient of the local Actor network based on the credit allocation signal, so as to generate the energy scheduling action for the next moment.

[0087] The calculations involved in the embodiments are all dimensionless numerical calculations, and the preset parameters and thresholds in the calculations are set by those skilled in the art according to the actual situation.

[0088] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.

[0089] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and inventive constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

Claims

1. A microgrid energy dispatching method based on multi-agent reinforcement learning, characterized in that, Includes the following steps: S1: Each agent acquires its local frequency value, active power, and frequency change rate, as well as the voltage phase angle difference change rate with neighboring nodes; S2: When the frequency deviation determined based on the frequency value exceeds the dead zone, and the voltage phase angle difference change rate exceeds the preset change rate threshold and is opposite in polarity to the local frequency change rate, it is determined that the transient coordinated adjustment state is entered. S3: Under transient coordinated regulation, the transient active power increment injected by this node into adjacent nodes before and after the moment of opposite polarity is calculated based on the active power. The ratio of the transient active power increment to the line reactance between this node and adjacent nodes is used as the estimated marginal contribution of this agent's action to frequency recovery, including: The active power of the measurement cycle before the moment of opposite polarity is taken as the first active power, and the active power of the measurement cycle after the moment of opposite polarity is taken as the second active power. Subtracting the first active power from the second active power yields the transient active power increment; Read the line reactance between this node and adjacent nodes; Dividing the transient active power increment by the line reactance yields the equivalent electromagnetic torque component, which is then used as an estimate of the marginal contribution of the agent's actions to frequency recovery. S4: The marginal contribution estimate is used as the correction term of the advantage function and fused with the value gradient output by the global Critic network in the centralized training to obtain the credit allocation signal for this agent. S5: Each agent updates the policy gradient of its local Actor network based on the credit allocation signal to generate the energy scheduling action for the next time step.

2. The microgrid energy dispatching method based on multi-agent reinforcement learning according to claim 1, characterized in that, Each agent acquires its local frequency value, active power, and frequency change rate, as well as the voltage phase angle difference rate with neighboring nodes, including: The local three-phase voltage is collected by a voltage transformer, and the local frequency value is obtained by performing a phase-locked loop operation on the local three-phase voltage. The frequency change rate is obtained by differentiating the local frequency value; the active power is calculated by collecting the local three-phase current through a current transformer and combining it with the local three-phase voltage. The three-phase voltages of adjacent nodes are collected by voltage transformers. The phase angles of the adjacent nodes are obtained by performing phase-locked loop operations on the three-phase voltages of the adjacent nodes. The voltage phase angle difference is obtained by subtracting the voltage phase angle of the adjacent nodes from the local voltage phase angle. The voltage phase angle difference is obtained by performing differential operations on the voltage phase angle difference to obtain the rate of change of the voltage phase angle difference.

3. The microgrid energy dispatching method based on multi-agent reinforcement learning according to claim 2, characterized in that, The local frequency value is obtained by performing a phase-locked loop operation on the local three-phase voltage, including: performing a Clarke transform on the local three-phase voltage to obtain the alpha-axis voltage component and beta-axis voltage component in a two-phase stationary coordinate system; performing a Park transform on the alpha-axis voltage component and beta-axis voltage component to obtain the direct-axis voltage component and quadrature-axis voltage component in a two-phase rotating coordinate system; inputting the quadrature-axis voltage component into a proportional-integral controller, the output of the proportional-integral controller as the correction amount for the local angular frequency, adding the correction amount to the rated angular frequency to obtain the local angular frequency, and dividing the local angular frequency by twice pi to obtain the local frequency value.

4. The microgrid energy dispatching method based on multi-agent reinforcement learning according to claim 1, characterized in that, When the frequency deviation determined based on the frequency value exceeds the dead zone, and the rate of change of the voltage phase angle exceeds a preset rate of change threshold and is opposite in polarity to the local frequency rate of change, it is determined that a transient coordinated adjustment state has been entered, including: The frequency deviation is obtained by subtracting the local frequency value from the rated frequency; it is then determined whether the absolute value of the frequency deviation exceeds the dead zone. If it does, the starting time when the frequency deviation exceeds the dead zone is recorded. The absolute value of the rate of change of voltage phase angle difference is compared with a preset rate of change threshold. Perform a sign bit XOR operation on the rate of change of voltage phase angle difference and the local rate of change of frequency. If the result is true, it is determined that the polarity of the rate of change of voltage phase angle difference and the local rate of change of frequency are opposite. When the frequency deviation exceeds the dead zone and the rate of change of voltage phase angle exceeds the preset rate of change threshold and is opposite in polarity to the local rate of change of frequency, it is determined that the transient coordinated adjustment state has been entered.

5. A microgrid energy dispatching method based on multi-agent reinforcement learning according to claim 4, characterized in that, The absolute value of the voltage phase angle difference change rate is compared with a preset change rate threshold, which is obtained by reading the line reactance between the current node and the adjacent node and the active power of the local load of the adjacent node; dividing the active power of the local load of the adjacent node by the line reactance between the current node and the adjacent node to obtain the sensitivity coefficient of the power angle of the adjacent node to frequency change; and multiplying the sensitivity coefficient by the allowable range of frequency deviation as the preset change rate threshold.

6. The microgrid energy dispatching method based on multi-agent reinforcement learning according to claim 1, characterized in that, The marginal contribution estimate is used as a correction term for the advantage function and fused with the value gradient output by the global Critic network during centralized training to obtain a credit allocation signal specific to this agent, including: The value gradient is obtained from the global Critic network trained in a centralized manner, with local frequency values ​​and frequency change rate as input conditions for the output. The marginal contribution estimate is multiplied by the attenuation factor to obtain the dominance function correction term. The attenuation factor is determined based on the ratio of the line reactance between the current node and adjacent nodes to the average line reactance of the entire network. The credit allocation signal for this agent is obtained by adding the advantage function correction term to the value gradient.

7. A microgrid energy dispatching method based on multi-agent reinforcement learning according to claim 6, characterized in that, The attenuation factor is determined based on the ratio of the line reactance between the current node and its adjacent nodes to the average line reactance of the entire network. This includes: obtaining the line reactance between each pair of adjacent nodes in the entire network; summing the line reactance between each pair of adjacent nodes in the entire network and dividing the sum by the total number of pairs of adjacent nodes in the entire network to obtain the average line reactance of the entire network; dividing the line reactance between the current node and its adjacent nodes by the average line reactance of the entire network to obtain the reactance ratio; and using the reactance ratio as the input of an exponential function, with the output of the exponential function serving as the attenuation factor.

8. The microgrid energy dispatching method based on multi-agent reinforcement learning according to claim 1, characterized in that, Each agent updates its local Actor network's policy gradient based on the credit allocation signal to generate the energy scheduling action for the next time step, including: Obtain the credit allocation signal; replace the value gradient of the global Critic network output in centralized training with the credit allocation signal, and calculate the policy gradient of the local Actor network; The policy gradient is used to update the parameters along the backpropagation path of the local Actor network parameters to obtain the updated local Actor network; the local frequency value and voltage phase angle difference change rate are input into the updated local Actor network to output the energy scheduling action for the next time step.

9. A microgrid energy dispatching system based on multi-agent reinforcement learning, used to implement the microgrid energy dispatching method based on multi-agent reinforcement learning as described in any one of claims 1-8, characterized in that, Includes the following modules: The electrical quantity acquisition module is used by each intelligent agent to acquire local frequency values, active power and frequency change rate, as well as the voltage phase angle difference change rate with adjacent nodes; The state determination module is used to determine that the transient coordinated adjustment state is entered when the frequency deviation determined based on the frequency value exceeds the dead zone and the voltage phase angle difference change rate exceeds the preset change rate threshold and is opposite to the polarity of the local frequency change rate. The contribution estimation module is used to calculate the transient active power increment injected by this node to the adjacent node before and after the moment of opposite polarity under transient coordinated adjustment state. The ratio of the transient active power increment to the line reactance between this node and the adjacent node is used as the marginal contribution estimate of this agent's action to frequency recovery. The allocation and fusion module is used to use the marginal contribution estimate as the advantage function correction term and fuse it with the value gradient output by the global Critic network in centralized training to obtain the credit allocation signal for this agent. The action generation module is used by each agent to update the policy gradient of the local Actor network based on the credit allocation signal, so as to generate the energy scheduling action for the next moment.