Price-dual coupling based multi-microgrid grid-connected neutral scheduling method, device and medium

By adopting a price-dual coupled multi-microgrid grid-connected neutral scheduling method, the problem of difficulty in simultaneously addressing grid connection deviations on both time scales in multi-microgrid cluster collaborative scheduling is solved. This method achieves rapid suppression of power spike impacts and elimination of persistent deviations, and price signals can be used for correction, thereby improving the economy and operational stability of multi-microgrids.

CN122225529APending Publication Date: 2026-06-16LIAONING UNIVERSITY OF PETROLEUM AND CHEMICAL TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LIAONING UNIVERSITY OF PETROLEUM AND CHEMICAL TECHNOLOGY
Filing Date
2026-03-05
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

In the collaborative scheduling of multi-microgrid clusters, it is difficult to balance the grid connection deviation across two time scales; the decoupling of price and constraints leads to insufficient correction and price fluctuations; and under strong constraints, multi-agent training is unstable and prone to outputting inactive actions.

Method used

A price-dual coupling-based multi-microgrid grid-connected neutral scheduling method is adopted. By collecting operation information, the net active power exchange and sliding window grid-connected neutral index are calculated. An asymmetric grid-connected neutral band is set, an instantaneous out-of-bounds suppression signal and sliding window scale cumulative deviation are constructed, dual variables are introduced for iterative updates, and an executable internal incentive price is generated. Combined with P2P transactions and power flow safety verification, online scheduling of multiple agents is realized.

Benefits of technology

It achieves rapid suppression of power spikes, continuous elimination of deviations, and price signal correction, thereby improving the economic efficiency and operational stability of multi-microgrids, reducing PCC net exchange fluctuations, and ensuring the execution and learning stability of strong physical constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122225529A_ABST
    Figure CN122225529A_ABST
Patent Text Reader

Abstract

The present application relates to the field of power distribution network grid-connected operation control and multi-agent reinforcement learning, and discloses a kind of multi-microgrid grid-connected neutral scheduling method based on price-dual coupling;The method collects multi-microgrid operation information and calculates core grid-connected index, sets asymmetric neutral zone and double-time scale deviation signal, iteratively updates dual variable;Combined with benchmark price, deviation and dual variable, executable internal incentive price is generated, forming a closed-loop regulation;Complete P2P transaction clearing and flow safety check and feedback reward;Through safe projection and improved multi-agent algorithm joint training scheduling strategy;The present application realizes the fast inhibition of grid-connected deviation + slow elimination precise regulation and control, and the price is upgraded from exogenous incentive to executable control quantity, and the deviation is corrected and the engineering acceptability is considered;P2P transaction and physical safety are deeply coupled, which improves the economy and safety;Algorithm training is stable under strong constraint, which improves renewable energy consumption rate and reduces grid regulation pressure and equipment loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of active distribution network operation control, microgrid energy management and multi-agent reinforcement learning technology. Specifically, it relates to a multi-microgrid grid-connected neutral scheduling method based on price-dual coupling, electronic equipment and computer-readable storage medium. In particular, it relates to a multi-microgrid collaborative scheduling technology that integrates sliding window grid-connected neutral deviation control, closed-loop coupling of dual variables and internal incentive prices, point-to-point (P2P) transaction settlement, power flow safety verification and safety action projection. Background Technology

[0002] With the large-scale grid connection of distributed energy sources such as distributed photovoltaic, distributed wind power, and electrochemical energy storage to the distribution network, the net power exchange between the distribution network's point of common coupling (PCC) and the main grid exhibits significant randomness and volatility. On the one hand, renewable energy output is affected by factors such as irradiance, temperature, and cloud cover, exhibiting strong short-term fluctuations. On the other hand, the uncertainty of residential and industrial electricity consumption behavior and the dynamic nature of flexible load dispatching lead to a continuous bias in power exchange on an hourly scale. This dual-timescale power fluctuation problem of "short-term peak exceeding limits + long-term bias accumulation" can easily trigger problems such as grid connection point power exceeding limits and reverse power flow. It can also trigger transformer and line thermal constraints, causing frequent operation of reactive power and voltage regulation devices, further resulting in voltage exceeding limits at distribution network nodes, increased line losses, and reduced service life of power equipment, while significantly increasing the power balance control pressure on the upstream grid. To mitigate the aforementioned risks, engineering practices often employ measures such as limiting renewable energy generation, expanding grid reserve capacity, tightening grid-connected power bandwidth, or increasing grid connection fees. However, these methods all come at the cost of sacrificing the benefits of renewable energy consumption and the economic viability of microgrids, making it difficult to achieve a balance between the safety and economy of distribution network operation.

[0003] For the problem of collaborative scheduling of multi-microgrid clusters, existing technologies mainly fall into two categories, both of which have significant technical shortcomings: The first category is based on methods such as model predictive control, robust / stochastic optimization, and distributed coordination (e.g., alternating direction multiplier method, augmented Lagrange method). These methods treat grid connection point power exchange and electrical constraints as explicit constraints for solution, but they have extremely high requirements for the accuracy of equipment models and prediction models and parameter calibration, and require online iterative solutions, resulting in significant computational and communication overhead. Prediction errors or model biases can significantly reduce the feasibility and performance of the scheduling scheme. The second category is based on collaborative control methods using multi-agent reinforcement learning. Although these methods have the advantage of adaptive learning in uncertain operating environments, in complex application scenarios with "strong physical constraints + P2P transaction settlement coupling + power flow security verification feedback", simply increasing the penalty coefficient to suppress out-of-bounds behavior can easily lead to problems such as training oscillations, infeasible strategy outputs, severe price signal jitter, and instability across random seeds. Meanwhile, within the framework of a unified price signal driven by multiple microgrids, without engineering-executable price update constraints (such as amplitude limiting, ramp limiting, smoothing, and projection) and a closed-loop correction mechanism coupled with grid-connection neutral deviation, price can only serve as an "exogenous incentive" and is unlikely to become an "executable control quantity." This not only makes it difficult to effectively fulfill the responsibility of correcting grid-connection deviations but may even amplify the collective oscillation problem caused by concerted group actions. Therefore, it is necessary to propose a unified collaborative scheduling scheme that can simultaneously consider "fast suppression, slow elimination, price executableness and acceptability, strong constraint feasibility, and training stability." Summary of the Invention

[0004] This invention aims to solve the following problems in multi-microgrid collaborative scheduling under grid-neutral constraints: difficulty in simultaneously addressing grid deviation across two time scales, insufficient correction and price fluctuations due to decoupling between price and constraints, and unstable multi-agent training with a strong constraint action space that easily outputs inactive actions.

[0005] To address the aforementioned technical problems, this invention proposes a neutral scheduling method for multi-microgrid interconnection based on price-dual coupling, comprising the following steps:

[0006] S1: Collect the operation information of the multi-microgrid cluster, and calculate the net active power exchange power and sliding window grid neutrality index of the cluster at the common coupling point.

[0007] S2: Set an asymmetric grid-connected neutral band, construct the instantaneous out-of-bounds suppression signal and the cumulative deviation of the sliding window scale, introduce a dual variable characterizing the tension of the grid-connected neutral constraint and update iteratively;

[0008] S3: Generate a candidate internal incentive price vector based on the benchmark electricity price, calculate the price correction increment by combining the sliding window grid-connected neutral index and dual variables and inject it into the candidate internal incentive price vector, apply an executable update operator to it, obtain the executable internal incentive price and publish it to the outside world;

[0009] S4: Each microgrid agent outputs local physical actions and P2P transaction intentions based on the published executable internal incentive price. After completing the centralized clearing of P2P transactions and the capacity verification of tie lines, the power flow safety verification of the distribution network is carried out, and the verification results are fed back to the multi-agent training reward.

[0010] S5: Under the centralized training-decentralized execution framework, the original actions of each micronet agent are safely projected. The multi-agent soft actor-critic algorithm is used to jointly train the coordinator and the scheduling strategy of each micronet to achieve neutral online scheduling of multiple micronets.

[0011] Preferably, the net active power exchange in step S1 is the sum of the power exchange between all microgrids in the multi-microgrid cluster and the main power grid; the sliding window grid neutrality index is the average value of the net active power exchange at the common coupling point in each discrete time period within a preset length sliding window, used to characterize the grid bias accumulation state of the multi-microgrid cluster in the continuous scheduling period corresponding to the sliding window.

[0012] Preferably, the instantaneous over-limit suppression signal in step S2 adopts the form of "dead zone + normalized double penalty + saturation". The penalty is triggered only when the absolute value of the net active power exchange at the common coupling point exceeds the instantaneous dead zone threshold. The penalty value increases twice with the excess amount and is limited by saturation to suppress power spike impact. The cumulative deviation of the sliding window scale is constructed based on the superband part of the grid neutral index of the sliding window that exceeds the grid neutral band, and is used to identify continuous grid deviation.

[0013] The dual variable is iteratively updated using the rule of "leakage pullback + deviation integral + interval projection". When the grid-connected neutral index of the sliding window exceeds the grid-connected neutral band for multiple consecutive scheduling cycles, the dual variable increases through the cumulative deviation integral to strengthen the correction intensity. When the index returns to the grid-connected neutral band, the dual variable gradually decreases through the leakage pullback term to avoid signal sticking. At the same time, the iterative update process is constrained by single-step amplitude limiting and interval projection to limit the unbounded growth of the dual variable.

[0014] Preferably, the candidate internal incentive price vector in step S3 is in multi-component form, defined by a three-component structure of "valley-peak-carbon term", and includes at least three types of executable settlement signals: valley electricity price, peak electricity price, and carbon or risk coefficient. The executable update operator includes at least an inertial smoothing operator, a single-step saturation operator, and an interval projection operator. The candidate internal incentive price vector with correction attribute is subjected to inertial smoothing, single-step ramp-up limiting, and interval projection operations in sequence through each operator to finally generate an executable internal incentive price that is released to the public.

[0015] Preferably, the centralized clearing of P2P transactions in step S4 follows a process of first determining the total clearable amount, then completing the bilateral power allocation, and finally performing tie-line capacity projection. The total clearable amount is the minimum of the total seller's declaration and the total buyer's declaration. After completing the bilateral power allocation, the coordinator constructs an antisymmetric bilateral clearing matrix that satisfies power conservation and applies a hard constraint on the upper limit of transmission capacity to each tie-line between microgrids. The power gap or surplus that is not absorbed by the internal P2P market is balanced at the physical layer through power exchange between each microgrid and the main grid.

[0016] Preferably, the power flow safety verification in step S4 involves mapping the P2P transaction settlement results with the local physical actions of each microgrid to electrical quantities, converting them into the injected power of each node in the distribution network, calling the distribution network power flow solver to solve for the node voltage amplitude and branch power flow data; if a node voltage exceeds the limit, a voltage limit soft penalty term is constructed and the degree of exceeding the limit is quantified and fed back to the reward term of the multi-agent training.

[0017] Preferably, the safe projection in step S5 involves first mapping the normalized original actions output by each microgrid agent into physical actions, and then projecting them onto the corresponding action feasible domain of each microgrid through the feasible domain projection operator to obtain feasible actions that satisfy rigid physical constraints such as energy storage power and state of charge, flexible load comfort range and tie line capacity.

[0018] Preferably, the multi-agent soft actor-critic algorithm described in step S5 introduces a projection alignment advantage regularization term during the policy update phase. This regularization term is formed by weighting the action deviation distance based on the difference in the evaluation function before and after projection, which is used to reduce the policy gradient estimation bias caused by action projection, encourage the policy to output actions in a direction with higher advantage and closer to the action feasible domain, and alleviate the actor-critic mismatch problem caused by "policy output-environment projection".

[0019] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the steps of the price-dual coupling-based multi-microgrid grid-connected neutral scheduling method as described in any of the preceding claims.

[0020] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the price-dual coupling-based multi-microgrid neutral scheduling method as described in any of the preceding claims.

[0021] Advantages of this invention:

[0022] Compared with existing technologies, this invention achieves synergistic innovation in four aspects: "dual-timescale governance of grid connection deviation, price-executable closed-loop correction, P2P settlement-power flow verification coupling, and strongly constrained multi-agent stable learning," and has at least the following beneficial effects:

[0023] 1. At the level of control mechanism, the present invention adopts a two-layer grid-connected neutral regulation structure of "instantaneous layer + sliding window accumulation layer": the instantaneous layer quickly suppresses the peak impact through the over-boundary suppression triggered by the dead zone, and the sliding window layer eliminates the continuous bias through dual integration and explicitly maps the constraint tension into an interpretable correction intensity, thereby taking into account both "fast suppression + slow elimination".

[0024] 2. At the price signal level, this invention elevates the internal settlement / incentive price from an exogenous market signal to an executable control input: by forming a closed loop through "deviation-dual variable-correction increment-executable price operator", the price can still assume the corrective responsibility under the engineering acceptability constraints such as limit, limit ramp, smoothing and projection, avoiding the instability of the strategy caused by drastic price fluctuations or simple penalty stacking.

[0025] 3. At the level of the collaborative mechanism, this invention embeds the centralized clearing of P2P transactions into a time-series closed loop: first, internal mutual assistance is cleared, and then the main network exchanges to make up for the gap. The capacity of the interconnection line is projected and constrained, so that the internal market will take the lead in absorbing the surplus / gap, reducing the fluctuation of PCC net exchange, and improving the overall economic efficiency of the multi-micro network.

[0026] 4. In terms of safety, feasibility, and learning stability, this invention introduces safe action projection and projection alignment advantage regularization under the centralized training-distributed execution framework: ensuring that the executed actions strictly meet strong physical constraints, while alleviating actor-critic mismatch and training oscillation, and improving consistency across random seeds and engineering deployability. Attached Figure Description

[0027] Figure 1 This is a flowchart of the neutral scheduling method for multi-microgrid interconnection based on price-dual coupling of the present invention;

[0028] Figure 2 This is a schematic diagram of the overall framework for neutral and collaborative energy management of multi-microgrid interconnection in this invention;

[0029] Figure 3 This is a schematic diagram of the algorithm framework of the multi-agent soft actor-commentator combined liquidation structure of the present invention;

[0030] Figure 4 This is a schematic diagram comparing the evolution of the dual variable, the grid-connected neutral indicator, and the price trajectory in this invention. In the diagram, (a) represents the dual variable. The training evolution curves, (b) are the sliding window grid-connected neutral index. (c) is the training evolution curve of the grid-connected neutral penalty term, (d) is the price trajectory comparison diagram, and (e) is the price-dual coupling dynamic response diagram.

[0031] Figure 5 This is a schematic diagram of the internal collaborative measurement and P2P transaction statistics of the present invention. In the figure, (a) is the P2P transaction ratio diagram and (b) is the average P2P power exchange amplitude diagram.

[0032] Figure 6 This is a schematic diagram of the typical daily multi-microgrid scheduling decomposition and energy storage charge state trajectory of the present invention;

[0033] Figure 7 This is a schematic diagram of the dynamic response of the grid-connected neutral index and price-dual coupling in this invention;

[0034] Figure 8 The diagram shows the P2P transaction structure and internal coordination measurement of the present invention. In the diagram, (a) is a typical daily P2P power exchange curve of each microgrid, (b) is an internal coordination measurement curve containing the proportion of P2P transactions and the internal transaction volume / power exchange amplitude, and (c) is a heat map of the average P2P transaction matrix. Detailed Implementation

[0035] An embodiment of the present invention will be further described below with reference to the accompanying drawings.

[0036] In this embodiment of the invention, a neutral scheduling method for multi-microgrid interconnection based on price-dual coupling is described in the flowchart below. Figure 1 As shown, the specific implementation process is as follows:

[0037] This embodiment focuses on a multi-microgrid cluster at the end of a distribution network. The underlying electrical topology uses typical distribution feeders (e.g., IEEE 13-node feeders), and an online power flow solver is used to perform power flow calculations to obtain safety feedback data such as node voltage and branch power flow. The system connects to multiple grid-connected microgrids (this embodiment uses three microgrids, MG-0 / 1 / 2, as an example). Each microgrid includes photovoltaic (PV), energy storage (ESS), and building flexible loads. Microgrids exchange energy bidirectionally via tie lines, with tie line power constrained by rated capacity. The net active power exchange between the grid connection point (PCC) and the main grid is the object of grid neutrality constraint management. The aforementioned three-layer coupling structure of "distribution network - multi-microgrid - transaction clearing - power flow security" is as follows: Figure 2 As shown, the application carrier and core control system constitute the scheduling method of the present invention.

[0038] like Figure 3 As shown, this invention forms a closed-loop process of "price release - decentralized response - centralized settlement - power flow verification - closed-loop correction - strategy training" within a single discrete scheduling period, specifically as follows:

[0039] S1: In each discrete scheduling period The coordinator / environment module collects operational information from each microgrid, including at least: energy storage state of charge, available photovoltaic output, uncontrollable load, flexible load status, power exchange between the microgrid and the main grid, and transaction intentions; among which, transaction intentions include at least the buy / sell direction and the declared quantity; flexible load status includes parameters such as comfort boundary, adjustable level or temperature status;

[0040] To meet the requirements of decentralized execution, each microgrid agent constructs its own decision state based solely on local observables and publicly available price signals, while the coordinator only obtains the global operating state and the joint actions of each microgrid during the centralized training phase.

[0041] Based on the data collected from each microgrid during the time period Grid-connected switching power The net active power exchange at the PCC is obtained by the coordinator. The formula is:

[0042]

[0043] in, The number of micronets in a multi-micronet cluster. This indicates that electricity is purchased from the main power grid (net power received). This indicates the supply of electricity to the main power grid (net power supply).

[0044] Based on the net active power exchange at PCC, a length of [length missing] is used. The moving average method is used to calculate the grid-connected neutral index of the sliding window. This indicator is used to characterize the medium- to long-term grid connection bias of multi-microgrid clusters, and the formula is:

[0045]

[0046] in, For the first Net active power switching power of multi-microgrid clusters at PCC during scheduling period This is the length of the sliding window.

[0047] S2: Set an asymmetric grid-connected neutral band, construct the instantaneous out-of-bounds suppression signal triggered by the dead zone and the cumulative deviation of the sliding window scale after exponential smoothing, and introduce a dual variable to represent the tension of the grid-connected neutral constraint, and complete the iterative update of the dual variable according to a preset rule, specifically:

[0048] S201: Set the grid-connected neutral band to an asymmetric interval form. ,in The neutral threshold for net internet access direction. The grid neutral threshold is the net power receiving direction. This asymmetric form can reflect the different tolerances for the "net power receiving / net power supply" sides in engineering.

[0049] Using sliding window grid connection neutral index As a criterion for determining grid neutrality in sliding window scale: when The sliding window size is determined to meet the grid neutrality requirement; otherwise, a continuous deviation is identified and correction is required. This embodiment uses... Enter and remain in the neutral zone of this asymmetric grid connection The goal is to achieve grid-connected neutral control.

[0050] S202: Based on the asymmetric grid neutral band, instantaneous out-of-bounds suppression signals and sliding window-scale cumulative deviations are constructed to achieve rapid suppression of short-term power spikes and accurate identification of medium- and long-term continuous grid connection deviations, forming a dual-time-scale hierarchical control of grid connection deviations.

[0051] (1) Regarding the net active power exchange at PCC To address the problem of short-term spike out-of-bounds crossings and severe oscillations, an instantaneous out-of-bounds suppression signal is constructed using an out-of-bounds metric with a deadband. The specific form is "dead zone + normalization secondary penalty + saturation" to avoid unnecessary adjustments caused by small power fluctuations, while avoiding the problem of unstable training values ​​through normalization and saturation limits.

[0052] when When within the dead zone, no penalty is triggered; when When the excess amount increases twice, the penalty value is normalized according to the bandwidth scale; when the excess amount is significantly greater than the preset bandwidth, the penalty value is saturated and limited, as shown in the formula:

[0053]

[0054] in, The instantaneous grid-connected neutral dead zone threshold. It is a very small positive number;

[0055] (2) The cumulative deviation of the sliding window scale is based on The superband portion exceeding the sliding window bandwidth is constructed, and the superband deviation is smoothed using an exponential smoothing coefficient. The smoothed deviation is used to drive the integral accumulation of the dual variable and form a stable correction signal.

[0056] S203: Introduce a dual variable to address the cumulative deviation of the sliding window scale. Characterizes the tension of grid-connected neutral constraints; for The iteration is completed using a composite update rule of "leakage pullback + deviation integration + interval projection" to obtain the dual variable for the next time period. This ensures accurate characterization of constraint tension while avoiding variable drift or unbounded growth. The formula is as follows:

[0057]

[0058] in, For projection operators, The leakage coefficient has a range of values. This is used to achieve a gradual decline in the dual variable after biased regression, avoiding long-term signal stagnation. A positive integration step size is used to adjust the accumulation rate of the dual variable. This is the maximum limit of the dual variable;

[0059] The update of the dual variable must simultaneously meet the following requirements: based on the sliding window superband bias smoothed by the exponential smoothing coefficient... Perform incremental updates, for Apply leakage pull-back item For incremental updates Apply single-step limit And project the updated results to the interval Inside;

[0060] The update property of the dual variable is: when When the load exceeds the neutral zone of asymmetric grid connection for an extended period, the integral term of the deviation accumulates continuously, driving... The continuous rise strengthens the subsequent price correction; when When returning to the grid neutral zone, the deviation integral term is 0, and the leakage term comes into play, making... As the iterations progress, the price gradually returns to the initial level, avoiding the problem of persistent price bias caused by the stickiness of the dual variable.

[0061] S3: Generate a candidate internal incentive price vector by combining the retail electricity price benchmark. Calculate the price correction increment based on the sliding window grid connection neutrality index and dual variables. Apply an executable update operator to complete the price constraint correction, generating an executable internal incentive price that balances the effectiveness of grid connection deviation correction with the acceptability of the engineering site. This forms a closed-loop regulation of physical deviation, dual variables, and price guidance, specifically:

[0062] S301: The coordinator uses retail electricity prices or the park's benchmark price as a reference to generate a multi-component candidate internal incentive price vector. This vector is defined in a three-component form: "valley-peak-carbon term". The standardized price vector structure must include at least three types of executable settlement signals: valley electricity price, peak electricity price, and carbon / risk coefficient. This structure is adaptable to the operating cost and revenue accounting needs of the microgrid at different times. The formula is as follows:

[0063]

[0064] in, for Off-peak electricity pricing for Peak electricity pricing for The carbon / risk coefficient for a given period is used to represent the carbon cost coefficient or the equipment operation risk cost coefficient calculated from the proxy amount of controllable failure factors.

[0065] S302: Coordinator based on sliding window grid neutrality index Dual variables At the same time, a deviation difference term is introduced. Capture the changing trend of grid connection deviation and calculate the incremental price correction. and according to The positive and negative directions are selected for the valley component. or peak component Targeted injection is performed, specifically when... When >0 (cluster net power reception exceeds grid neutral band), for peak segment components Injecting correction increments to increase peak-hour electricity purchase costs and guide microgrids to reduce their electricity purchases from the main grid; when When <0 (the cluster's net internet access exceeds the neutral band), the valley segment component... Injecting correction increments to reduce off-peak electricity sales revenue and guide microgrids to reduce power transmission to the main grid, thereby achieving a precise return of cluster net exchange power to the grid neutral zone;

[0066] Corrective increment Injected into the baseline candidate price vector, resulting in candidate prices with corrective properties. The formula is:

[0067]

[0068] in, for The benchmark executable internal incentive price vector published by the time-sharing coordinator. This is the adjustment coefficient for the grid-connected neutral index of the sliding window. The adjustment coefficient for the dual variable. This is the adjustment coefficient for the deviation difference term;

[0069] The magnitude of the price correction increment and degree of deviation The degree of constraint tension is positively correlated; the more significant the deviation and the tighter the constraint, the greater the increment of price correction and the stronger the price correction guidance.

[0070] S303: For candidate prices with correction attributes First, obtain the benchmark executable price using the price executable update operator, and then inject price correction increments under executable constraints. Finally, an executable update operator is applied. The coordinator publicly released the executable internal incentive price. ;

[0071] The price update operator includes at least an inertial smoothing operator. Single-step saturation operator With interval projection operator The calculation of the baseline executable price and the final executable internal incentive price satisfies:

[0072] (7)

[0073] (8)

[0074] in,

[0075] (9)

[0076] in, As the benchmark executable internal incentive price vector, This is the amplitude limit value for the single-step saturation operator. For inertial smoothing operators The smoothness coefficient, For time period The peak and valley attribute indicator, with a value of 0 or 1, where 0 represents a valley segment and 1 represents a peak segment;

[0077] Price constraint correction can also be accomplished using a simplified executable update operator. The three constraint operations—inertia smoothing, single-step ramp limiting, and interval projection—are executed sequentially to directly correct the final executable internal incentive price. The formula is publicly released by the coordinator:

[0078]

[0079] in, Index for discrete scheduling periods ( =1,2,…,N, where N is the total number of scheduling periods). , These are the lower and upper limits of the internal incentive price, respectively. The inertial smoothing coefficients are related to the inertial smoothing operator. The smoothing coefficients are consistent. The upper limit of single-step climbing, and the single-step saturation operator. amplitude limit Consistent For single-step saturation operators, and Equivalent implementations are both used to limit the single-step update range of variables;

[0080] Through the complete process of "benchmark generation + correction increment injection + executable operator constraint", a price regulation closed loop based on price-dual coupling is constructed. This breaks away from the limitation of traditional prices as only exogenous incentive signals, and transforms the internal incentive price into an executable control input that can actively regulate the grid connection status. It can guide the subsequent action decisions and transaction intention output of the microgrid agent through price, realize the effective correction of grid connection deviation, and avoid the strategy instability problem caused by drastic price fluctuations. It takes into account both the effectiveness of grid connection regulation and the feasibility of engineering execution.

[0081] S4: Each microgrid agent outputs its local physical action and P2P transaction intent based on the published executable internal incentive price. First, it performs centralized clearing of P2P transactions and verifies tie-line capacity. Then, it combines the clearing results with the physical actions to conduct a distribution network power flow safety verification. The verification results are fed back to the multi-agent training reward, specifically:

[0082] S401: Each microgrid agent receives the executable internal incentive price publicly released by the coordinator. Subsequently, relying solely on local observation information (energy storage state of charge, photovoltaic available output, load operating status, etc.), distributed autonomous decision-making is achieved, independently outputting original local physical actions and P2P transaction intentions. :

[0083] The local physical actions include energy storage charging / discharging decisions (including power and direction), flexible load adjustment decisions (such as HVAC settings, peak shaving and valley filling of interruptible loads, etc.), and photovoltaic consumption / curtailment decisions; P2P transaction intent Clearly define the buy / sell direction and the declared transaction power quantity; each microgrid action first outputs the "original action", and then performs security projection correction into an actionable action;

[0084] S402: The coordinator gathers the P2P transaction intentions of all microgrids and completes the P2P energy transaction clearing process according to the standardized centralized clearing process of "first determining the total liquidation volume, then making bilateral allocations, and finally projecting the tie-line capacity". It also performs hard constraint verification on the tie-line capacity, specifically:

[0085] (1) Distinguish between the participants in the transaction and define the seller set as Buyer group The declared quantities of each seller are The quantities declared by each buyer are The total liquidable amount is the minimum of the total amounts declared by both the buyer and seller. The formula is:

[0086]

[0087] (2) Pre-defined liquidation rules for proportional or priority allocation will determine the total liquidable amount. The volume of transactions is allocated to each microgrid and the transaction volume between each microgrid is determined. And construct a bilateral clearing matrix based on the transaction volume. The bilateral clearing matrix must satisfy the power conservation requirement and be an antisymmetric matrix, i.e. This clarifies the actual transaction efficiency between each micronetwork.

[0088] (3) Rated transmission capacity of each interconnect line between microgrids Perform amplitude-limited projection on the bilateral clearing matrix. Implement constraint adjustments to ensure that the actual power of any bilateral transaction does not exceed the tie-line capacity limit;

[0089] (4) The power gap / surplus that is not absorbed after the internal P2P market liquidation is the power exchanged between each microgrid and the main grid. Balancing is achieved at the physical layer, thereby reducing the net active power exchange fluctuation at the common coupling point (PCC) and improving the neutrality level of multi-microgrid clusters.

[0090] S403: The coordinator maps the P2P transaction settlement results to the local physical actions of each microgrid using electrical quantities, conducts online power flow safety verification of the distribution network, and feeds the verification results back to the multi-agent training reward. This constructs a market-physical closed-loop constraint verification mechanism with dual dimensions of grid neutrality and electrical safety, while mitigating two types of operational risks: distribution network electrical boundary violations and excessive grid connection deviations. Specifically:

[0091] (1) P2P bilateral clearing matrix It is integrated with the local physical actions of each microgrid (energy storage charging and discharging, photovoltaic power output, load power, etc.) and converted into the injected power of each node of the distribution network. The calculation of the node injected power needs to be adapted to the wiring rules of the underlying electrical topology of the distribution network (such as the IEEE 13-node feeder).

[0092] (2) Call the power flow solver of the distribution network to perform online power flow calculation on the injected power of the nodes and obtain the voltage amplitude of each node. Data on the flow of each branch;

[0093] (3) If a node voltage exceeds the limit, construct a normalized quadratic form voltage over-limit soft penalty term. To quantify the degree of voltage exceeding the limit and avoid training oscillations caused by hard constraints, the formula is:

[0094]

[0095] in, For the distribution network The voltage amplitude at each node, , These are the upper and lower threshold values ​​for the node voltage, respectively.

[0096] (4) Soft penalty for voltage exceeding limit The branch power flow over-limit penalty and the grid neutrality penalty (Exponential Moving Average, EMA) including instantaneous over-limit suppression penalty and sliding window scale cumulative deviation penalty are jointly included in the reward / constraint penalty of multi-agent reinforcement learning. By deducting the agent's training reward and increasing the penalty value, safety feedback is achieved, providing accurate safety feedback signals for the subsequent policy training of the coordinator and microgrid agents. From the perspective of reward and punishment mechanism, the electrical safety and grid neutrality requirements of the distribution network operation are guaranteed at the same time.

[0097] S5: Under the centralized training-distributed execution framework, the original actions of the microgrid agents are subjected to secure projection to satisfy various strong physical constraints. A multi-agent soft actor-critic (MASAC) reinforcement learning algorithm is used to jointly train the scheduling strategies of the coordinator and each microgrid. After training, the optimal strategy is deployed to each microgrid agent and the coordinator to carry out neutral online scheduling of multiple microgrids. Specifically:

[0098] S501: Normalized raw actions output by each micronet agent First, it is mapped to a physical action, and then projected onto the corresponding action-feasible domain of the microgrid to obtain the actionable actions. This action satisfies strong physical constraints such as energy storage power / energy storage state of charge, flexible load comfort range, and tie line capacity, completely avoiding inactive outputs at the execution level. The formula is:

[0099]

[0100] in, For feasible region projection operators, For micro-network The feasible domain of actions;

[0101] S502: Employs the Multi-Agent Soft Actor-Critic (MASAC) algorithm to jointly train the scheduling policies of the coordinator and each micro-network agent, balancing global stability of training with distributed execution characteristics; simultaneously, it introduces a Projection Alignment Advantage (PAA) regularization term in policy updates to optimize training performance.

[0102] (1) Centralized training phase: A centralized commentator calls upon the global operating state and the joint actions of each agent to stably evaluate the value and provide accurate basis for policy updates;

[0103] (2) Decentralized execution phase: Each microgrid agent makes autonomous decisions based solely on local observation information and price signals publicly released by the coordinator, reducing system communication and computing overhead;

[0104] (3) Regularization term optimization: The projection alignment advantage regularization term is used to reduce the bias caused by the deviation of the action before and after projection on the policy gradient estimation, and the action deviation distance is weighted according to the difference of the evaluation function before and after projection. This regularization term alleviates the actor-critic mismatch problem caused by "policy output-environment projection", encourages the policy to output actions closer to the feasible domain in the direction with higher advantage, and improves the stability of training under strong constraint action space and cross random seed consistency.

[0105] S503: After training, the optimal strategy is deployed to each microgrid agent and coordinator to carry out neutral online scheduling of multi-microgrid interconnection, and the strategy is continuously iterated and optimized based on the operation log. At the same time, the grid connection operation status is monitored in real time and anomaly correction is performed.

[0106] To further clarify the closed-loop correction mechanism and operational effect of the present invention, combined with Figures 4-8 The explanation is as follows: Figure 4 As shown, dual variables Neutral index of grid connection with sliding window During training and operation, the price remained stable within a safe range, and was adjusted up or down accordingly during the correction phase. After returning to the grid neutral zone, the price quickly fell back, demonstrating the precise, efficient, and executable correction capability of the "price-dual closed loop" for grid deviation. Figure 5 By statistically analyzing the proportion of P2P transactions and the average P2P exchange power amplitude, the operational advantages of this invention are intuitively reflected: significantly improved energy mutual assistance level within the multi-microgrid, optimized tie-line utilization, and effective reduction of external grid-connected power fluctuations. Figure 6 The decomposition results of multi-microgrid scheduling and the energy storage state-of-charge trajectory of a typical day are presented, verifying that energy storage and load can achieve efficient coordinated peak shifting under the guidance of price signals, and that all kinds of rigid physical constraints can be met throughout the process, fully demonstrating the feasibility and economic advantages of the scheduling strategy of this invention; Figure 7 Demonstrated grid-connected neutral indicators The intraday dynamic evolution of the price-dual coupled response verifies the excellent performance of this invention, which features rapid regulation response, outstanding dynamic tracking capability, and stable maintenance of grid-connected neutral state. Figure 8 The invention demonstrates the technical advantages of a reasonable internal transaction structure and excellent collaborative scheduling effect in multi-microgrid clusters through P2P power exchange curves, internal coordination metrics, and average transaction matrix.

[0107] This embodiment can be implemented in an electronic device through a combination of hardware and software: the "data acquisition and state construction module, price-dual closed-loop module, P2P clearing and flow verification module, security projection and strategy inference, training and execution module" are encapsulated as a program that can run on a processor and stored in memory or computer-readable storage media (such as solid-state drives, USB flash drives, cloud storage, industrial control memory cards, etc.). During online operation, the device cyclically executes the above steps according to the scheduling cycle, outputs control commands and published prices for each microgrid, and records operation logs for continuous training or offline retraining.

[0108] The method provided by this invention has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are merely for the purpose of helping to understand the core ideas of this invention. It should be noted that those skilled in the art can make various improvements and modifications to this invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this invention.

Claims

1. A neutral scheduling method for multi-microgrid interconnection based on price-dual coupling, characterized in that, Includes the following steps: S1: Collect the operation information of the multi-microgrid cluster, and calculate the net active power exchange power and sliding window grid neutrality index of the cluster at the common coupling point. S2: Set an asymmetric grid-connected neutral band, construct the instantaneous out-of-bounds suppression signal and the cumulative deviation of the sliding window scale, introduce a dual variable characterizing the tension of the grid-connected neutral constraint and update iteratively; S3: Generate a candidate internal incentive price vector based on the benchmark electricity price, calculate the price correction increment by combining the sliding window grid-connected neutral index and dual variables and inject it into the candidate internal incentive price vector, apply an executable update operator to it, obtain the executable internal incentive price and publish it to the outside world; S4: Each microgrid agent outputs local physical actions and P2P transaction intentions based on the published executable internal incentive price. After completing the centralized clearing of P2P transactions and the capacity verification of tie lines, the power flow safety verification of the distribution network is carried out, and the verification results are fed back to the multi-agent training reward. S5: Under the centralized training-decentralized execution framework, the original actions of each micronet agent are safely projected. The multi-agent soft actor-critic algorithm is used to jointly train the coordinator and the scheduling strategy of each micronet to achieve neutral online scheduling of multiple micronets.

2. The method according to claim 1, characterized in that, The net active power exchange in step S1 is the sum of the power exchange between all microgrids in the multi-microgrid cluster and the main power grid; the sliding window grid neutrality index is the average value of the net active power exchange at the common coupling point in each discrete time period within the preset length sliding window, which is used to characterize the grid bias accumulation state of the multi-microgrid cluster in the continuous scheduling period corresponding to the sliding window.

3. The method according to claim 1, characterized in that, The instantaneous over-limit suppression signal in step S2 adopts the form of "dead zone + normalized double penalty + saturation". The penalty is triggered only when the absolute value of the net active power exchange at the common coupling point exceeds the instantaneous dead zone threshold. The penalty value increases twice with the excess amount and is limited by saturation to suppress power spike impact. The cumulative deviation of the sliding window scale is constructed based on the overband part of the grid neutral index of the sliding window that exceeds the grid neutral band, and is used to identify continuous grid deviation. The dual variable is iteratively updated using the rule of "leakage pullback + deviation integral + interval projection". When the grid-connected neutral index of the sliding window exceeds the grid-connected neutral band for multiple consecutive scheduling cycles, the dual variable increases through the cumulative deviation integral to strengthen the correction intensity. When the index returns to the grid-connected neutral band, the dual variable gradually decreases through the leakage pullback term to avoid signal sticking. At the same time, the iterative update process is constrained by single-step amplitude limiting and interval projection to limit the unbounded growth of the dual variable.

4. The method according to claim 1, characterized in that, The candidate internal incentive price vector in step S3 is in multi-component form, defined by a three-component structure of "valley-peak-carbon term", and includes at least three types of executable settlement signals: valley electricity price, peak electricity price, and carbon or risk coefficient. The executable update operator includes at least an inertial smoothing operator, a single-step saturation operator, and an interval projection operator. The candidate internal incentive price vector with correction attributes is subjected to inertial smoothing, single-step ramp-up limiting, and interval projection operations in sequence through each operator, and finally an executable internal incentive price is generated for external release.

5. The method according to claim 1, characterized in that, The P2P transaction centralized clearing in step S4 follows a process of first determining the total clearable amount, then completing the bilateral power allocation, and finally executing the tie-line capacity projection. The total clearable amount is the minimum of the total seller's declaration amount and the total buyer's declaration amount. After completing the bilateral power allocation, the coordinator constructs an antisymmetric bilateral clearing matrix that satisfies power conservation and imposes a hard constraint on the upper limit of transmission capacity for each interconnection line between microgrids; the power gap or surplus that is not absorbed by the internal P2P market is balanced at the physical layer through power exchange between each microgrid and the main grid.

6. The method according to claim 1, characterized in that, The power flow safety verification in step S4 involves mapping the P2P transaction settlement results with the local physical actions of each microgrid to electrical quantities, converting them into the injected power of each node in the distribution network, and then calling the distribution network power flow solver to solve the problem and obtain the node voltage amplitude and branch power flow data. If a node voltage exceeds the limit, a soft penalty term for voltage exceeding the limit is constructed, and the degree of exceeding the limit is quantified and fed back to the reward term for multi-agent training.

7. The method according to claim 1, characterized in that, The safe projection in step S5 involves first mapping the normalized original actions output by each microgrid agent into physical actions, and then projecting them onto the corresponding action feasible domain of each microgrid through the feasible domain projection operator to obtain feasible actions that satisfy rigid physical constraints such as energy storage power and state of charge, flexible load comfort range and tie line capacity.

8. The method according to claim 1, characterized in that, In step S5, the multi-agent soft actor-critic algorithm introduces a projection alignment advantage regularization term during the policy update phase. This regularization term is formed by weighting the action deviation distance based on the difference in the evaluation function before and after projection. It is used to reduce the policy gradient estimation bias caused by action projection, encourage the policy to output actions in a direction with higher advantage and closer to the action feasible domain, and alleviate the actor-critic mismatch problem caused by "policy output-environment projection".

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the steps of the multi-microgrid neutral scheduling method based on price-dual coupling as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of the price-dual coupling-based multi-microgrid neutral scheduling method as described in any one of claims 1 to 8.