A Multi-Microgrid Cooperative Scheduling Method and System Based on Game Theory

Through the multi-micronet collaborative scheduling method based on game theory, a space-time coupled game feature space is constructed, dynamic game initial strategy set is generated, distributed negotiation and reinforcement learning is carried out, efficient and flexible resource allocation and scheduling of the multi-micronet system is realized, and the coordination and scheduling problems exist in the multi-micronet system are solved.

CN120280930BActive Publication Date: 2025-08-01HANGZHOU KGOOER ELECTRONIC TECH CO LTD

Patent Information

Application Number
CN202510764664.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-08-01
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

In the coordinated scheduling, the multi-micronet system has problems such as delay in information transmission, single-point system failure, inflexible scheduling and unbalanced resource allocation, making it difficult to achieve optimized dynamic adjustment.

Method used

The multi-micronet collaborative scheduling method based on game theory is adopted, and the game feature space that is coupled with space-time coupled, and a multi-agent strategy network is used to generate a dynamic game initial strategy set, combined with the asymmetric Nash bargaining model for distributed negotiation, and a equilibrium benefit allocation scheme is generated, and a time-decay reinforcement learning algorithm is used to iteratively correct the energy storage priority indicators, and finally, energy storage resources are allocated in real time through the decentralized gradient consensus mechanism.

Benefits of technology

It improves the operating efficiency and adaptability of the multi-micro-grid system, meets the needs of future intelligent distribution networks, and realizes the flexibility of efficient allocation and scheduling of resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120280930B_ABST
    Figure CN120280930B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-microgrid collaborative scheduling method and system based on game theory. The method includes: generating a dynamic game initial strategy set through a multi-agent policy network according to the state of charge, charge and discharge efficiency of user energy storage devices, and the real-time scheduling requirements of the power grid; based on the dynamic game initial strategy set, using an asymmetric Nash bargaining model for distributed negotiation to generate an equilibrium interest distribution scheme; according to the equilibrium interest distribution scheme, using a time-decaying reinforcement learning algorithm to iteratively correct the energy storage priority index to generate a dynamic bidding rule including supply and demand elasticity coefficients and risk compensation; based on the dynamic bidding rule, real-time allocating energy storage resources through a decentralized gradient consensus mechanism to generate a final collaborative scheduling instruction and synchronize it to each microgrid terminal. By using the embodiments of the present invention, the operation efficiency and adaptive ability of the multi-microgrid system can be improved to meet the requirements of future intelligent distribution networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of microgrids, and particularly relates to a multi-microgrid collaborative scheduling method and system based on game theory. Background Art

[0002] With the rapid popularization of new energy and the development of distributed generation technology, multi-microgrid systems, as an important part of the power system, exhibit high flexibility and scalability. Microgrids can achieve autonomous operation by integrating distributed energy sources, energy storage devices, and control systems, effectively alleviating the grid load pressure and improving energy utilization efficiency. However, due to the autonomy and heterogeneity of microgrids, their coordinated scheduling faces many challenges.

[0003] Traditional scheduling methods mostly adopt centralized schemes, relying on a single control center for decision-making, which suffer from problems such as information transmission delay, single-point system failure, and inflexible scheduling. On the other hand, the diverse demands and unbalanced resource allocation among microgrids also increase the complexity of coordinated scheduling. In addition, the state of charge, charge and discharge efficiency changes of microgrid energy storage devices, and the dynamic changes in users' energy demands all require dynamic adjustment of scheduling strategies to achieve optimization. Summary of the Invention

[0004] The purpose of the present invention is to provide a multi-microgrid collaborative scheduling method and system based on game theory to solve the deficiencies in the prior art, which can improve the operation efficiency and adaptive ability of multi-microgrid systems and meet the requirements of future intelligent distribution grids.

[0005] An embodiment of the present application provides a multi-microgrid collaborative scheduling method based on game theory, the method comprising:

[0006] Construct a spatio-temporal coupled game feature space according to the state of charge, charge and discharge efficiency of the user's energy storage device and the real-time scheduling requirements of the power grid, and generate a dynamic game initial strategy set through a multi-agent strategy network, wherein the dynamic game initial strategy set includes the charge and discharge amounts and bid combinations of each microgrid;

[0007] Based on the dynamic game initial strategy set, adopt an asymmetric Nash bargaining model for distributed negotiation, and dynamically adjust the weight factors of each microgrid through a virtual resource exchange market to generate an equilibrium benefit distribution scheme, wherein the weight factors are jointly calculated by the user credit rating and the power grid emergency level;

[0008] According to the equilibrium benefit distribution scheme, use a time-decaying reinforcement learning algorithm to iteratively correct the energy storage priority index and generate a dynamic bidding rule including a supply-demand elasticity coefficient and risk compensation;

[0009] Based on the dynamic bidding rules, the energy storage resources are allocated in real time through a decentralized gradient consensus mechanism, and the verifiable random function is combined to verify the effectiveness of the allocation, generating the final collaborative scheduling instruction and synchronizing it to each microgrid terminal.

[0010] Optionally, according to the state of charge, charge and discharge efficiency of the user's energy storage device and the real-time scheduling requirements of the power grid, a spatio-temporal coupled game feature space is constructed, and a dynamic game initial strategy set is generated through a multi-agent policy network, where the dynamic game initial strategy set includes the charge and discharge amounts and bid combinations of each microgrid, including:

[0011] Collect the real-time state of charge data and historical charge and discharge efficiency curves of each microgrid energy storage device, combine the regional load forecast and electricity price fluctuation signals issued by the power grid dispatching center, and perform spatio-temporal domain interpolation processing to generate a spatio-temporal coupled state matrix with a preset time granularity;

[0012] Input the spatio-temporal coupled state matrix into a convolutional recurrent hybrid neural network, extract the charge and discharge correlation features between adjacent microgrids through a time sliding window, and output a game relationship graph including the strength of supply and demand relationship and capacity constraints;

[0013] Based on the game relationship graph, construct an adversarial training environment for the multi-agent policy network, and use a double-layer policy gradient algorithm to synchronously optimize the charge and discharge amount strategies and bid strategies of each microgrid, generating an initial strategy set including Nash equilibrium constraints;

[0014] Perform Pareto front screening on the initial strategy set, eliminate the strategy combinations that violate the power grid safety and stable operation constraints, and output the dynamic game initial strategy set.

[0015] Optionally, based on the dynamic game initial strategy set, an asymmetric Nash bargaining model is used for distributed negotiation, and the weight factors of each microgrid are dynamically adjusted through a virtual resource exchange market, generating an equilibrium interest distribution scheme, where the weight factors are jointly calculated by the user credit rating and the power grid emergency level, including:

[0016] Analyze the charge and discharge amounts and bid combinations in the initial strategy set, initialize the threat point and bargaining space of the asymmetric Nash bargaining model, and set the threat point as the minimum revenue threshold when the microgrid operates independently;

[0017] According to the historical compliance records of microgrid users and the probability of grid node voltage over-limit, calculate the credit score and emergency coefficient respectively, and generate an initial weight factor by fusing the two types of parameters through the hyperbolic tangent function;

[0018] Deploy a distributed negotiation protocol in the virtual resource exchange market for distributed negotiation. Each microgrid calculates the interest utility of the strategy set based on the current weight factor, and iteratively updates the bargaining solution through the alternating direction multiplier method;

[0019] When the difference in the Shapley value between two adjacent iterations is less than the preset threshold, a dynamic adjustment mechanism for the weight factor is triggered. Among them, the weights of the microgrids with a decreased credit rating decay exponentially, and the weights of the microgrids within the emergency area of the power grid increase linearly.

[0020] Generate an equilibrium benefit distribution plan including the transfer payment plan and the charging and discharging timing plan according to the final negotiation result, and verify the individual rationality and coalition stability of the plan through zero-knowledge proof technology.

[0021] Optionally, according to the equilibrium benefit distribution plan, use the time-decaying reinforcement learning algorithm to iteratively correct the energy storage priority index, and generate a dynamic bidding rule including the supply-demand elasticity coefficient and risk compensation, including:

[0022] Extract the historical winning bid rate and revenue fluctuation data of each microgrid from the equilibrium benefit distribution plan, and initialize the energy storage priority index as a weighted function of the supply-demand ratio and the risk value;

[0023] Construct a state transition model of time-decaying reinforcement learning, define the state as the difference between the real-time load of the power grid and the available capacity of the energy storage, the action as the priority adjustment amplitude, and the reward function includes the electricity price arbitrage revenue and the power grid regulation contribution degree;

[0024] Adopt a decay factor to dynamically adjust the exploration rate of the Q-learning algorithm, enhance the weight of historical experience during peak load periods, increase the random exploration probability during valley periods, and update the probability distribution of the energy storage priority index;

[0025] According to the updated priority index and its probability distribution, calculate the ratio of the price change rate to the trading volume change rate as the supply-demand elasticity coefficient, and quantify the risk compensation value of energy storage invocation based on the conditional value-at-risk model;

[0026] Encode the supply-demand elasticity coefficient and the risk compensation value of energy storage invocation as a piecewise linear function, generate a dynamic bidding rule including the ladder quotation upper limit and penalty clauses, and write it into the consensus verification contract of the blockchain.

[0027] Optionally, based on the dynamic bidding rule, allocate energy storage resources in real time through a decentralized gradient consensus mechanism, verify the allocation effectiveness by combining a verifiable random function, generate a final coordinated scheduling instruction and synchronize it to each microgrid terminal, including:

[0028] Broadcast the dynamic bidding rule in the decentralized network, each microgrid node generates a candidate allocation plan based on local constraint conditions, and calculates the gradient vector of the plan for the global objective function;

[0029] Adopt the gradient consensus algorithm to aggregate the gradient information of each node, and solve the optimal allocation solution that satisfies the power grid power flow constraint through the projected subgradient descent method to generate a preliminary coordinated scheduling instruction;

[0030] Call a verifiable random function to generate a distributed random number seed. Each node performs a pre-commitment operation based on the collaborative scheduling instruction and generates an operation hash. After aggregating all hash values through threshold signature, they are compared with the random number to verify the consistency.

[0031] When the verification pass rate exceeds the preset threshold, write the collaborative scheduling instruction into the trusted execution environment of each microgrid terminal. Otherwise, trigger the elastic reallocation mechanism and return to execute the steps of broadcasting dynamic bidding rules in the decentralized network. Each microgrid node generates a candidate allocation plan based on local constraint conditions and calculates the gradient vector of the plan for the global objective function for iterative optimization until convergence.

[0032] Another embodiment of the present application provides a multi-microgrid collaborative scheduling system based on game theory, which includes:

[0033] A construction module, configured to construct a spatio-temporal coupled game feature space according to the state of charge, charge and discharge efficiency of the user's energy storage device and the real-time scheduling requirements of the power grid, and generate an initial dynamic game strategy set through a multi-agent policy network. Among them, the initial dynamic game strategy set includes the charge and discharge amounts and bid combinations of each microgrid.

[0034] A negotiation module, configured to perform distributed negotiation based on the initial dynamic game strategy set, adopt an asymmetric Nash bargaining model, dynamically adjust the weight factors of each microgrid through a virtual resource exchange market, and generate an equilibrium interest distribution plan. Among them, the weight factors are jointly calculated by the user credit rating and the power grid emergency level.

[0035] A correction module, configured to iteratively correct the energy storage priority index according to the equilibrium interest distribution plan by using a time-decaying reinforcement learning algorithm, and generate a dynamic bidding rule including a supply-demand elasticity coefficient and risk compensation.

[0036] An allocation module, configured to allocate energy storage resources in real time based on the dynamic bidding rule through a decentralized gradient consensus mechanism, verify the allocation effectiveness in combination with a verifiable random function, generate a final collaborative scheduling instruction and synchronize it to each microgrid terminal.

[0037] Another embodiment of the present application provides a storage medium, in which a computer program is stored. Among them, the computer program is set to execute the method described in any one of the above when running.

[0038] Another embodiment of the present application provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is set to run the computer program to execute the method described in any one of the above.

[0039] Compared with the prior art, a multi-microgrid collaborative scheduling method based on game theory provided by the present invention generates a dynamic game initial strategy set according to the state of charge, charge and discharge efficiency of user energy storage devices, and the real-time scheduling requirements of the power grid; based on the dynamic game initial strategy set, an asymmetric Nash bargaining model is used for distributed negotiation to generate an equilibrium benefit distribution plan; according to the equilibrium benefit distribution plan, a time-decaying reinforcement learning algorithm is used to iteratively correct the energy storage priority index to generate a dynamic bidding rule including supply-demand elasticity coefficients and risk compensation; based on the dynamic bidding rule, the energy storage resources are allocated in real time through a decentralized gradient consensus mechanism to generate a final collaborative scheduling instruction and synchronize it to each microgrid terminal, so as to improve the operation efficiency and adaptive ability of the multi-microgrid system and meet the requirements of future intelligent distribution networks. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a hardware structure block diagram of a computer terminal for a multi-microgrid collaborative scheduling method based on game theory provided by an embodiment of the present invention;

[0041] Figure 2 It is a schematic flowchart of a multi-microgrid collaborative scheduling method based on game theory provided by an embodiment of the present invention;

[0042] Figure 3 It is a schematic structural diagram of a multi-microgrid collaborative scheduling system based on game theory provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0044] An embodiment of the present invention first provides a multi-microgrid collaborative scheduling method based on game theory, which can be applied to electronic devices, such as computer terminals, specifically ordinary computers, etc.

[0045] The following takes running on a computer terminal as an example for a detailed description. Figure 1 It is a hardware structure block diagram of a computer terminal for a multi-microgrid collaborative scheduling method based on game theory provided by an embodiment of the present invention. As Figure 1 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory may include a non-volatile storage medium and an internal memory.

[0046] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any multi-microgrid collaborative scheduling method based on game theory.

[0047] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.

[0048] The internal memory provides an environment for the operation of a computer program in a non-volatile storage medium. When the computer program is executed by the processor, the processor can be caused to execute any multi-microgrid collaborative scheduling method based on game theory.

[0049] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 1 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0050] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0051] See Figure 2 , an embodiment of the present invention provides a multi-microgrid collaborative scheduling method based on game theory, which may include the following steps:

[0052] S201, according to the state of charge, charge and discharge efficiency of the user's energy storage device and the real-time scheduling requirements of the power grid, construct a spatio-temporal coupled game feature space, and generate a dynamic game initial strategy set through a multi-agent strategy network. Among them, the dynamic game initial strategy set includes the charge and discharge amounts and price quotation combinations of each microgrid; specifically, it may include:

[0053] S2011, collect the real-time state of charge data and historical charge and discharge efficiency curves of each microgrid energy storage device, combine the regional load forecast and electricity price fluctuation signals issued by the power grid dispatching center, and perform spatio-temporal domain interpolation processing to generate a spatio-temporal coupled state matrix with a preset time interval as the granularity; the preset time interval is, for example, 15 minutes.

[0054] In this step, a high-precision state matrix is constructed through multi-source data fusion and spatio-temporal alignment techniques, providing a structured input for subsequent game modeling.

[0055] ‌Data collection and preprocessing‌:

[0056] ‌Energy storage device data collection‌:

[0057] ‌Real-time state of charge (SOC)‌: Read the SOC value (0 - 100%) of the lithium battery management system through the CAN bus protocol, with a sampling frequency of 1 Hz. Use moving average filtering (window length 60 seconds) to eliminate noise. For example, the SOC of the energy storage in Microgrid A is 72.3% at 14:00, with an accuracy of ±0.5%.

[0058] ‌Historical charge and discharge efficiency curve‌: Extract the charge and discharge data of the past 30 days from the SCADA system and calculate the efficiency η = actual discharge amount / theoretical capacity. For example, the average efficiency of Microgrid B at 20°C is η = 93.2%.

[0059] ‌Grid dispatching data access‌:

[0060] ‌Regional load forecasting‌: Receive 24-hour load forecasting data (CSV format) released by the grid dispatching center, with a time resolution of 5 minutes and a power range of 0 - 100 MW. For example, the regional load forecast at 15:00 is 58.7 MW, with a confidence interval of ±3%.

[0061] ‌Electricity price fluctuation signal‌: Analyze the difference between the day-ahead electricity price and the real-time electricity price in the electricity trading market to generate a fluctuation index (-1 to +1). A negative value indicates oversupply. For example, the real-time electricity price at 14:30 is 12% lower than the day-ahead price, and the fluctuation index is -0.4.

[0062] ‌Spatio-temporal domain interpolation processing‌:

[0063] ‌Spatial interpolation‌: For the discrete SOC data of 12 microgrids, use Kriging interpolation to generate a SOC distribution map of a 500m × 500m grid. The semi-variogram selects the Gaussian model, with the nugget value set to 0.05 and the range set to 2 km.

[0064] ‌Time alignment‌: Unify data with different sampling frequencies to a 15-minute granularity. For the 1 Hz SOC data, take the median within each 15-minute window; for the 5-minute load data, generate a 15-minute sequence through cubic spline interpolation.

[0065] Matrix construction: Organize the interpolated data into a four-dimensional tensor (spatial rows × spatial columns × time slices × feature dimensions), with a size of 12×12×96×8 (a 12-row × 12-column grid, 96 15-minute time slices, and 8 features: SOC, efficiency, load, electricity price, etc.). For example, the matrix position (5,7,32,3) represents that the load value of the grid at the 5th row and 7th column at 08:00 is 45.6 MW.

[0066] In S2012, input the spatio-temporal coupling state matrix into the convolutional recurrent hybrid neural network, extract the charge-discharge correlation features between adjacent microgrids through a time-sliding window, and output a game relationship graph containing the strength of the supply-demand relationship and capacity constraints.

[0067] This step mines spatio-temporal correlation features through a hybrid neural network and constructs a topological representation of the dynamic game relationship between microgrids.

[0068] Convolutional recurrent hybrid neural network (CRHNN) design:

[0069] Spatial feature extraction: Use a 3D convolutional layer (Conv3d) to process the spatial-temporal dimensions. The convolutional kernel size is 3×3×3 (height × width × time), and the number of channels is expanded from 8 to 32. For example, the input tensor (12×12×96×8) outputs (10×10×94×32) after passing through Conv3d, capturing the charge-discharge spatial patterns of local microgrid groups.

[0070] Temporal dependence modeling: Connect a bidirectional LSTM (Bi-LSTM) after the convolutional layer. The hidden layer dimension of each LSTM cell is 256, and the time window sliding step is 4 (i.e., one step per hour). For example, for 94 time slices, they are divided into 23 windows with a step of 4, and each window processes 4 consecutive time slices (i.e., 1 hour of data).

[0071] Correlation feature fusion: Calculate the correlation weights between adjacent microgrids through the multi-head attention mechanism (Multi-Head Attention, number of heads = 8). For example, the attention score α_ij of microgrid i to microgrid j is 0.73, indicating that the charge-discharge behavior of i is greatly affected by j.

[0072] Game relationship graph generation:

[0073] Calculation of the strength of the supply-demand relationship: Define the strength S_ij = α_ij × (SOC_i - SOC_j) × η_i, where η_i is the charge-discharge efficiency of microgrid i. For example, S_AB of microgrid A to B is 0.73×(72% - 65%)×93% = 0.047, indicating the potential for A to supply power to B.

[0074] Capacity Constraint Encoding: According to the transmission limit of the lines between microgrids, generate a hard constraint matrix \(C_{ij}=\min(\text{line capacity}, \text{remaining energy storage capacity})\). For example, if the line capacity from A to B is 2 MW and the current discharge capacity of A is 1.8 MW, then \(C_{AB} = 1.8\) MW.

[0075] Graph Construction: Encode \(S_{ij}\) and \(C_{ij}\) into a weighted adjacency matrix, where the nodes represent microgrids and the edge weights are \([S_{ij}, C_{ij}]\). For example, the adjacency matrix element from A to B is \([0.047, 1.8]\), which is stored as a JSON-formatted topology file.

[0076] S2013, based on the game relationship graph, construct an adversarial training environment for the multi-agent policy network, and use the double-layer policy gradient algorithm to synchronously optimize the charge and discharge strategies and bidding strategies of each microgrid, generating an initial strategy set containing Nash equilibrium constraints;

[0077] In this step, through adversarial training and game equilibrium constraints, generate an optimal strategy candidate set that satisfies the interests of multiple parties.

[0078] Multi-Agent Policy Network Architecture:

[0079] Policy Network (Actor): Each microgrid independently has a policy network. The input is its own state (SOC, efficiency) and the features of adjacent nodes in the game graph, and the output is the charge and discharge amount \(a_t\) (-1 to +1, normalized value) and the bid price \(p_t\) (0 to 1, corresponding to 80% - 120% of the market price). The network structure is a 4-layer fully connected network (256→128→64→2), and the activation function uses LeakyReLU (negative slope 0.01) except for the last layer which uses Tanh.

[0080] Value Network (Critic): The global value network receives the states and actions of all microgrids and outputs the joint Q value. Use a graph attention network (GAT) to aggregate neighbor information, with 4 attention heads and a hidden layer dimension of 128.

[0081] Double-Layer Policy Gradient Algorithm:

[0082] Inner Layer Optimization: Fix the parameters of the Critic network and update the Actor network through the deterministic policy gradient (DPG). Gradient calculation formula: \(\nabla_{\theta}J\approx E[\nabla_{\theta}Q(s,a)| \{a = \mu(s)\}]\), learning rate is set to \(1e - 4\), and for the Adam optimizer, \(\beta_1 = 0.9\), \(\beta_2 = 0.999\).

[0083] Outer layer equilibrium constraint: Introduce the Nash equilibrium constraint condition, requiring that the marginal revenue difference of each microgrid is less than the threshold ε = 0.05. Add the constraint to the loss function through the Lagrange multiplier method, with the initial value of the multiplier λ = 0.1, which doubles every 10 rounds of training until the constraint is satisfied.

[0084] Adversarial training mechanism: Set up a virtual opponent policy network to generate adversarial bid perturbations (±5%) in each round of training, forcing the main network to improve its robustness. The frequency of generating adversarial samples is once per batch (batch size = 64).

[0085] Initial policy set generation:

[0086] Policy sampling: After the training converges, sample 1000 groups of policies (charge and discharge power + bid price) from the policy network to form the initial set. For example, a certain policy is {Microgrid A discharges 0.8 MW and bids 0.52 yuan / kWh, Microgrid B charges 1.2 MW and bids 0.48 yuan / kWh}.

[0087] Nash equilibrium verification: Calculate the Nash convergence degree index δ = Σ|revenue_i - maximum possible revenue_i| / N of each policy, and retain the policies with δ < 0.1. For example, if the δ of a certain policy is 0.07, it indicates that it is close to the equilibrium state.

[0088] In S2014, perform Pareto front screening on the initial policy set, eliminate the policy combinations that violate the constraints of the safe and stable operation of the power grid, and output the initial policy set of the dynamic game.

[0089] This step ensures that the policy set takes into account both economy and grid security through multi-objective optimization and security constraint filtering.

[0090] Pareto front screening:

[0091] Optimization objective definition:

[0092] Objective 1: Minimize the total operating cost (including power purchase cost and equipment loss);

[0093] Objective 2: Maximize the grid regulation contribution degree (calculated according to the response speed and regulation amount);

[0094] Objective 3: Optimize the revenue fairness (minimize the Gini coefficient).

[0095] Application of NSGA-III algorithm:

[0096] Reference point generation: Use the Das-Dennis method to generate 100 uniform reference points in the three-dimensional objective space;

[0097] Non-dominated sorting: Sort 1,000 initial strategies and retain the top 50 Pareto optimal solutions;

[0098] Diversity maintenance: Calculate the correlation degree between the strategies and the reference point to ensure a uniform distribution of the frontier.

[0099] Power grid security constraint handling:

[0100] Power flow constraint verification: Call the Matpower toolkit to perform fast DC power flow calculations and verify whether the node voltages are within the range of 0.95 - 1.05 p.u. For example, if a certain strategy causes the node voltage to drop to 0.93 p.u., it is marked as a violation.

[0101] Energy storage constraint filtering: Eliminate strategies whose charge and discharge rates exceed the rated value of the equipment (such as the maximum 2C rate for lithium batteries). For example, if a certain strategy requires the microgrid C to discharge at 2.5C, the filtering mechanism is triggered.

[0102] N - 1 security criterion: When simulating any single line fault, verify whether the remaining strategies still meet the load demand. Use parallel computing (MPI multi - process) to accelerate the verification, and the fault scenario library contains 12 preset topologies.

[0103] Dynamic strategy set generation:

[0104] After screening by NSGA - III and filtering by security constraints, the number of initial strategies is reduced from 1,000 to 32;

[0105] Each strategy is labeled with multi - dimensional attribute tags, for example:

[0106] {

[0107] "strategy_id": "S23",

[0108] "cost": 45670, / / Unit: yuan

[0109] "contribution": 0.82, / / Normalized score

[0110] "gini": 0.15,

[0111] "violation_flag": false

[0112] }。

[0113] Build a strategy retrieval index to support fast query according to conditions such as cost range and contribution degree threshold.

[0114] In this step, by integrating the real-time status of energy storage devices (such as state of charge, charge-discharge efficiency) and the dynamic demands of the power grid (such as load forecasting, electricity price fluctuations), a game feature space in the spatio-temporal dimension is constructed, and the multi-agent strategy network is used to simulate the interest game behaviors of each microgrid. Specifically, the spatio-temporal correlation features between microgrids (such as the charge-discharge dependence relationship between adjacent microgrids) are extracted through a convolutional recurrent hybrid neural network, and an initial strategy set that satisfies the Nash equilibrium is generated based on the double-layer policy gradient algorithm, ensuring that the strategy maximizes the charge-discharge flexibility and bidding competitiveness of each microgrid on the premise of meeting the power grid security constraints, solving the strategy conflict problem caused by spatio-temporal differences in multi-microgrid collaborative scheduling, unifying the decentralized microgrid behaviors into a globally optimized initial strategy set through game modeling, providing a basic support for subsequent interest distribution and dynamic bidding, and avoiding the risks of uneven resource allocation or local overload.

[0115] S202. Based on the initial dynamic game strategy set, an asymmetric Nash bargaining model is used for distributed negotiation, and the weight factors of each microgrid are dynamically adjusted through a virtual resource exchange market to generate an equilibrium interest distribution plan, where the weight factors are jointly calculated by the user credit rating and the power grid emergency level; specifically, it may include:

[0116] S2021. Analyze the charge-discharge amount and bidding combination in the initial strategy set, initialize the threat point and bargaining space of the asymmetric Nash bargaining model, and set the threat point as the minimum revenue threshold when the microgrid operates independently;

[0117] In this step, by analyzing the strategy data of multiple microgrids, the bargaining starting point and feasible space of the game are defined, laying a foundation for subsequent distributed negotiation.

[0118] ‌ Strategy data analysis:

[0119] ‌ Charge-discharge amount analysis: Extract the charge-discharge amount strategies of each microgrid from the initial dynamic game strategy set (in JSON format). For example, the strategy of microgrid A is "discharge 0.8 MW (14:00 - 15:00)" and "charge 1.2 MW (18:00 - 19:00)", with a power resolution of 0.1 MW.

[0120] ‌ Bidding combination analysis: Analyze the price range in the bidding strategy. For example, the bid of microgrid B is "discharge price from RMB 0.52 yuan / kWh (minimum price limit) to RMB 0.68 yuan / kWh (maximum price limit)", with a time granularity of 15 minutes, aligned with the real-time power market trading platform.

[0121] ‌ Data verification: Check the matching of the charge-discharge amount and the energy storage capacity. If a certain strategy requires microgrid C to discharge 2 MW when the SOC = 20% (exceeding its maximum discharge rate of 1.5C), it is marked as an invalid strategy and excluded.

[0122] Threat point initialization:

[0123] Minimum revenue threshold calculation: Based on the historical data during the independent operation of the microgrid (such as the revenue quantile in the past 30 days), set the threat point as the 10% quantile of the revenue distribution. For example, when the microgrid D operates independently, the daily average revenue is between 1200 yuan and 1800 yuan, and the 10% quantile is 950 yuan, so the threat point is set at 950 yuan.

[0124] Bargaining space construction: Define the feasible bargaining region as the difference between the total revenue of all microgrids during joint operation and the total revenue during independent operation. For example, if the total independent revenue of three microgrids is 3000 yuan and the predicted revenue during joint operation is 4200 yuan, then the bargaining space is 1200 yuan.

[0125] Constraint condition encoding: Encode the power grid security constraints (such as node voltage deviation <5%) and energy storage life constraints (charge-discharge cycle times <5000 times) as linear inequalities and write them into the constraint matrix of the bargaining model.

[0126] In S2022, according to the historical compliance records of microgrid users and the probability of grid node voltage overlimit, calculate the credit score and emergency coefficient respectively, and generate the initial weight factor by fusing the two types of parameters through the hyperbolic tangent function;

[0127] This step dynamically adjusts the microgrid's right to speak in bargaining by quantifying user credit and grid emergency level, ensuring the fairness of benefit distribution and the stability of the power grid.

[0128] Credit score calculation:

[0129] Historical compliance record analysis: Extract the execution records of dispatching instructions of the microgrid in the past 90 days, and define the compliance rate λ = actual completion volume / committed volume. For example, microgrid E fulfilled 28 times out of 30 dispatches, so λ = 93.3%.

[0130] Time decay weighting: Use the exponential decay model to assign higher weights to recent compliance. The decay factor α = 0.9, and calculate the weighted compliance rate: λ_weighted = Σ(α^{t}·λ_t) / Σα^{t}, where t is the historical number of days and λ_t is the compliance rate in the historical t days. For example, for microgrid F with a compliance rate of 100% in the past 7 days, λ_weighted = 98.6%.

[0131] Credit score mapping: Map λ_weighted to 0 - 100 points through a piecewise linear function: when λ≥95%, get 100 points; when 80%≤λ<95%, get (λ - 80) / 0.15×30 + 70 points; when λ<80%, directly get 0 points.

[0132] Emergency coefficient calculation:

[0133] Voltage Out-of-Limit Probability Prediction: Based on the real-time state estimation of the power grid (SCADA data), a random forest model is used to predict the voltage out-of-limit probability P_voltage of each node within the next 1 hour. The input features include load factor, reactive power compensation, power exchange volume with adjacent microgrids, etc. For example, the current load factor of node G is 85%, and the predicted P_voltage = 23%.

[0134] Emergency Degree Classification: The emergency level is classified according to P_voltage: P≥30% is "Red Emergency" (coefficient 1.0), 20%≤P<30% is "Orange Emergency" (coefficient 0.7), and P<20% is "Normal" (coefficient 0.3).

[0135] Regional Aggregation: If a microgrid is connected to multiple nodes, the maximum value is taken as its emergency degree coefficient. For example, microgrid H is connected to node J (P = 25%) and node K (P = 18%), then the coefficient is 0.7.

[0136] Weight Factor Fusion:

[0137] Application of Hyperbolic Tangent Function: The credit score S (0 - 100) and the emergency degree coefficient E (0 - 1) are normalized and then input into the tanh function: W_initial = tanh(0.1S + 2E). For example, when S = 80 and E = 0.7, W_initial≈0.86.

[0138] Weight Normalization: Softmax processing is performed on W_initial of all microgrids within the alliance to ensure ΣW_i = 1. For example, the W_initial of three microgrids are 0.86, 0.64, and 0.50 respectively, and the normalized weights are 0.43, 0.32, and 0.25.

[0139] In S2023, a distributed negotiation protocol is deployed in the virtual resource exchange market for distributed negotiation. Each microgrid calculates the interest utility of the strategy set based on the current weight factor, and iteratively updates the bargaining solution through the alternating direction multiplier method;

[0140] This step realizes decentralized interest negotiation through a distributed optimization algorithm, avoiding single-point failures of central nodes and protecting data privacy.

[0141] Distributed Negotiation Protocol Design:

[0142] Communication Topology Definition: The Gossip protocol is used to construct a P2P network, and each microgrid directly communicates only with 3 - 5 geographically adjacent microgrids. For example, microgrid J forms a neighbor group with microgrids K, L, and M, and the communication delay < 50ms.

[0143] Message Format Standardization: Define the JSON Schema for negotiation messages, including fields:

[0144] {

[0145] "timestamp": "2023-08-20T14:30:00Z",

[0146] "microgrid_id": "MG-005",

[0147] "proposed_strategy": {"discharge": 0.6, "price": 0.55},

[0148] "weight": 0.35,

[0149] "utility_value": 1280.5

[0150] }。

[0151] Benefit Utility Calculation: Each microgrid calculates the strategy utility U_i = W_i × (Revenue_i - Threat Point_i) + (1 - W_i) × Σ(Revenue_j - Threat Point_j) according to the weight W_i. For example, for microgrid N, W = 0.4, Revenue is 1300 yuan, Threat Point is 900 yuan, and the total revenue difference of neighbors is 2000 yuan, then U_N = 0.4 × 400 + 0.6 × 2000 = 160 + 1200 = 1360.

[0152] Alternating Direction Method of Multipliers (ADMM) Iteration:

[0153] Variable Splitting: Decompose the global optimization problem into local sub-problems for each microgrid, and define the consistency constraint variable z_ij to represent the consensus value of microgrid i and j on the common variable. For example, microgrids O and P need to reach an agreement on the power exchange volume of the tie line, and z_OP is initially set to the average of their proposals.

[0154] Parameter Setting: Penalty factor ρ = 1.2 (controls the convergence speed), maximum number of iterations T = 100, convergence threshold ε = 0.01 (the change rate of the objective function between two iterations < 1%).

[0155] In S2024, when the difference in Shapley values between two adjacent iterations is less than the preset threshold, trigger the dynamic adjustment mechanism of the weight factor. Among them, the weight of the microgrid with a decreased credit rating decays exponentially, and the weight of the microgrid in the emergency area of the power grid increases linearly;

[0156] This step responds to credit risks and grid emergency states through the dynamic weight adjustment mechanism to achieve adaptive optimization of benefit distribution.

[0157] Shapley Value Monitoring:

[0158] Contribution Calculation: For each microgrid \(i\), calculate its Shapley value \(\varphi_i=\sum_{S\subseteq N\setminus\{i\}}\frac{|S|!(|N|-|S|-1)!}{|N|!}(v(S\cup\{i\}) - v(S))\), where \(v(S)\) is the total revenue of coalition \(S\). For example, for 3 microgrids, the \(\varphi_i\) needs to calculate \(2^2 = 4\) subset combinations.

[0159] Difference Evaluation: Define the difference \(\Delta=\max\frac{|\varphi_i^{k}-\varphi_i^{k - 1}|}{\varphi_i^{k - 1}}\), and the preset threshold \(\Delta_{threshold}=5\%\). When \(\Delta < 5\%\), it is considered that the coalition structure tends to be stable and the weight adjustment is triggered.

[0160] Weight Dynamic Adjustment Rules:

[0161] Credit Decrease Penalty: If the fulfillment rate of a certain microgrid drops by more than 10% in the past 24 hours, its weight is calculated as \(W_i^{new}=W_i^{old}\times e^{-\beta t}\), where \(\beta = 0.1\) (decay rate) and \(t\) is the number of hours of continuous credit decrease. For example, if the fulfillment rate of microgrid \(R\) drops by 15% for 3 consecutive hours, the weight drops from 0.3 to \(0.3\times e^{-0.1\times3}=0.222\).

[0162] Emergency Area Reward: For microgrids in areas where the probability of voltage violation \(> 25\%\), the weight increases linearly: \(W_i^{new}=W_i^{old}+\gamma\times(P_{voltage}-0.25)\), where \(\gamma = 0.5\) (adjustment coefficient). For example, if \(P_{voltage}=30\%\) in the area where microgrid \(S\) is located, the weight increases by \(0.5\times(0.3 - 0.25)=0.025\).

[0163] Weight Rebalancing: After adjustment, re - normalization is required to ensure \(\sum W_i = 1\). For example, the original weights \([0.3, 0.4, 0.3]\) become \([0.25, 0.45, 0.3]\) after adjustment, and after normalization, they are \([0.25 / 1.0, 0.45 / 1.0, 0.3 / 1.0]\).

[0164] In S2025, generate an equilibrium benefit distribution plan containing the transfer payment plan and the charging and discharging time sequence plan according to the final negotiation result, and verify the individual rationality and coalition stability of the plan through zero - knowledge proof technology.

[0165] This step transforms the negotiation result into an executable plan and ensures the fairness and immutability of the plan through cryptographic technology.

[0166] ‌Transfer payment plan generation‌:

[0167] ‌Revenue Difference Calculation‌: For each microgrid i, calculate the difference between the actual revenue π_i and the Shapley value φ_i, Δ_i = π_i - φ_i. If Δ_i > 0, Δ_i must be paid to other members; if Δ_i < 0, |Δ_i| compensation is received. For example, if microgrid V receives φ_V = RMB 1500 and the actual revenue π_V = RMB 1400, it must receive RMB 100 in compensation.

[0168] Payment path optimization: Dijkstra's algorithm is used to find the minimum transaction path. For example, if microgrid W needs to pay RMB 200 to microgrid X, and microgrid Y needs to pay RMB 150 to X, the paths are merged into a W→Y→X path to reduce the number of transactions.

[0169] Smart Contract Deployment: The payment plan is written into a blockchain smart contract, which is triggered to execute automatically after the microgrid completes the dispatch instruction. For example, after microgrid Z completes 1.5MW of discharge at 3:00 PM, the on-chain contract transfers RMB 320 to its address.

[0170] ‌Charge and discharge timing plan‌:

[0171] Time Window Optimization: Based on the grid load forecast curve, discharge is prioritized during peak electricity price periods (e.g., 2:00 PM to 4:00 PM) and charging is scheduled during off-peak periods (e.g., 2:00 AM to 4:00 AM). For example, Microgrid α's discharge schedule is 0.9 MW from 2:30 PM to 3:30 PM, and its charging schedule is 1.1 MW from 3:00 AM to 4:00 AM.

[0172] Energy storage life balancing: Limit the number of charge-discharge cycles per storage unit to ≤ 1 within three consecutive time slices to avoid frequent switching that can degrade battery life. For example, after completing a charge-discharge cycle between 10:00 AM and 11:30 AM, the energy storage system in microgrid β must rest for at least 45 minutes.

[0173] ‌Zero-knowledge proof verification‌:

[0174] Individual Rational Verification: Using zk-SNARKs, a proof π_IR is generated, showing that for any microgrid i, π_i ≥ threat_i, without revealing the specific value of the payoff. For example, the verifier only needs to confirm the validity of π_IR and does not need to know the actual payoff of microgrid γ, which is RMB 1,650.

[0175] Alliance Stability Verification: A Merkle tree is constructed to store the Shapley values of each microgrid. The integrity of φ_i is verified through the root hash to ensure that no member can arbitrarily modify the distribution ratio. For example, an attacker cannot forge the φ value of microgrid δ because its hash value is inconsistent with that of other nodes.

[0176] Taking the initial strategy set as the input, the bargaining power (weight factor) of each microgrid is defined through the asymmetric Nash bargaining model. The weight factor combines the user credit score (such as the historical performance rate) and the grid emergency level (such as the probability of node voltage violation). In the virtual resource exchange market, the alternating direction multiplier method is used to iteratively optimize the benefit distribution scheme. The weight of the microgrid with a low credit rating decays, and the weight of the microgrid in the emergency area increases. Finally, the fairness and stability of the scheme are verified through zero-knowledge proof, ensuring that the benefit distribution not only satisfies individual rationality (the revenue of a single microgrid is not lower than the independent operation threshold) but also guarantees the stability of the coalition (optimal global resource utilization), balances the interest conflicts among multiple microgrids, and preferentially guarantees the grid emergency needs and the rights and interests of high-credit users through the dynamic weight adjustment mechanism, improving the fairness and enforceability of collaborative scheduling and avoiding the "free-rider" or "resource occupation" problems in traditional centralized scheduling.

[0177] S203. According to the equilibrium benefit distribution scheme, use the time-decay type reinforcement learning algorithm to iteratively correct the energy storage priority index and generate dynamic bidding rules including the supply-demand elasticity coefficient and risk compensation; specifically, it may include:

[0178] S2031. Extract the historical winning rate and revenue fluctuation data of each microgrid from the equilibrium benefit distribution scheme, and initialize the energy storage priority index as a weighted function of the supply-demand ratio and risk value;

[0179] In this step, by mining historical transaction data and revenue characteristics, an energy storage priority evaluation system is constructed to provide a quantitative basis for dynamic bidding rules.

[0180] ‌Data extraction and cleaning‌:

[0181] ‌Calculation of historical winning rate‌: Query the bidding records of each microgrid in the past 30 trading periods from the blockchain database and count the proportion of the number of winning bids to the total number of bids. For example, microgrid A participates in the day-ahead market bidding 20 times and wins 12 times. The historical winning rate = 12 / 20 = 60%.

[0182] ‌Quantification of revenue volatility‌: Extract the actual revenue data (unit: yuan / MWh) of each trading period and calculate the ratio of the standard deviation to the mean as the volatility. For example, the peak-time revenue sequence of microgrid B is [520 yuan, 580 yuan, 490 yuan], the mean is 530 yuan, and the standard deviation is 45 yuan. The volatility = 45 / 530 ≈ 8.5%.

[0183] Outlier handling: The Tukey's fences method is used to eliminate extreme values. If the revenue in a certain period exceeds Q3 + 1.5IQR (Q3 is the third quartile and IQR is the interquartile range), it is marked as an outlier. For example, in the revenue sequence of Microgrid C, 850 yuan exceeds Q3 (620 yuan) + 1.5×IQR (200), so it is regarded as an outlier and replaced with the mean value of adjacent periods.

[0184] Supply-demand ratio and risk value calculation:

[0185] Real-time supply-demand ratio calculation: The supply-demand ratio is defined as the ratio of the total regional load (MW) to the total available energy storage capacity (MWh) in the current period. For example, at 14:00 in a certain region, the load is 150 MW and the available energy storage capacity is 120 MWh (SOC≥30%), then the supply-demand ratio = 150 / 120 = 1.25.

[0186] Risk value assessment: Based on the CVaR (Conditional Value at Risk) model, the expected loss at the tail of the revenue distribution (such as the 5% quantile) is calculated. For example, in the revenue distribution of Microgrid D, the average revenue in the worst 5% scenarios is 360 yuan, so the risk value = 360 yuan.

[0187] Priority index initialization: The supply-demand ratio (S) and the risk value (R) are normalized and weighted. The weight coefficients are α = 0.6 (emphasizing supply-demand balance) and β = 0.4 (emphasizing risk aversion): Priority = α×(S / S_max) + β×(1 - R / R_max). Here, S_max is the historical maximum supply-demand ratio (such as 2.0), and R_max is the maximum risk value (such as 500 yuan). For example, if S = 1.25 and R = 360 yuan, then Priority = 0.6×(1.25 / 2.0)+0.4×(1 - 360 / 500)=0.487.

[0188] S2032, construct a state transition model for time-decaying reinforcement learning. Define the state as the difference between the real-time grid load and the available energy storage capacity, the action as the adjustment range of the priority, and the reward function includes the revenue from electricity price arbitrage and the contribution degree of grid regulation;

[0189] This step dynamically adapts to market environment changes through reinforcement learning, optimizing the real-time performance and robustness of the energy storage scheduling strategy.

[0190] State space modeling:

[0191] Definition of state variables: The state s_t = real-time load L_t (MW) - available energy storage capacity C_t (MWh). For example, at 15:00, the load is 160 MW and the energy storage capacity is 130 MWh, then s_t = 160 - 130 = 30.

[0192] State discretization: The continuous state is divided into 10 intervals, such as [-∞, -50), [-50, -30), [-30, -10), [-10,10), [10,30), [30,50), [50,70), [70,90), [90,110), [110, +∞]. Each interval corresponds to a discrete state code (e.g., s = 30 corresponds to code 5).

[0193] State transition probability: Based on historical data, the state transition law is statistically analyzed. For example, when s_t = 30 (code 5), the probability that the next state transitions to s_{t+1} = 50 (code 6) is 40%, and the probability of transitioning to s_{t+1} = 10 (code 4) is 35%.

[0194] Action space design:

[0195] Definition of adjustment amplitude: The action a ∈ {-0.2, -0.1, 0, +0.1, +0.2}, which represents the adjustment amount of the priority index. For example, when the current priority is 0.5, after executing the action +0.1, it becomes 0.6.

[0196] Action constraint conditions: Set the maximum adjustment step size to prevent sudden changes in priority. For example, the single - time adjustment amplitude shall not exceed ±0.2, and after 3 consecutive same - direction adjustments, a forced reverse adjustment is required once.

[0197] Reward function design:

[0198] Profit from electricity price arbitrage: Calculate the arbitrage profit R_arbitrage of the micro - grid in period t = (discharge electricity price - charging electricity price) × actual transaction volume. For example, the discharge price is 0.68 yuan / kWh, the charging price is 0.42 yuan / kWh, and the transaction volume is 10 MWh, then R_arbitrage=(0.68 - 0.42)×10000 = 2600 yuan.

[0199] Grid regulation contribution degree Contribution: Calculate the contribution degree according to the regulation amount (unit: MW) of the micro - grid to the grid frequency deviation. For example, if the micro - grid G provides 5 MW of frequency regulation service, the contribution degree = 5 / total regional frequency regulation demand (20 MW)=25%.

[0200] Comprehensive reward calculation: Weighted aggregation of two types of rewards, with weights γ = 0.7 (arbitrage profit) and δ = 0.3 (regulation contribution): Reward = γ×R_arbitrage / R1_max + δ×Contribution.

[0201] Among them, R1_max is the historical maximum arbitrage profit (such as RMB 5,000). For example, if R_arbitrage = RMB 2,600 and Contribution = 25%, then Reward = 0.7×2,600 / 5,000 + 0.3×25% = 0.439.

[0202] S2033, adopt a decay factor to dynamically adjust the exploration rate of the Q-learning algorithm, enhance the weight of historical experience during peak load periods, increase the probability of random exploration during valley load periods, and update the probability distribution of energy storage priority indicators;

[0203] This step balances experience utilization and unknown exploration through an adaptive learning strategy, improving the decision-making efficiency of the model during different load periods.

[0204] ‌Decay factor mechanism‌:

[0205] ‌Time decay function‌: Define the decay factor λ(t)=λ_base × e^(-kt), where λ_base = 0.9 (base decay rate), k = 0.05 (decay speed coefficient), and t is the number of consecutive explorations. For example, after a certain strategy has not been selected for 3 consecutive times, λ(3)=0.9×e^(-0.05×3)≈0.9×0.861 = 0.775.

[0206] ‌Load period division‌: Define peak periods (such as 10:00 - 12:00, 18:00 - 20:00), flat periods (7: O0 - 10:00, 12:00 - 18:00), and valley periods (0:00 - 7:00, 20:00 - 24:00) according to the historical load curve.

[0207] ‌Dynamic adjustment of exploration rate‌:

[0208] ‌Peak period strategy‌: When the load ≥ 120% of the regional average load, set the exploration rate ε = 0.1 (10% random exploration), and preferentially select the action with the highest Q value. For example, the probability that Microgrid I selects the action +0.2 according to the Q-table during peak periods is 90%, and the probability of randomly exploring other actions is 10%.

[0209] ‌Valley period strategy‌: When the load ≤ 80% of the regional average load, increase the exploration rate to ε = 0.4 to encourage the discovery of new strategies. For example, Microgrid J has a 40% probability of randomly trying the actions -0.1 or 0 at 2:00 am to test the feasibility of low-price charging.

[0210] ‌Flat period transition strategy‌: Adjust the exploration rate using linear interpolation. For example, when the load is 100% of the average load, ε = 0.25.

[0211] ‌Update of priority probability distribution‌:

[0212] Q - value update rule: Q(s,a) ← Q(s,a) + α×[R + γ×max Q(s',a') - Q(s,a)], where α = 0.2 (learning rate) and γ = 0.9 (discount factor). For example, the original Q - value of action +0.1 in state s = 30 (encoded as 5) is 1.2, the new reward R = 0.5, and the maximum Q - value of the next state = 1.5. Then the updated Q = 1.2+0.2×(0.5 + 0.9×1.5 - 1.2)=1.2+0.2×(0.5 + 1.35 - 1.2)=1.2+0.2×0.65 = 1.33.

[0213] Probability distribution generation: For each state s, calculate the Softmax probability for all actions a: P(a|s)=e^(Q(s,a) / τ) / Σe^(Q(s,a') / τ), where τ = 0.5 (temperature coefficient). For example, for three actions with Q - values of 1.2, 1.0, and 0.8 respectively, their probabilities are e^(1.2 / 0.5)=e^2.4≈11.02, e^2≈7.39, e^1.6≈4.95, the sum is approximately 23.36, and the normalized probabilities are 47.1%, 31.6%, and 21.3% respectively.

[0214] S2034. Calculate the ratio of the price change rate to the volume change rate as the supply - demand elasticity coefficient based on the updated priority index and its probability distribution, and quantify the risk compensation value for energy storage invocation based on the conditional value - at - risk model;

[0215] This step realizes the dynamic balance of market supply and demand and the reasonable sharing of risks through the elasticity coefficient and the risk compensation mechanism.

[0216] Calculation of supply - demand elasticity coefficient:

[0217] Calculation of price change rate: ΔPrice=(Average electricity price this period - Average electricity price last period) / Average electricity price last period. For example, the average electricity price of micro - grid L at time t is 0.58 yuan / kWh, and at time t + 1 is 0.63 yuan / kWh. Then ΔPrice=(0.63 - 0.58) / 0.58≈8.62%.

[0218] Calculation of volume change rate: ΔVolume=(Volume this period - Volume last period) / Volume last period. For example, the volume traded at time t is 10 MWh, and at time t + 1 is 13 MWh. Then ΔVolume=(13 - 10) / 10 = 30%.

[0219] Elasticity coefficient calculation: E = ΔVolume / ΔPrice. For example, E = 30% / 8.62% ≈ 3.48, which means that for every 1% increase in price, the trading volume increases by 3.48%, indicating a relatively high demand elasticity.

[0220] Risk compensation value quantification:

[0221] CVaR model application: Set the confidence level β = 95%, and calculate the conditional expected value of the returns distribution below VaR (Value at Risk). For example, in the returns distribution of microgrid M, the 5th percentile corresponds to VaR = RMB 280, and CVaR is the average of all returns below RMB 280 (such as RMB 250).

[0222] Risk compensation calculation: Compensation value = CVaR × risk aversion coefficient η, where η = 0.3 (set by user preference). For example, if CVaR = RMB 250, then the compensation value = 250 × 0.3 = RMB 75 / MWh.

[0223] Dynamic adjustment mechanism: When the actual returns of the microgrid are lower than CVaR for two consecutive periods, η is increased to 0.5; conversely, if the returns are higher than CVaR for three consecutive periods, η is decreased to 0.2.

[0224] In S2035, encode the supply-demand elasticity coefficient and the energy storage call risk compensation value as a piecewise linear function, generate a dynamic bidding rule including the step price ceiling and penalty clauses, and write it into the consensus verification contract of the blockchain.

[0225] This step ensures the transparency and immutability of the bidding process through structured rule design and blockchain technology.

[0226] Piecewise linear function design:

[0227] Elasticity coefficient segmentation: Divide it into three grades according to the elasticity value E:

[0228] E ≥ 2.0 (high elasticity): Price ceiling = benchmark price × 1.2 + risk compensation;

[0229] 1.0 ≤ E < 2.0 (medium elasticity): Price ceiling = benchmark price × 1.0 + risk compensation;

[0230] E < 1.0 (low elasticity): Price ceiling = benchmark price × 0.8 + risk compensation.

[0231] For example, the benchmark price is RMB 0.55 / kWh, E = 2.5, and the risk compensation is RMB 0.10 / kWh, then the price ceiling = 0.55 × 1.2 + 0.10 = RMB 0.76 / kWh.

[0232] Penalty Clause Setting: For microgrids with actual transaction volume lower than 80% of the committed volume, the margin will be deducted according to the difference ratio. For example, if the committed discharge is 10 MWh and the actual completion is 8 MWh, with a difference of 20%, then the deducted margin = 20% × 10 MWh × upper limit of the quoted price × 0.5.

[0233] Blockchain Contract Deployment:

[0234] Smart Contract Logic: Write a Solidity contract to achieve the following functions:

[0235] Verify whether the quoted price complies with the stepped upper limit rule;

[0236] Automatically calculate the penalty amount according to the transaction result;

[0237] Call the off-chain oracle to obtain real-time electricity price and load data.

[0238] For example, when Microgrid O quotes 0.75 yuan / kWh (exceeding the upper limit of 0.72 yuan / kWh corresponding to its elasticity coefficient E = 1.8), the contract automatically rejects the quote.

[0239] Consensus Verification Process: Each node reaches a consensus on the bidding rules through the PBFT (Practical Byzantine Fault Tolerance) protocol, and synchronizes the rule hash value every 15 minutes to ensure network-wide consistency.

[0240] Based on the balanced interest distribution result, use the time-decaying reinforcement learning algorithm to dynamically adjust the energy storage priority index. Strengthen historical experience (such as high arbitrage profit strategies) during peak load periods, increase random exploration (such as risk diversification strategies) during low load periods, combine the supply-demand elasticity coefficient (the ratio of the price change rate to the transaction volume change rate) and the conditional value-at-risk model to quantify the energy storage call risk (such as the default probability caused by insufficient capacity), and finally generate stepped bidding rules and risk compensation clauses, and write the rules into the blockchain contract to ensure transparency and immutability. Adapt to the real-time supply-demand changes of the power grid through dynamic bidding rules, use reinforcement learning to balance short-term benefits and long-term risks, improve the market response speed and economy of energy storage resources, and at the same time reduce the impact of uncertainty in the dispatching process through the risk compensation mechanism.

[0241] S204, based on the dynamic bidding rules, allocate energy storage resources in real time through the decentralized gradient consensus mechanism, verify the allocation effectiveness by combining verifiable random functions, and generate the final coordinated dispatching instructions and synchronize them to each microgrid terminal. Specifically, it can include:

[0242] S2041, Broadcast the dynamic bidding rules in the decentralized network, each microgrid node generates a candidate allocation plan based on local constraint conditions, and calculates the gradient vector of the plan for the global objective function;

[0243] The dynamic bidding rules are broadcast to all microgrid nodes participating in collaborative scheduling (such as 10 microgrids) through the blockchain network. After each node receives the rules, it generates a candidate allocation plan based on the parameters of local energy storage devices (such as state of charge SOC = 65%, maximum charge and discharge power 50kW) and the real-time demand of the power grid (such as the need to increase energy storage discharge by 200kWh in the region). The candidate plan needs to meet the following constraints:

[0244] Charge and discharge power constraint: not exceeding ± the discharge power of the device rating (such as ± power);

[0245] SOC safety constraint: the SOC remains between 20% and 80% after charge and discharge;

[0246] Quotation constraint: according to the step quotation upper limit in the dynamic bidding rules (such as during peak hours ≤ price constraint: yuan / kWh).

[0247] The gradient vector calculation adopts a distributed optimization framework:

[0248] Global objective function: minimize the grid regulation cost + maximize the total revenue of the microgrid alliance;

[0249] Local objective function: each microgrid constructs a revenue function according to its own quotation and SOC status (such as revenue = discharge amount × quotation - energy storage loss cost).

[0250] The gradient of the candidate plan with respect to the global objective function is calculated by automatic differentiation technology (such as Autograd in PyTorch). For example, a certain microgrid node calculates that if the discharge amount is increased from 30kW to 40kW, the global cost gradient is -0.2 (a negative gradient indicates a cost reduction), then a gradient vector [-0.2, 0.1,...] (the dimension is the same as the number of variables) is generated.

[0251] Example: The candidate plan of Microgrid A is to discharge 35kW and the quotation is 0.12 yuan / kWh. Calculate its gradient with respect to the global objective as [-0.15, 0.08], indicating that increasing the discharge amount can reduce the global cost, but increasing the quotation will slightly increase the cost.

[0252] S2042, the gradient consensus algorithm is used to aggregate the gradient information of each node, and the optimal allocation solution that satisfies the power grid power flow constraint is solved by the projected subgradient descent method to generate a preliminary collaborative scheduling instruction;

[0253] The gradient consensus algorithm realizes decentralized aggregation through the Gossip protocol:

[0254] Gradient exchange: each node randomly selects 3 neighbor nodes (such as Microgrid B, C, D) and sends its local gradient vector;

[0255] Weighted average: After a node receives the gradients of its neighbors, it calculates the average gradient according to the weights (e.g., its own weight is 0.6, and each neighbor's weight is 0.2).

[0256] Iterative convergence: After repeating 10 rounds, the gradient difference rate of each node is < 1% (regarded as convergence).

[0257] The projected subgradient descent method is used to handle the power flow constraints of the power grid:

[0258] Variable update: Update the charge and discharge amount and the quotation according to the aggregated gradient, and the step size is α;

[0259] Constraint projection: If the updated variable exceeds the boundary (e.g., SOC > 80%), project it into the feasible region (e.g., force SOC = 79%);

[0260] Power grid power flow verification: Call the Matpower toolbox for power flow calculation to ensure that the line load rate is < 95%.

[0261] Example: After 5 iterations, each node reaches a consensus. The optimal solution is the total discharge amount of 180 kW (Microgrid A: 40 kW, Microgrid B: 35 kW...), the average quotation is 0.13 yuan / kWh, and the power flow load rate is 92%. Generate a preliminary instruction: "Microgrid A discharges 40 kW from 14:00 to 14:15, and the quotation is 0.13 yuan".

[0262] S2043, call the verifiable random function to generate distributed random number seeds. Each node performs a pre-submission operation based on the collaborative scheduling instruction and generates an operation hash. After all hash values are aggregated by threshold signature, they are compared with the random numbers to verify the consistency;

[0263] The verifiable random function (VRF) adopts the ECVRF-ED25519-SHA512 algorithm:

[0264] Random number generation: Each node uses its private key to generate a random number seed for the current block height (e.g., Block#7821), and the output is a 512-bit hash value;

[0265] Seed aggregation: Generate a global random number through threshold signature (Threshold Signature, which requires signatures from more than 2 / 3 of the nodes), for example, the signature aggregation result is 0x3a7d...;

[0266] Hash pre-submission: Before each node executes the scheduling instruction, calculate the operation hash (e.g., SHA-3 hash the instruction content "Microgrid A discharges 40 kW"), and sign and submit it to the blockchain.

[0267] Consistency verification:

[0268] Hash comparison: The pre-committed hashes of all nodes are aggregated through a Merkle tree. The root hash is bitwise XORed with the random number generated by the VRF. If the results are the same, the instruction is considered valid.

[0269] Threshold verification: If the hashes of 7 out of 10 nodes are the same (pass rate 70%), but lower than the 95% threshold, elastic reallocation is triggered.

[0270] Example: Among 10 nodes, 9 submit the same hash (root hash 0x7f2c...), and the XOR result with the VRF random number 0x7f2c... is 0, so the verification passes.

[0271] S2044, when the verification pass rate exceeds the preset threshold, the collaborative scheduling instruction is written into the trusted execution environment of each microgrid terminal; otherwise, the elastic reallocation mechanism is triggered, and the steps of broadcasting the dynamic bidding rules in the decentralized network are returned. Each microgrid node generates a candidate allocation plan based on local constraint conditions and calculates the gradient vector of the plan for the global objective function to perform iterative optimization until convergence. The preset threshold is, for example, 95%.

[0272] Trusted Execution Environment (TEE) adopts Intel SGX technology to ensure the secure execution of instructions:

[0273] Instruction writing: The instruction is encrypted and transmitted to the SGX enclave of each microgrid through a secure channel;

[0274] Secure execution: The instruction is decrypted inside the enclave and the energy storage device is controlled to prevent malicious tampering (such as the charge and discharge amount being maliciously modified).

[0275] Elastic reallocation mechanism:

[0276] Trigger condition: The consistency rate of pre-committed hashes < 95% (e.g., 8 out of 10 nodes are consistent);

[0277] Cause analysis: Identify the inconsistent node (such as microgrid E submitting an abnormal hash), and reduce its weight factor by 50%;

[0278] Iterative optimization: Return to step S2041, adjust the gradient aggregation weights (normal node weight +10%, abnormal node weight -30%), and regenerate the candidate plan;

[0279] Termination condition: The pass rate in three consecutive iterations ≥ the iteration pass rate or the maximum number of iterations is 20 times.

[0280] Example: The first verification pass rate is 90%. After identifying microgrid E as an abnormal node and reducing its weight, the pass rate in the second iteration is increased to 96%, and the instruction is successfully written into the TEE and executed.

[0281] In a decentralized network, each microgrid node generates a candidate allocation plan based on dynamic bidding rules, aggregates global gradient information through a gradient consensus algorithm, and uses the projected subgradient descent method to solve the optimal solution that satisfies the power grid power flow constraints. A verifiable random function (VRF) is used to generate a random number seed to verify the consistency of the pre-allocation operation hash values submitted by each node (threshold signature aggregation comparison). If the verification pass rate meets the standard, it is written into the trusted execution environment for execution; otherwise, the elastic reallocation mechanism is triggered for iterative optimization to ensure that the allocation result satisfies economy, security, and verifiability simultaneously, realizing efficient resource allocation in a decentralized environment. Through the gradient consensus and VRF verification mechanisms, the global optimality and anti-tampering of scheduling instructions are guaranteed, the anti-attack ability and collaborative response efficiency of the multi-microgrid system are improved, and the real-time scheduling requirements in a complex power grid environment are met.

[0282] It can be seen that according to the state of charge, charge and discharge efficiency of the user's energy storage device, and the real-time scheduling requirements of the power grid, a dynamic game initial strategy set is generated through a multi-agent policy network; based on the dynamic game initial strategy set, an asymmetric Nash bargaining model is used for distributed negotiation to generate an equilibrium benefit allocation plan; according to the equilibrium benefit allocation plan, the time decay reinforcement learning algorithm is used to iteratively correct the energy storage priority index to generate dynamic bidding rules including supply-demand elasticity coefficients and risk compensation; based on the dynamic bidding rules, energy storage resources are allocated in real time through a decentralized gradient consensus mechanism, and the final collaborative scheduling instructions are generated and synchronized to each microgrid terminal, thereby improving the operation efficiency and adaptive ability of the multi-microgrid system and meeting the requirements of future intelligent distribution networks.

[0283] Another embodiment of the present invention provides a multi-microgrid collaborative scheduling system based on game theory. Refer to Figure 3 , the system may include:

[0284] A construction module 301, configured to construct a spatio-temporal coupled game feature space according to the state of charge, charge and discharge efficiency of the user's energy storage device, and the real-time scheduling requirements of the power grid, and generate a dynamic game initial strategy set through a multi-agent policy network, wherein the dynamic game initial strategy set includes the charge and discharge amounts and quotation combinations of each microgrid;

[0285] A negotiation module 302, configured to perform distributed negotiation based on the dynamic game initial strategy set, use an asymmetric Nash bargaining model to dynamically adjust the weight factors of each microgrid through a virtual resource exchange market, and generate an equilibrium benefit allocation plan, wherein the weight factors are jointly calculated by the user credit rating and the power grid emergency level;

[0286] A correction module 303, configured to iteratively correct the energy storage priority index according to the equilibrium benefit allocation plan by using a time decay reinforcement learning algorithm, and generate dynamic bidding rules including supply-demand elasticity coefficients and risk compensation;

[0287] An allocation module 304 is configured to allocate energy storage resources in real time based on the dynamic bidding rules through a decentralized gradient consensus mechanism, verify the effectiveness of the allocation by combining a verifiable random function, generate a final collaborative scheduling instruction, and synchronize it to each microgrid terminal.

[0288] It can be seen that according to the state of charge, charge and discharge efficiency of the user's energy storage device, and the real-time scheduling requirements of the power grid, a dynamic game initial strategy set is generated through a multi-agent policy network; based on the dynamic game initial strategy set, an asymmetric Nash bargaining model is used for distributed negotiation to generate an equilibrium benefit allocation plan; according to the equilibrium benefit allocation plan, a time-decaying reinforcement learning algorithm is used to iteratively correct the energy storage priority index to generate dynamic bidding rules including supply-demand elasticity coefficients and risk compensation; based on the dynamic bidding rules, energy storage resources are allocated in real time through a decentralized gradient consensus mechanism, a final collaborative scheduling instruction is generated and synchronized to each microgrid terminal, thereby improving the operation efficiency and adaptive ability of the multi-microgrid system and meeting the requirements of future intelligent distribution networks.

[0289] An embodiment of the present invention also provides a storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0290] Specifically, in this embodiment, the above storage medium may be configured to store a computer program for executing the following steps:

[0291] S201, according to the state of charge, charge and discharge efficiency of the user's energy storage device, and the real-time scheduling requirements of the power grid, construct a spatio-temporal coupled game feature space, and generate a dynamic game initial strategy set through a multi-agent policy network, where the dynamic game initial strategy set includes the charge and discharge amounts and bid combinations of each microgrid;

[0292] S202, based on the dynamic game initial strategy set, use an asymmetric Nash bargaining model for distributed negotiation, dynamically adjust the weight factors of each microgrid through a virtual resource exchange market, and generate an equilibrium benefit allocation plan, where the weight factors are jointly calculated by the user credit rating and the power grid emergency level;

[0293] S203, according to the equilibrium benefit allocation plan, use a time-decaying reinforcement learning algorithm to iteratively correct the energy storage priority index, and generate dynamic bidding rules including supply-demand elasticity coefficients and risk compensation;

[0294] S204, based on the dynamic bidding rules, allocate energy storage resources in real time through a decentralized gradient consensus mechanism, verify the effectiveness of the allocation by combining a verifiable random function, generate a final collaborative scheduling instruction, and synchronize it to each microgrid terminal.

[0295] It can be seen that, according to the state of charge, charge and discharge efficiency of the user's energy storage device, and the real-time scheduling requirements of the power grid, a dynamic game initial strategy set is generated through a multi-agent policy network; based on the dynamic game initial strategy set, an asymmetric Nash bargaining model is used for distributed negotiation to generate an equilibrium benefit distribution plan; according to the equilibrium benefit distribution plan, a time-decaying reinforcement learning algorithm is used to iteratively correct the energy storage priority index to generate a dynamic bidding rule including supply-demand elasticity coefficients and risk compensation; based on the dynamic bidding rule, energy storage resources are allocated in real time through a decentralized gradient consensus mechanism, and a final collaborative scheduling instruction is generated and synchronized to each microgrid terminal, thereby being able to improve the operation efficiency and adaptive ability of the multi-microgrid system and meet the requirements of future intelligent distribution grids.

[0296] An embodiment of the present invention also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0297] Specifically, the above electronic device may further include a transmission device and an input / output device. Among them, the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0298] Specifically, in this embodiment, the above processor may be configured to execute the following steps through a computer program:

[0299] S201, according to the state of charge, charge and discharge efficiency of the user's energy storage device, and the real-time scheduling requirements of the power grid, construct a spatio-temporal coupled game feature space, and generate a dynamic game initial strategy set through a multi-agent policy network, where the dynamic game initial strategy set includes the charge and discharge amounts and bid combinations of each microgrid;

[0300] S202, based on the dynamic game initial strategy set, use an asymmetric Nash bargaining model for distributed negotiation, and dynamically adjust the weight factors of each microgrid through a virtual resource exchange market to generate an equilibrium benefit distribution plan, where the weight factors are jointly calculated by the user credit rating and the power grid emergency level;

[0301] S203, according to the equilibrium benefit distribution plan, use a time-decaying reinforcement learning algorithm to iteratively correct the energy storage priority index to generate a dynamic bidding rule including supply-demand elasticity coefficients and risk compensation;

[0302] S204, based on the dynamic bidding rule, allocate energy storage resources in real time through a decentralized gradient consensus mechanism, verify the allocation effectiveness in combination with a verifiable random function, and generate a final collaborative scheduling instruction and synchronize it to each microgrid terminal.

[0303] It can be seen that according to the state of charge, charge-discharge efficiency of the user energy storage device and the real-time dispatching requirements of the power grid, an initial strategy set of dynamic game is generated through a multi-agent policy network; based on the initial strategy set of dynamic game, an asymmetric Nash bargaining model is used for distributed negotiation to generate an equilibrium benefit distribution scheme; according to the equilibrium benefit distribution scheme, a time-decaying reinforcement learning algorithm is used to iteratively correct the energy storage priority index to generate a dynamic bidding rule including the supply-demand elasticity coefficient and risk compensation; based on the dynamic bidding rule, the energy storage resources are allocated in real time through a decentralized gradient consensus mechanism to generate a final coordinated dispatching instruction and synchronize it to each microgrid terminal, so as to improve the operation efficiency and adaptive ability of the multi-microgrid system and meet the requirements of the future intelligent distribution network.

[0304] The structure, features and function effects of the present invention have been described in detail based on the embodiments shown in the drawings. The above is only the preferred embodiment of the present invention, but the present invention is not limited to the scope defined by the drawings. Any changes made according to the concept of the present invention, or equivalent embodiments modified to equivalent changes, still within the spirit covered by the specification and drawings, shall be within the protection scope of the present invention.

Claims

1. A multi - microgrid collaborative scheduling method based on game theory, characterized in that, The method includes: According to the state of charge, charge-discharge efficiency of the user's energy storage device, and the real-time dispatching requirements of the power grid, construct a spatio-temporal coupled game feature space, and generate an initial dynamic game strategy set through a multi-agent strategy network. Among them, the initial dynamic game strategy set includes the charge-discharge amounts and bid combinations of each microgrid; among them, collect the real-time state of charge data and historical charge-discharge efficiency curves of each microgrid energy storage device, combine the regional load forecast and electricity price fluctuation signals issued by the power grid dispatching center, and perform spatio-temporal domain interpolation processing to generate a spatio-temporal coupled state matrix with a granularity of 15 minutes; input the spatio-temporal coupled state matrix into a convolutional recurrent hybrid neural network, extract the charge-discharge correlation features between adjacent microgrids through a time-sliding window, and output a game relationship graph including the strength of supply-demand relationship and capacity constraints; Based on the game relationship graph, construct an adversarial training environment for the multi-agent strategy network, and use a double-layer policy gradient algorithm to synchronously optimize the charge-discharge amount strategy and bid strategy of each microgrid to generate an initial strategy set including Nash equilibrium constraints; perform Pareto front screening on the initial strategy set, eliminate the strategy combinations that violate the power grid security and stable operation constraints, and output the initial dynamic game strategy set; Based on the initial dynamic game strategy set, use an asymmetric Nash bargaining model for distributed negotiation, and dynamically adjust the weight factors of each microgrid through a virtual resource exchange market to generate an equilibrium benefit distribution scheme. Among them, the weight factors are jointly calculated by the user credit rating and the power grid emergency level; According to the equilibrium benefit distribution scheme, use a time-decaying reinforcement learning algorithm to iteratively correct the energy storage priority index to generate a dynamic bidding rule including supply-demand elasticity coefficients and risk compensation; Based on the dynamic bidding rule, allocate energy storage resources in real time through a decentralized gradient consensus mechanism, combine a verifiable random function to verify the effectiveness of the allocation, and generate a final collaborative dispatching instruction and synchronize it to each microgrid terminal.

2. The method according to claim 1, wherein The method of using an asymmetric Nash bargaining model for distributed negotiation based on the initial dynamic game strategy set, dynamically adjusting the weight factors of each microgrid through a virtual resource exchange market to generate an equilibrium benefit distribution scheme, where the weight factors are jointly calculated by the user credit rating and the power grid emergency level, includes: Analyze the charge-discharge amounts and bid combinations in the initial strategy set, initialize the threat point and bargaining space of the asymmetric Nash bargaining model, and set the threat point as the minimum revenue threshold when the microgrid operates independently; According to the historical compliance records of microgrid users and the probability of grid node voltage over-limit, calculate the credit score and emergency coefficient respectively, and generate an initial weight factor by fusing the two types of parameters through a hyperbolic tangent function; Deploy a distributed negotiation protocol in the virtual resource exchange market for distributed negotiation. Each microgrid calculates the benefit utility of the strategy set based on the current weight factor, and iteratively updates the bargaining solution through the alternating direction multiplier method; When the difference in Shapley values between two adjacent iterations is less than a preset threshold, trigger the weight factor dynamic adjustment mechanism, where the weight of the microgrid with a decreasing credit rating decays exponentially, and the weight of the microgrid in the power grid emergency area increases linearly; Generate an equilibrium benefit distribution plan that includes a transfer payment plan and a charging and discharging time sequence plan based on the final negotiation results, and verify the individual rationality and coalition stability of the plan through zero-knowledge proof technology.

3. The method according to claim 2, wherein According to the equilibrium benefit distribution plan, use the time-decaying reinforcement learning algorithm to iteratively correct the energy storage priority index, and generate a dynamic bidding rule that includes the supply-demand elasticity coefficient and risk compensation, including: Extract the historical winning bid rate and revenue fluctuation data of each microgrid from the equilibrium benefit distribution plan, and initialize the energy storage priority index as a weighted function of the supply-demand ratio and risk value; Construct a state transition model for time-decaying reinforcement learning, define the state as the difference between the real-time grid load and the available capacity of the energy storage, the action as the priority adjustment range, and the reward function includes the electricity price arbitrage revenue and the grid regulation contribution; Use the decay factor to dynamically adjust the exploration rate of the Q-learning algorithm, enhance the weight of historical experience during peak load periods, increase the probability of random exploration during low load periods, and update the probability distribution of the energy storage priority index; According to the updated priority index and its probability distribution, calculate the ratio of the price change rate to the volume change rate as the supply-demand elasticity coefficient, and quantify the risk compensation value of energy storage call based on the conditional value-at-risk model; Encode the supply-demand elasticity coefficient and the risk compensation value of energy storage call as a piecewise linear function, generate a dynamic bidding rule that includes the step quotation upper limit and penalty clauses, and write it into the consensus verification contract of the blockchain.

4. The method according to claim 3, characterized in that, Based on the dynamic bidding rule, allocate energy storage resources in real time through a decentralized gradient consensus mechanism, verify the effectiveness of the allocation by combining a verifiable random function, and generate a final coordinated scheduling instruction and synchronize it to each microgrid terminal, including: Broadcast the dynamic bidding rule in the decentralized network, each microgrid node generates a candidate allocation plan based on local constraint conditions, and calculates the gradient vector of the plan for the global objective function; Use the gradient consensus algorithm to aggregate the gradient information of each node, and solve the optimal allocation solution that satisfies the grid power flow constraint through the projected subgradient descent method to generate a preliminary coordinated scheduling instruction; Call the verifiable random function to generate a distributed random number seed, each node performs a pre-submission operation based on the coordinated scheduling instruction and generates an operation hash, and all hash values are aggregated by threshold signature and compared with the random number to verify the consistency; When the verification pass rate exceeds 95%, write the coordinated scheduling instruction into the trusted execution environment of each microgrid terminal, otherwise trigger the elastic reallocation mechanism, and return to execute the step of broadcasting the dynamic bidding rule in the decentralized network, each microgrid node generates a candidate allocation plan based on local constraint conditions, and calculates the gradient vector of the plan for the global objective function to perform iterative optimization until convergence.

5. A multi - microgrid collaborative scheduling system based on game theory, characterized in that, The system includes: A building block is used to construct a spatio-temporal coupled game feature space according to the state of charge, charge-discharge efficiency of the user's energy storage device, and the real-time grid dispatching requirements, and generate an initial dynamic game strategy set through a multi-agent strategy network. Among them, the initial dynamic game strategy set includes the charge-discharge amounts and bid combinations of each microgrid; among them, real-time state-of-charge data and historical charge-discharge efficiency curves of each microgrid energy storage device are collected, combined with the regional load forecast and electricity price fluctuation signals issued by the grid dispatching center, and spatio-temporal domain interpolation processing is performed to generate a spatio-temporal coupled state matrix with a granularity of 15 minutes; the spatio-temporal coupled state matrix is input into a convolutional recurrent hybrid neural network, and the charge-discharge correlation features between adjacent microgrids are extracted through a time sliding window, and a game relationship graph including the strength of supply-demand relationship and capacity constraints is output. Based on the game relationship graph, an adversarial training environment for the multi-agent strategy network is constructed, and the double-layer policy gradient algorithm is used to synchronously optimize the charge-discharge amount strategy and bid strategy of each microgrid to generate an initial strategy set including Nash equilibrium constraints; the initial strategy set is screened by the Pareto front, and the strategy combinations that violate the grid security and stable operation constraints are excluded, and the initial dynamic game strategy set is output. A negotiation module is used to perform distributed negotiation based on the initial dynamic game strategy set by using an asymmetric Nash bargaining model, and dynamically adjust the weight factors of each microgrid through a virtual resource exchange market to generate an equilibrium benefit distribution plan, where the weight factors are jointly calculated by the user credit rating and the grid urgency. A correction module is used to iteratively correct the energy storage priority index by using a time-decaying reinforcement learning algorithm according to the equilibrium benefit distribution plan to generate a dynamic bidding rule including the supply-demand elasticity coefficient and risk compensation. An allocation module is used to allocate energy storage resources in real time based on the dynamic bidding rule through a decentralized gradient consensus mechanism, verify the allocation effectiveness by combining a verifiable random function, and generate a final collaborative scheduling instruction and synchronize it to each microgrid terminal.

6. The system according to claim 5, wherein, The negotiation module is specifically used for: Parse the charge-discharge amounts and bid combinations in the initial strategy set, initialize the threat point and bargaining space of the asymmetric Nash bargaining model, and set the threat point as the minimum revenue threshold when the microgrid operates independently. Calculate the credit score and urgency coefficient respectively according to the historical compliance records of microgrid users and the probability of grid node voltage over-limit, and generate an initial weight factor by fusing the two types of parameters through a hyperbolic tangent function. Deploy a distributed negotiation protocol in the virtual resource exchange market for distributed negotiation. Each microgrid calculates the interest utility of the strategy set based on the current weight factor, and iteratively updates the bargaining solution through the alternating direction multiplier method. When the difference in the Shapley values between two adjacent iterations is less than a preset threshold, trigger the weight factor dynamic adjustment mechanism, where the weight of the microgrid with a decreased credit rating decays exponentially, and the weight of the microgrid in the grid emergency area increases linearly. Generate an equilibrium benefit distribution plan including the transfer payment plan and the charge-discharge time sequence plan according to the final negotiation result, and verify the individual rationality and coalition stability of the plan through zero-knowledge proof technology.

7. A storage medium, characterized in that, A computer program is stored in the storage medium, wherein the computer program is configured to execute the method according to any one of claims 1-4 when running.

8. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Game theory-based multi-micro-grid interconnection running optimization method

    CN107545325A

  • Multi-microgrid collaborative optimization scheduling method based on cooperative game and consistency algorithm

    CN113065696A

Cited By

  • Port virtual power plant interval game method and device considering transaction uncertainty

    CN121120133A