Multi-microgrid cooperative scheduling method and system based on game theory
Through the multi-micronet collaborative scheduling method based on game theory, a game feature space with time-space coupled is constructed, and distributed negotiation is used using the multi-agent strategy network and asymmetric Nash bargaining model. Combining the time-decay reinforcement learning and the decentralized gradient consensus mechanism, the coordination and scheduling problem of the multi-micronet system is solved, and efficient resource allocation and adaptive capabilities are achieved.
Patent Information
- Application Number
- CN202510764664.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
In the coordinated scheduling, the multi-micronet system has problems such as delay in information transmission, single-point system failure, inflexible scheduling and unbalanced resource allocation, making it difficult to achieve optimized dynamic adjustment.
The multi-micronet collaborative scheduling method based on game theory is adopted to construct a space-time-coupled game feature space, and a multi-agent strategy network is used to generate dynamic game initial strategy sets, combine asymmetric Nash bargaining model for distributed negotiation, and use a time-decay reinforcement learning algorithm to iteratively correct the energy storage priority indicators, and allocate energy storage resources in real time through a decentralized gradient consensus mechanism.
It improves the operating efficiency and adaptability of the multi-micro grid system, meets the needs of future intelligent distribution networks, and realizes dynamic optimized allocation and coordinated scheduling of resources.
Smart Images

Figure CN120280930A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of microgrids, and particularly relates to a multi-microgrid collaborative scheduling method and system based on game theory. Background Art
[0002] With the rapid popularization of new energy and the development of distributed generation technology, multi-microgrid systems, as an important part of the power system, exhibit high flexibility and scalability. Microgrids can achieve autonomous operation by integrating distributed energy sources, energy storage devices, and control systems, effectively alleviating the grid load pressure and improving energy utilization efficiency. However, due to the autonomy and heterogeneity of microgrids, their coordinated scheduling faces many challenges.
[0003] Traditional scheduling methods mostly adopt centralized schemes, relying on a single control center for decision-making, which suffer from problems such as information transmission delay, single-point system failure, and inflexible scheduling. On the other hand, the diverse demands and unbalanced resource allocation among microgrids also increase the complexity of coordinated scheduling. In addition, the state of charge, charge and discharge efficiency changes of microgrid energy storage devices, and the dynamic changes in user energy demands all require dynamic adjustment of scheduling strategies to achieve optimization. Summary of the Invention
[0004] The purpose of the present invention is to provide a multi-microgrid collaborative scheduling method and system based on game theory to solve the deficiencies in the prior art, which can improve the operation efficiency and adaptive ability of multi-microgrid systems and meet the requirements of future intelligent distribution networks.
[0005] An embodiment of the present application provides a multi-microgrid collaborative scheduling method based on game theory, the method comprising: Constructing a spatio-temporal coupled game feature space according to the state of charge, charge and discharge efficiency of user energy storage devices and the real-time scheduling requirements of the power grid, and generating a dynamic game initial strategy set through a multi-agent strategy network, wherein the dynamic game initial strategy set includes the charge and discharge amounts and quotation combinations of each microgrid; Based on the dynamic game initial strategy set, using an asymmetric Nash bargaining model for distributed negotiation, dynamically adjusting the weight factors of each microgrid through a virtual resource exchange market, and generating an equilibrium benefit distribution scheme, wherein the weight factors are jointly calculated by user credit ratings and grid urgency; According to the equilibrium benefit distribution scheme, using a time-decaying reinforcement learning algorithm to iteratively correct the energy storage priority index, and generating a dynamic bidding rule including supply-demand elasticity coefficients and risk compensation; Based on the dynamic bidding rule, allocating energy storage resources in real time through a decentralized gradient consensus mechanism, verifying the effectiveness of the allocation in combination with a verifiable random function, and generating a final collaborative scheduling instruction and synchronizing it to each microgrid terminal.
[0006] Optionally, based on the state of charge, charge and discharge efficiency of the user's energy storage device, and the real-time grid scheduling requirements, a spatio-temporal coupling game feature space is constructed, and a dynamic game initial strategy set is generated through a multi-agent strategy network. The dynamic game initial strategy set includes the charge and discharge amounts and bid combinations of each microgrid, and specifically includes: Collect the real-time state of charge data and historical charge and discharge efficiency curves of each microgrid energy storage device, combine the regional load forecast and electricity price fluctuation signals issued by the grid dispatching center, and perform spatio-temporal domain interpolation processing to generate a spatio-temporal coupling state matrix with a preset time period as the granularity; Input the spatio-temporal coupling state matrix into a convolutional recurrent hybrid neural network, extract the charge and discharge correlation features between adjacent microgrids through a time sliding window, and output a game relationship graph including the intensity of supply-demand relationship and capacity constraints; Based on the game relationship graph, construct an adversarial training environment for the multi-agent strategy network, and use a double-layer policy gradient algorithm to synchronously optimize the charge and discharge amount strategies and bid strategies of each microgrid to generate an initial strategy set including Nash equilibrium constraints; Perform Pareto front screening on the initial strategy set, eliminate the strategy combinations that violate the grid security and stable operation constraints, and output the dynamic game initial strategy set.
[0007] Optionally, based on the dynamic game initial strategy set, an asymmetric Nash bargaining model is used for distributed negotiation, and the weight factors of each microgrid are dynamically adjusted through a virtual resource exchange market. The weight factors are jointly calculated by the user credit rating and the grid emergency level, and specifically include: Analyze the charge and discharge amounts and bid combinations in the initial strategy set, initialize the threat point and bargaining space of the asymmetric Nash bargaining model, and set the threat point as the minimum revenue threshold when the microgrid operates independently; According to the historical compliance records of microgrid users and the probability of grid node voltage violation, calculate the credit score and emergency coefficient respectively, and generate the initial weight factor by fusing the two types of parameters through the hyperbolic tangent function; Deploy a distributed negotiation protocol in the virtual resource exchange market for distributed negotiation. Each microgrid calculates the interest utility of the strategy set based on the current weight factor, and iteratively updates the bargaining solution through the alternating direction multiplier method; When the difference in Shapley values between two adjacent iterations is less than the preset threshold, trigger the weight factor dynamic adjustment mechanism. Among them, the weight of the microgrid with a decreased credit rating decays exponentially, and the weight of the microgrid in the grid emergency area increases linearly; Generate an equilibrium interest distribution plan including the transfer payment plan and the charge and discharge time sequence plan according to the final negotiation result, and verify the individual rationality and coalition stability of the plan through zero-knowledge proof technology.
[0008] Optionally, according to the balanced benefit distribution scheme, the time decay type reinforcement learning algorithm is used to iteratively correct the energy storage priority index to generate a dynamic bidding rule including a supply-demand elasticity coefficient and risk compensation, including: Extract the historical winning bid rate and revenue fluctuation data of each microgrid from the balanced benefit distribution scheme, and initialize the energy storage priority index as a weighted function of the supply-demand ratio and risk value; Construct a state transition model of time decay type reinforcement learning, define the state as the difference between the real-time load of the power grid and the available capacity of the energy storage, the action as the priority adjustment range, and the reward function includes the electricity price arbitrage revenue and the power grid regulation contribution degree; Adopt a decay factor to dynamically adjust the exploration rate of the Q-learning algorithm, enhance the weight of historical experience during peak load periods, increase the probability of random exploration during low load periods, and update the probability distribution of the energy storage priority index; According to the updated priority index and its probability distribution, calculate the ratio of the price change rate to the trading volume change rate as the supply-demand elasticity coefficient, and quantify the risk compensation value of energy storage call based on the conditional value at risk model; Encode the supply-demand elasticity coefficient and the risk compensation value of energy storage call as a piecewise linear function, generate a dynamic bidding rule including the ladder quotation upper limit and penalty clauses, and write it into the consensus verification contract of the blockchain.
[0009] Optionally, based on the dynamic bidding rule, the energy storage resources are allocated in real time through a decentralized gradient consensus mechanism, combined with a verifiable random function to verify the effectiveness of the allocation, generate a final coordinated scheduling instruction and synchronize it to each microgrid terminal, including: Broadcast the dynamic bidding rule in the decentralized network, each microgrid node generates a candidate allocation scheme based on local constraint conditions, and calculates the gradient vector of the scheme for the global objective function; Adopt a gradient consensus algorithm to aggregate the gradient information of each node, and solve the optimal allocation solution that satisfies the power grid power flow constraint through the projected subgradient descent method to generate a preliminary coordinated scheduling instruction; Call the verifiable random function to generate a distributed random number seed, each node performs a pre-submission operation based on the coordinated scheduling instruction and generates an operation hash, and all hash values are aggregated by threshold signature and compared with the random number to verify the consistency; When the verification passing rate exceeds the preset threshold, write the coordinated scheduling instruction into the trusted execution environment of each microgrid terminal, otherwise trigger the elastic reallocation mechanism, and return to execute the steps of broadcasting the dynamic bidding rule in the decentralized network, each microgrid node generating a candidate allocation scheme based on local constraint conditions, and calculating the gradient vector of the scheme for the global objective function for iterative optimization until convergence.
[0010] Another embodiment of the present application provides a multi-microgrid coordinated scheduling system based on game theory, the system includes: A building module, configured to construct a spatio-temporal coupled game feature space according to the state of charge, charge and discharge efficiency of the user's energy storage device, and the real-time scheduling requirements of the power grid, and generate an initial dynamic game strategy set through a multi-agent strategy network, wherein the initial dynamic game strategy set includes the charge and discharge amounts and bid combinations of each microgrid; A negotiation module, configured to perform distributed negotiation based on the initial dynamic game strategy set by using an asymmetric Nash bargaining model, dynamically adjust the weight factors of each microgrid through a virtual resource exchange market, and generate an equilibrium benefit distribution scheme, wherein the weight factors are jointly calculated by the user credit rating and the power grid urgency; A correction module, configured to iteratively correct the energy storage priority index by using a time-decaying reinforcement learning algorithm according to the equilibrium benefit distribution scheme, and generate a dynamic bidding rule including a supply-demand elasticity coefficient and a risk compensation; An allocation module, configured to allocate energy storage resources in real time based on the dynamic bidding rule through a decentralized gradient consensus mechanism, verify the allocation effectiveness by combining a verifiable random function, generate a final collaborative scheduling instruction, and synchronize it to each microgrid terminal.
[0011] Another embodiment of the present application provides a storage medium, in which a computer program is stored, wherein the computer program is set to execute the method described in any one of the above when running.
[0012] Another embodiment of the present application provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is set to run the computer program to execute the method described in any one of the above.
[0013] Compared with the prior art, a multi-microgrid collaborative scheduling method based on game theory provided by the present invention constructs an initial dynamic game strategy set through a multi-agent strategy network according to the state of charge, charge and discharge efficiency of the user's energy storage device, and the real-time scheduling requirements of the power grid; based on the initial dynamic game strategy set, an asymmetric Nash bargaining model is used for distributed negotiation to generate an equilibrium benefit distribution scheme; according to the equilibrium benefit distribution scheme, a time-decaying reinforcement learning algorithm is used to iteratively correct the energy storage priority index to generate a dynamic bidding rule including a supply-demand elasticity coefficient and a risk compensation; based on the dynamic bidding rule, energy storage resources are allocated in real time through a decentralized gradient consensus mechanism to generate a final collaborative scheduling instruction and synchronize it to each microgrid terminal, so as to improve the operation efficiency and adaptive ability of the multi-microgrid system and meet the requirements of future intelligent distribution networks. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a hardware structure block diagram of a computer terminal of a multi-microgrid collaborative scheduling method based on game theory provided by an embodiment of the present invention; Figure 2Schematic flow chart of a multi - microgrid collaborative scheduling method provided by an embodiment of the present invention; Figure 3 Schematic structural diagram of a multi - microgrid collaborative scheduling system provided by an embodiment of the present invention. Detailed implementation manners
[0015] The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0016] An embodiment of the present invention first provides a multi - microgrid collaborative scheduling method based on game theory. This method can be applied to electronic devices, such as computer terminals, specifically ordinary computers, etc.
[0017] The following takes running on a computer terminal as an example to explain it in detail. Figure 1 Hardware structure block diagram of a computer terminal for a multi - microgrid collaborative scheduling method provided by an embodiment of the present invention. As Figure 1 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory may include a non - volatile storage medium and an internal memory.
[0018] The non - volatile storage medium can store an operating system and a computer program. The computer program includes program instructions. When the program instructions are executed, the processor can execute any multi - microgrid collaborative scheduling method based on game theory.
[0019] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.
[0020] The internal memory provides an environment for the operation of the computer program in the non - volatile storage medium. When the computer program is executed by the processor, the processor can execute any multi - microgrid collaborative scheduling method based on game theory.
[0021] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 1 the structure shown in
[0022] It should be understood that the processor can be a Central Processing Unit (CPU), and the processor can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0023] See Figure 2 , an embodiment of the present invention provides a multi-microgrid collaborative scheduling method based on game theory, which may include the following steps: S201. According to the state of charge, charge and discharge efficiency of the user's energy storage device, and the real-time grid scheduling requirements, construct a spatio-temporal coupled game feature space, and generate a dynamic game initial strategy set through a multi-agent strategy network. Among them, the dynamic game initial strategy set includes the charge and discharge amounts and quotation combinations of each microgrid; specifically, it may include: S2011. Collect the real-time state of charge data and historical charge and discharge efficiency curves of each microgrid energy storage device, combine the regional load forecast and electricity price fluctuation signals released by the grid dispatching center, and perform spatio-temporal domain interpolation processing to generate a spatio-temporal coupled state matrix with a preset time period as the granularity; the preset time period is, for example, 15 minutes.
[0024] This step constructs a high-precision state matrix through multi-source data fusion and spatio-temporal alignment technology, providing a structured input for subsequent game modeling.
[0025] Data collection and preprocessing: Energy storage device data collection: Real-time state of charge (SOC): Read the SOC value (0-100%) of the lithium battery management system through the CAN bus protocol, with a sampling frequency of 1Hz, and use a moving average filter (window length 60 seconds) to eliminate noise. For example, the SOC of the energy storage of Microgrid A is 72.3% at 14:00, with an accuracy of ±0.5%.
[0026] Historical charge and discharge efficiency curve: Extract the charge and discharge data of the past 30 days from the SCADA system, and calculate the efficiency η = actual discharge amount / theoretical capacity. For example, the average efficiency of Microgrid B at 20°C is η = 93.2%.
[0027] Grid scheduling data access: Regional load forecasting: Receive 24-hour load forecasting data (in CSV format) released by the power grid dispatching center, with a time resolution of 5 minutes and a power range of 0 - 100 MW. For example, the regional load forecast at 15:00 is 58.7 MW, with a confidence interval of ±3%.
[0028] Electricity price fluctuation signal: Analyze the difference between the day-ahead electricity price and the real-time electricity price in the power trading market to generate a fluctuation index (-1 to +1), where a negative value indicates supply exceeding demand. For example, the real-time electricity price at 14:30 has dropped by 12% compared to the day-ahead price, and the fluctuation index is -0.4.
[0029] Spatio-temporal interpolation processing: Spatial interpolation: For the discrete SOC data of 12 microgrids, use Kriging interpolation to generate a SOC distribution map of a 500m×500m grid. The semi-variogram selects the Gaussian model, with the nugget value set to 0.05 and the range set to 2 km.
[0030] Time alignment: Unify data with different sampling frequencies to a 15-minute granularity. For 1Hz SOC data, take the median within each 15-minute window; for 5-minute load data, generate a 15-minute sequence through cubic spline interpolation.
[0031] Matrix construction: Organize the interpolated data into a four-dimensional tensor (spatial rows × spatial columns × time slices × feature dimensions), with a size of 12×12×96×8 (a 12-row × 12-column grid, 96 15-minute time slices, and 8 features: SOC, efficiency, load, electricity price, etc.). For example, the matrix position (5,7,32,3) represents the load value of 45.6 MW at 08:00 for the grid at row 5 and column 7.
[0032] S2012, input the spatio-temporal coupling state matrix into a convolutional recurrent hybrid neural network, extract the charge-discharge correlation features between adjacent microgrids through a time-sliding window, and output a game relationship graph containing the strength of the supply-demand relationship and capacity constraints; This step mines spatio-temporal correlation features through a hybrid neural network to construct a topological representation of the dynamic game relationship between microgrids.
[0033] Design of convolutional recurrent hybrid neural network (CRHNN): Spatial feature extraction: Use a 3D convolutional layer (Conv3d) to process the spatial-temporal dimension, with a convolutional kernel size of 3×3×3 (height × width × time), and the number of channels expanded from 8 to 32. For example, the input tensor (12×12×96×8) outputs (10×10×94×32) after passing through Conv3d, capturing the charge-discharge spatial patterns of local microgrid groups.
[0034] Time-dependent modeling: A bidirectional LSTM (Bi-LSTM) is connected after the convolutional layer. The hidden layer dimension of each LSTM unit is 256, and the sliding step size of the time window is 4 (i.e., one step per hour). For example, for 94 time slices, they are divided into 23 windows with a step size of 4, and each window processes 4 consecutive time slices (i.e., 1-hour data).
[0035] Associated feature fusion: The associated weights between adjacent microgrids are calculated through the multi-head attention mechanism (Multi-Head Attention, number of heads = 8). For example, the attention score α_ij of microgrid i to microgrid j is 0.73, indicating that the charging and discharging behavior of i is greatly affected by j.
[0036] Game relationship graph generation: Calculation of supply-demand relationship intensity: Define the intensity S_ij = α_ij × (SOC_i - SOC_j) × η_i, where η_i is the charging and discharging efficiency of microgrid i. For example, S_AB of microgrid A to B is 0.73×(72% - 65%)×93% = 0.047, indicating the potential for A to supply power to B.
[0037] Capacity constraint encoding: According to the transmission limit of the lines between microgrids, a hard constraint matrix C_ij = min(line capacity, remaining energy storage capacity) is generated. For example, the line capacity from A to B is 2MW, and the current dischargeable amount of A is 1.8MW, then C_AB = 1.8MW.
[0038] Graph construction: Encode S_ij and C_ij into a weighted adjacency matrix, where the nodes represent microgrids and the edge weights are [S_ij, C_ij]. For example, the element A→B of the adjacency matrix is [0.047, 1.8], which is stored as a JSON-format topology file.
[0039] S2013, based on the game relationship graph, construct an adversarial training environment for the multi-agent policy network, and use the double-layer policy gradient algorithm to synchronously optimize the charging and discharging quantity strategies and bidding strategies of each microgrid to generate an initial strategy set containing Nash equilibrium constraints; In this step, through adversarial training and game equilibrium constraints, an optimal strategy candidate set that satisfies the interests of multiple parties is generated.
[0040] Multi-agent policy network architecture: Actor (Policy Network): Each microgrid independently has a policy network. The inputs are its own state (SOC, efficiency) and the features of adjacent nodes in the game graph. The outputs are the charge / discharge amount \(a_t\) (-1 to +1, normalized value) and the bid price \(p_t\) (0 to 1, corresponding to 80% - 120% of the market price). The network structure is a 4 - layer fully connected network (256→128→64→2). The activation function uses LeakyReLU (negative slope 0.01) for all layers except the last layer, where Tanh is used.
[0041] Critic (Value Network): The global value network receives the states and actions of all microgrids and outputs the joint Q - value. A Graph Attention Network (GAT) is used to aggregate neighbor information, with 4 attention heads and a hidden layer dimension of 128.
[0042] Double - layer Policy Gradient Algorithm: Inner - layer Optimization: Fix the parameters of the Critic network and update the Actor network through the Deterministic Policy Gradient (DPG). The gradient calculation formula is: \(\nabla_{\theta}J\approx E[\nabla_{\theta}Q(s,a)| {a = \mu(s)\}]\). The learning rate is set to \(1e - 4\), and for the Adam optimizer, \(\beta_1 = 0.9\) and \(\beta_2 = 0.999\).
[0043] Outer - layer Equilibrium Constraint: Introduce the Nash equilibrium constraint condition, requiring that the marginal revenue difference of each microgrid is less than the threshold \(\varepsilon = 0.05\). The constraint is added to the loss function through the Lagrange multiplier method. The initial value of the multiplier is \(\lambda = 0.1\), and it doubles every 10 training rounds until the constraint is satisfied.
[0044] Adversarial Training Mechanism: Set up a virtual opponent policy network to generate adversarial bid perturbations (±5%) in each training round, forcing the main network to improve its robustness. The adversarial sample generation frequency is once per batch (batch size = 64).
[0045] Initial Policy Set Generation: Policy Sampling: After training convergence, sample 1000 groups of policies (charge / discharge amount + bid price) from the policy network to form the initial set. For example, a certain policy is {Microgrid A discharges 0.8 MW and bids 0.52 yuan / kWh, Microgrid B charges 1.2 MW and bids 0.48 yuan / kWh}.
[0046] Nash Equilibrium Verification: Calculate the Nash convergence degree index \(\delta=\sum_{i = 1}^{N}|Revenue_i - Maximum\ Possible\ Revenue_i| / N\), and retain the policies with \(\delta\lt0.1\). For example, if \(\delta = 0.07\) for a certain policy, it indicates that it is close to the equilibrium state.
[0047] In S2014, perform Pareto - front screening on the initial policy set, eliminate the policy combinations that violate the power grid security and stable operation constraints, and output the initial dynamic game policy set.
[0048] This step ensures that the policy set takes into account both economy and grid security through multi-objective optimization and security constraint filtering.
[0049] Pareto front screening: Optimization objective definition: Objective 1: Minimize the total operating cost (including power purchase cost and equipment loss); Objective 2: Maximize the grid regulation contribution (calculated based on response speed and regulation amount); Objective 3: Optimize the fairness of benefits (minimize the Gini coefficient).
[0050] Application of NSGA-III algorithm: Reference point generation: Use the Das-Dennis method to generate 100 uniform reference points in the three-dimensional objective space; Non-dominated sorting: Sort 1000 initial policies and retain the top 50 Pareto optimal solutions; Diversity maintenance: Calculate the correlation between policies and reference points to ensure a uniform distribution of the front.
[0051] Treatment of grid security constraints: Power flow constraint verification: Call the Matpower toolkit for fast DC power flow calculation to verify whether the node voltage is within the range of 0.95 - 1.05 p.u. For example, if a certain policy causes the node voltage to drop to 0.93 p.u., it is marked as a violation.
[0052] Energy storage constraint filtering: Eliminate policies with charge and discharge rates exceeding the equipment rated value (such as the maximum 2C rate for lithium batteries). For example, if a certain policy requires the microgrid C to discharge at 2.5C, the filtering mechanism is triggered.
[0053] N - 1 security criterion: When simulating any single-line fault, verify whether the remaining policies still meet the load demand. Use parallel computing (MPI multi-process) to accelerate the verification, and the fault scenario library contains 12 preset topologies.
[0054] Generation of dynamic policy set: After screening by NSGA-III and filtering by security constraints, the initial policies are reduced from 1000 to 32; Each policy is labeled with multi-dimensional attribute tags, for example: { "strategy_id": "S23", "cost": 45670, / / Unit: yuan "contribution": 0.82, / / Normalized score "gini": 0.15, "violation_flag": false }。
[0055] Build a policy retrieval index to support fast query by conditions such as cost range and contribution threshold.
[0056] In this step, by integrating the real-time status of energy storage devices (such as state of charge, charge-discharge efficiency) and the dynamic demands of the power grid (such as load forecasting, electricity price fluctuations), a game feature space in the spatio-temporal dimension is constructed, and the multi-agent policy network is used to simulate the interest game behaviors of each microgrid. Specifically, the spatio-temporal correlation features between microgrids (such as the charge-discharge dependence relationship between adjacent microgrids) are extracted through a convolutional recurrent hybrid neural network, and an initial policy set that satisfies the Nash equilibrium is generated based on the double-layer policy gradient algorithm, ensuring that the policy maximizes the charge-discharge flexibility and bidding competitiveness of each microgrid on the premise of meeting the power grid security constraints, solving the policy conflict problem caused by spatio-temporal differences in multi-microgrid coordinated scheduling, unifying the decentralized microgrid behaviors into a globally optimized initial policy set through game modeling, providing a basic support for subsequent interest distribution and dynamic bidding, and avoiding the risks of uneven resource allocation or local overload.
[0057] S202. Based on the initial dynamic game policy set, use the asymmetric Nash bargaining model for distributed negotiation, and dynamically adjust the weight factors of each microgrid through the virtual resource exchange market to generate an equilibrium interest distribution plan, where the weight factors are jointly calculated by the user credit rating and the power grid emergency level; specifically, it may include: S2021. Analyze the charge-discharge amount and bidding combination in the initial policy set, initialize the threat point and bargaining space of the asymmetric Nash bargaining model, and set the threat point as the minimum revenue threshold when the microgrid operates independently; In this step, by analyzing the policy data of multiple microgrids, the bargaining starting point and feasible space of the game are defined, laying a foundation for subsequent distributed negotiation.
[0058] Policy data analysis : Charge-discharge amount analysis : Extract the charge-discharge amount policy of each microgrid from the initial dynamic game policy set (JSON format). For example, the policy of microgrid A is "discharge 0.8 MW (14:00 - 15:00)" and "charge 1.2 MW (18:00 - 19:00)", and the power resolution is 0.1 MW.
[0059] Bidding combination analysis : Analyze the price range in the bidding policy. For example, the bid of microgrid B is "discharge price from RMB 0.52 yuan / kWh (minimum limit price) to RMB 0.68 yuan / kWh (maximum limit price)", with a time granularity of 15 minutes, aligned with the real-time power market trading platform.
[0060] Data verification: Verify the matching of the charge and discharge amount and the energy storage capacity. If a certain strategy requires Microgrid C to discharge 2 MW at SOC = 20% (exceeding its maximum discharge rate of 1.5C), it is marked as an invalid strategy and excluded.
[0061] Threat point initialization: Calculation of the minimum benefit threshold: Based on the historical data during the independent operation of the microgrid (such as the income quantile in the past 30 days), set the threat point as the 10% quantile of the income distribution. For example, when Microgrid D operates independently, the daily average income is from RMB 1200 to RMB 1800, and the 10% quantile is RMB 950, that is, the threat point is set to RMB 950.
[0062] Bargaining space construction: Define the bargaining feasible region as the difference between the total income during the combined operation of all microgrids and the total income during independent operation. For example, if the total independent income of three microgrids is RMB 3000 and the predicted income during combined operation is RMB 4200, then the bargaining space is RMB 1200.
[0063] Encoding of constraint conditions: Encode the grid security constraints (such as node voltage deviation < 5%) and the energy storage life constraints (charge and discharge cycle times < 5000 times) into linear inequalities and write them into the constraint matrix of the bargaining model.
[0064] In S2022, according to the historical compliance records of microgrid users and the probability of grid node voltage over-limit, calculate the credit score and the emergency coefficient respectively, and generate the initial weight factor by fusing the two types of parameters through the hyperbolic tangent function; This step dynamically adjusts the voice of the microgrid in bargaining by quantifying user credit and grid emergency level, ensuring the fairness of benefit distribution and the stability of the grid.
[0065] Calculation of credit score: Analysis of historical compliance records: Extract the execution records of the dispatching instructions of the microgrid in the past 90 days, and define the compliance rate λ = actual completion volume / committed volume. For example, Microgrid E fulfilled 28 times out of 30 dispatches, λ = 93.3%.
[0066] Time decay weighting: Use the exponential decay model to assign higher weights to recent compliance. The decay factor α = 0.9, and calculate the weighted compliance rate: λ_weighted = Σ(α^{t}·λ_t) / Σα^{t}, where t is the historical number of days and λ_t is the compliance rate in the historical t days. For example, for Microgrid F with a compliance rate of 100% in the recent 7 days, λ_weighted = 98.6%.
[0067] Credit score mapping: Map λ_weighted to a score range of 0 - 100 through a piecewise linear function: 100 points are obtained when λ ≥ 95%, (λ - 80) / 0.15×30 + 70 points are obtained when 80% ≤ λ < 95%, and 0 points are directly obtained when λ < 80%.
[0068] Calculation of urgency coefficient: Prediction of voltage violation probability: Based on the real-time state estimation of the power grid (SCADA data), use a random forest model to predict the voltage violation probability P_voltage of each node within the next 1 hour. Input features include load rate, reactive power compensation amount, adjacent microgrid power exchange amount, etc. For example, the current load rate of node G is 85%, and the predicted P_voltage = 23%.
[0069] Urgency classification: Divide the emergency level according to P_voltage: "Red emergency" (coefficient 1.0) when P ≥ 30%, "Orange emergency" (coefficient 0.7) when 20% ≤ P < 30%, and "Normal" (coefficient 0.3) when P < 20%.
[0070] Regional aggregation: If a microgrid is connected to multiple nodes, take the maximum value as its urgency coefficient. For example, microgrid H is connected to node J (P = 25%) and node K (P = 18%), then the coefficient is 0.7.
[0071] Weight factor fusion: Application of hyperbolic tangent function: Normalize the credit score S (0 - 100) and the urgency coefficient E (0 - 1) and then input them into the tanh function: W_initial = tanh(0.1S + 2E). For example, when S = 80 and E = 0.7, W_initial ≈ 0.86.
[0072] Weight normalization: Perform Softmax processing on W_initial of all microgrids within the alliance to ensure ΣW_i = 1. For example, if the W_initial values of three microgrids are 0.86, 0.64, and 0.50 respectively, the normalized weights are 0.43, 0.32, and 0.25.
[0073] In S2023, deploy a distributed negotiation protocol in the virtual resource exchange market for distributed negotiation. Each microgrid calculates the interest utility of the strategy set based on the current weight factor, and iteratively updates the bargaining solution through the alternating direction multiplier method; This step realizes decentralized interest negotiation through a distributed optimization algorithm, avoiding single-point failures of the central node and protecting data privacy.
[0074] Design of distributed negotiation protocol: Communication topology definition: The Gossip protocol is used to build a P2P network, and each microgrid communicates directly with only 3-5 geographically adjacent microgrids. For example, microgrid J forms a neighbor group with microgrids K, L, and M, and the communication delay < 50ms.
[0075] Message format standardization: Define the JSON Schema for negotiation messages, including fields: { "timestamp": "2023-08-20T14:30:00Z", "microgrid_id": "MG-005", "proposed_strategy": {"discharge": 0.6, "price": 0.55}, "weight": 0.35, "utility_value": 1280.5 }。
[0076] Benefit utility calculation: Each microgrid calculates the strategy utility U_i = W_i × (benefit_i - threat point_i) + (1 - W_i) × Σ(benefit_j - threat point_j) according to the weight W_i. For example, for microgrid N, W = 0.4, the benefit is 1300 yuan, the threat point is 900 yuan, and the total benefit difference of neighbors is 2000 yuan, then U_N = 0.4 × 400 + 0.6 × 2000 = 160 + 1200 = 1360.
[0077] Alternating direction method of multipliers (ADMM) iteration: Variable splitting: Decompose the global optimization problem into local sub-problems of each microgrid, and define the consensus constraint variable z_ij to represent the consensus value of microgrid i and j on the common variable. For example, microgrids O and P need to reach an agreement on the power exchange volume of the tie line, and z_OP is initially set to the average of their proposals.
[0078] Parameter setting: Penalty factor ρ = 1.2 (controlling the convergence speed), maximum number of iterations T = 100, convergence threshold ε = 0.01 (the change rate of the objective function between two iterations < 1%).
[0079] S2024, when the difference in Shapley values between two adjacent iterations is less than the preset threshold, trigger the dynamic adjustment mechanism of the weight factor. Among them, the weight of the microgrid with a decreased credit rating decays exponentially, and the weight of the microgrid in the emergency area of the power grid increases linearly; This step responds to credit risks and grid emergency states through a dynamic weight adjustment mechanism to achieve adaptive optimization of benefit distribution.
[0080] Shapley Value Monitoring: Contribution Calculation: For each microgrid i, calculate its Shapley value φ_i = Σ_{S⊆N{i}} [|S|!(|N|-|S|-1)!) / |N|!] (v(S∪{i}) - v(S)), where v(S) is the total revenue of coalition S. For example, for 3 microgrids, 2^2 = 4 subset combinations need to be calculated for φ_i.
[0081] Difference Evaluation: Define the difference Δ = max |φ_i^{k} - φ_i^{k-1}| / φ_i^{k-1}, and the preset threshold Δ_threshold = 5%. When Δ < 5%, it is considered that the coalition structure tends to be stable and the weight adjustment is triggered.
[0082] Weight Dynamic Adjustment Rule: Credit Decrease Penalty: If the compliance rate of a certain microgrid drops by more than 10% in the past 24 hours, its weight is calculated as W_i^{new} =W_i^{old} × e^{-βt}, where β = 0.1 (decay rate) and t is the number of hours of continuous credit decrease. For example, if the compliance rate of microgrid R drops by 15% for 3 consecutive hours, the weight drops from 0.3 to 0.3×e^{-0.1×3}=0.222.
[0083] Emergency Area Reward: For microgrids in areas with a voltage violation probability > 25%, the weight increases linearly: W_i^{new} = W_i^{old} + γ×(P_voltage - 0.25), where γ = 0.5 (adjustment coefficient). For example, if P_voltage = 30% in the area where microgrid S is located, the weight increases by 0.5×(0.3 - 0.25)=0.025.
[0084] Weight Rebalancing: After adjustment, re-normalize to ensure ΣW_i = 1. For example, the original weights [0.3, 0.4, 0.3] become [0.25, 0.45, 0.3] after adjustment, and after normalization, they are [0.25 / 1.0, 0.45 / 1.0, 0.3 / 1.0].
[0085] In S2025, generate an equilibrium benefit distribution plan including the transfer payment plan and the charging and discharging timing plan according to the final negotiation result, and verify the individual rationality and coalition stability of the plan through zero-knowledge proof technology.
[0086] This step converts the negotiation result into an executable plan and ensures the fairness and immutability of the plan through cryptographic technology.
[0087] Transfer Payment Plan Generation: Calculation of income difference : For each microgrid i, calculate the difference Δ_i = π_i - φ_i between the actual income π_i and the Shapley value φ_i. If Δ_i>0, Δ_i needs to be paid to other members; if Δ_i<0, |Δ_i| compensation can be obtained. For example, microgrid V obtains φ_V = RMB 1500, and the actual income π_V = RMB 1400, then it needs to receive RMB 100 compensation.
[0088] Payment path optimization: Dijkstra algorithm is used to find the minimum transaction path. For example, microgrid W needs to pay RMB 200 to microgrid X, and microgrid Y needs to pay RMB 150 to X. The paths are merged into W→Y→X to reduce the number of transactions.
[0089] Smart Contract Deployment: The payment plan is written into the blockchain smart contract, and the trigger condition is that the microgrid will automatically execute after completing the dispatch instruction. For example, after microgrid Z completes 1.5MW discharge at 15:00, the on-chain contract transfers RMB 320 to its address.
[0090] Charge and discharge timing plan: Time window optimization: Based on the grid load forecast curve, discharge is prioritized during peak electricity price periods (such as 14:00-16:00), and charging is arranged during valley periods (such as 02:00-04:00). For example, the discharge plan of microgrid α is to discharge 0.9MW from 14:30 to 15:30, and the charging plan is to charge 1.1MW from 03:00 to 04:00.
[0091] Energy storage life balance: Limit the number of charge and discharge switching of the same energy storage unit to ≤ 1 time within 3 consecutive time slices to avoid frequent switching and loss of battery life. For example, after the energy storage system of microgrid β completes a charge and discharge cycle from 10:00 to 11:30, it needs to be left idle for at least 45 minutes.
[0092] Zero-knowledge proof verification: Individual rational verification: Using zk-SNARKs technology, generate proof π_IR, indicating that for any microgrid i, there is π_i ≥ threat point_i, without revealing the specific value of the benefit. For example, the verifier only needs to confirm the validity of π_IR, without knowing the actual benefit of microgrid γ, RMB 1,650.
[0093] Alliance stability verification: Construct a Merkle tree to store the Shapley value of each microgrid, and verify the integrity of φ_i through the root hash to ensure that no member modifies the allocation ratio without authorization. For example, an attacker cannot forge the φ value of microgrid δ because its hash value is inconsistent with other nodes.
[0094] Taking the initial strategy set as input, the bargaining power (weight factor) of each microgrid is defined through the asymmetric Nash bargaining model. The weight factor is combined with the user credit score (such as the historical fulfillment rate) and the grid urgency (such as the probability of node voltage exceeding the limit). The alternating direction multiplier method is used to iteratively optimize the benefit distribution scheme in the virtual resource exchange market. The weight of microgrids with low credit ratings is attenuated, and the weight of microgrids in emergency areas is increased. Finally, the fairness and stability of the scheme are verified through zero-knowledge proof to ensure that the benefit distribution satisfies both individual rationality (the income of a single microgrid is not lower than the independent operation threshold) and alliance stability (global resource utilization is optimal), balance the conflicts of interest between multiple microgrids, and prioritize the rights and interests of grid emergency needs and high-quality credit users through a dynamic weight adjustment mechanism, improve the fairness and executability of coordinated scheduling, and avoid the problems of "free riding" or "resource crowding" in traditional centralized scheduling.
[0095] S203, according to the balanced benefit distribution plan, using the time-decayed reinforcement learning algorithm to iteratively correct the energy storage priority index, and generate a dynamic bidding rule including supply and demand elasticity coefficient and risk compensation; specifically, it may include: S2031, extracting the historical bid winning rate and revenue fluctuation data of each microgrid from the balanced benefit distribution plan, and initializing the energy storage priority index as a weighted function of the supply-demand ratio and the risk value; This step constructs an energy storage priority evaluation system by mining historical transaction data and revenue characteristics, providing a quantitative basis for dynamic bidding rules.
[0096] Data extraction and cleaning: Calculation of historical winning bid rate: Query the bidding records of each microgrid in the past 30 trading periods from the blockchain database, and calculate the proportion of winning bids to the total number of bids. For example, microgrid A participated in the bidding 20 times in the day-ahead market and won 12 times, and the historical winning bid rate = 12 / 20 = 60%.
[0097] Quantification of return volatility: Extract the actual return data of each trading period (unit: RMB / MWh), and calculate the ratio of standard deviation to mean as volatility. For example, the peak return sequence of microgrid B is [RMB 520, RMB 580, RMB 490], with a mean of RMB 530, a standard deviation of RMB 45, and a volatility of 45 / 530≈8.5%.
[0098] Outlier processing: Tukey's fences method is used to remove extreme values. If the revenue of a period exceeds Q3+1.5IQR (Q3 is the third quartile, IQR is the interquartile range), it is marked as an outlier. For example, in the revenue sequence of microgrid C, RMB 850 exceeds Q3 (RMB 620) + 1.5×IQR (200), it is regarded as an outlier and replaced by the mean of the adjacent period.
[0099] Calculation of supply - demand ratio and risk value: Calculation of real - time supply - demand ratio: Define the supply - demand ratio as the ratio of the total regional load (MW) to the total available energy storage capacity (MWh) in the current period. For example, at 14:00 in a certain region, the load is 150 MW and the available energy storage capacity is 120 MWh (SOC≥30%), then the supply - demand ratio = 150 / 120 = 1.25.
[0100] Risk value assessment: Based on the CVaR (Conditional Value at Risk) model, calculate the expected loss at the tail of the profit distribution (such as the 5% quantile). For example, in the profit distribution of Microgrid D, the average profit in the worst 5% scenarios is 360 yuan, and the risk value = 360 yuan.
[0101] Initialization of priority index: Normalize and weight the supply - demand ratio (S) and risk value (R). The weight coefficients are α = 0.6 (emphasizing supply - demand balance) and β = 0.4 (emphasizing risk aversion): Priority = α×(S / S_max)+β×(1 - R / R_max). Where S_max is the historical maximum supply - demand ratio (such as 2.0), and R_max is the maximum risk value (such as 500 yuan). For example, if S = 1.25 and R = 360 yuan, then Priority = 0.6×(1.25 / 2.0)+0.4×(1 - 360 / 500)=0.487.
[0102] S2032, construct a state - transition model of time - decaying reinforcement learning. Define the state as the difference between the real - time grid load and the available energy storage capacity, the action as the adjustment amplitude of the priority, and the reward function includes the electricity price arbitrage profit and the grid regulation contribution; This step dynamically adapts to market environment changes through reinforcement learning, optimizing the real - time performance and robustness of the energy storage scheduling strategy.
[0103] State - space modeling: Definition of state variables: The state s_t = real - time load L_t (MW)-available energy storage capacity C_t (MWh). For example, at 15:00, the load is 160 MW and the energy storage capacity is 130 MWh, then s_t = 160 - 130 = 30.
[0104] State discretization: Divide the continuous state into 10 intervals, such as [-∞, - 50), [-50, - 30), [-30, - 10), [-10,10), [10,30), [30,50), [50,70), [70,90), [90,110), [110, +∞]. Each interval corresponds to a discrete state code (such as s = 30 corresponding to code 5).
[0105] State transition probability: Based on historical data statistics, the state transition law is obtained. For example, when s_t = 30 (encoded as 5), the probability of the next state transitioning to s_{t + 1}=50 (encoded as 6) is 40%, and the probability of transitioning to s_{t + 1}=10 (encoded as 4) is 35%.
[0106] Action space design: Definition of adjustment amplitude: The action a ∈ {-0.2, -0.1, 0, +0.1, +0.2}, representing the adjustment amount of the priority index. For example, when the current priority is 0.5, it becomes 0.6 after executing the action +0.1.
[0107] Action constraint conditions: Set the maximum adjustment step size to prevent sudden changes in priority. For example, the single - time adjustment amplitude shall not exceed ±0.2, and after three consecutive same - direction adjustments, a forced reverse adjustment is required once.
[0108] Reward function design: Electricity price arbitrage profit: Calculate the arbitrage profit R_arbitrage of the micro - grid in period t = (discharge electricity price - charge electricity price)×actual transaction volume. For example, the discharge price is 0.68 yuan / kWh, the charge price is 0.42 yuan / kWh, and the transaction volume is 10 MWh, then R_arbitrage=(0.68 - 0.42)×10000 = 2600 yuan.
[0109] Grid regulation contribution degree Contribution: Calculate the contribution degree according to the regulation amount (unit: MW) of the micro - grid to the grid frequency deviation. For example, if the micro - grid G provides 5 MW of frequency regulation service, the contribution degree = 5 / total regional frequency regulation demand (20 MW)=25%.
[0110] Comprehensive reward calculation: Aggregate the two types of rewards with weights γ = 0.7 (arbitrage profit) and δ = 0.3 (regulation contribution): Reward = γ×R_arbitrage / R1_max+δ×Contribution.
[0111] Among them, R1_max is the historical maximum arbitrage profit (such as 5000 yuan). For example, when R_arbitrage = 2600 yuan and Contribution = 25%, then Reward = 0.7×2600 / 5000 + 0.3×25% = 0.439.
[0112] S2033, adopt a decay factor to dynamically adjust the exploration rate of the Q - learning algorithm, enhance the weight of historical experience during peak load periods, increase the probability of random exploration during off - peak periods, and update the probability distribution of the energy storage priority index; This step balances experience exploitation and unknown exploration through an adaptive learning strategy, improving the decision-making efficiency of the model during different load periods.
[0113] Decay factor mechanism: Time decay function: Define the decay factor λ(t)=λ_base × e^(-kt), where λ_base = 0.9 (base decay rate), k = 0.05 (decay speed coefficient), and t is the number of consecutive explorations. For example, after a certain strategy has not been selected for 3 consecutive times, λ(3)=0.9×e^(-0.05×3)≈0.9×0.861 = 0.775.
[0114] Load period division: Define peak periods (e.g., 10:00 - 12:00, 18:00 - 20:00), flat periods (7:00 - 10:00, 12:00 - 18:00), and low - valley periods (0:00 - 7:00, 20:00 - 24:00) according to the historical load curve.
[0115] Dynamic adjustment of exploration rate: Peak period strategy: When the load ≥ 120% of the regional average load, set the exploration rate ε = 0.1 (10% random exploration), and preferentially select the action with the highest Q value. For example, for Microgrid I during the peak period, the probability of selecting the action +0.2 according to the Q - table is 90%, and the probability of randomly exploring other actions is 10%.
[0116] Low - valley period strategy: When the load ≤ 80% of the regional average load, increase the exploration rate to ε = 0.4 to encourage the discovery of new strategies. For example, at 2:00 am, Microgrid J has a 40% probability of randomly trying the actions -0.1 or 0 to test the feasibility of low - price charging.
[0117] Flat - period transition strategy: Adjust the exploration rate using linear interpolation. For example, when the load is 100% of the average load, ε = 0.25.
[0118] Update of priority probability distribution: Q - value update rule: Q(s,a) ← Q(s,a) + α×[R + γ×max Q(s',a') - Q(s,a)], where α = 0.2 (learning rate), γ = 0.9 (discount factor). For example, the original Q - value of the action +0.1 in state s = 30 (encoded as 5) is 1.2, the new reward R = 0.5, and the maximum Q - value of the next state = 1.5. Then the updated Q = 1.2 + 0.2×(0.5 + 0.9×1.5 - 1.2)=1.2 + 0.2×(0.5 + 1.35 - 1.2)=1.2 + 0.2×0.65 = 1.33.
[0119] Probability distribution generation: For each state s, calculate the Softmax probability of all actions a: P(a|s) = e^(Q(s,a) / τ) / Σe^(Q(s,a') / τ), where τ = 0.5 (temperature coefficient). For example, for three actions with Q values of 1.2, 1.0, and 0.8 respectively, their probabilities are e^(1.2 / 0.5) = e^2.4 ≈ 11.02, e^2 ≈ 7.39, e^1.6 ≈ 4.95, the sum is ≈ 23.36, and the normalized probabilities are 47.1%, 31.6%, and 21.3% respectively.
[0120] S2034, according to the updated priority index and its probability distribution, calculate the ratio of the price change rate to the trading volume change rate as the supply-demand elasticity coefficient, and quantify the risk compensation value of energy storage call based on the conditional value-at-risk model; This step realizes the dynamic balance of market supply and demand and the reasonable sharing of risks through the elasticity coefficient and the risk compensation mechanism.
[0121] Calculation of supply-demand elasticity coefficient: Calculation of price change rate: ΔPrice = (Average electricity price in this period - Average electricity price in the previous period) / Average electricity price in the previous period. For example, the average electricity price of Microgrid L at time t is 0.58 yuan / kWh, and at time t+1 it is 0.63 yuan / kWh, so ΔPrice = (0.63 - 0.58) / 0.58 ≈ 8.62%.
[0122] Calculation of trading volume change rate: ΔVolume = (Trading volume in this period - Trading volume in the previous period) / Trading volume in the previous period. For example, the trading volume at time t is 10 MWh, and at time t+1 it is 13 MWh, so ΔVolume = (13 - 10) / 10 = 30%.
[0123] Calculation of elasticity coefficient: E = ΔVolume / ΔPrice. For example, E = 30% / 8.62% ≈ 3.48, indicating that for every 1% increase in price, the trading volume increases by 3.48%, showing a relatively high demand elasticity.
[0124] Quantification of risk compensation value: Application of CVaR model: Set the confidence level β = 95%, and calculate the conditional expected value below VaR (value at risk) in the profit distribution. For example, in the profit distribution of Microgrid M, the 5% quantile corresponds to VaR = 280 yuan, and CVaR is the average of all profits below 280 yuan (such as 250 yuan).
[0125] Risk compensation calculation: Compensation value = CVaR × risk aversion coefficient η, η = 0.3 (set by user preference). For example, if CVaR = RMB 250 yuan, the compensation value = 250 × 0.3 = RMB 75 yuan / MWh.
[0126] Dynamic adjustment mechanism: When the actual revenue of the microgrid is lower than CVaR for two consecutive periods, η is increased to 0.5; conversely, if the revenue is higher than CVaR for three consecutive periods, η is decreased to 0.2.
[0127] In S2035, encode the supply-demand elasticity coefficient and the energy storage call risk compensation value as a piecewise linear function, generate a dynamic bidding rule including the stepped bid ceiling and penalty clauses, and write it into the consensus verification contract of the blockchain.
[0128] This step ensures the transparency and immutability of the bidding process through structured rule design and blockchain technology.
[0129] Piecewise linear function design: Segmentation of elasticity coefficient: Divide it into three grades according to the elasticity value E: E ≥ 2.0 (high elasticity): Bid ceiling = benchmark price × 1.2 + risk compensation; 1.0 ≤ E < 2.0 (medium elasticity): Bid ceiling = benchmark price × 1.0 + risk compensation; E < 1.0 (low elasticity): Bid ceiling = benchmark price × 0.8 + risk compensation.
[0130] For example, the benchmark price is RMB 0.55 yuan / kWh, E = 2.5, and the risk compensation is RMB 0.10 yuan / kWh, then the bid ceiling = 0.55 × 1.2 + 0.10 = RMB 0.76 yuan / kWh.
[0131] Penalty clause setting: For a microgrid with an actual transaction volume less than 80% of the committed volume, deduct the margin according to the difference ratio. For example, if the committed discharge is 10 MWh and the actual completion is 8 MWh, the difference is 20%, then the deducted margin = 20% × 10 MWh × bid ceiling × 0.5.
[0132] Blockchain contract deployment: Smart contract logic: Write a Solidity contract to implement the following functions: Verify whether the bid conforms to the stepped ceiling rule; Automatically calculate the penalty amount according to the transaction result; Call the off-chain oracle to obtain real-time electricity prices and load data.
[0133] For example, when the microgrid O quotes 0.75 yuan / kWh (exceeding the upper limit of 0.72 yuan / kWh corresponding to its elasticity coefficient E = 1.8), the contract automatically rejects the quote.
[0134] Consensus verification process: Each node reaches a consensus on the bidding rules through the PBFT (Practical Byzantine Fault Tolerance) protocol, and synchronizes the rule hash value every 15 minutes to ensure network-wide consistency.
[0135] Based on the balanced interest distribution result, a time-decaying reinforcement learning algorithm is used to dynamically adjust the energy storage priority index. Strengthen historical experience (such as high arbitrage revenue strategies) during peak load periods, increase random exploration (such as risk diversification strategies) during low load periods, combine the supply-demand elasticity coefficient (the ratio of the price change rate to the volume change rate) and the conditional value-at-risk model to quantify the risk of energy storage invocation (such as the default probability caused by insufficient capacity), and finally generate a stepped bidding rule and a risk compensation clause, and write the rule into the blockchain contract to ensure transparency and immutability. Adapt to the real-time supply-demand changes of the power grid through dynamic bidding rules, use reinforcement learning to balance short-term benefits and long-term risks, improve the market response speed and economy of energy storage resources, and at the same time reduce the impact of uncertainty in the dispatching process through the risk compensation mechanism.
[0136] S204. Based on the dynamic bidding rules, the energy storage resources are allocated in real time through a decentralized gradient consensus mechanism, and the validity of the allocation is verified by combining a verifiable random function, and the final coordinated dispatching instruction is generated and synchronized to each microgrid terminal. Specifically, it may include: S2041. Broadcast the dynamic bidding rules in a decentralized network. Each microgrid node generates a candidate allocation plan based on local constraint conditions and calculates the gradient vector of the plan for the global objective function; The dynamic bidding rules are broadcast to all microgrid nodes participating in coordinated dispatching (such as 10 microgrids) through the blockchain network. After each node receives the rules, it generates a candidate allocation plan based on local energy storage device parameters (such as state of charge SOC = 65%, maximum charge-discharge power 50kW) and the real-time demand of the power grid (such as the need to increase 200kWh of energy storage discharge in the region). The candidate plan needs to meet the following constraints: Charge-discharge power constraint: not exceeding ± the discharge power of the device rated value (such as ± power); SOC safety constraint: The SOC remains between 20% and 80% after charge and discharge; Quotation constraint: according to the stepped quotation upper limit in the dynamic bidding rules (such as ≤ yuan / kWh during peak hours).
[0137] The gradient vector calculation uses a distributed optimization framework: Global objective function: minimize the power grid regulation cost + maximize the total revenue of the microgrid alliance; Local objective function: Each microgrid constructs a revenue function based on its own bid price and SOC status (e.g., revenue = discharge amount × bid price - energy storage loss cost).
[0138] Calculate the gradient of the candidate solution with respect to the global objective function through automatic differentiation techniques (such as Autograd in PyTorch). For example, a certain microgrid node calculates that if the discharge amount is increased from 30 kW to 40 kW, the global cost gradient is -0.2 (a negative gradient indicates a cost reduction), then a gradient vector [-0.2, 0.1,...] (with the same dimension as the number of variables) is generated.
[0139] Example: The candidate solution for Microgrid A is to discharge 35 kW at a bid price of 0.12 yuan / kWh. Calculate its gradient with respect to the global objective as [-0.15, 0.08], indicating that increasing the discharge amount can reduce the global cost, but increasing the bid price will slightly increase the cost.
[0140] S2042, Use the gradient consensus algorithm to aggregate the gradient information of each node, and solve for the optimal allocation solution that satisfies the power grid flow constraints through the projected subgradient descent method to generate preliminary coordinated scheduling instructions; The gradient consensus algorithm achieves decentralized aggregation through the Gossip protocol: Gradient exchange: Each node randomly selects 3 neighbor nodes (such as Microgrids B, C, and D) and sends its local gradient vector; Weighted average: After a node receives the neighbor gradients, it calculates the average gradient according to the weights (such as its own weight of 0.6 and 0.2 for each neighbor); Iterative convergence: After repeating 10 rounds, the gradient difference rate of each node < 1% (regarded as convergence).
[0141] The projected subgradient descent method is used to handle the power grid flow constraints: Variable update: Update the charge and discharge amount and bid price according to the aggregated gradient, with a step size of α; Constraint projection: If the updated variable exceeds the boundary (such as SOC > 80%), project it onto the feasible region (such as force SOC = 79%); Power grid flow verification: Call the Matpower toolkit for power flow calculation to ensure that the line load rate < 95%.
[0142] Example: After 5 iterations, each node reaches a consensus. The optimal solution is a total discharge amount of 180 kW (Microgrid A: 40 kW, Microgrid B: 35 kW...), with an average bid price of 0.13 yuan / kWh and a power flow load rate of 92%. Generate preliminary instructions: "Microgrid A discharges 40 kW from 14:00 to 14:15 at a bid price of 0.13 yuan".
[0143] In S2043, a verifiable random function is called to generate a distributed random number seed. Each node performs a pre-commit operation based on the collaborative scheduling instruction and generates an operation hash. After aggregating all hash values through threshold signature, they are compared with the random number to verify consistency. The verifiable random function (VRF) adopts the ECVRF-ED25519-SHA512 algorithm: Random number generation: Each node uses its private key to generate a random number seed for the current block height (such as Block#7821), and the output is a 512-bit hash value. Seed aggregation: A global random number is generated through threshold signature (Threshold Signature, which requires signatures from more than 2 / 3 of the nodes). For example, the signature aggregation result is 0x3a7d... Hash pre-commit: Before each node executes the scheduling instruction, it calculates the operation hash (such as the SHA-3 hash of the instruction content "Microgrid A discharges 40kW"), and signs and submits it to the blockchain.
[0144] Consistency verification: Hash comparison: All nodes' pre-committed hashes are aggregated through a Merkle tree. The root hash is exclusive-ORed with the random number generated by VRF bit by bit. If the result is consistent, the instruction is considered valid. Threshold verification: If the hashes of 7 / 10 nodes are consistent (the passing rate is 70%), but it is lower than the 95% threshold, then elastic reallocation is triggered.
[0145] Example: Among 10 nodes, 9 submit the same hash (root hash 0x7f2c...), and the exclusive-OR result with the VRF random number 0x7f2c... is 0, and the verification passes.
[0146] In S2044, when the verification passing rate exceeds the preset threshold, the collaborative scheduling instruction is written into the trusted execution environment of each microgrid terminal. Otherwise, the elastic reallocation mechanism is triggered, and the process returns to execute the step of broadcasting dynamic bidding rules in the decentralized network. Each microgrid node generates a candidate allocation plan based on local constraint conditions and calculates the gradient vector of the plan for the global objective function for iterative optimization until convergence. The preset threshold is, for example, 95%.
[0147] The trusted execution environment (TEE) adopts Intel SGX technology to ensure the secure execution of instructions: Instruction writing: The instruction is encrypted and transmitted to the SGX enclave of each microgrid through a secure channel. Secure execution: The instruction is decrypted within the enclave and used to control the energy storage device to prevent malicious tampering (such as the charge and discharge amount being maliciously modified).
[0148] Elastic reallocation mechanism: Trigger condition: The pre-submission hash consistency rate < 95% (e.g., 8 / 10 nodes are consistent); Cause analysis: Identify inconsistent nodes (e.g., abnormal hash submitted by Microgrid E), and reduce its weight factor by 50%; Iterative optimization: Return to step S2041, adjust the gradient aggregation weights (normal node weight +10%, abnormal node weight -30%), and regenerate the candidate solution; Termination condition: The passing rate of three consecutive iterations ≥ the iteration pass rate or the maximum number of iterations is 20 times.
[0149] Example: The passing rate of the first verification is 90%. After identifying Microgrid E as an abnormal node and reducing its weight, the passing rate of the second iteration is increased to 96%, and the instruction is successfully written into the TEE and executed.
[0150] In a decentralized network, each microgrid node generates a candidate allocation solution based on the dynamic bidding rule, aggregates the global gradient information through the gradient consensus algorithm, and uses the projected subgradient descent method to solve the optimal solution that satisfies the power grid power flow constraint. Use a verifiable random function (VRF) to generate a random number seed, verify the consistency of the pre-allocation operation hash values submitted by each node (threshold signature aggregation comparison). If the verification passing rate meets the standard, write it into the trusted execution environment for execution. Otherwise, trigger the elastic reallocation mechanism for iterative optimization to ensure that the allocation result simultaneously meets economy, security, and verifiability, realize efficient resource allocation in a decentralized environment, ensure the global optimality and anti-tampering of the scheduling instruction through the gradient consensus and VRF verification mechanisms, improve the anti-attack ability and collaborative response efficiency of the multi-microgrid system, and adapt to the real-time scheduling requirements in a complex power grid environment.
[0151] It can be seen that according to the state of charge, charge and discharge efficiency of the user's energy storage device and the real-time scheduling requirements of the power grid, a dynamic game initial strategy set is generated through a multi-agent strategy network; based on the dynamic game initial strategy set, an asymmetric Nash bargaining model is used for distributed negotiation to generate an equilibrium interest distribution plan; according to the equilibrium interest distribution plan, the time-decaying reinforcement learning algorithm is used to iteratively correct the energy storage priority index to generate a dynamic bidding rule including supply and demand elasticity coefficients and risk compensation; based on the dynamic bidding rule, the energy storage resources are allocated in real time through a decentralized gradient consensus mechanism, and the final collaborative scheduling instruction is generated and synchronized to each microgrid terminal, so as to improve the operation efficiency and adaptive ability of the multi-microgrid system and meet the requirements of future intelligent distribution networks.
[0152] Another embodiment of the present invention provides a multi-microgrid collaborative scheduling system based on game theory, see Figure 3 , the system may include: The construction module 301 is configured to construct a spatio-temporal coupled game feature space according to the state of charge, charge and discharge efficiency of the user's energy storage device and the real-time grid scheduling requirements, and generate an initial dynamic game strategy set through a multi-agent strategy network. The initial dynamic game strategy set includes the charge and discharge amounts and bid combinations of each microgrid. The negotiation module 302 is configured to perform distributed negotiation based on the initial dynamic game strategy set by using an asymmetric Nash bargaining model, dynamically adjust the weight factors of each microgrid through a virtual resource exchange market, and generate an equilibrium benefit distribution scheme. The weight factors are jointly calculated by the user credit rating and the grid emergency level. The correction module 303 is configured to iteratively correct the energy storage priority index according to the equilibrium benefit distribution scheme by using a time-decaying reinforcement learning algorithm, and generate a dynamic bidding rule including the supply-demand elasticity coefficient and risk compensation. The allocation module 304 is configured to allocate energy storage resources in real time based on the dynamic bidding rule through a decentralized gradient consensus mechanism, verify the effectiveness of the allocation by combining a verifiable random function, and generate a final collaborative scheduling instruction and synchronize it to each microgrid terminal.
[0153] It can be seen that according to the state of charge, charge and discharge efficiency of the user's energy storage device and the real-time grid scheduling requirements, an initial dynamic game strategy set is generated through a multi-agent strategy network; based on the initial dynamic game strategy set, an asymmetric Nash bargaining model is used for distributed negotiation to generate an equilibrium benefit distribution scheme; according to the equilibrium benefit distribution scheme, a time-decaying reinforcement learning algorithm is used to iteratively correct the energy storage priority index to generate a dynamic bidding rule including the supply-demand elasticity coefficient and risk compensation; based on the dynamic bidding rule, energy storage resources are allocated in real time through a decentralized gradient consensus mechanism to generate a final collaborative scheduling instruction and synchronize it to each microgrid terminal, thereby improving the operation efficiency and adaptive ability of the multi-microgrid system and meeting the requirements of the future intelligent distribution network.
[0154] The embodiment of the present invention also provides a storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above method embodiments when running.
[0155] Specifically, in this embodiment, the above storage medium can be configured to store a computer program for executing the following steps: S201, according to the state of charge, charge and discharge efficiency of the user's energy storage device and the real-time grid scheduling requirements, construct a spatio-temporal coupled game feature space, and generate an initial dynamic game strategy set through a multi-agent strategy network. The initial dynamic game strategy set includes the charge and discharge amounts and bid combinations of each microgrid. S202. Based on the initial dynamic game strategy set, use the asymmetric Nash bargaining model for distributed negotiation, and dynamically adjust the weight factors of each microgrid through a virtual resource exchange market to generate an equilibrium benefit distribution scheme, where the weight factors are jointly calculated by the user credit rating and the grid emergency level; S203. According to the equilibrium benefit distribution scheme, use the time-decaying reinforcement learning algorithm to iteratively correct the energy storage priority index to generate a dynamic bidding rule including the supply-demand elasticity coefficient and risk compensation; S204. Based on the dynamic bidding rule, allocate energy storage resources in real time through a decentralized gradient consensus mechanism, verify the effectiveness of the allocation in combination with a verifiable random function, and generate a final collaborative scheduling instruction and synchronize it to each microgrid terminal.
[0156] It can be seen that according to the state of charge, charge and discharge efficiency of the user's energy storage device and the real-time scheduling requirements of the power grid, a dynamic game initial strategy set is generated through a multi-agent policy network; based on the dynamic game initial strategy set, the asymmetric Nash bargaining model is used for distributed negotiation to generate an equilibrium benefit distribution scheme; according to the equilibrium benefit distribution scheme, the time-decaying reinforcement learning algorithm is used to iteratively correct the energy storage priority index to generate a dynamic bidding rule including the supply-demand elasticity coefficient and risk compensation; based on the dynamic bidding rule, the energy storage resources are allocated in real time through a decentralized gradient consensus mechanism to generate a final collaborative scheduling instruction and synchronize it to each microgrid terminal, thereby improving the operation efficiency and adaptive ability of the multi-microgrid system and meeting the requirements of future intelligent distribution networks.
[0157] The embodiment of the present invention also provides an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0158] Specifically, the above electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0159] Specifically, in this embodiment, the above processor may be configured to execute the following steps through a computer program: S201. According to the state of charge, charge and discharge efficiency of the user's energy storage device and the real-time scheduling requirements of the power grid, construct a spatio-temporal coupled game feature space, and generate an initial dynamic game strategy set through a multi-agent policy network, where the initial dynamic game strategy set includes the charge and discharge amounts and bid combinations of each microgrid; S202. Based on the initial dynamic game strategy set, use the asymmetric Nash bargaining model for distributed negotiation, and dynamically adjust the weight factors of each microgrid through a virtual resource exchange market to generate an equilibrium benefit distribution scheme, where the weight factors are jointly calculated by the user credit rating and the grid emergency level; S203. According to the equilibrium benefit distribution scheme, use the time-decaying reinforcement learning algorithm to iteratively correct the energy storage priority index and generate a dynamic bidding rule that includes the supply-demand elasticity coefficient and risk compensation; S204. Based on the dynamic bidding rule, allocate energy storage resources in real time through a decentralized gradient consensus mechanism, verify the effectiveness of the allocation by combining a verifiable random function, and generate a final collaborative scheduling instruction and synchronize it to each microgrid terminal.
[0160] It can be seen that according to the state of charge, charge and discharge efficiency of the user's energy storage device, and the real-time scheduling requirements of the power grid, an initial dynamic game strategy set is generated through a multi-agent policy network; based on the initial dynamic game strategy set, the asymmetric Nash bargaining model is used for distributed negotiation to generate an equilibrium benefit distribution scheme; according to the equilibrium benefit distribution scheme, the time-decaying reinforcement learning algorithm is used to iteratively correct the energy storage priority index to generate a dynamic bidding rule that includes the supply-demand elasticity coefficient and risk compensation; based on the dynamic bidding rule, the energy storage resources are allocated in real time through a decentralized gradient consensus mechanism to generate a final collaborative scheduling instruction and synchronize it to each microgrid terminal, thereby improving the operation efficiency and adaptive ability of the multi-microgrid system and meeting the requirements of future intelligent distribution networks.
[0161] The above has detailed the structure, features and function effects of the present invention according to the illustrated embodiments. The above is only a preferred embodiment of the present invention, but the present invention is not limited to the scope defined by the drawings. Any changes made according to the concept of the present invention, or equivalent embodiments modified to equivalent changes, still within the spirit covered by the specification and drawings, shall be within the protection scope of the present invention.
Claims
1. A multi-microgrid collaborative scheduling method based on game theory, characterized in that, The method comprises: According to the state of charge, charging and discharging efficiency and real-time dispatching requirements of the power grid of the user's energy storage equipment, a time-space coupled game feature space is constructed, and a dynamic game initial strategy set is generated through a multi-agent strategy network, wherein the dynamic game initial strategy set includes the charging and discharging amount and quotation combination of each microgrid; Based on the initial strategy set of the dynamic game, an asymmetric Nash bargaining model is used for distributed negotiation, and the weight factors of each microgrid are dynamically adjusted through the virtual resource exchange market to generate a balanced benefit distribution plan, wherein the weight factors are jointly calculated by the user's credit rating and the urgency of the power grid; According to the balanced benefit distribution scheme, the energy storage priority index is iteratively corrected using a time-decayed reinforcement learning algorithm to generate a dynamic bidding rule including supply and demand elasticity coefficients and risk compensation; Based on the dynamic bidding rules, energy storage resources are allocated in real time through a decentralized gradient consensus mechanism, and the validity of the allocation is verified in combination with a verifiable random function to generate the final coordinated scheduling instructions and synchronize them to each microgrid terminal.
2. The method according to claim 1, wherein According to the state of charge, charging and discharging efficiency and real-time dispatching requirements of the power grid of the user's energy storage equipment, a time-space coupled game feature space is constructed, and a dynamic game initial strategy set is generated through a multi-agent strategy network, wherein the dynamic game initial strategy set includes the charging and discharging amount and quotation combination of each microgrid, including: Collect the real-time state of charge data and historical charge and discharge efficiency curves of each microgrid energy storage device, combine the regional load forecast and electricity price fluctuation signals released by the power grid dispatching center, perform spatiotemporal interpolation processing, and generate a spatiotemporal coupling state matrix with a preset time length as the granularity; The spatiotemporal coupling state matrix is input into a convolutional recursive hybrid neural network, and the charge-discharge correlation features between adjacent microgrids are extracted through a time sliding window, and a game relationship map containing the strength of the supply-demand relationship and capacity constraints is output; Based on the game relationship graph, an adversarial training environment for the multi-agent strategy network is constructed. The double-layer strategy gradient algorithm is used to simultaneously optimize the charging and discharging strategies and bidding strategies of each microgrid to generate an initial strategy set containing Nash equilibrium constraints. The initial strategy set is screened by Pareto frontier, the strategy combinations that violate the constraints of safe and stable operation of the power grid are eliminated, and the initial strategy set of the dynamic game is output.
3. The method according to claim 2, characterized in that, Based on the initial strategy set of the dynamic game, an asymmetric Nash bargaining model is used for distributed negotiation, and the weight factors of each microgrid are dynamically adjusted through the virtual resource exchange market to generate a balanced benefit distribution plan, wherein the weight factor is jointly calculated by the user's credit rating and the grid urgency, including: Analyze the charge and discharge volume and quotation combination in the initial strategy set, initialize the threat point and bargaining space of the asymmetric Nash bargaining model, and set the threat point as the minimum profit threshold when the microgrid operates independently; According to the historical performance records of microgrid users and the voltage over-limit probability of grid nodes, the credit score and urgency coefficient are calculated respectively, and the initial weight factor is generated by fusing the two types of parameters through the hyperbolic tangent function. A distributed negotiation protocol is deployed in the virtual resource exchange market to conduct distributed negotiation. Each microgrid calculates the benefit utility of the strategy set based on the current weight factor and iteratively updates the bargaining solution through the alternating direction multiplier method. When the difference in Shapley values between two adjacent iterations is less than the preset threshold, a dynamic adjustment mechanism for the weight factor is triggered. Among them, the weight of the microgrid with a decreased credit rating decays exponentially, and the weight of the microgrid within the emergency area of the power grid increases linearly. Generate an equilibrium benefit distribution plan containing the transfer payment plan and the charging and discharging timing plan according to the final negotiation result, and verify the individual rationality and coalition stability of the plan through zero-knowledge proof technology.
4. The method according to claim 3, characterized in that, According to the equilibrium benefit distribution plan, use the time-decaying reinforcement learning algorithm to iteratively correct the energy storage priority index, and generate dynamic bidding rules including the supply-demand elasticity coefficient and risk compensation, including: Extract the historical winning bid rate and revenue fluctuation data of each microgrid from the equilibrium benefit distribution plan, and initialize the energy storage priority index as a weighted function of the supply-demand ratio and the risk value. Construct a state transition model for time-decaying reinforcement learning, define the state as the difference between the real-time load of the power grid and the available capacity of the energy storage, the action as the priority adjustment range, and the reward function includes the electricity price arbitrage income and the power grid regulation contribution degree. Adopt a dynamic adjustment of the decay factor to explore the rate of the Q-learning algorithm, enhance the weight of historical experience during peak load periods, increase the probability of random exploration during low load periods, and update the probability distribution of the energy storage priority index. According to the updated priority index and its probability distribution, calculate the ratio of the price change rate to the volume change rate as the supply-demand elasticity coefficient, and quantify the risk compensation value of energy storage invocation based on the conditional value-at-risk model. Encode the supply-demand elasticity coefficient and the risk compensation value of energy storage invocation as a piecewise linear function, generate dynamic bidding rules including the ladder bid price ceiling and penalty clauses, and write them into the consensus verification contract of the blockchain.
5. The method according to claim 4, wherein Based on the dynamic bidding rules, allocate energy storage resources in real time through the decentralized gradient consensus mechanism, verify the effectiveness of the allocation by combining with the verifiable random function, and generate the final coordinated scheduling instruction and synchronize it to each microgrid terminal, including: Broadcast the dynamic bidding rules in the decentralized network, each microgrid node generates a candidate allocation plan based on local constraint conditions, and calculates the gradient vector of the plan for the global objective function. Adopt the gradient consensus algorithm to aggregate the gradient information of each node, and solve the optimal allocation solution that satisfies the power grid power flow constraint through the projected subgradient descent method to generate a preliminary coordinated scheduling instruction. Call the verifiable random function to generate a distributed random number seed, each node performs a pre-submission operation based on the coordinated scheduling instruction and generates an operation hash, and all hash values are aggregated by threshold signature and compared with the random number to verify the consistency. When the verification pass rate exceeds the preset threshold, write the coordinated scheduling instruction into the trusted execution environment of each microgrid terminal, otherwise trigger the elastic reallocation mechanism, and return to execute the step of broadcasting the dynamic bidding rules in the decentralized network, each microgrid node generates a candidate allocation plan based on local constraint conditions, and calculates the gradient vector of the plan for the global objective function, so as to perform iterative optimization until convergence.
6. A multi-microgrid collaborative scheduling system based on game theory, characterized in that, The system includes: A construction module, which is used to construct a spatio-temporal coupled game feature space according to the state of charge, charge and discharge efficiency of the user's energy storage device, and the real-time grid scheduling requirements, and generate an initial dynamic game strategy set through a multi-agent strategy network, where the initial dynamic game strategy set includes the charge and discharge amounts and bid combinations of each microgrid; A negotiation module, which is used to perform distributed negotiation based on the initial dynamic game strategy set by using an asymmetric Nash bargaining model, and dynamically adjust the weight factors of each microgrid through a virtual resource exchange market to generate an equilibrium benefit distribution scheme, where the weight factors are jointly calculated by the user credit rating and the grid urgency; A correction module, which is used to iteratively correct the energy storage priority index according to the equilibrium benefit distribution scheme by using a time-decaying reinforcement learning algorithm, and generate a dynamic bidding rule including the supply-demand elasticity coefficient and risk compensation; An allocation module, which is used to allocate energy storage resources in real time based on the dynamic bidding rule through a decentralized gradient consensus mechanism, verify the effectiveness of the allocation by combining a verifiable random function, and generate a final coordinated scheduling instruction and synchronize it to each microgrid terminal.
7. The system according to claim 6, wherein The construction module is specifically used for: Collecting the real-time state of charge data and historical charge and discharge efficiency curves of each microgrid energy storage device, combining the regional load forecast and electricity price fluctuation signals released by the grid dispatching center, and performing spatio-temporal domain interpolation processing to generate a spatio-temporal coupled state matrix with a preset time interval as the granularity; Inputting the spatio-temporal coupled state matrix into a convolutional recurrent hybrid neural network, extracting the charge and discharge correlation features between adjacent microgrids through a time sliding window, and outputting a game relationship graph including the supply-demand relationship intensity and capacity constraints; Based on the game relationship graph, constructing an adversarial training environment for the multi-agent strategy network, and synchronously optimizing the charge and discharge amount strategies and bid strategies of each microgrid by using a double-layer policy gradient algorithm to generate an initial strategy set including Nash equilibrium constraints; Performing Pareto front screening on the initial strategy set, eliminating the strategy combinations that violate the grid safety and stable operation constraints, and outputting the initial dynamic game strategy set.
8. The system according to claim 7, wherein The negotiation module is specifically used for: Analyzing the charge and discharge amounts and bid combinations in the initial strategy set, initializing the threat point and bargaining space of the asymmetric Nash bargaining model, and setting the threat point as the minimum revenue threshold when the microgrid operates independently; Calculating the credit score and urgency coefficient respectively according to the historical performance records of microgrid users and the probability of grid node voltage violation, and generating an initial weight factor by fusing the two types of parameters through a hyperbolic tangent function; Deploying a distributed negotiation protocol in the virtual resource exchange market for distributed negotiation, each microgrid calculates the interest utility of the strategy set based on the current weight factor, and iteratively updates the bargaining solution through the alternating direction multiplier method; When the difference in the Shapley values of two adjacent iterations is less than a preset threshold, triggering a dynamic weight factor adjustment mechanism, where the weight of the microgrid with a decreased credit rating decays exponentially, and the weight of the microgrid in the grid emergency area increases linearly; Generating an equilibrium benefit distribution scheme including a transfer payment scheme and a charge and discharge time sequence plan according to the final negotiation result, and verifying the individual rationality and coalition stability of the scheme through zero-knowledge proof technology.
9. A storage medium, characterized in that, A computer program is stored in the storage medium, wherein the computer program is configured to execute the method according to any one of claims 1-5 when running.
10. An electronic device, comprising a memory and a processor, characterized in that A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method according to any one of claims 1-5.
Citation Information
Patent Citations
Game theory-based multi-micro-grid interconnection running optimization method
CN107545325A
Multi-microgrid collaborative optimization scheduling method based on cooperative game and consistency algorithm
CN113065696A
Large-scale electric vehicle charging and discharging optimization scheduling method based on multi-main-body double-layer game
CN114662759A
Optimized operation strategy of multi-microgrid shared energy storage in power distribution network based on mixed game
CN117875479A
Distribution network distributed optimization scheduling method and system under multi-interest subject game
CN119051038A
Cited By
Intelligent park source network load storage and charging integrated scheduling method based on AI
CN120474006A
Virtual machine network address dynamic management method and system based on neural network
CN120475016A
A method and system for dynamic management of virtual machine network addresses based on neural network
CN120475016B
Energy storage power station SOC automatic calibration and inter-cluster dynamic balance cooperative control method
CN120722263A
Generative artificial intelligence-based power grid resource allocation method and system
CN120746209A