Microgrid optimal scheduling method and system combining double-layer game and a3c

By combining two-level game theory with the A3C algorithm, a microgrid optimization scheduling method is proposed, which solves the problems of response lag and training stability of traditional systems in complex environments. This method enables real-time, robust, and low-carbon operation of the microgrid, improving the system's economy and deployability.

CN122118966AActive Publication Date: 2026-05-29SHANGHAI UNIVERSITY OF ELECTRIC POWER

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI UNIVERSITY OF ELECTRIC POWER
Filing Date
2026-04-22
Publication Date
2026-05-29

Smart Images

  • Figure CN122118966A_ABST
    Figure CN122118966A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of microgrid optimization scheduling method and system combining double-layer game and A3C, comprising the following steps: constructing source-grid-load-storage integrated microgrid model;With the total income of maximum microgrid operation as the goal, combined with the operation constraint of microgrid, the constraint of different external environmental conditions factors to the on-site renewable energy power output, power constraint and microgrid power supply and demand balance constraint, establish the microgrid optimization scheduling model based on A3C;Extract state space feedback from environment, make pricing decision through upper Stackelberg game, as the leader of large power grid, the decision of microgrid as follower obtains current power purchase and V2G usage, updates price signal;Solve the trained microgrid scheduling model, obtain the strategy output by lower reinforcement learning at current time step.The present application has the advantages of realizing dynamic price signal closed-loop linkage, suppressing early convergence through adaptive exploration, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of microgrid scheduling technology, and in particular to a microgrid optimization scheduling method and system that combines two-level game theory and A3C. Background Technology

[0002] With the large-scale integration of distributed photovoltaic and other renewable energy sources, microgrids are rapidly developing in scenarios such as industrial parks, islands, and remote areas. These systems exhibit significant time-varying and volatile characteristics on the load side, and randomness and intermittency on the source side. Simultaneously, the widespread participation of electric vehicles makes vehicle-to-grid (V2G) interaction a crucial factor affecting power balance and economic benefits. Traditional energy management systems (EMS) largely rely on fixed time-of-use pricing and preset control strategies. Their response to external electricity prices, carbon trading costs, and V2G supply and demand changes is often delayed, easily leading to curtailment of solar power, increased peak-hour electricity purchase costs, and supply-demand instability. This makes it difficult to meet the stable, economical, and low-carbon operation requirements of large-capacity microgrids in complex coupled environments.

[0003] Chinese patent application publication number CN118473021A discloses a microgrid optimization scheduling method and system combining CMA-ES and DDPG algorithms. By establishing mathematical models for each unit of the microgrid, selecting optimization indices with the goal of maximizing total revenue or minimizing total cost, and combining evolutionary algorithms and deep reinforcement learning for policy optimization, the method improves the adaptability and scheduling efficiency of the strategy to a certain extent. However, with the increasing marketization and the introduction of carbon trading mechanisms, existing solutions still have shortcomings in the following aspects: First, the coupling between environmental factors and electric vehicle factors is insufficient; carbon emissions and environmental protection costs are often handled retroactively or in a simplified manner, making it difficult to price and settle V2G behavior based on user-side state of charge (SOC) and market supply and demand. Second, there is insufficient real-time linkage with the electricity market; electricity prices, carbon prices, and V2G prices often participate in optimization offline or with fixed parameters. Third, reinforcement learning suffers from poor training stability and sample efficiency in non-stationary, high-noise environments, lacking adaptive control of policy exploration intensity and a mechanism for prioritizing key samples, resulting in slow convergence and weak online update capabilities.

[0004] Therefore, there is an urgent need for a microgrid optimization scheduling technology solution that integrates source-grid-load-storage and simultaneously considers carbon trading and V2G markets. This solution should be able to achieve closed-loop linkage updates of price and scheduling at the time step level, and optimize stability and efficiency through reinforcement learning training to meet the real-time, robust and low-carbon operation requirements of large-capacity microgrids in complex coupled environments. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the existing technology and provide a microgrid optimization scheduling method and system that combines two-layer game theory and A3C. It aims to achieve coordinated solution and stable convergence of strategies in continuous time and multi-constraint operating environments, taking into account both economy and emission reduction, and improving deployability and robustness in real-world scenarios.

[0006] The objective of this invention can be achieved through the following technical solutions: One aspect of the present invention provides a microgrid optimal scheduling method combining two-level game theory and A3C, comprising the following steps: Construct an integrated microgrid model encompassing power generation, grid, load, and energy storage; With the goal of maximizing the total revenue during microgrid operation, and taking into account the microgrid's operational constraints, the constraints on renewable energy power output, power constraints, and the power supply and demand balance constraints of the microgrid, a dynamic scheduling model for microgrids based on A3C is established. The state space feedback is extracted from the environment. Through the upper-level Stackelberg game, the large power grid, as the leader, makes pricing decisions, while the microgrid, as the follower, makes decisions to obtain the current electricity purchase and V2G usage, and updates the price signal. Based on the price signal, the trained microgrid scheduling model is solved to obtain the strategy output by the lower-level reinforcement learning at the current time step.

[0007] As a preferred technical solution, in the A3C algorithm of the microgrid optimization scheduling model, the state space of the agent in the environment includes load demand, photovoltaic power generation, energy storage system SOC value, energy storage system charging power, energy storage system discharging power, energy storage system remaining capacity, gas micro turbine power, fuel cell power, V2G dynamic price, time-of-use price, electric vehicle (EV) charging amount, EV discharging amount, and charging pile activity ratio. The action space of the agent includes energy storage charging and discharging actions, V2G discharging actions, V2G price adjustment actions, and demand response actions.

[0008] As a preferred technical solution, during the training process of the microgrid optimal scheduling model, the Actor network and / or Critic network are trained based on importance sampling weights. The process of obtaining the importance sampling weights includes the following steps: Based on time difference error and preset control priority Calculate experience priority ; Based on experience priority and the minimum sampling probability in the preset experience pool Calculate importance sampling weights .

[0009] As a preferred technical solution, the Actor network update process during the training of the microgrid optimal scheduling model includes the following steps: Target value is calculated using reverse rolling based on intraday time steps. The time difference error is obtained. ; With time difference error As an advantage, by blocking gradient flow to prevent inverse dependence on the Critic network, the product of the policy log probability and the advantage is calculated as the core expectation term. ; Using policy entropy as the regularization term, based on the recent moving average of policy entropy and target policy entropy Calculate adaptive weights Among them, the moving average of policy entropy With adaptive weights Inverse correlation; Based on the adaptive weights and the importance sampling weight The policy entropy is weighted and combined with the core expectation term. We obtain the loss function value of the Actor network and then perform training through policy gradient ascent optimization.

[0010] As a preferred technical solution, the Critic network update process during the training of the microgrid optimal scheduling model includes the following steps: Calculate the difference between the target value and the actual value at each step, and use the importance sampling weights. The weighted loss function value of the Critic network is obtained, and the network is updated based on the loss function value to achieve training.

[0011] As a preferred technical solution, the microgrid model constructs a deep reinforcement learning-based simulation environment within the microgrid's operating environment, based on an objective function and constraints, to reflect key factors and dynamic changes in microgrid operation. The objective function is: In the formula, Let be the objective function. for Carbon benefits from EV charging The electricity generation costs are respectively for gas turbines and fuel cells. The carbon costs of gas turbines and fuel cells, respectively. For the cost of V2G, The cost of purchasing electricity and the carbon cost of microgrids The cost of responding to demand.

[0012] As a preferred technical solution, in the microgrid optimized scheduling model, the components on the power supply side include solar photovoltaic, energy storage systems and generator sets, and the components on the load side include electric vehicle loads and residential loads.

[0013] As a preferred technical solution, the modeling process of each component in the microgrid optimized scheduling model includes: The power generation cost of gas turbines is modeled based on the active power output of gas turbines, natural gas prices, gas turbine efficiency, and low calorific value of natural gas, and upper and lower limits of output power and ramp rate constraints are constructed. Based on the power generation cost, active power output, and efficiency of fuel cells, the power generation cost of fuel cells is modeled, and upper and lower limits of output power and ramp rate constraints are constructed. Based on the state of charge ratio, previous state of charge, state of charge, charge and discharge, charge and discharge efficiency, charge and discharge power, and rated capacity of the energy storage system, the state of charge of the energy storage system is modeled, and charge and discharge power constraints and state of charge boundary constraints are constructed. Construct constraints on the balance between power supply and demand in microgrids.

[0014] As a preferred technical solution, the strategy includes energy storage charging and discharging, V2G trading, and electricity purchase and sale decisions.

[0015] Another aspect of the present invention provides a microgrid optimal scheduling system combining two-level game theory and A3C, for implementing the aforementioned microgrid optimal scheduling method, the system comprising: The upper-level leader pricing module is used to jointly optimize electricity and carbon prices based on the state-space feedback of microgrid operation and market supply and demand information, and output price signals. The mid-level microgrid follower pricing module is used to obtain the V2G electricity price and V2G carbon price based on the price signal, combined with the energy storage status and V2G supply and demand ratio, and to obtain the current electricity purchase and V2G usage from the grid. The lower-level A3C scheduling module is used for policy learning based on target entropy-driven adaptive entropy coefficients, exponential decay of learning rate, and gradient clipping. The microgrid environment module is used to inject the price signal into the environment and collect new state feedback in real time after each scheduling step and A3C strategy update, driving the next round of pricing and decision-making in the upper and middle layers to form an iterative closed loop.

[0016] Compared with the prior art, the present invention has at least one of the following beneficial effects: (1) Realize dynamic price signal closed-loop linkage: This invention extracts state space feedback from the environment, and through the upper-level Stackelberg game, the large power grid, as the leader, makes pricing decisions, and the microgrid, as the follower, makes decisions to obtain the current electricity purchase and V2G usage, and updates the price signal. Based on the price signal, the trained microgrid scheduling model is solved to obtain the strategy output by the lower-level reinforcement learning at the current time step. The upper layer gives the electricity price / carbon price / V2G price through the extended Stackelberg game, and the lower-level A3C executes at the device layer and feeds back the real state space to the upper layer, forming a closed-loop linkage of price, scheduling, and feedback, and improving the consistency between price and operation decisions.

[0017] (2) Suppressing early convergence through adaptive exploration: During the training process of the A3C Actor network in this invention, early convergence is suppressed based on adaptive weights. Importance sampling weights Weight the policy entropy and combine it with the core expectation term. The loss function value of the Actor network is obtained, and training is achieved through policy gradient ascent optimization, where the policy entropy moving mean is used. With adaptive weights Anticorrelation, when hour, To enhance exploration, when When it is large, To enhance utilization and convergence.

[0018] (3) Improve training effectiveness: This invention is based on time difference error and preset control priority Calculate experience priority Based on experience priority and the minimum sampling probability in the preset experience pool Calculate importance sampling weights It also participates in the training process of the Actor and Critic networks to improve the training effect.

[0019] (4) Achieve two-layer decoupling of price and strategy: In the two-layer architecture of this invention, the upper-layer price is solved through the leader-follower structure to ensure price boundary and convergence, while the lower-layer A3C is responsible for device and behavior optimization to avoid strong coupling leading to instability. Attached Figure Description

[0020] Figure 1 The flowchart shows the microgrid optimal scheduling method combining two-level game theory and A3C in the embodiment. Figure 2 This is a schematic diagram of the microgrid structure in the embodiment; Figure 3This is a flowchart of microgrid optimization scheduling combined with the A3C algorithm in the embodiment. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0022] Example 1 To address the problems existing in the aforementioned technologies, this embodiment provides a microgrid optimization scheduling method that combines two-layer game theory and A3C. It aims to achieve coordinated solution and stable convergence of strategies such as electricity price, carbon price, vehicle-grid interaction price, and energy storage charging and discharging under continuous time and multi-constraint operating environments, taking into account both economic efficiency and emission reduction, and improving deployability and robustness in real-world scenarios.

[0023] See Figure 1 The method includes the following steps: Step S1: Establish mathematical models of some microgrid units and load PV, EV and other data.

[0024] like Figure 2 As shown, the constructed microgrid model is an integrated source-grid-load-storage system. The power source side includes solar photovoltaic (PV) systems, energy storage systems, and generator sets (including fuel cells and gas turbines), while the load side includes electric vehicle (EV) loads and residential loads. It also includes a microgrid management system and a higher-level grid. This method can be deployed in a microgrid management center. In this step, the first state space is generated by establishing mathematical models of some microgrid units and loading PV, EV, and other data.

[0025] Establish the formula for the power generation cost of gas micro turbines: in For time step Operating costs of gas turbines For the number of gas turbines, For the first Active power output of a gas turbine (unit: kW). The price is for natural gas (unit: yuan / cubic meter). For the first The efficiency of a gas turbine. The low calorific value of natural gas (unit: kWh / cubic meter).

[0026] The output power upper and lower limits that gas micro turbines need to meet are: in and These are the minimum and maximum output power of the gas turbine (unit: kW).

[0027] The ramp rate constraint that the gas micro turbine needs to meet: in To limit the ramp rate, the power variation between adjacent time steps is restricted.

[0028] Establish the formula for fuel cell power generation cost: in For time step The cost of generating electricity from fuel cells, For the number of fuel cells, For the first Active power output of a fuel cell (unit: kW). For the first The efficiency of Taiwan's fuel cells.

[0029] Fuel cells need to meet the following upper and lower limits of output power constraints: in and These represent the minimum and maximum output power of the fuel cell.

[0030] Fuel cells need to meet the following ramp rate constraints: in The ramp rate is limited for fuel cells.

[0031] Establish a dynamic SOC model for an Energy Storage System (ESS): in For time step The state of charge ratio of the energy storage system (between 0 and 1). For the previous time step The state of charge; and This is a charging / discharging status indicator variable, which is binary and takes the value 0 or 1, and they are mutually exclusive; and For charging and discharging efficiency, energy loss must be considered; and Charge / discharge power (unit: kW); Rated capacity of the energy storage system (unit: kWh); For time step.

[0032] Energy storage systems must meet the following charging and discharging power constraints: in and These represent the maximum charge / discharge power and mutual exclusion constraints, respectively. To ensure that charging and discharging cannot occur simultaneously, the state variable is binary: .

[0033] The SOC (State of Charge) of an energy storage system must meet the SOC boundary constraints: ,in and For SOC at time step The minimum and maximum values ​​at that time.

[0034] Establish microgrid power supply and demand balance constraints: The left side represents the total power supply, including: Photovoltaic power generation capacity Gas turbine output power Fuel cell output power Power purchased from the upper-level power grid Discharge power of energy storage system Electric vehicle discharge power (V2G mode); The right side shows the total power consumption, including: Electric vehicle charging power Energy storage system charging power Transferable load power, Actual load after implementing demand response.

[0035] Step S2: With the goal of maximizing the total revenue of the microgrid during system operation, and combining the operational constraints of each unit of the microgrid system, the constraints of different external environmental conditions on the renewable energy power output on site, power-related constraints, and the power supply and demand balance constraints of the microgrid, a dynamic scheduling model for the microgrid based on the A3C algorithm is established.

[0036] In this embodiment, within the microgrid's operating environment, a deep reinforcement learning model guides the agent's decision-making by constructing a simulation environment. The specific steps for establishing the deep reinforcement learning model are as follows: Within the microgrid's operating environment, a deep reinforcement learning-based simulation environment is constructed based on the objective function and constraints to ensure a comprehensive reflection of the key factors and dynamic changes in the microgrid's operation. This includes the following steps: Configure the state space of the agent in the environment: The state space provides the agent with comprehensive information about the current state of the microgrid, including economic and energy factors, so that the agent can make optimal decisions.

[0037] The meanings of each dimension are as follows: For current load demand, Photovoltaic power generation capacity, This refers to the SOC value of the energy storage system. The charging power for the energy storage system, This refers to the discharge power of the energy storage system. For the remaining capacity of the energy storage system, For the power of the gas micro turbine, For fuel cell power, For V2G dynamic pricing, For time-of-use electricity pricing, Charge for EVs, For EV discharge capacity, This represents the percentage of active charging stations.

[0038] Define the action space of the agent in the environment: The action space is the set of all possible actions that the agent can choose. These actions represent the ways in which the agent interacts with the environment. This defines the action space for the microgrid optimization scheduling problem. The actions that an intelligent agent can take are defined as follows: in, For energy storage charging and discharging operations, This is a V2G discharge operation. This is a V2G price adjustment. This is a demand response action.

[0039] The action space reflects the actions that an agent can take, which will affect the operation and performance of the microgrid.

[0040] 3. Establish an agent reward function. The reward function is defined as the immediate reward received by the agent after taking a corresponding action in a preset state. Feedback guides the agent's behavior, with the goal of maximizing the total reward obtained throughout the learning process. Represented as: in, Let $\frac{ ...

[0041] After the above steps are completed, the A3C neural network is trained based on the agent's state, actions, and reward function to obtain the target microgrid optimal scheduling model. This model takes into account external electricity prices, carbon prices, V2G price signals, and renewable energy output in each scheduling cycle, and outputs control commands for energy storage charging and discharging, generator output, V2G trading, and electricity purchase and sale.

[0042] The training method for the microgrid optimal scheduling model includes the following steps: During the initialization phase, the microgrid environment is created and loaded. Collect the data to generate the first state space. Construct the global network. With 4 local Initialize optimizer Exponentially decaying learning rate. Establish a priority replay experience pool. Used for storage and sample transfer.

[0043] host Initialize the two-level game solver and periodically execute the upper-level game: extract state space feedback from the environment to construct market information. Leaders (large power grids) and followers (microgrids) set their own pricing and power purchase agreements. Decision-making involves updating price signals based on inertia weighting and supply-demand fine-tuning. The updated version Injected into the environment, it serves as "effective electricity price / carbon price" in subsequent interactions. The word "price" was used.

[0044] each In the current state The policy network outputs the action distribution and samples the actions. It supports returning policy entropy for recording and interacting with the environment to obtain the next state. With rewards , will transfer Enter the experience pool and local buffer.

[0045] Record discount factors Calculate the target value by reverse rolling based on intraday time steps. At the end of the day The time difference error is obtained. This is used as an advantage approximation to directly drive policy improvement and value function fitting, and is applied to advantage approximation in Actor networks and regression in Critic networks. Relative to the total round return, It can transmit reward signals to the current action more quickly, significantly reduce variance and improve credit allocation for long sequences, making it suitable for daily-round microgrid scheduling tasks.

[0046] The Actor network updates optimize parameters through policy gradient ascent. The core expectation term is the product of the policy log probability and the advantage: ,in As an advantage approximation and through Block the inverse dependency on Critic. To encourage exploration and prevent premature policy convergence, policy entropy is introduced. As a regularization term, and with adaptive weights ,in The recent entropy moving average, The target entropy. This mechanism makes... Inversely changing to the entropy level: when policy uncertainty is low, i.e. hour, To enhance exploration; when strategy uncertainty is high, i.e. When it is large, To enhance utilization and convergence. The final Actor loss is defined as the expected negative value with importance sampling weights: Policy gradient ascent is achieved by minimizing this loss; where Importance sampling derived from priority playback is used to eliminate estimation bias caused by non-uniform sampling.

[0047] Priority playback achieves efficient storage and sampling through a SumTree structure, defining empirical priorities as follows: ,in For TD error, Setting it as a constant ensures that all samples can be sampled. Control priority; sampling probability is This allows high-priority samples to be accessed more frequently. To correct for the bias introduced by non-uniform sampling, importance sampling weights are used. ,in The minimum sampling probability, The error is gradually increased to 1 during training to progressively eliminate corrections. The absolute value of the TD error is used when updating priorities in batches. This mechanism aims to focus training on sparse key events, such as high-price windows, constraint boundary triggers, and termination steps. Larger sample sizes significantly improve sample efficiency.

[0048] The Critic network updates the value function by minimizing the weighted mean square error. The loss function is defined as Its purpose is to provide a stable state-value baseline for policy gradients. The motivation is that an accurate baseline allows the advantage function, i.e., The error is close to zero mean, which significantly reduces the variance of the policy gradient and avoids drastic fluctuations in policy updates. Especially in complex reward function scenarios involving multiple cost and benefit decompositions, this variance reduction is crucial for training stability and ensures that the gradient direction accurately reflects the direction of policy improvement.

[0049] Expected items With policy entropy Together they constitute the update signal, using an adaptive entropy coefficient. Controlled exploration—utilizing trade-offs; The value function is fitted using a weighted mean square error. The final loss is: Experience Write During sampling Obtain batches and calculate Correcting non-uniform sampling bias; batch use Update priority. The sampling batch is used for a single playback update, which is performed in parallel with the main interactive update.

[0050] The local network constructs gradients and performs norm clipping, then calls the optimizer to update global parameters with a global step size. ), then pull ( This ensures that local parameters are consistent with global parameters; the learning rate decays exponentially step by step to guarantee stable convergence in the later stages of iteration.

[0051] The objective function of the microgrid optimal scheduling model is: in, Carbon benefits from charging EVs The power generation cost and carbon cost of the generator set, For V2G cost, The costs of purchasing electricity from the main grid and carbon costs for microgrids. Cost of responding to demand.

[0052] in, The following method for quantifying EV carbon emission reduction based on fuel substitution effect is used for calculation: in For fuel carbon intensity, For the thermal efficiency of internal combustion engines, For the carbon intensity of the power grid, The substitution efficiency coefficient (0–1) represents the proportion of fuel-powered travel that can be replaced by charging. Charging battery cars, The charging energy of electric vehicles over a time step The EV charging power within the current time step. for EV charging power at all times.

[0053] By introducing a quantitative method for EV carbon emission reduction based on the fuel substitution effect, electric vehicles are transformed from simple electricity consumers into environmentally beneficial prosumers. By dynamically embedding real-time carbon trading costs and benefits into the payoff function of a two-layer game, an economic coupling mechanism between carbon price signals and electricity price signals is established. This achieves spontaneous guidance for low-carbon operation while pursuing the minimum system cost. The payoff function refers to the payoff / reward function of each participant in the two-layer Stackelberg game, which is used to describe the impact of electricity price, carbon price, and V2G trading on the benefits of each party and drive pricing and decision-making iterations.

[0054] Step S3: Initialize the two-layer game solver and periodically execute the upper-layer game: extract state space feedback from the environment and construct market information. Leaders (large power grids) and followers (microgrids) set their own pricing and power purchase agreements. Decision-making involves updating price signals based on inertia weighting and supply-demand fine-tuning. The updated version Injected into the environment, serving as the electricity price / carbon price / in subsequent interactions. The price was used.

[0055] Specifically, in the two-level Stackelberg game, a Stackelberg two-level game structure is used to collaboratively optimize electricity prices, carbon prices, and V2G prices. The upper level consists of large grid operators as leaders, the middle level consists of microgrid operators as followers, and market feedback from active V2G user groups is also considered. Price signals are uniformly represented as quaternions. It is managed by the price signal structure and updated iteratively.

[0056] The market status is composed of actual operational feedback, denoted as Including total demand Internal power generation Net purchased electricity V2G capacity and supply and demand indicators, energy storage status and unit power, dynamic electricity price and time-of-use base price, and the proportion of renewable energy. This state space drives the decisions of each participant, ensuring that price evolution aligns with physical operations.

[0057] Leaders' electricity pricing features a markup structure anchored to time-of-use pricing: in For profit margin, Weights are used for dynamic electricity price adjustments. Electricity price boundary constraints are... .

[0058] Leaders adjust carbon prices based on net carbon trading demand and the proportion of renewable energy. This reduces carbon emissions from heat engines. V2G Carbon emissions from electricity purchases Net carbon demand Then carbon price and constraints Total Leader Rewards ,in This is equivalent to the electricity purchased by the microgrid.

[0059] Followers make power balancing and electricity purchase decisions based on a state space. (Electricity purchase amount) ,in Net supply for V2G.

[0060] V2G electricity and carbon prices are intelligently set by the microgrid based on grid prices and equipment status. V2G electricity price The constraints are as follows: V2G carbon price .

[0061] Microgrid revenue is calculated by subtracting the costs of purchasing electricity, internal generation, and V2G electricity from user electricity sales revenue, and also includes V2G carbon revenue. in , The cost of internal power generation, This represents the carbon gain that the microgrid obtains from the electric vehicle side.

[0062] Convergence determination and benefit assessment. The step difference for all four price categories is less than the tolerance. The convergence is determined at the time of convergence; at the convergence or iteration upper limit, the revenue of leaders, followers and V2G users is calculated, and energy efficiency, carbon efficiency and economic efficiency indicators are evaluated based on the real state space as a verification of the price strategy effect.

[0063] Step S4: Solve the microgrid optimization scheduling model based on the A3C algorithm: Using external electricity price, carbon price and V2G price signals and renewable energy output as inputs, solve the trained microgrid scheduling model and output the strategy for the current time step; this strategy can dynamically update energy storage charging and discharging, V2G trading and power purchase and sale decisions according to the external environment and renewable energy fluctuations, taking into account both energy supply stability and economy.

[0064] like Figure 3 As shown, the training steps of the deep reinforcement learning model used in this embodiment are an iterative process, addressing the dynamics and stochasticity of microgrid operation. Through the coordinated interaction of the improved A3C algorithm and the electricity price, carbon price, and V2G price signals generated by the two-layer Stackelberg game, the Actor network parameters and scheduling strategy are continuously optimized, achieving a synergistic optimal balance between comprehensive economic benefits and low-carbon constraints under conditions of renewable energy fluctuations and load changes. The training method for the microgrid optimal scheduling model includes the following steps: Step 1: Modeling and Initialization: Establish microgrid environment and equipment models, and configure the A3C trainer.

[0065] First, a microgrid environment and equipment model is established to clarify the state and action space and operational constraints. Simultaneously, an improved A3C trainer and multiple working threads are created, and stabilization parameters such as learning rate, entropy target, and gradient clipping are set.

[0066] Step 2: Price and Constraint Injection: The upper-level game generates electricity price / carbon price / V2G price, which is then injected into the environment.

[0067] Subsequently, the upper-level Stackelberg game generates current electricity price, carbon price, and V2G price signals, which are injected into the environment as external constraints and take effect immediately, forming a training scenario that closely resembles the market.

[0068] Step 3: Interactive sampling: Input state, output action, calculate reward, return to the next state, and store in the experience pool.

[0069] Next, at each time step, the agent reads the current state and outputs a scheduling action; the environment sequentially executes demand response, photovoltaic consumption, energy storage charging and discharging, and V2G trading, settles the purchase and sale of electricity and carbon costs, and returns to the next state and reward. At the same time, the experience is stored in the experience pool.

[0070] Step 4: Update network parameters.

[0071] Based on priority-based experience replay sampling batches, the target value and TD error are calculated and the Actor and Critic are jointly updated according to importance sampling weights. During the policy update process, dynamic entropy is used to adaptively adjust the exploration intensity, the learning rate decays smoothly with training steps, and gradient pruning suppresses oscillations. The local network pushes the gradient to the global network and periodically pulls the latest parameters to maintain the consistency of parallel training.

[0072] Step 5: Price feedback.

[0073] Through leader-follower game theory, feedback from lower-level scheduling is used to update electricity and V2G prices and determine convergence. The updated market signals are injected and take effect in the next training cycle, forming a closed loop of collaborative optimization between price and scheduling.

[0074] Step 6: Multiplication scheduling scheme.

[0075] Finally, a convergent microgrid scheduling scheme is obtained.

[0076] To verify the effectiveness of this method, refer to Table 1. Through comparative experiments, the performance differences in revenue and V2G between the game-theoretic model combined with the A3C (Game-A3C) strategy in this embodiment and the traditional A3C combined with the TOU strategy were verified. The experiments aimed to quantify the specific performance of the two schemes in key indicators such as microgrid revenue, V2G user revenue, and carbon emission reduction benefits.

[0077] Microgrid benefits: The experimental group's microgrid benefits reached 2.66 × 10⁻⁶. 7 Compared with the control group (2.17×10), 7 The growth rate was 22.63%, which verifies the effectiveness of game-theoretic pricing in ensuring profitability.

[0078] On the user side, EV carbon benefits also steadily increased by 5.09%. Overall, although users are still in a net expenditure state, mainly due to charging costs, the total loss of users in the experimental group decreased by about 4.15%, indicating that the mechanism effectively takes into account and improves user benefits while improving the overall efficiency of the system.

[0079] Table 1 Comparison of Game Theory Model + A3C and A3C + TOU Strategy Furthermore, referring to Table 2, a comparative experiment was conducted to evaluate the comprehensive impact of introducing an EV Carbon Trading mechanism in a microgrid system on the benefits for various stakeholders, including microgrid operators, the main grid, and V2G users. The experiment compared the benefits under two scenarios: enabling EV Carbon Trading (experimental group) and not enabling EV Carbon Trading (control group), aiming to verify the effectiveness of the carbon trading mechanism in promoting overall system efficiency and optimizing user-side costs. The total microgrid revenue in the experimental group increased by approximately 16.65% compared to the control group. The net expenditure situation on the V2G user side was significantly improved, with losses reduced by approximately 17.95%. The improvement in user-side revenue (approximately 602,200 RMB) highly matched the total EV Carbon revenue generated by the system (approximately 602,200 RMB). This indicates that, under the current mechanism design, EV Carbon revenue is almost completely and effectively transmitted to the user side, directly compensating for users' charging costs and opportunity costs of participating in V2G interaction. The carbon trading mechanism has become an important way to alleviate users' vehicle usage costs and incentivize users to participate in vehicle-grid interaction.

[0080] Table 2 Simulation Comparison of Carbon Trading Mechanisms In summary, this method improves upon the traditional A3C method to adapt to the highly stochastic and dynamically changing operating environment of microgrids. Based on the parallel exploration of multiple worker threads in A3C, mechanisms such as dynamic entropy adjustment of exploration intensity, adaptive learning rate decay, and gradient pruning are added to improve policy stability and global search capability. Simultaneously, the electricity price and carbon price signals generated by a two-layer Stackelberg game ensure that training maintains convergence and interpretability even in fluctuating external environments. Therefore, when renewable energy output continuously changes, the model can update the energy allocation and scheduling of the microgrid in real time, balancing energy supply stability and economy. By combining the advantages of the two-layer Stackelberg game and the A3C algorithm, the asynchronous push / pull mechanism of A3C's multiple workers and shared global step size significantly accelerates convergence and reduces policy variance, improving sample efficiency through asynchronous parallel reinforcement learning. A dynamic entropy coefficient is introduced, increasing the entropy weight when exploration is insufficient and automatically converging after policy stability, suppressing early convergence through adaptive exploration. Samples are prioritized based on their temporal difference error, and experiences with higher priority are sampled first, through the introduction of a priority experience replay mechanism. The upper-level price is solved using a leader-follower structure to ensure price boundaries and convergence; the lower-level A3C is responsible for device and behavior optimization to avoid instability caused by strong coupling.

[0081] Example 2 Building upon Example 1, this example provides a microgrid optimal scheduling system combining two-layer game theory and A3C (Automatic Three-Channel) to implement the aforementioned microgrid optimal scheduling method. The system includes an upper-layer leader pricing module, a middle-layer microgrid follower pricing module, a lower-layer A3C scheduling module, and a microgrid environment module. The method uses time-of-use pricing as an anchor point, drives price updates through state-space feedback, and A3C learns its operating strategy in an extended state space, forming a closed loop of "pricing-scheduling-feedback." (1) Upper-level leader pricing module. It is used to jointly optimize electricity price and carbon price based on state-space feedback and market supply and demand information of microgrid operation, and output price signal.

[0082] The upper-level leader pricing module is used to jointly optimize electricity and carbon prices based on state-space feedback and market supply and demand information from microgrid operations. Electricity prices consist of a base time-of-use price, profit margin, demand pressure, supply-demand balance, and dynamic factors, ensuring price stability and interpretability. Carbon prices are adjusted based on net carbon trading demand and the proportion of renewable energy, with reasonable boundaries set. The module outputs price signals including electricity price, carbon price, V2G electricity price, and V2G carbon price, which drive subsequent decision-making and training in the middle and lower layers.

[0083] (2) Mid-level microgrid follower pricing module, which is used to obtain V2G electricity price and V2G carbon price based on price signal, combined with energy storage status and V2G supply and demand ratio, and to obtain the current period's electricity purchase and V2G usage with the grid.

[0084] The mid-level microgrid follower pricing module is used to set V2G electricity and carbon prices based on the availability of electricity and carbon prices, combined with energy storage status and V2G supply-demand ratio, and to determine the current period's electricity purchase and V2G usage from the grid. This module achieves a balance between economic benefits and emission reduction by coordinating the internal power generation, energy storage charging and discharging, and vehicle-to-grid interaction.

[0085] (3) The lower-level A3C scheduling module is used for policy learning based on the target entropy-driven adaptive entropy coefficient, learning rate exponential decay and gradient clipping.

[0086] The lower-level A3C scheduling module is used for policy learning in a thirteen-dimensional state space that includes load demand, photovoltaic output, energy storage SOC and charging / discharging power, unit power, EV charging / discharging capacity, dynamic electricity price, time-of-use base price, and charging pile activity ratio. This module introduces an adaptive entropy coefficient driven by target entropy, exponential decay of the learning rate, and gradient pruning to improve the convergence stability and exploration efficiency of the training process, thereby obtaining an executable operating strategy under multiple constraints.

[0087] (4) Microgrid environment module, used to inject price signals into the environment and collect new state feedback in real time after each scheduling step and A3C strategy update, driving the next round of pricing and decision-making in the upper and middle layers to form an iterative closed loop.

[0088] The microgrid environment module is used to inject the leader's pricing results into the environment and collect new status feedback in real time after each scheduling step and A3C policy update, including power balance, equipment power boundary, carbon emission estimation and V2G supply and demand indicators, to drive the next round of pricing and decision-making in the upper and middle layers, forming an iterative closed loop.

[0089] In summary, this system, through a multi-layered collaborative mechanism of upper-level game theory pricing, mid-level microgrid pricing, lower-level A3C scheduling, and environmental feedback, achieves the coordinated convergence of electricity prices, carbon prices, and V2G prices, as well as the stable optimization of operating strategies, while ensuring the interpretability of key physical quantities and the controllability of parameter ranges. This significantly improves the economy, low-carbon nature, and engineering feasibility of microgrids in real-world scenarios.

[0090] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A microgrid optimal scheduling method combining two-level game theory and A3C, characterized in that, Includes the following steps: Construct an integrated microgrid model encompassing power generation, grid, load, and energy storage; With the goal of maximizing the total revenue during microgrid operation, and taking into account the microgrid's operational constraints, the constraints on renewable energy power output, power constraints, and the power supply and demand balance constraints of the microgrid, a dynamic scheduling model for microgrids based on A3C is established. The state space feedback is extracted from the environment. Through the upper-level Stackelberg game, the large power grid, as the leader, makes pricing decisions, while the microgrid, as the follower, makes decisions to obtain the current electricity purchase and V2G usage, and updates the price signal. Based on the price signal, the trained microgrid scheduling model is solved to obtain the strategy output by the lower-level reinforcement learning at the current time step.

2. The microgrid optimal scheduling method combining two-level game theory and A3C as described in claim 1, characterized in that, In the A3C algorithm of the microgrid optimization scheduling model, the state space of the agent in the environment includes load demand, photovoltaic power generation, energy storage system SOC value, energy storage system charging power, energy storage system discharging power, energy storage system remaining capacity, gas micro-turbine power, fuel cell power, V2G dynamic price, time-of-use price, EV charging amount, EV discharging amount, and charging pile activity ratio. The action space of the agent includes energy storage charging and discharging actions, V2G discharging actions, V2G price adjustment actions, and demand response actions.

3. The microgrid optimal scheduling method combining two-level game theory and A3C as described in claim 1, characterized in that, During the training process of the microgrid optimal scheduling model, the Actor network and / or Critic network are trained based on importance sampling weights. The process of obtaining the importance sampling weights includes the following steps: Based on time difference error and preset control priority Calculate experience priority ; Based on experience priority and the minimum sampling probability in the preset experience pool Calculate importance sampling weights .

4. The microgrid optimal scheduling method combining two-level game theory and A3C as described in claim 3, characterized in that, During the training process of the microgrid optimal scheduling model, the update process of the Actor network includes the following steps: Target value is calculated using reverse rolling based on intraday time steps. The time difference error is obtained. ; With time difference error As an advantage, by blocking gradient flow to prevent inverse dependence on the Critic network, the product of the policy log probability and the advantage is calculated as the core expectation term. ; Using policy entropy as the regularization term, based on the recent moving average of policy entropy and target policy entropy Calculate adaptive weights Among them, the moving average of policy entropy With adaptive weights Inverse correlation; Based on the adaptive weights and the importance sampling weight The policy entropy is weighted and combined with the core expectation term. We obtain the loss function value of the Actor network and then perform training through policy gradient ascent optimization.

5. The microgrid optimal scheduling method combining two-level game theory and A3C as described in claim 3, characterized in that, During the training process of the microgrid optimal scheduling model, the update process of the Critic network includes the following steps: Calculate the difference between the target value and the actual value at each step, and use the importance sampling weights. The weighted loss function value of the Critic network is obtained, and the network is updated based on the loss function value to achieve training.

6. The microgrid optimal scheduling method combining two-level game theory and A3C as described in claim 1, characterized in that, The microgrid model constructs a deep reinforcement learning-based simulation environment within the microgrid's operating environment, based on an objective function and constraints, to reflect key factors and dynamic changes in microgrid operation. The objective function is: In the formula, Let be the objective function. for Carbon benefits from EV charging The electricity generation costs are respectively for gas turbines and fuel cells. The carbon costs of gas turbines and fuel cells, respectively. For the cost of V2G, The cost of purchasing electricity and the carbon cost of microgrids The cost of responding to demand.

7. The microgrid optimal scheduling method combining two-level game theory and A3C as described in claim 1, characterized in that, In the microgrid optimization scheduling model, the power supply components include solar photovoltaic, energy storage systems and generator sets, while the load components include electric vehicle loads and residential loads.

8. A microgrid optimal scheduling method combining two-level game theory and A3C as described in claim 7, characterized in that, The modeling process for each component in the microgrid optimal scheduling model includes: The power generation cost of gas turbines is modeled based on the active power output of gas turbines, natural gas prices, gas turbine efficiency, and low calorific value of natural gas, and upper and lower limits of output power and ramp rate constraints are constructed. Based on the power generation cost, active power output, and efficiency of fuel cells, the power generation cost of fuel cells is modeled, and upper and lower limits of output power and ramp rate constraints are constructed. Based on the state of charge ratio, previous state of charge, state of charge, charge and discharge, charge and discharge efficiency, charge and discharge power, and rated capacity of the energy storage system, the state of charge of the energy storage system is modeled, and charge and discharge power constraints and state of charge boundary constraints are constructed. Construct constraints on the balance between power supply and demand in microgrids.

9. A microgrid optimal scheduling method combining two-level game theory and A3C as described in claim 1, characterized in that, The strategy includes energy storage charging and discharging, V2G trading, and power purchase and sale decisions.

10. A microgrid optimal scheduling system combining two-level game theory and A3C, characterized in that, For implementing the microgrid optimal scheduling method as described in any one of claims 1-9, the system includes: The upper-level leader pricing module is used to jointly optimize electricity and carbon prices based on the state-space feedback of microgrid operation and market supply and demand information, and output price signals. The mid-level microgrid follower pricing module is used to obtain the V2G electricity price and V2G carbon price based on the price signal, combined with the energy storage status and V2G supply and demand ratio, and to obtain the current electricity purchase and V2G usage from the grid. The lower-level A3C scheduling module is used for policy learning based on target entropy-driven adaptive entropy coefficients, exponential decay of learning rate, and gradient clipping. The microgrid environment module is used to inject the price signal into the environment and collect new state feedback in real time after each scheduling step and A3C strategy update, driving the next round of pricing and decision-making in the upper and middle layers to form an iterative closed loop.