Virtual power plant scheduling method for evolutionary deep reinforcement learning

Through the evolution of the virtual power plant scheduling method with deep reinforcement learning, combined with hydrogen energy green certificate trading and ladder carbon trading, the problem of power supply and demand balance in virtual power plants and the problem of unstable hydrogen energy supply is solved, and efficient and low-carbon multi-energy system optimization scheduling is achieved.

CN120494429AInactive Publication Date: 2025-08-15STATE GRID JIANGXI ELECTRIC POWER CO LTD RES INST

Patent Information

Application Number
CN202510928762.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-08-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The intermittent and volatility of distributed energy in virtual power plants lead to increased difficulty in balance between power supply and demand, insufficient traditional regulation capabilities, and existing optimization algorithms are inefficient in high-dimensional scenarios, which are prone to falling into local optimal solutions, unstable hydrogen energy supply, and failure to fully tap the potential of hydrogen energy application.

Method used

The virtual power plant scheduling method with evolutionary deep reinforcement learning is adopted, combined with hydrogen energy green certificate trading and ladder carbon trading, and optimize scheduling through the evolutionary soft action-judge (ESAC) algorithm to build a multi-purpose utilization model for hydrogen-containing virtual power plants, expand the search range, avoid local optimal solutions, and improve the algorithm convergence speed and stability.

Benefits of technology

It improves the optimization accuracy and robustness of virtual power plant scheduling, reduces operating costs, promotes the low-carbon operation of multi-energy systems, and improves the application efficiency and supply stability of hydrogen energy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494429A_ABST
    Figure CN120494429A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual power plant scheduling method for evolutionary deep reinforcement learning, and belongs to the technical field of power control. According to the method, a hydrogen-containing virtual power plant low-carbon economic dispatching model is constructed, hydrogen energy gradient utilization is realized through electric hydrogen production, electric-to-gas and hydrogen-to-ammonia technologies, a stepped carbon trading model is improved, a hydrogen energy green certificate trading model is introduced, and an evolutionary soft action device-evaluator (ESAC) algorithm is adopted to carry out optimized dispatching. According to the algorithm, population optimization and deep reinforcement learning of an evolutionary strategy are fused, and the mutation rate is dynamically adjusted through an automatic mutation adjustment mechanism. According to the method, the optimization precision and robustness of virtual power plant scheduling are remarkably improved, the operation cost is effectively reduced, and low-carbon operation of a multi-energy system is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of power control technology, and specifically relates to a virtual power plant scheduling method based on evolutionary deep reinforcement learning. Background Art

[0002] The intermittent and volatile nature of distributed energy resources in virtual power plants significantly increases the difficulty of balancing power supply and demand, and traditional power grid regulation capabilities are unable to meet these demands. Hydrogen energy, as an efficient means of energy coupling, storage, and low-carbon development, can be combined with virtual power plants to form hydrogen-based virtual power plants. Hydrogen-based virtual power plants can effectively address the challenges posed by the increasing proportion of renewable energy and provide new ideas for enhancing the grid's ability to accommodate and absorb fluctuating renewable energy sources under virtual power plants.

[0003] Existing research has effectively improved carbon emissions from hydrogen-based virtual power plants by introducing carbon capture systems and adopting a combined wind power-carbon capture-power-to-gas (P2G) operation model. Considering the combined operation of two-stage P2G (power-to-gas) and gas-hydrogen blending units reduces energy cascade losses and VPP carbon emissions. Furthermore, a wide-power adaptation model for electrolyzers has been proposed in wind-solar-hydrogen-thermal virtual power plants, enhancing the electrolyzer's ability to cope with fluctuations in wind and solar output while absorbing more excess wind and solar power. Most of the aforementioned existing research focuses on technologies such as renewable energy hydrogen production and storage. However, there is a supply-demand imbalance between renewable energy hydrogen production and the actual hydrogen energy demand of hydrogen-based virtual power plants, making it difficult for hydrogen-based virtual power plants to ensure a stable hydrogen supply. Furthermore, the existing virtual battery control model uses a relatively simple approach to hydrogen production and utilization, failing to fully tap the potential of hydrogen energy applications in virtual power plants.

[0004] Currently, most virtual power plant optimization scheduling models are solved using CPLEX, YALMIP solvers, or intelligent optimization algorithms. However, due to the uncertainty of renewable energy and load, virtual power plant scheduling often involves a large number of renewable energy resources and energy storage systems, resulting in high uncertainty and dynamics. This is particularly true in high-dimensional scenarios, where the performance of solvers and intelligent optimization algorithms is often significantly affected, and they are prone to becoming trapped in local optimal solutions, resulting in inefficient search. In recent years, deep reinforcement learning has been gradually applied to virtual power plant optimization scheduling due to its ability to process high-dimensional data and flexible policy adjustments in complex decision-making scenarios. Although these algorithms have the ability to handle continuous action spaces, in practice, the agent's policy exploration process is overly conservative when exploring the environment and its exploration efficiency in high-dimensional spaces is poor, which significantly affects the agent's training efficiency. Summary of the Invention

[0005] In order to overcome the shortcomings of the existing technology, the present invention provides a virtual power plant scheduling method based on evolutionary deep reinforcement learning, which takes into account carbon trading and hydrogen green certificate trading, establishes a diversified utilization model of hydrogen-containing virtual power plants, and further explores the carbon reduction potential while reducing the operating costs of virtual power plants; considering the advantages of evolutionary strategies in global search and jumping out of local optimal solutions, the diversity of strategy exploration is improved, and an evolutionary soft actor-critic (ESAC) algorithm is adopted for scheduling. The evolutionary soft actor-critic algorithm uses evolutionary strategies to explore the weight space of the actor network, expands the search range, and can largely avoid local optimal solutions, thereby improving the convergence speed and stability of the algorithm.

[0006] The present invention is implemented through the following technical solutions: A virtual power plant scheduling method based on evolutionary deep reinforcement learning, comprising the following steps: S1. Initialize the environment and parameters: Based on the hydrogen virtual power plant scheduling model and related constraints, build a deep reinforcement learning environment, input historical wind power and load data, and initialize the parameters of the evolutionary soft actor-evaluator algorithm; S2. Offline Training: The agent obtains state variables for the current period from the environment. These state variables include electrical load, thermal load, wind power output, real-time natural gas prices, and actions from the previous period. The agent then outputs actions based on these state variables, interacts with the environment, provides feedback on rewards, and updates the state for the next period. The agent is a deep reinforcement learning model optimized using an evolutionary soft actor-critic algorithm. S3. Maintain Model: Maintain convergence of the deep reinforcement learning model optimized by the evolving soft actor-critic algorithm after reaching a pre-defined maximum number of iterations; S4. Online Dispatch: Obtain real-time data on the electric load, thermal load, wind power output, and natural gas prices during the forecast period, and combine this with the actual output data of each unit to form state variables. S5. Generate a scheduling plan: Input the state variables into the deep reinforcement learning model optimized by the trained evolutionary soft actor-critic algorithm to generate the optimal scheduling plan for time period t, ensuring that the constraints are met and the operating cost is minimized.

[0007] It is further preferred to construct a hydrogen-containing virtual power plant scheduling model based on the multi-energy coupling operation mode and carbon emission model. The multi-energy coupling operation mode takes hydrogen energy as the core, integrates electricity-to-hydrogen, electricity-to-gas, gas-to-electricity, and hydrogen-to-ammonia technologies to achieve the cascade utilization of hydrogen energy: hydrogen is first directly supplied to the mixed hydrogen gas pipeline network, and then used for synthetic natural gas, and the remaining part is used for ammonia production or hydrogen storage; when the hydrogen-containing virtual power plant has insufficient hydrogen energy supply, hydrogen is purchased through hydrogen energy green certificate transactions; when the hydrogen-containing virtual power plant has excess hydrogen energy, hydrogen is sold through hydrogen energy green certificate transactions; the hydrogen-containing virtual power plant scheduling model includes: a hydrogen-containing virtual power plant operation model composed of various equipment models and various gas production and consumption models on the source side, a ladder carbon trading model, a hydrogen energy green certificate trading model, an objective function and constraints.

[0008] Further preferably, the objective function of the hydrogen-containing virtual power plant scheduling model is to minimize the total operating cost of the hydrogen-containing virtual power plant.

[0009] Further optimization, the ladder carbon trading model introduces a ladder carbon trading mechanism, including: Allocation of free carbon quotas: The baseline method is used to allocate carbon emission quotas for thermal power units, hydrogen-mixed gas turbines, and hydrogen-mixed gas boilers; Tiered penalty pricing: When actual carbon emissions exceed the quota threshold, the excess amount will be subject to tiered incremental pricing; Emission reduction compensation coefficient: When emissions are lower than the quota benchmark, the market-based trading benefits are enhanced through the compensation coefficient.

[0010] Further optimization is performed to construct the action space of the deep reinforcement learning model based on the hydrogen virtual power plant scheduling model and state space .

[0011] Further preferably, the evolutionary soft actor-evaluator algorithm includes: Criterion: Uses a Q-function neural network to estimate the soft Q-function. The training goal is to minimize the temporal difference error, and the network parameters are updated using exponential moving average. Actor: optimizes the policy to maximize the expected soft value; Evolutionary strategies: generate populations, evaluate fitness, perform soft winner selection and crossover, and update baseline strategy parameters; Automatic mutation adjustment mechanism: used to dynamically adjust the mutation rate during the random perturbation process of the evolution strategy.

[0012] The present invention also provides an electronic device, including a memory and a processor, wherein the memory stores computer-readable instructions, and when the instructions are executed by the processor, the processor implements the above-mentioned virtual power plant scheduling method.

[0013] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the above-mentioned virtual power plant scheduling method is implemented.

[0014] Based on a multi-energy coupling operation mode with hydrogen energy as the core, the present invention constructs a low-carbon economic dispatch model for a hydrogen-containing virtual power plant. At the operational level, the model adopts a variety of technologies such as electricity to hydrogen, electricity to gas, gas to electricity, and hydrogen to ammonia to broaden the application scenarios of hydrogen energy and improve the comprehensive utilization efficiency of energy; at the low-carbon mechanism level, it introduces ladder carbon trading and hydrogen energy green certificate trading to strengthen the carbon reduction effect on the source side. In view of the characteristics of the low-carbon economic dispatch model of the hydrogen-containing virtual power plant with high-dimensional variables and complex constraints, the present invention designs a deep reinforcement learning environment, action space, state space and reward function, and uses the evolutionary soft actor-critic (ESAC) algorithm for training. By generating the optimal dispatch strategy offline, the low-carbon economic operation of the hydrogen-containing virtual power plant is finally achieved. While improving the dispatch robustness and optimization accuracy, the present invention provides a new solution for the low-carbonization of hydrogen-containing virtual power plants. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 Flow chart of the method of the present invention.

[0016] Figure 2 This is the ESAC algorithm training flowchart. DETAILED DESCRIPTION

[0017] The present invention is further described below with reference to the embodiments. It is necessary to point out that the following embodiments are only used to further illustrate the present invention and are not to be construed as limiting the scope of protection of the present invention. Non-essential improvements and adjustments made by persons skilled in the art based on the above-mentioned invention contents still fall within the scope of protection of the present invention.

[0018] This embodiment is explained from three aspects: multi-energy coupling operation mode and carbon emission model, hydrogen-containing virtual power plant scheduling model, and deep reinforcement learning model optimized by evolutionary soft actor-evaluator algorithm. Finally, the steps of a virtual power plant scheduling method based on evolutionary deep reinforcement learning are explained.

[0019] 1. Multi-energy coupling operation mode and carbon emission model The energy production side of a hydrogen virtual power plant (HVPP) consists of a modular electric hydrogen production unit, integrating five heterogeneous energy sources: electricity, heat, hydrogen, natural gas, and ammonia. Core components include thermal power units, a hydrogen-based gas turbine (HGT), a hydrogen-based gas boiler (HGB), electric hydrogen production equipment, a power-to-ammonia conversion unit, a methane reactor, an electric boiler (EB), and energy storage devices, enabling efficient conversion and coordinated utilization of multiple energy flows. The electric hydrogen production unit leverages renewable energy generation to achieve large-scale hydrogen production, while the energy storage system enhances the flexibility of the hydrogen virtual power plant by storing hydrogen, heat, and electricity.

[0020] 1.1 Multi-energy coupling operation mode The present invention proposes a multi-energy coupling operation mode based on hydrogen energy in the HVPP framework, the core links of which include a water electrolysis hydrogen production unit and a hydrogen energy multi-conversion module.

[0021] Traditional P2G technology has a low overall energy conversion efficiency. To improve hydrogen utilization efficiency and promote its multi-path utilization, P2G is split into two steps: electricity-to-hydrogen production and methanation. The electricity-to-hydrogen production process is powered by renewable energy, and the produced hydrogen is utilized in a cascade through various pathways: first, it is directly supplied to the hydrogen-mixed gas pipeline network for use in gas equipment; second, the hydrogen is transported to the methanation unit to synthesize natural gas, achieving carbon recycling; a small amount is used to provide feedstock for electricity-to-ammonia units to produce ammonia-blended fuel for injection into coal-fired units; and the remainder is stored in hydrogen storage tanks to meet the flexible hydrogen demand of hydrogen-containing virtual power plants. When the hydrogen supply of a hydrogen-containing virtual power plant is insufficient, hydrogen is purchased through hydrogen green certificate trading to maintain the hydrogen supply of the hydrogen-containing virtual power plant. If the hydrogen demand of the hydrogen-containing virtual power plant is lower than the hydrogen production, the surplus hydrogen is sold through hydrogen green certificate trading, significantly improving the economic efficiency of energy utilization and reducing the operating costs of the hydrogen-containing virtual power plant.

[0022] 1.2 Carbon Emission Model Hydrogen virtual power plants The equipment includes thermal power units, hydrogen-mixed gas turbines and hydrogen-mixed gas boilers. In order to promote low-carbon transformation, carbon capture technology has been introduced on the production capacity side. Specifically, the carbon generated by hydrogen-mixed gas turbines, hydrogen-mixed gas boilers and thermal power units in the hydrogen-containing virtual power plant is The flue gas is collected through the flue gas bypass, part of which is introduced into the carbon capture unit for treatment through a multi-stage diversion device, and the rest is discharged into the atmosphere. , using amine solution absorption-steam regeneration process, part After chemical absorption in the absorption tower, thermal desorption in the regeneration tower, and supercritical compression, the methane is used as fuel replacement in the methane tank, and the remaining part is isolated for a long time through geological storage or mineral storage. The carbon emission model of each part of the hydrogen virtual power plant is shown below: (1); Where: is the actual carbon emissions of hydrogen-containing virtual power plants, 、 and are the carbon emissions of thermal power units, hydrogen-mixed gas turbines, and hydrogen-mixed gas boilers during period t; is the carbon storage amount in period t, is the methane tank at time t Utilization, T is the total number of time periods.

[0023] 2. Hydrogen Virtual Power Plant Scheduling Model 2.1 Hydrogen Virtual Power Plant Operation Model In this invention, the electric hydrogen production device utilizes the surplus wind energy of the hydrogen-containing virtual power plant to produce hydrogen energy to meet the hydrogen energy demand of the hydrogen-containing virtual power plant. Simultaneously, the hydrogen-mixed gas turbine and hydrogen-mixed gas boiler utilize the hydrogen-mixed gas to output electricity and heat. The electric boiler consumes electricity to output heat. The electric-to-ammonia device releases energy during the ammonia production process through the Hubble ammonia synthesis reaction, and the generated ammonia is supplied to the thermal power unit. The modeling of each device is shown below: (2); Where: and are the power supply power and heating power of the hydrogen-mixed gas turbine in period t respectively; and are the natural gas and hydrogen power consumed by the hydrogen-blended gas turbine during period t, respectively; and The electrical and thermal efficiency of the hydrogen-mixed gas turbine. is the total output of the hydrogen-gas-fired boiler in period t; and is the natural gas and hydrogen power consumed by the hydrogen-mixed gas boiler during period t; is the thermal efficiency of the hydrogen-gas boiler; is the thermal power of the hydrogen-gas boiler during period t; is the net electrical power output of the hydrogen-mixed gas turbine during period t; is the energy consumption of carbon capture operation during period t; is the fixed energy consumption for carbon capture during period t; Carbon capture unit Energy consumption of carbon capture; The carbon capture process during time t the amount; is the power consumption of the hydrogen production unit during period t; and are the heat generation power and power consumption of the electric boiler during period t respectively; is the electric-to-heat conversion efficiency of the electric boiler; is the thermal power provided by the power-to-ammonia unit to the hydrogen virtual power plant during period t; The heat release ratio of the power-to-ammonia unit used for heating; The thermal power released to generate unit NH3; is the quality of ammonia produced by the power-to-ammonia unit during period t.

[0024] The carbon capture device will capture part of It is transported to the methane reaction tank as high-quality raw material for synthesis of natural gas. The ammonia-blended combustion technology is used in traditional thermal power units. By mixing ammonia with pulverized coal for combustion, the zero-carbon characteristics of ammonia fuel can be used to reduce carbon emissions per unit of power generation. The combustion products of ammonia are mainly nitrogen and water, and no CO2 is produced. The carbon emission reduction rate is equal to the calorific value replacement rate of ammonia fuel. To ensure the safe operation of hydrogen-mixed gas turbines and hydrogen-mixed gas boilers, the hydrogen mixing ratio of hydrogen-mixed gas turbines is generally 10%-20%, and the burner can burn safely and stably. When hydrogen-mixed gas boilers use hydrogen-mixed gas, the molar mass ratio of hydrogen needs to be maintained at 2%~20%. The production and consumption models of each gas on the source side are shown below: (3); (4); Where: is the hydrogen power consumed by the methane tank during period t; The amount of CO2 required to generate natural gas per unit of power; The efficiency of natural gas production for methane tanks; is the amount of ammonia produced by the power-to-ammonia unit during period t; is the power consumption of the power-to-ammonia unit during period t; The unit energy consumption of ammonia production by the power-to-ammonia unit; for carbon capture efficiency; is the flue gas split ratio during period t; Carbon emissions from ammonia-doped thermal power units; is the carbon emission per unit of coal combustion; is the coal consumption during period t; and They are the carbon emissions of hydrogen-mixed gas turbines and hydrogen-mixed gas boilers; and are the carbon emission coefficients of the hydrogen-mixed gas turbine and hydrogen-mixed gas boiler, respectively; is the consumption characteristic coefficient of the thermal power unit; is the power generation capacity of the thermal power unit during period t; and are the lower heating values of ammonia and coal, respectively; is the hydrogen power produced by hydrogen production during period t; The hydrogen production efficiency of electric hydrogen production; is the hydrogen power produced by electro-ammonia conversion during period t; The hydrogen production efficiency of electric hydrogen production; is the calorific value of hydrogen; is the calorific value of natural gas; and They are The hydrogen mixing ratio of hydrogen-mixed gas turbine and hydrogen-mixed gas boiler during the period; is the volume of natural gas produced by the methane tank during period t; The efficiency of natural gas production in methane tanks.

[0025] The models of energy storage devices such as batteries, heat storage tanks, and hydrogen storage tanks are similar, so unified modeling of energy storage devices can be achieved using modeling methods disclosed in the prior art.

[0026] 2.2 Tiered carbon trading model and hydrogen green certificate trading model 2.2.1 Tiered carbon trading model Carbon quota allocation methods are mainly divided into two types: paid allocation and free allocation. The carbon quota in this paper adopts the free allocation adopted by most literature, and uses the baseline method as the allocation method of carbon emission quotas. The carbon quota baseline value can be determined by referring to the existing technology. The objects of carbon quota allocation include thermal power units, hydrogen-mixed gas turbines and hydrogen-mixed gas boilers. The ladder carbon trading model is as follows: (5); Where: 、 、 and They are the carbon emission quotas for hydrogen-containing virtual power plants, thermal power units, hydrogen-mixed gas turbines and hydrogen-mixed gas boilers; 、 and These are the carbon quota benchmark values for thermal power units, hydrogen-mixed gas turbines and hydrogen-mixed gas boilers respectively.

[0027] To more effectively control carbon emissions, this paper introduces a tiered carbon trading mechanism based on a unified carbon trading framework. This mechanism determines transaction costs by dividing carbon emissions into intervals. When a hydrogen-based virtual power plant's actual carbon emissions exceed a preset quota threshold, the excess is subject to escalating punitive pricing. If emissions fall below the quota benchmark, the hydrogen-based virtual power plant can trade its excess carbon emission rights in a market-based transaction, thereby generating economic benefits. Furthermore, to further incentivize emission reductions, a compensation factor is introduced to increase incentives for emission reductions, thereby strengthening the disciplining effect of the carbon market.

[0028] 2.2.2 Hydrogen Energy Green Certificate Trading The current production cost of green hydrogen is relatively high, but its environmental value is significant. To this end, this invention introduces the hydrogen energy green certificate trading mechanism in the electricity market into the hydrogen energy market and formulates a trading price based on the supply and demand relationship of hydrogen energy green certificates to promote the economic development and environmental benefits of green hydrogen. The hydrogen energy green certificate trading model is as follows: (6); Where: is the clearing price of hydrogen green certificates; The maximum acceptable price for hydrogen green certificates; The demand ratio for hydrogen green certificates.

[0029] 2.3 Objective Function and Constraints 2.3.1 Objective Function For hydrogen-containing virtual power plants that consider carbon trading and hydrogen energy green certificate trading, in order to achieve economic benefits while taking into account environmental benefits, the objective function of the hydrogen-containing virtual power plant scheduling model is to minimize the total operating cost of the hydrogen-containing virtual power plant, as shown in the following formula: (7); Where: is the total operating cost of the hydrogen virtual power plant; The unit operating cost; is the tiered carbon trading cost; for carbon sequestration costs; for the cost of carbon capture equipment; Penalty costs for wind curtailment; The cost of purchasing gas; The start-up and shutdown and coal consumption costs of thermal power units; The specific expression of each part is as follows: 1) Unit operating costs: (8); Where: is the operation and maintenance cost of the u-th equipment in period t; is the output of the uth type of equipment other than thermal power units during period t, and N is the number of equipment types.

[0030] 2) Tiered carbon trading costs: (9); Where: for Carbon trading costs of hydrogen virtual power plants during this period.

[0031] 3) Carbon sequestration costs: (10); Where: is the unit cost of carbon sequestration.

[0032] 4) Carbon capture equipment cost: (11); Where: is the depreciation cost of carbon capture equipment during period t; and are the total cost and depreciation period of carbon capture equipment respectively; The discount rate for carbon capture power plant projects.

[0033] 5) Wind curtailment penalty costs: (12); Where: is the wind curtailment penalty coefficient; is the amount of abandoned wind after absorption in period t.

[0034] 6) Gas purchase cost: (13); Where: is the gas purchase volume during period t, is the purchase price of natural gas during period t.

[0035] 7) Thermal power unit start-up and shutdown and coal consumption costs: (14); Where: is the start-up and shutdown cost of thermal power units; is the coal consumption cost of thermal power units; is the start-up and shutdown cost coefficient of thermal power units; Indicates the state of the thermal power unit during period t, which is a 0 or 1 variable; Indicates the status of the thermal power unit during the t-1 period.

[0036] 8) Hydrogen Energy Green Certificate Trading Profits: (15); Where: It is the hydrogen power sold through hydrogen green certificate trading during period t.

[0037] 2.3.2 Constraints The constraints of the hydrogen-containing virtual power plant scheduling model of the present invention include conventional unit constraints, safe operation constraints and balance constraints.

[0038] 1) Conventional unit constraints According to the operating requirements of hydrogen-containing virtual power plants, various energy conversion equipment must meet the following operating output constraints: (16); Where: and They are the upper and lower limits of thermal power unit output respectively; and They are the upper and lower limits of the electrical output of the hydrogen-mixed gas turbine respectively; and They are the upper and lower limits of thermal output of hydrogen-mixed gas turbine respectively; and They are the upper and lower limits of heating for the hydrogen-gas-fired boiler unit respectively; and They are the upper and lower limits of hydrogen consumption in the methane tank; and They are the upper and lower limits of power consumption of the electric hydrogen production unit; and They are the upper and lower limits of heating of electric boilers respectively; is the wind power forecast value during period t; is the wind power output during period t; It is the upper limit of power consumption of the power-to-ammonia unit during period t.

[0039] Parameter constraints of the unit: (17); Where: is the maximum flue gas split ratio; is the ammonia blending ratio of thermal power units, It is the upper limit of ammonia blending ratio for thermal power units; and These are the upper and lower limits of hydrogen blending ratio for hydrogen-mixed gas turbines and hydrogen-mixed gas boilers, respectively.

[0040] Unit climbing constraints: (18); Where: and are the climbing rate and sliding rate of thermal power units respectively; and are the ramp rate and ramp rate of the hydrogen-blended gas turbine, respectively; and They are the ramp rate and ramp rate of the hydrogen-gas-fired boiler respectively; and are the climbing rate and sliding rate of the methane tank, respectively; and are the climbing rate and sliding rate of the electric boiler respectively; and They are the climbing rate and sliding rate of the hydrogen production unit respectively.

[0041] 2) Safe operation constraints Since the hydrogen-containing virtual power plant is connected to the upper-level distribution network and the natural gas network, it is necessary to ensure the safety of natural gas transmission to meet the safety requirements of scheduling operation. The present invention adopts the safety constraint model in the existing technology.

[0042] 3) Balance Constraints (19); Where: is the battery discharge power during period t; is the battery charging power during period t; is the electric load demand during period t; is the heat release power of the heat storage tank during period t; is the heating power of the heat storage tank during period t; is the heat load demand during period t; is the volume of natural gas consumed by the hydrogen-blended gas turbine during period t; is the volume of natural gas consumed by the hydrogen-fired gas boiler during period t; is the hydrogen power released by the hydrogen storage tank during period t; is the hydrogen power stored in the hydrogen storage tank during period t, is the volume of natural gas produced by the methane tank during period t.

[0043] 3. Deep reinforcement learning model optimized by evolutionary soft actor-critic algorithm Deep reinforcement learning is a cutting-edge AI technology that combines the representational capabilities of deep learning with the decision-making framework of reinforcement learning. Deep learning automatically extracts complex data features through multi-layer neural networks, eliminating the need for manual feature engineering and achieving efficient representation learning. Reinforcement learning, on the other hand, optimizes strategies through trial-and-error and reward-based feedback through the interaction between an agent and its environment, maximizing long-term cumulative returns. This interaction can be described as a Markov decision process, enabling the agent to adjust its behavior based on environmental feedback. It is suitable for scenarios such as robotic control and game play. Deep reinforcement learning combines the strengths of both, utilizing deep networks to process high-dimensional states and optimizing strategies through dynamic programming to overcome the curse of dimensionality. This empowers agents with goal-oriented decision-making capabilities, enabling them to autonomously evolve optimal strategies in dynamic environments.

[0044] 3.1 Reward Function Design The model constructed by this invention takes reducing the operating cost of the virtual power plant as the optimization goal. The agent adjusts its action strategy based on the reward function returned by the environment to maximize the reward value. The reward function is as follows: (20); Where: is the weight coefficient of the operating cost of the hydrogen-containing virtual power plant; To adjust the parameters; is the state during period t; is the action during period t; is the reward function for period t; is the total operating cost of the hydrogen virtual power plant.

[0045] 3.2 Action Space Design The HVPP constructed by the present invention has a wide variety of energy equipment. In order to simplify the action space design and reduce the pressure of model training, the present invention adopts an optimization design method. The gas purchase volume of the gas station, the heating power of the hydrogen-mixed gas boiler and the power of the electric boiler are determined by the power balance constraint equation (19), while other equipment such as wind power, thermal power, methane reaction tank, hydrogen-mixed gas turbine, electric hydrogen production device, energy storage system, carbon capture equipment, gas hydrogen mixing ratio and thermal power ammonia mixing ratio are all action variables. for: (twenty one); in, is the wind power output during period t, is the power generation capacity of the thermal power unit during period t, is the power supply of the hydrogen-mixed gas turbine during period t, is the hydrogen power consumed by the methane tank during period t, is the power consumption of the hydrogen production unit during period t, is the action variable of the energy storage system during period t, is the flue gas split ratio during period t, and They are The hydrogen mixing ratio of the hydrogen-mixed gas turbine and the hydrogen-mixed gas boiler during the period, is the ammonia blending ratio of the thermal power unit. In order to ensure that the action of the designed environment can adapt to the actual operation requirements of the hydrogen virtual power plant scheduling model, the action must meet various operating constraints in the model.

[0046] 3.3 State Space Design In the hydrogen virtual power plant discussed in this paper, the information transmitted by the environment to the deep reinforcement learning agent generally includes electrical load, thermal load, wind power output, and real-time natural gas prices. In addition, the action value of the previous period is also a state variable, specifically expressed as: (twenty two); Where: is the state space, is the electricity load demand during period t; is the heat load demand during period t, is the predicted value of wind power in period t; is the unit price of natural gas purchased during period t.

[0047] 3.4 Evolving Soft Actor-Criter Algorithm Evolutionary strategies (ES) are renowned for their efficient parallel search capabilities and ability to optimize continuous parameter spaces. The SAC algorithm, a deep reinforcement learning algorithm based on the maximum entropy principle, excels in handling high-dimensional continuous action spaces, emphasizing the balance between exploration and exploitation. However, SAC uses a dual Q-network to estimate the lower bound of Q-values. Because the estimated Q-values are approximate and can deviate significantly from the true Q-values, this leads to overly centralized strategies, making the agent prone to local optima, conservative exploration, and difficulty expanding to a wider range of action spaces. To address these issues, the present invention employs an evolutionary soft actor-critic (ESAC) algorithm, incorporating the population optimization principle of ES to enhance SAC's global search capabilities. ESAC leverages ES's diverse exploration mechanisms to address SAC's shortcomings in local optima and exploration efficiency, avoiding the oversampling problem and achieving more robust and efficient learning in complex environments.

[0048] 3.4.1 Algorithm Objective The goal of the ESAC algorithm is to find an optimal target strategy , so that the expected value of the sum of the cumulative returns and the entropy regularization term of the strategy is maximize: (twenty three); Where: represents the expectation function, Indicated by strategy Generated environment interaction trajectories; is the state during period t; is the action during period t; It is the immediate reward returned by the environment; is the discount factor; It's a strategy In state Entropy under is the temperature parameter , T is the total number of time periods.

[0049] 3.4.2 Core Parameter Update Rules The ESAC algorithm network structure mainly includes an actor and a judge. The judge contains two Q-function neural networks for estimating Q values and a target Q-function neural network to enhance the training stability of the algorithm.

[0050] The task of the judge is to estimate the soft Q function In practice, ESAC uses a Q-function neural network To approximate the Q value, is the network parameter. The training goal of the judge is to minimize the temporal difference error, and its loss function is Defined as: (twenty four); Where: It is an experience pool that stores historical interaction data ; is the reward value in period t, is the action in period t+1, is the state in period t+1; is the target value; is the target Q function of the neural network.

[0051] The network parameters are updated using exponential moving average: (25); Where: is the smoothing factor, are the updated network parameters.

[0052] The actor is responsible for optimizing the strategy ,in is the parameter of the policy network. The goal of ESAC is to maximize the expected soft value, and its objective function as follows: (26); Where: and are the mean and standard deviation of the policy network output, respectively.

[0053] Optimization via gradient ascent , the policy network gradually adjusts and , so that the action distribution maintains appropriate entropy while obtaining a high Q value.

[0054] Temperature parameters The role of SAC is to dynamically adjust the balance between exploration and utilization. By defining the target entropy , and optimize the loss function , as follows: (27); By updating with gradient descent, the actual entropy of the policy is close to the target entropy .

[0055] 3.4.3 Evolutionary Strategy Evolutionary strategies are a heuristic search process that involves mutating a population of offspring using random perturbations. After mutation, the fitness metrics corresponding to each member of the population are evaluated, and the offspring with the highest scores are recombined to form the next generation of the population. This paper improves on these characteristics by combining deep reinforcement learning with deep reinforcement learning.

[0056] Assume that the baseline strategy parameters of the evolution strategy are , n is the number of offspring in the population, then the nth individual can be expressed as There are n random perturbations in total . Random perturbations are drawn from a Gaussian distribution Sampling in order to mutate And evaluate the fitness target. The fitness value after a single noise disturbance is highly random, so the expected value of the reward value is taken as the final fitness, as shown in the following formula: (28); Where, is the fitness of the nth individual, is the mutation rate; is the reward value, is the nth random perturbation.

[0057] After calculating the fitness of all individuals, soft winner selection is performed, that is, the fitness of all individuals is normalized and sorted. The first w winner individuals are selected to form the winner individual set W: (29); Where: e is the proportion of dominant individuals.

[0058] In traditional ES, the mutation rate It is an important hyperparameter that directly affects the perturbation amplitude of the strategy parameters. Too small may lead to insufficient population diversity and fall into local optimality; Excessive perturbation noise will mask the effective improvement of the strategy, resulting in slow convergence. Therefore, the automatic mutation tuning mechanism (AMT) is introduced to dynamically adjust the mutation rate. , which enables ESAC to adaptively adjust the disturbance amplitude and reduce the sensitivity of hyperparameters according to the current performance of the population, thereby improving the robustness of the algorithm. The AMT update rule is as follows: (30); Where: is the mutation rate during period t+1; is the mutation rate during period t; is the shear function; is the learning rate of the evolution strategy; is the loss function for automatic mutation adjustment; is the maximum fitness of the population; is the average fitness of the population; is the adjustment factor, limiting The maximum change.

[0059] Loss Function Combining the characteristics of absolute loss and square loss is the key to AMT's balanced exploration and utilization. The mathematical form is as follows: (31); The purpose of the loss function is to achieve smooth gradients and establish an adaptive response mechanism. For smooth gradients, a quadratic function is used when the error is small, ensuring that the gradient changes linearly with the error, avoiding drastic fluctuations and maintaining fine-tuning stability. When the error is large, a linear function is used, maintaining a constant gradient to prevent gradient explosion. Regarding the adaptive response mechanism, when the winner is only slightly better than the population average (e.g., small fluctuations in renewable energy output), the smoothness of the loss function allows for small parameter adjustments to avoid excessive perturbations. However, when the winner significantly outperforms other individuals (e.g., sudden load peaks or sudden changes in wind power output), a linear response allows for rapid parameter increases, driving the population to explore new strategies.

[0060] Baseline strategy parameters of evolution strategy After executing the soft winner selection, the update is performed. The update formula is as follows: (32); Where: and are the baseline strategy parameters for the current and next generation populations respectively; is the learning rate of the evolution strategy; and Respectively The fitness and random disturbance value of each individual.

[0061] 3.4.4 Evolutionary Soft Actor-Criterion Algorithm Update Process The above article introduces the core contents of the ESAC algorithm's goals, network structure and evolutionary strategy in detail. In order to explain the algorithm's implementation process more comprehensively and in-depth, the following will focus on the algorithm's update process and elaborate on its specific steps and mechanisms. The ESAC algorithm training process is as follows: Figure 2 As shown, the training steps are as follows: (1) Initialization parameters and population Parameter initialization: state value function , objective value function , network parameters , baseline strategy parameters , parameters of the policy network , experience pool D.

[0062] Hyperparameter setting: SAC algorithm agent learning rate , learning rate of evolution strategy , mutation rate , SAC algorithm update probability , adjustment factor , by strategy Generated environment interaction trajectory , the proportion of dominant individuals e and the fixed interval g.

[0063] Let the number of iterations k=0.

[0064] Population initialization: population , each individual passes generate.

[0065] Initialize the system state from the environment, let t=0.

[0066] (2) Evaluating individual fitness Each individual according to its own strategy and current state Output Action The environment moves to a new state based on actions and random source load scenarios. , and calculate the reward value , Store it in the experience pool D, and calculate the fitness according to formula (28) , forming a fitness set , is the fitness of the nth individual, and determines whether t≥T is satisfied. If not, set t=t+1, re-output the action and continue updating. If yes, go to step (3); (3) Soft winner selection, crossover, and mutation rate update Perform soft winner selection, select fitness before A victorious individual, , the winners are selected according to fitness order to form a winner set , a crossover operation is performed between the winner and the offspring to retain the dominant features. The mutation rate is updated by equation (30), and the baseline strategy parameters are updated by equation (32).

[0067] (4) SAC algorithm gradient update and experience replay When the number of iterations k is a multiple of the fixed interval g (k mod g = 0), the random sample value μ∼N(0,1) is less than , randomly sample from the experience pool D at a fixed interval g , the judge, actor and temperature parameters are updated by equations (24), (26) and (27) respectively.

[0068] (5) Population update and iteration By the winner individual after crossover , network parameters and baseline strategy parameters Together they form a new population.

[0069] (6) If the algorithm training has reached the maximum number of iterations K (k ≥ K), the ESAC algorithm training process ends. If not, return to step (2) and start a new round of training.

[0070] 3.5 Optimizing the Scheduling Process Based on Evolutionary Deep Reinforcement Learning The present invention adopts the technical route of "model training, online optimization" to transform the constructed hydrogen virtual power plant scheduling model into a training environment adapted to the ESAC algorithm. After the model training is completed, it is used for online optimization scheduling. Figure 1 The specific implementation process is as follows: S1. Initialize the environment and parameters: Based on the hydrogen virtual power plant scheduling model and related constraints, build a deep reinforcement learning environment, input historical wind power and load data, and initialize the parameters of the Evolutionary Soft Actor-Criterion (ESAC) algorithm. S2. Offline Training: The agent obtains state variables for the current period from the environment. These state variables include electrical load, thermal load, wind power output, real-time natural gas prices, and actions from the previous period. The agent then outputs actions based on these state variables, interacts with the environment, provides feedback on rewards, and updates the state for the next period. The agent is a deep reinforcement learning model optimized using an evolutionary soft actor-critic algorithm. S3. Maintain Model: Maintain convergence of the deep reinforcement learning model optimized by the evolving soft actor-critic algorithm after reaching a pre-defined maximum number of iterations; S4. Online Dispatch: Obtain real-time data on the electric load, thermal load, wind power output, and natural gas prices during the forecast period, and combine this with the actual output data of each unit to form state variables. S5. Generate a scheduling plan: Input the state variables into the deep reinforcement learning model optimized by the trained evolutionary soft actor-critic algorithm to generate the optimal scheduling plan for time period t, ensuring that the constraints are met and the operating cost is minimized.

[0071] Another embodiment of the present invention provides an electronic device, including a memory and a processor, wherein the memory stores computer-readable instructions, and when the instructions are executed by the processor, the processor implements the above-mentioned virtual power plant scheduling method.

[0072] Yet another embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above-mentioned virtual power plant scheduling method when executed by a processor.

[0073] The above description merely represents preferred embodiments of the present invention and is not intended to limit the present invention in any other manner. Any person skilled in the art may utilize the above disclosure to modify or modify the present invention into equivalent embodiments. However, any simple modifications, equivalent variations, and modifications to the above embodiments that do not depart from the technical content of the present invention and are based on the technical essence of the present invention remain within the scope of protection of the present invention.

Claims

1. A virtual power plant scheduling method based on evolutionary deep reinforcement learning, characterized in that: The following steps are involved: S1. Initialize the environment and parameters: Based on the hydrogen virtual power plant scheduling model and related constraints, build a deep reinforcement learning environment, input historical wind power and load data, and initialize the parameters of the evolutionary soft actor-evaluator algorithm; S2. Offline Training: The agent obtains state variables for the current period from the environment. These state variables include electrical load, thermal load, wind power output, real-time natural gas prices, and actions from the previous period. The agent then outputs actions based on these state variables, interacts with the environment, provides feedback on rewards, and updates the state for the next period. The agent is a deep reinforcement learning model optimized using an evolutionary soft actor-critic algorithm. S3. Maintain Model: Maintain convergence of the deep reinforcement learning model optimized by the evolving soft actor-critic algorithm after reaching a pre-defined maximum number of iterations; S4. Online Dispatch: Obtain real-time data on the electric load, thermal load, wind power output, and natural gas prices during the forecast period, and combine this with the actual output data of each unit to form state variables. S5. Generate a scheduling plan: Input the state variables into the deep reinforcement learning model optimized by the trained evolutionary soft actor-critic algorithm to generate the optimal scheduling plan for time period t, ensuring that the constraints are met and the operating cost is minimized.

2. The virtual power plant scheduling method according to claim 1, characterized in that: A hydrogen-containing virtual power plant scheduling model is constructed based on the multi-energy coupling operation mode and carbon emission model. The multi-energy coupling operation mode takes hydrogen energy as the core, integrates electricity-to-hydrogen, electricity-to-gas, gas-to-electricity, and hydrogen-to-ammonia technologies to achieve the cascade utilization of hydrogen energy: hydrogen is first directly supplied to the mixed hydrogen gas pipeline network, then used for synthetic natural gas, and the remainder is used for ammonia production or hydrogen storage; when the hydrogen-containing virtual power plant has insufficient hydrogen energy supply, hydrogen is purchased through hydrogen energy green certificate trading; when the hydrogen-containing virtual power plant has excess hydrogen energy, hydrogen is sold through hydrogen energy green certificate trading; the hydrogen-containing virtual power plant scheduling model includes: a hydrogen-containing virtual power plant operation model composed of various equipment models and various gas production and consumption models on the source side, a ladder carbon trading model, a hydrogen energy green certificate trading model, an objective function and constraints.

3. The virtual power plant scheduling method according to claim 2, characterized in that: The objective function of the hydrogen virtual power plant scheduling model is to minimize the total operating cost of the hydrogen virtual power plant, as shown in the following formula: ; Where: is the total operating cost of the hydrogen virtual power plant; The unit operating cost; is the tiered carbon trading cost; for carbon sequestration costs; for the cost of carbon capture equipment; Penalty costs for wind curtailment; The cost of purchasing gas; The start-up and shutdown and coal consumption costs of thermal power units; The proceeds from hydrogen green certificate trading.

4. The virtual power plant scheduling method according to claim 3, characterized in that: The ladder carbon trading model introduces a ladder carbon trading mechanism, including: Allocation of free carbon quotas: The baseline method is used to allocate carbon emission quotas for thermal power units, hydrogen-mixed gas turbines, and hydrogen-mixed gas boilers; Tiered penalty pricing: When actual carbon emissions exceed the quota threshold, the excess amount will be subject to tiered incremental pricing; Emission reduction compensation coefficient: When emissions are lower than the quota benchmark, the market-based trading benefits are enhanced through the compensation coefficient.

5. The virtual power plant scheduling method according to claim 3, characterized in that: The hydrogen green certificate trading model is as follows: ; Where: is the clearing price of hydrogen green certificates; The maximum acceptable price for hydrogen green certificates; The demand ratio of hydrogen energy green certificates is is the hydrogen power produced by hydrogen production during period t; T is the total number of periods.

6. The virtual power plant scheduling method according to claim 1, characterized in that: Constructing the action space of a deep reinforcement learning model based on a hydrogen-containing virtual power plant dispatch model and state space , action space for: ; in, is the wind power output during period t, is the power generation capacity of the thermal power unit during period t, is the power supply of the hydrogen-mixed gas turbine during period t, is the hydrogen power consumed by the methane tank during period t, is the power consumption of the hydrogen production unit during period t, is the action variable of the energy storage system during period t, is the flue gas split ratio during period t, and They are The hydrogen mixing ratio of the hydrogen-mixed gas turbine and the hydrogen-mixed gas boiler during the period, is the ammonia blending ratio of thermal power units, and each action must meet the constraints of the hydrogen virtual power plant scheduling model; State Space Expressed as: ; Where: is the electricity load demand during period t; is the heat load demand during period t, is the predicted value of wind power in period t; is the unit price of natural gas purchased during period t.

7. The virtual power plant scheduling method according to claim 1, characterized in that: The evolutionary soft actor-critic algorithm includes: Criterion: Uses a Q-function neural network to estimate the soft Q-function. The training goal is to minimize the temporal difference error, and the network parameters are updated using exponential moving average. Actor: optimizes the policy to maximize the expected soft value; Evolutionary strategies: generate populations, evaluate fitness, perform soft winner selection and crossover, and update baseline strategy parameters; Automatic mutation adjustment mechanism: used to dynamically adjust the mutation rate during the random perturbation process of the evolution strategy.

8. The virtual power plant scheduling method according to claim 7, characterized in that: The automatic mutation adjustment mechanism is expressed as follows: ; ; Where: is the mutation rate during period t+1; is the mutation rate during period t; is the shear function; is the learning rate of the evolution strategy; is the loss function for automatic mutation adjustment; is the maximum fitness of the population; is the average fitness of the population; is the adjustment factor.

9. An electronic device comprising a memory and a processor, wherein the memory stores computer-readable instructions, wherein: When the instruction is executed by the processor, the processor implements the virtual power plant scheduling method described in any one of claims 1-8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the virtual power plant scheduling method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Virtual power plant optimal scheduling method considering carbon transaction and green certificate transaction

    CN115081715A

  • Deep reinforcement learning low-carbon scheduling method and system for integrated energy system

    CN116993128A

  • Scheduling decision model establishment method based on SumTree-TD3 algorithm

    CN117291390A

  • Virtual power plant scheduling method considering demand response under carbon-green certificate transaction mechanism

    CN118378828A

  • Virtual power plant low-carbon optimization scheduling method, device, equipment, medium and product

    CN118967358A

Cited By

  • Intelligent scheduling method and system for carbon capture of thermal power plant

    CN121052459A