Optimal Scheduling Method for Integrated Electric-Hydrogen-Thermal Energy Microgrid Based on Improved TD3 Algorithm

Through the improved TD3 algorithm and priority experience playback mechanism, a mathematical model of hydrogen energy storage is constructed, which solves the accuracy and real-time scheduling results in the electro-hydrogen-thermal integrated energy microgrid, improves the environmental friendliness and scheduling efficiency of the energy storage system, achieves lower operation and environmental costs, and verifies the migratory ability of the strategy.

CN119787352BActive Publication Date: 2025-05-27SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510272622.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-05-27
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the time-varying operation characteristics and complex constraints of various power components in the electro-hydrogen-thermal integrated energy microgrid, which makes it difficult to ensure the accuracy and real-time nature of the scheduling results. At the same time, the experience playback process of the deep reinforcement learning algorithm fails to effectively utilize the importance or priority of experience, resulting in low sampling efficiency.

Method used

The improved TD3 algorithm is used to construct a mathematical model of hydrogen energy storage, explore the electrothermal coupling characteristics between the hydrogen energy storage system and the microgrid, and use the priority experience replay mechanism to train the scheduling strategy aimed at minimizing environmental costs, equipment life loss costs and large grid interaction costs.

Benefits of technology

The environmental friendliness and scheduling strategy solution speed of the microgrid energy storage system are improved, the daily operating costs and environmental costs of the system are achieved, and the migration of the strategy is verified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119787352B_ABST
    Figure CN119787352B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of integrated energy microgrid scheduling optimization, and discloses an optimized scheduling method for an electric-hydrogen-thermal integrated energy microgrid based on an improved TD3 algorithm, aiming to minimize the sum of environmental costs, equipment life depreciation costs, and interaction costs with the large power grid. The improved TD3 algorithm is used to study the optimized scheduling strategy of the electric-hydrogen-thermal integrated energy microgrid system, construct a mathematical model of hydrogen energy storage, explore the electro-thermal coupling characteristics between the hydrogen energy storage system and the microgrid, transform the optimized scheduling process of the microgrid into a Markov decision process, and introduce an experience replay mechanism to improve the performance of the algorithm. The present invention improves the environmental friendliness of the microgrid energy storage system and the solving speed of the scheduling strategy, and uses the improved prioritized experience replay reinforcement learning algorithm to accelerate the solving speed of the optimal scheduling strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of integrated energy microgrid dispatching optimization, and specifically to an optimized dispatching method for an electric-hydrogen-thermal integrated energy microgrid based on an improved TD3 (Twin Delayed Deep Deterministic Policy Gradient) algorithm. Background Technique

[0002] Under the background of carbon peaking and carbon neutrality, renewable energy represented by photovoltaic and wind power has developed rapidly in recent years due to its advantages of cleanliness and sustainability, and its penetration rate in the power grid has been continuously increasing.

[0003] At present, most renewable energy sources are connected to the power grid through power electronic devices such as inverters to form distributed power sources in the power system. Compared with the centralized power supply form of fossil energy, distributed power sources have two operating modes: grid-connected and off-grid, which can improve the flexibility of the power supply method and reduce the environmental pressure caused by fossil energy to a certain extent. However, their output has obvious regional distribution differences and spatio-temporal volatility, and when they are connected to the grid on a large scale, they will cause impacts and reduce the stability of the power system.

[0004] In order to effectively solve the contradiction between traditional power grids and distributed power sources, scholars such as R.H. Lasserter proposed the concept of a microgrid. A microgrid refers to a small power system in which a high proportion of distributed power sources are the main energy supply method for loads, and at the same time, energy storage and energy conversion devices, and data acquisition and monitoring control systems are equipped. The operation mode of the microgrid has great flexibility. It can operate in parallel with the grid or operate independently from the main grid. It can be both a load of the main grid and supply electric energy to the main grid. The microgrid is regarded as an important part of the future new power system because it has six major advantages: improving the two-way flow of bearing information, improving power quality, power system reliability and resilience, energy comprehensive utilization efficiency, enhancing energy coordination and interaction between regions, and realizing active control of loads.

[0005] As a flexible, efficient and clean secondary energy, introducing hydrogen energy into the microgrid and coupling it with renewable energy can effectively improve the flexibility, renewable energy consumption capacity and cleanliness of the microgrid.

[0006] However, for the optimized dispatching of electric-hydrogen-thermal integrated energy microgrids, the following technical problems still exist:

[0007] 1. It is necessary to establish mathematical models for each power component in the microgrid to ensure the accuracy of the dispatching results. However, the operating characteristics and constraint conditions of each component during the entire dispatching process are time-varying. Establishing an accurate mathematical model requires comprehensively considering complex factors such as the dynamic characteristics, energy consumption, and energy conversion of each component, and it is difficult to establish a real-time parameter model during operation.

[0008] 2. The experience replay process based on the scheduling optimization of the deep reinforcement learning algorithm is based on the "first in, first out" principle and does not consider the importance or priority of experiences, which may lead to some important experiences being frequently overwritten or forgotten, thus resulting in the problem of low sampling efficiency. Summary of the Invention

[0009] In view of the above problems, the purpose of the present invention is to provide an optimized scheduling method for an electric-hydrogen-thermal integrated energy microgrid based on an improved TD3 algorithm. By constructing a mathematical model of hydrogen energy storage, exploring the electro-thermal coupling characteristics between the hydrogen energy storage system and the microgrid, the environmental friendliness of the microgrid energy storage system and the solution speed of the scheduling strategy are improved, and an improved prioritized experience replay reinforcement learning algorithm is used to accelerate the solution speed of the optimal scheduling strategy. The technical solution is as follows:

[0010] The optimized scheduling method for an electric-hydrogen-thermal integrated energy microgrid based on an improved TD3 algorithm includes the following steps:

[0011] Step 1: Establish corresponding mathematical models for the components of the electric-hydrogen-thermal integrated energy system: electrolyzer, fuel cell, hydrogen storage tank, and compressor;

[0012] Step 2: Based on the Markov decision process, establish a Markov decision process for the optimized scheduling of the electric-thermal-hydrogen integrated energy microgrid: ; where, is the possible state space of each power component in the microgrid environment; is the action space, including various operations and scheduling strategies adopted by the microgrid system; is the state transition probability, describing the probability distribution of the microgrid system transitioning from one state to another after taking a certain action in state ; is the immediate reward, representing the reward obtained by the agent after taking a certain action in state ; is a discount factor between 0 and 1, indicating the degree of emphasis on future rewards;

[0013] Step 3: Train the agent based on the prioritized experience replay TD3 algorithm to obtain a day-ahead scheduling strategy with the goal of minimizing the sum of the life loss costs of the electrolyzer and lead-acid battery, the carbon cost of the microgrid system, and the natural gas cost purchased by the microgrid system from the natural gas supplier.

[0014] The beneficial effects of the present invention are:

[0015] 1. The present invention improves the environmental friendliness and the solution speed of the scheduling strategy of the microgrid energy storage system: By introducing a hydrogen energy system and deeply studying its electro-thermal coupling characteristics, not only the energy storage efficiency of the microgrid is improved, but also the environmental friendliness of the system is significantly enhanced. At the same time, by using the improved Prioritized Experience Replay Reinforcement Learning Algorithm (TD3-PER, Twin Delayed Deep Deterministic Policy Gradient with Prioritized Experience Replay, that is, the improved TD3 algorithm), the solution speed of the optimal scheduling strategy is greatly accelerated.

[0016] 2. The present invention achieves lower daily operating costs and environmental costs of the system and verifies the transferability of the strategy: Compared with the traditional TD3 algorithm, the TD3-PER algorithm can more efficiently utilize the scheduling strategy experience and find a better action decision space, thereby reducing the daily operating costs and environmental costs of the system while ensuring the system stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 FIG. is a schematic flow chart of the optimized scheduling method for the integrated electricity-hydrogen-thermal energy microgrid based on the improved TD3 algorithm of the present invention.

[0018] Figure 2 FIG. is a structural diagram of the integrated electricity-hydrogen-thermal energy microgrid system of the simulation example of the present invention.

[0019] Figure 3 FIG. is a curve graph showing the change of the reward value during the training process of the intelligent agent with the number of training rounds.

[0020] FIG. 4(a) shows the optimized scheduling result for a typical summer day.

[0021] FIG. 4(b) shows the optimized scheduling result for a typical winter day.

[0022] Figure 5 FIG. is the calculation result of the similarity between power grids.

[0023] Figure 6 FIG. is the convergence result of the reward value of the day-ahead optimized scheduling. DETAILED DESCRIPTION OF THE INVENTION

[0024] The present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0025] The present invention provides an optimized scheduling method for an integrated electricity-hydrogen-thermal energy microgrid based on an improved TD3 algorithm, aiming to minimize the sum of environmental cost, equipment life depreciation cost, and interaction cost with the large power grid. The improved TD3 algorithm is used to study the optimized scheduling strategy of the integrated electricity-hydrogen-thermal energy microgrid system, construct a mathematical model of hydrogen energy storage, explore the electro-thermal coupling characteristics between the hydrogen energy storage system and the microgrid, transform the optimized scheduling process of the microgrid into a Markov decision process, introduce an experience replay mechanism to improve the algorithm performance, and verify that transfer learning can effectively improve the training rate of the agent reward value.

[0026] As Figure 1 shown, the optimized scheduling method for the integrated electricity-hydrogen-thermal energy microgrid based on the improved TD3 algorithm of the present invention includes the following steps:

[0027] Step 1: Establish corresponding mathematical models for the components of the integrated electricity-hydrogen-thermal energy system, namely, electrolyzers, fuel cells, hydrogen storage tanks, and compressors. The specific steps are as follows:

[0028] The structure of the integrated electricity-hydrogen-thermal energy microgrid system studied in the present invention is as Figure 2 shown. Considering cost, technological maturity, and the power demand of the microgrid, an alkaline electrolyzer is selected as the hydrogen production device for the integrated electricity-hydrogen-thermal energy microgrid. An alkaline electrolyzer usually includes an anode and a cathode, both of which are made of inert materials such as platinum. The anode and the cathode are connected to each other through an electrolyte solution or an alkaline substance in the electrolyzer, and the electrolysis reaction mainly occurs on the anode and the cathode.

[0029] On the anode, water undergoes an oxidation reaction to produce oxygen and hydrogen ions, expressed as:

[0030] ;

[0031] On the cathode, water undergoes a reduction reaction to generate hydrogen and hydroxide ions, expressed as:

[0032] ;

[0033] Passing direct current through the alkaline electrolyzer can electrolyze water to produce hydrogen. The relationship between its hydrogen production rate and the external current is expressed as:

[0034] ;

[0035] In the formula, represents the hydrogen production rate of the electrolyzer at time is the Faraday constant, is the number of electrons generated per reaction, is the external current of the electrolyzer at time is the number of electrolyzers, is the Faraday efficiency, and its calculation formula is as follows:

[0036] ;

[0037] In the formula, is the area of the electrolysis module, is the operating temperature of the electrolytic cell; is the Faraday efficiency coefficient.

[0038] The formula for calculating the voltage of a single electrolytic cell is expressed as:

[0039] ;

[0040] In the formula, is the standard free energy of liquid water, is the empirical temperature coefficient, , are the Ohmic resistance parameters of the electrolytic cell, and are the empirical fitting parameters of the overvoltage model of the electrolytic cell, and and are the basic activation overvoltage constants, and are the temperature linear correction coefficients, and are the temperature quadratic correction coefficients.

[0041] The output power of the electrolytic cell is expressed as:

[0042] ;

[0043] In the formula, is the efficiency of the electrolytic cell.

[0044] In the present invention, it is combined with a lead-acid battery to form a microgrid composite energy storage system. The voltage of the proton exchange membrane fuel cell stack is expressed as:

[0045] ;

[0046] In the formula, is the terminal voltage of the fuel cell, is the internal equivalent impedance of the fuel cell, is the fuel cell current, , and are empirical coefficients, is the number of fuel cells connected in series.

[0047] The hydrogen consumption rate of the fuel cell is expressed as:

[0048] ;

[0049] In the formula, is the hydrogen utilization efficiency of the fuel cell.

[0050] This invention selects compressed hydrogen storage technology for hydrogen storage. The internal pressure of the hydrogen storage tank is expressed as:

[0051] ;

[0052] In the formula, is the internal pressure of the hydrogen storage tank at time is the gas constant, is the temperature of the hydrogen storage tank, is the volume of the hydrogen storage tank, represents the hydrogen consumption rate of the heat load.

[0053] Since the pressure in the hydrogen storage tank is proportional to the amount of stored hydrogen, the relative capacity of hydrogen at this time can be obtained by the following formula:

[0054] ;

[0055] In the formula, is the relative capacity of hydrogen at time is the maximum pressure that the hydrogen storage tank can withstand.

[0056] Step 2: Based on the Markov decision process, establish the Markov decision process for the dispatching optimization of the integrated electric, thermal, and hydrogen energy microgrid: ; where is the possible state space of each power component in the microgrid environment; is the action space, including various operations and dispatching strategies adopted by the microgrid system; is the state transition probability, describing the probability distribution of the microgrid system transferring from one state to another after taking a certain action in state ; is the immediate reward, representing the reward obtained by the agent after taking a certain action in state ; is a discount factor between 0 and 1, indicating the degree of emphasis on future rewards. The specific steps include:

[0057] (1) Possible state space :

[0058] Each time of the microgrid state variable of the integrated energy microgrid system

[0059] ;

[0060] In the formula, , and respectively represent the output power of photovoltaic and wind power and the electrical load within the scheduling moment, and respectively represent the price of natural gas and the carbon trading price within the scheduling moment; is the state of charge of the lead-acid battery; represents the capacity of the fuel cell at the moment; represents the demand for the heat load of the microgrid at the moment.

[0061] (2) Action space :

[0062] For each moment, the charging and discharging actions of the composite energy storage system of the microgrid control the supply-demand balance of the microgrid. At the same time, the electrolyzer is controlled to produce hydrogen to provide raw materials for the fuel cell and meet the demand for the heat load. Then the action variable of the system is expressed as:

[0063] ;

[0064] In the formula, is the power of the lead-acid battery at the moment; represents the power of the fuel cell at the moment;

[0065] (3) Immediate reward :

[0066] The goal in reinforcement learning is to maximize the cumulative reward. However, the objective function of the integrated energy microgrid of the present invention is to minimize the daily operating cost. Therefore, the reward function of reinforcement learning is set to the negative of the objective function to achieve direction consistency, which is expressed as:

[0067] ;

[0068] In the formula, is the objective function of the day-ahead optimal scheduling of the microgrid.

[0069] (4) State transition probability :

[0070] ;

[0071] In the formula, is the state space at time is the action space at time

[0072] Step 3: Train the agent based on the prioritized experience replay TD3 algorithm to obtain a day-ahead scheduling strategy with the goal of minimizing the sum of the life loss costs of the electrolyzer and the composite energy storage system, the carbon cost of the microgrid system, and the natural gas cost purchased by the microgrid system from the natural gas supplier.

[0073] The specific steps are as follows:

[0074] When the hydrogen in the system cannot meet the demand of the heat load, the microgrid needs to purchase natural gas from the natural gas supplier to fill the gap. Carbon dioxide will be generated during the use of natural gas. Therefore, the environmental cost of the microgrid needs to be considered. At the same time, the life of the electrolyzer and energy storage equipment will also be damaged during operation. In summary, the day-ahead optimal scheduling objective function of the microgrid is constructed as follows:

[0075] ;

[0076] In the formula, and respectively represent the life loss costs of the electrolyzer and the composite energy storage system at time represents the carbon cost of the microgrid system, is the natural gas cost purchased by the microgrid system from the natural gas supplier, is the interaction power between the microgrid and the main grid at time is the real-time electricity price at time

[0077] At the same time, the optimal scheduling problem of the integrated electricity-hydrogen-heat energy microgrid should satisfy the following constraints:

[0078] (1) Electric load balance:

[0079] ;

[0080] In the formula, represents the fuel cell power at time

[0081] (2) Heat load balance:

[0082] ;

[0083] In the formula, and respectively represent the thermal energy provided by hydrogen and natural gas at a moment, represents the thermal load demand of the microgrid at a moment.

[0084] (3) Power constraints of each power component:

[0085] ;

[0086] In the formula, and respectively represent the maximum and minimum state of charge of the lead-acid battery; and respectively represent the maximum and minimum output of the lead-acid battery; represents the capacity of the fuel cell at a moment, and respectively represent the upper and lower limits of the fuel cell capacity, and respectively represent the upper and lower limits of the fuel cell power, and respectively represent the upper and lower limits of the electrolyzer power, represents the volume of hydrogen gas passing through the compressor at a moment, represents the upper limit of the volume of hydrogen gas passing through the compressor, represents the volume of hydrogen gas output by the compressor at a moment, represents the upper limit of the volume of hydrogen gas output by the compressor.

[0087] According to a possible implementation manner, during the process of training the intelligent agent based on the prioritized experience replay TD3 algorithm, the algorithm alternately performs policy evaluation, policy improvement, and prioritized experience replay; among them,

[0088] In the policy evaluation stage, it is necessary to calculate the state-action value, that is, , and this Q function can be represented by the Bellman equation as:

[0089] ;

[0090] In the formula, represents the state-action value in the value evaluation Q network; represents the immediate reward at a moment and the state-action value with the discount factor at a moment the sum of expectations.

[0091] After parameterizing the Q - function using a neural network, the Q - function is approximated by minimizing the Bellman residual:

[0092] ;

[0093] where, is the minimized Bellman residual, and represent the parameters of the value - evaluation Q - network and the TargetQ - network respectively.

[0094] In the policy improvement stage, after parameterizing the Q - function using a neural network, the objective function is used to update the gradient of the network parameters, that is:

[0095] ;

[0096] where, represents the derivative with respect to the policy network parameters ; is the expectation, represents the derivative with respect to the action in the state - action value of the policy network value - evaluation network in the th training episode, represents the derivative with respect to the parameter in the policy network in the th training episode.

[0097] According to a possible implementation, during the process of training an agent using the improved TD3 algorithm, target value evaluation is also performed through two Target Q - networks with different initial parameters, and the smaller value is selected as the target value. Therefore, the Bellman residual to be minimized is modified to:

[0098] ;

[0099] where, is the reward at time , is the smaller value of the value evaluation at time ; is the action after adding perturbation at time ; k represents the number of value - evaluation Q - networks; is the value evaluation at time .

[0100] Use target - policy smoothing regularization to enhance the stability of the policy and smooth the Q - function; that is, when calculating the Bellman residual, for the action taken in the next state will be selected as:

[0101] ;

[0102] where μ represents the Target policy network; ε is the added noise, usually selected as Gaussian noise, and its amplitude is clipped to be limited within a small range.

[0103] Prioritized experience replay measures the importance of each sample by its corresponding Temporal Difference Error (TD-error). The agent preferentially selects samples with higher values for training, thus reducing the exploration time. Specifically, samples with larger TD-errors (i.e., larger deviations between the actual Q-value and the estimated Q-value of the sample) are considered more important for the network's learning, so they are given higher probabilities and are preferentially used for training.

[0104] Introduce the concept of sample priority: Define each sample in the experience pool the sampling probability by the agent is positively correlated with the absolute value of the TD-error and is expressed as:

[0105] ;

[0106] where represents the sample sampling probability matrix, represents the probability of each sample being sampled, , is a hyperparameter used to solve the extreme value problem that occurs in the TD-error calculation process (degenerates to uniform sampling when ), is a small positive constant used to ensure that each sample has at least a small priority, even if their TD-error is zero;

[0107] Weight importance sampling rate correction: When using priority sampling, high-priority samples will be selected more frequently. However, this may lead to uneven sample distribution. By introducing the weight importance sampling rate, this unevenness can be balanced to ensure that both high-priority and low-priority samples can receive appropriate attention. Define the importance sampling coefficient as follows:

[0108] ;

[0109] where is the total number of samples in the experience pool; is a hyperparameter used to control the balance between sampling priority and weights in experience replay and smooth the high-variance importance sampling weights.

[0110] Figure 3 The curve of the reward value changing with the number of training episodes during the agent training process is given. It can be observed from the figure that after 900 episodes of training, the DDPG (Deep Deterministic Policy Gradient) algorithm converges near a reward value of 180. However, compared with the TD3 algorithm, the stability of the optimal scheduling strategy obtained by the DDPG algorithm is weaker. At the same time, it can be seen from the figure that in the initial stage, due to insufficient training of the agent, the number of samples in the experience pool is small, resulting in low accuracy of priority calculation, which limits the ability of TD3-PER to make full use of the priority sampling mechanism. The probability that the agent extracts low-reward value experiences is relatively high. As the agent continuously interacts with the environment, the accumulation of experience and the optimization of priority calculation enable TD3-PER to use high-reward value experiences to improve the strategy at a high frequency, and can make better scheduling decisions and obtain higher reward values in the subsequent training process.

[0111] The optimized scheduling results for the selected typical summer and winter days are shown in Figures 4(a) and 4(b). It can be seen from the figures that in summer, the temperature is relatively high, the photovoltaic power supply is large while the internal heat load demand in the power grid is small. During the day, on the basis of meeting the electrical load demand, most of the surplus electrical energy is supplied to the electrolyzer for hydrogen production for use when the heat load demand is large in winter. At the same time, the battery system is charged for use at night. The electrical load and heat load demands at night are relatively higher than those during the day. At this time, the battery discharges, and a small part of the hydrogen in the hydrogen storage tank is used for the fuel cell to release electrical energy to assist the lithium battery to achieve load balance, thereby reducing the operating cost of the microgrid. In winter, the temperature of the typical day is relatively low, the photovoltaic output is low while the heat load demand in the microgrid is high. Excluding the power supplied to the fuel cell, the net power of the electrolyzer for hydrogen production is greatly reduced and can no longer meet the heat load demand, resulting in a gap in heat demand. At this time, the microgrid needs to purchase natural gas from the natural gas supplier to meet the heat load consumption.

[0112] Step 4: Analyze the impact of transfer learning on the training rate of improving the agent's reward value.

[0113] The specific steps are as follows:

[0114] To verify the improvement effect of transfer learning on the generalization ability of the optimized scheduling strategy, taking the above-mentioned microgrid as the source domain, 5 microgrids with similar geographical locations are selected as the target domains. The geographical locations are shown in Table 1. Using the MMD function to calculate the distribution difference between domains with the wind-solar-load data as the standard , taking as the similarity, the calculation results are as Figure 5 shown. It can be seen from the figure that the wind-solar-load data of the No. 3 microgrid within one year has the highest similarity with the source domain microgrid. Therefore, it is regarded as the target domain for migrating the optimized scheduling strategy.

[0115] Table 1 Coordinates of 5 target-domain microgrids

[0116] Power station number Longitude Latitude Source domain -120.6867 37.5929 Target domain microgrid No. 1 -120.6658 37.5795 Target domain microgrid No. 2 -120.6987 37.5800 Target domain microgrid No. 3 -120.6889 37.5961 Target domain microgrid No. 4 -120.6735 36.5821 Target domain microgrid No. 5 -120.6797 37.5885

[0117] When the reward value of the optimal scheduling task of the source-domain microgrid converges, the network parameters are "frozen" and used as the initial values of the network parameters for the optimal scheduling task of the target-domain microgrid for policy transfer. The convergence result of the day-ahead optimal scheduling reward value is as Figure 6 shown, and the agent is retrained to solve the optimum.

[0118] The present invention provides strong support for the wide application and flexible scheduling of microgrids by verifying the transferability of the optimal scheduling strategies between microgrids.

Claims

1. An optimized dispatching method for electric, hydrogen and heat integrated energy microgrid based on improved TD3 algorithm, characterized in that: The following steps are involved: Step 1: Establish corresponding mathematical models for the components of the electric-hydrogen-heat integrated energy system: electrolyzer, fuel cell, hydrogen storage tank and compressor; Step 2: Based on the Markov decision process, establish the Markov decision process for dispatch optimization of electric, thermal and hydrogen integrated energy microgrid: ;in, It is the possible state space of each power component in the microgrid environment; The action space includes various operations and dispatch strategies adopted by the microgrid system; is the state transition probability, describing the state Take an action After that, the probability distribution of the microgrid system transferring from one state to another; is an immediate reward, representing the Take an action After that, the agent gets the reward; is a discount factor between 0 and 1, indicating the importance of future rewards; Step 3: Train the agent based on the priority experience replay TD3 algorithm to obtain a day-ahead dispatch strategy that minimizes the sum of the life depreciation costs of the electrolyzer and lead-acid battery, the carbon cost and electricity transaction cost of the microgrid system, and the cost of natural gas purchased by the microgrid system from the natural gas supplier.

2. The method for optimizing the dispatching of electric, hydrogen and heat integrated energy microgrid based on the improved TD3 algorithm according to claim 1 is characterized in that: In step 1, the mathematical model is as follows: 1) Electrolyzer mathematical model: The relationship between the hydrogen production rate of the alkaline electrolyzer and the external current is expressed as: ; In the formula, express The hydrogen production rate of the electrolyzer at the moment, is the Faraday constant, represents the number of electrons produced in each reaction, express The external current of the electrolytic cell at the moment, Indicates the number of electrolytic cells, represents the Faraday efficiency; The calculation formula of the cell voltage is: ; In the formula, represents the standard free energy of liquid water, represents the empirical temperature coefficient, , They represent the ohmic resistance parameters of the electrolytic cell, and are the empirical fitting parameters of the electrolyzer overvoltage model, and and is the basic activation overvoltage constant, and is the temperature linear correction coefficient, and is the temperature secondary correction coefficient; represents the electrolysis module area, Indicates the working temperature of the electrolytic cell; The output power of the electrolyzer is expressed as: ; In the formula, is the efficiency of the electrolyzer; 2) Fuel cell mathematical model: The electrolyzer and lead-acid battery form a microgrid composite energy storage system, and the voltage of the proton exchange membrane fuel cell stack is expressed as: ; In the formula, is the terminal voltage of the fuel cell, is the equivalent impedance inside the fuel cell, is the fuel cell current, , and is the empirical coefficient, is the number of fuel cells connected in series; is the electromotive force of the fuel cell; The fuel cell hydrogen consumption rate is expressed as: ; In the formula, The hydrogen utilization efficiency of the fuel cell; 3) Mathematical model of hydrogen storage tank and compressor: The internal pressure of the hydrogen storage tank is expressed as: ; In the formula, express The internal pressure of the hydrogen storage tank at the moment, is the gas constant, Indicates the temperature of the hydrogen storage tank. Indicates the volume of the hydrogen storage tank, Indicates the hydrogen consumption rate of the heat load; The relative capacity of hydrogen at this moment is obtained by the following formula: ; In the formula, for The relative capacity of hydrogen at the time, The maximum pressure that the hydrogen storage tank can withstand; The compressor model is expressed as: ; In the formula, express The volume of hydrogen output by the compressor at the moment, express The volume of hydrogen passing through the compressor at any given moment, and Represent the compressor output and input hydrogen pressure respectively, Indicates the compressor efficiency.

3. The method for optimizing the dispatching of electric, hydrogen and heat integrated energy microgrid based on the improved TD3 algorithm according to claim 2 is characterized in that: In step 2, the Markov decision process The specific expressions are as follows: 1) State Space : Microgrids The state of the microgrid system of electric, hydrogen and heat integrated energy at the moment , expressed as: ; In the formula, , and Respectively indicate scheduling Photovoltaic and wind power output power and electrical load at all times, and Respectively, in scheduling The price of natural gas and carbon trading prices at the moment; Indicates the state of charge of the lead-acid battery; express The capacity of the fuel cell at any moment; express Microgrid heat load demand at every moment; express Real-time electricity prices at the moment; 2) Action Space : Operation of the composite energy storage system , expressed as: ; In the formula, express Lead-acid battery power at the moment; express Fuel cell power at the time; express Electrolyzer power at the moment; 3) Instant Rewards : The reward function of reinforcement learning is set to the opposite of the microgrid day-ahead optimization scheduling objective function to achieve directional consistency, which is expressed as: ; In the formula, represents the objective function of the day-ahead optimization dispatch of the microgrid; 4) State transition probability : In Status Take an action After that, the probability distribution of the microgrid system transferring from one state to another is expressed as: ; In the formula, express The state space at time, express Action space at any moment.

4. The method for optimizing the dispatching of electric, hydrogen and heat integrated energy microgrid based on the improved TD3 algorithm according to claim 3 is characterized in that: In step 3, the objective of minimizing the sum of the life depreciation cost of the microgrid composite energy storage system, the carbon cost and electricity transaction cost of the microgrid system, and the cost of natural gas purchased by the microgrid system from the natural gas supplier is to configure the microgrid day-ahead optimization scheduling objective function: ; In the formula, and Respectively The life loss cost of electrolyzers and lead-acid batteries at all times, represents the carbon cost of the microgrid system, represents the natural gas cost purchased by the microgrid system from the natural gas supplier, express The interaction power between the microgrid and the main grid at any time, express Real-time electricity prices at any given moment.

5. The method for optimizing the dispatching of electric, hydrogen and heat integrated energy microgrid based on the improved TD3 algorithm according to claim 3 is characterized in that: In step 3, the optimization scheduling problem of the electric, hydrogen and heat integrated energy microgrid satisfies the following constraints: 1) Electrical load balance: ; In the formula, express Fuel cell power at the time; 2) Heat load balance: ; In the formula, and Respectively The heat energy provided by hydrogen and natural gas at all times, express Microgrid heat load demand at every moment; 3) Power constraints of each power component: ; In the formula, and Respectively represent the maximum and minimum state of charge of the lead-acid battery; and Respectively represent the maximum and minimum output of lead-acid batteries; express The capacity of the fuel cell at the moment, and Respectively represent the upper and lower limits of fuel cell capacity, and Respectively represent the upper and lower limits of fuel cell power, and Respectively represent the upper and lower limits of the electrolytic cell power, express The volume of hydrogen passing through the compressor at any given moment, Indicates the upper limit of the volume of hydrogen that can pass through the compressor. express The volume of hydrogen output by the compressor at the moment, Indicates the upper limit of the hydrogen volume output by the compressor.

6. The method for optimizing the dispatching of electric, hydrogen and heat integrated energy microgrid based on the improved TD3 algorithm according to claim 3 is characterized in that: In step 3, the agent training based on the priority experience playback TD3 algorithm is specifically as follows: In the process of training the agent based on the Priority Experience Replay TD3 algorithm, the algorithm alternately performs strategy evaluation, strategy improvement, and Priority Experience Replay; Among them, in the strategy evaluation stage, the state-action value is calculated, and its Q function is expressed by the Bellman equation as follows: ; In the formula, Represents the state-action value in the value evaluation Q network; express Instant rewards at the moment and Discount factor at time State-action value The expectation of the sum; After parameterizing the Q function using a neural network, the Q function is approximated by minimizing the Bellman residual: ; In the formula, is the minimized Bellman residual, and Represent the parameters of the value assessment Q network and the Target Q network respectively; In the strategy improvement stage, the Q function is parameterized using a neural network to minimize the objective function The gradient used to update the network parameters is: ; In the formula, Represents the policy network parameters Derivation; For expectations, Indicates The state-action value in the Q network in the training rounds Actions in Seeking guidance, Indicates Policy network in training rounds Parameters in Derivation; In the process of training the agent with the improved TD3 algorithm, the target value is evaluated through two Target Q networks with different initial parameters, and the smaller value is selected as the target value, so the minimized Bellman residual is corrected to: ; In the formula, express The rewards of the moment, express The smaller value of the moment value assessment; express The action after adding disturbance at every moment; represents the number of value evaluation Q networks; express The value assessment of the moment; Use target policy smoothing regularization to enhance the stability of the policy and smooth the Q function; that is, when calculating the Bellman residual, the state at the next moment Actions taken Will be selected as: ; In the formula, Represents the Target policy network; For the added noise, is the action after adding disturbance.

7. The method for optimizing the dispatching of electric, hydrogen and heat integrated energy microgrid based on the improved TD3 algorithm according to claim 6 is characterized in that: In the Priority Experience Replay TD3 algorithm, the concept of sample priority is introduced: define each sample in the experience pool The probability of being sampled by the agent and the absolute value of the time difference error The correct sample sampling probability matrix for positive correlation and immediate reward replenishment is expressed as: ; In the formula, represents the sample sampling probability matrix, represents the probability of each sample being sampled, , is a hyperparameter; is a small positive constant; is the time difference error; Indicates that there is Samples; Weighted importance sampling rate correction: The weighted importance sampling rate is introduced to balance the uneven distribution of samples and ensure that both high-priority and low-priority samples receive appropriate attention. The importance sampling coefficient is defined as follows: ; In the formula, is the total number of samples in the experience pool, is a hyperparameter that controls the balance between sampling priority and weights in experience replay and smoothes high-variance importance sampling weights.

Citation Information

Patent Citations

  • Scheduling decision model establishment method based on SumTree-TD3 algorithm

    CN117291390A

  • Power distribution network region voltage control method based on multi-agent deep reinforcement learning

    CN119070315A