A reinforcement learning-consensus theory combined wind-solar-hydrogen storage microgrid regulation method
By combining reinforcement learning and consistency theory to develop a microgrid control method for wind, solar, hydrogen, and energy storage, the problem of low control accuracy of green hydrogen systems in new energy power generation systems has been solved. This method achieves efficient and stable new energy consumption and green hydrogen production, improving the system's flexibility and economy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies have failed to effectively utilize the wide operating conditions and continuous, rapid, and adjustable potential of green hydrogen systems in new energy power generation systems, resulting in low control accuracy. Furthermore, deep reinforcement learning is limited in training efficiency and engineering applicability when faced with fluctuations in wind and solar power output and equipment heterogeneity.
A wind-solar-hydrogen-storage microgrid control method combining reinforcement learning and consensus theory is adopted. By establishing a multi-physics coupling model and designing an optimized control model, combined with the improved deep reinforcement learning algorithm of TD3 and the improved consensus algorithm with gradient compensation, multi-device collaborative optimization and hydrogen storage state equilibrium are achieved.
It has improved the utilization rate of new energy sources, ensured the reliability and economy of green hydrogen production, enhanced the overall optimization capability of the control strategy and the stability of system operation, and reduced operating costs.
Smart Images

Figure CN120914886B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of new energy and micro-grid regulation, and particularly relates to a wind-solar-hydrogen storage micro-grid regulation method combining reinforcement learning and consensus theory. BACKGROUND
[0002] New energy generation has the characteristics of randomness, intermittency and volatility, and large-scale grid connection poses a serious challenge to the flexibility of grid regulation and power supply stability. In recent years, hydrogen energy is becoming an important part of future strategic energy. It not only can be used as a carrier of clean energy, but also has excellent "source-load interaction" characteristics. Through the new energy green hydrogen production project, the local consumption of new energy can be realized, thereby significantly improving the flexibility of new power systems.
[0003] The current green hydrogen system regulation technology has obvious limitations. The mainstream research regards green hydrogen production as a stable load, ignoring its wide operating condition and continuous rapid adjustable potential, which leads to less consideration of the participation of green hydrogen system in grid interaction. Existing models are often based on static or steady-state assumptions, ignoring the dynamic characteristics of hydrogen production and storage devices during operation, which seriously affects the regulation accuracy, and the overall optimal coordination of multiple performance indicators of the system, including power supply reliability, economy and environmental friendliness, is difficult to achieve.
[0004] Although deep reinforcement learning (DRL) has shown potential in power dispatching, the pure data-driven scheme has inherent bottlenecks. In the face of dramatic fluctuations in wind and light output and device heterogeneity, DRL needs to deal with a high-dimensional action space, which seriously restricts the training efficiency and engineering applicability of the strategy. SUMMARY
[0005] Based on the above deficiencies, the present application provides a wind-solar-hydrogen storage micro-grid regulation method combining reinforcement learning and consensus theory, which provides intelligent and efficient regulation technology for large-scale green hydrogen production projects, improves new energy consumption rate, guarantees green hydrogen production demand, and enhances the reliability and economy of electrolytic hydrogen production projects, thereby solving the deficiencies described in the background art.
[0006] To achieve the above purpose, the technical scheme adopted by the present application is as follows: a wind-solar-hydrogen storage micro-grid regulation method combining reinforcement learning and consensus theory, comprising the following steps:
[0007] Step 1: Establish a model of wind-solar-hydrogen storage micro-grid, which covers electrochemical process model, thermodynamic process model and mass transfer process model, and the wind-solar-hydrogen storage micro-grid includes wind power generation equipment, photovoltaic power generation equipment, electrolytic tank, hydrogen storage tank, fuel cell and energy storage equipment;
[0008] Step two: An optimization control model of the wind-solar-hydrogen storage micro-grid is designed with the optimization objectives of minimizing the operation cost and maximizing the hydrogen production, and considering the safe operation boundary constraints of the equipment, including the power balance constraint, the working power constraint, the working temperature constraint, the climbing power constraint of the electrolyzer, the output power constraint, the climbing power constraint of the fuel cell, and the hydrogen storage state constraint of the hydrogen storage tank;
[0009] Step three: The improved consistency algorithm based on gradient compensation is used to dynamically allocate the power of multiple groups of heterogeneous electrolyzers, so as to realize the balanced control of the storage states of multiple hydrogen storage tanks.
[0010] Step four: The deep reinforcement learning algorithm based on improved TD3 is used to dynamically coordinate the electrolyzer, hydrogen storage tank, fuel cell and energy storage equipment, and the optimization result of the micro-grid operation is obtained through learning iteration.
[0011] Among them, step three and step four constitute a hierarchical control architecture, step four is the upper layer control, which realizes the collaborative optimization of multiple devices, and step three is the lower layer control, which realizes the power allocation and hydrogen storage state balance.
[0012] Further, in step one,
[0013] The electrochemical process model includes: the voltage of the alkaline electrolyzer stack is equal to the sum of the reversible voltage, ohmic loss, activation loss and diffusion loss, and its expression is:
[0014] V el = N el V cell = N el (V rev + V ohm + V act + V diff ) (1)
[0015] In the formula, N el is the number of series of alkaline electrolyzers in the stack; V cell is the voltage of a single electrolyzer; V rev is the reversible voltage; V ohm is the ohmic loss; V act is the activation loss; and V diff is the diffusion loss.
[0016] The reversible voltage V rev is obtained from the Nernst equation:
[0017]
[0018] In the formula, z is the number of electrons transferred by each hydrogen molecule, which is 2; F is the Faraday constant; R is the universal gas constant; P is the working pressure; and T is the working temperature, unit: ℃. V is the temperature-dependent reversible voltage at standard pressure; P v,KOH V is the vapor pressure of the electrolyte solution; a is the water activity in the electrolyte;
[0019] V is the ohmic loss ohm as shown in the following formula:
[0020]
[0021] wherein A is the electrode surface area of the electrolytic cell; r1 and r2 are empirical parameters, r2 simulates the increase of the resistance value with temperature, and I is the current of the electrolytic cell;
[0022] V is the activation loss act is caused by the oxidation-reduction reaction occurring on the anode and cathode, and the relationship between the current in the electrode and the activation overvoltage is an approximate formula covering the entire current process under high activation voltage conditions:
[0023] V act = V act,a + V act,c (4)
[0024]
[0025] wherein s1, s2, s3, t1, t2, t3, v1, v2, v3, w1, w2, w3 are empirical parameters, reflecting the dependence relationship between the current of the anode and cathode current sources and the activation overvoltage;
[0026] V is the diffusion loss diff : Since the working current density of the alkaline electrolytic cell is relatively low, its influence is negligible, so it is not included in the model,
[0027] Therefore, the total power consumption of the alkaline electrolytic cell is:
[0028]
[0029] wherein I i and V el,i are the working current and working voltage of the i-th electrolytic cell stack, respectively;
[0030] The thermodynamic process mathematical model comprises:
[0031] The overall heat balance expression of the electrolytic cell is:
[0032]
[0033] wherein the left side term describes the change of the electrolytic cell temperature with time, which depends on the total heat capacity C t ; The heat generated internally is the part of the electric power input to the electrolyzer that is lost as heat. When the electric power input to the electrolyzer exceeds the thermodynamic energy requirement, heat is generated, which is expressed as:
[0034]
[0035] V tn = V Δh + V irrev (10)
[0036]
[0037] where V tn is the heat-neutral voltage; V Δh is the enthalpy voltage calculated from the Faraday's law of electrolysis; and V irrev exists in the actual electrolysis process, where y is the number of moles of water vapor generated per 1 mole of hydrogen gas electrolyzed; is the molar enthalpy of water vapor generated by the electrolyte at the working temperature and pressure; is the molar enthalpy of liquid water under standard conditions; V Δh , is an empirical parameter obtained from experiments;
[0038] Only the heat transfer process between the overall surface of the electrolyzer and the environment is considered, which is expressed as:
[0039]
[0040] where h is the convective-radiative heat transfer coefficient; A stack is the outer surface of the electrolyzer cell stack; A sep is the total outer surface of the two gas separators; T amb is the ambient temperature;
[0041] is the heat removed by the cooling water, which is expressed as:
[0042]
[0043] where C cw is the heat capacity of the cooling water; T cw,i and T cw,o are the inlet and outlet temperatures of the cooling water, respectively. The inlet temperature of the cooling water is known and remains constant, and the outlet temperature is obtained from the following equation:
[0044]
[0045] UA HX = h cond + h convI (16)
[0046] where UA HX is the heat transfer coefficient of the heat exchanger obtained by an empirical formula; h cond and h conv represent the conduction and convection heat transfer coefficients, respectively;
[0047] related to the material leaving the system and the water loss, which is much smaller than the other values, so this term is neglected;
[0048] The mathematical model of the mass transfer process includes:
[0049] According to Faraday's law, the molar flow rate of hydrogen generated at the electrode is related to the activation current of the cathode:
[0050]
[0051] where η F represents the Faraday efficiency, defined as the ratio of the actual gas flow rate to the theoretical flow rate, caused by parasitic current losses; I act,c and I act,a are the activation currents of the cathode and anode, respectively;
[0052] The dissolved substances in the separator of the electrolytic cell undergo mutual diffusion, and the expression for the molar flow rate through the separator is:
[0053]
[0054] where D is the molecular diffusion coefficient of hydrogen; d, ε, and τ are the thickness, porosity, and tortuosity of the separator, respectively; A cell is the surface area of the electrolytic cell compartment; are the molar concentrations of hydrogen at the cathode and anode, respectively;
[0055] A device is installed between the separators in the electrolytic cell to balance the OH - charges consumed / generated in the electrochemical reaction process, and the flow rate of hydrogen transferred to the opposite separator through the mixing tube is expressed as a fraction of the net flow rate into its original separator, and its expression is:
[0056]
[0057] where the empirical expressions for λ and μ are defined in terms of the current, temperature, and pressure:
[0058]
[0059] Finally, the expression for the available hydrogen flow rate at the outlet of the electrolytic cell is obtained:
[0060]
[0061] The hydrogen produced by the electrolyzer enters the hydrogen storage tank for storage. The expression of the hydrogen storage pressure of the hydrogen storage tank is:
[0062]
[0063] In the formula, n sto is the hydrogen storage amount of the hydrogen storage tank; n sto (t0) is the hydrogen storage amount of the hydrogen storage tank at t0; is the hydrogen usage flow rate from the hydrogen storage tank; p sto is the hydrogen storage tank pressure; T sto is the hydrogen storage tank temperature, in K; and V sto is the hydrogen storage tank volume.
[0064] Further, in step two, the control variable vector u and the state variable vector s of the wind-solar-hydrogen storage microgrid model are respectively expressed as:
[0065]
[0066] In the formula, I AEL,m T is the current reference value of the m alkaline electrolyzers, P FC is the fuel cell power reference value; P ES , P new T , P AEL,m T are the current power values of the battery energy storage, the wind-solar stations and the conventional electrical load, and the m alkaline electrolyzers; T AEL,m T are the internal temperatures of the m electrolyzers; SOH m T are the internal hydrogen storage states of the m hydrogen storage tanks; are the hydrogen production rates of the electrolyzers and the hydrogen flow rates from the hydrogen storage tanks for fuel cell power generation, respectively;
[0067] The microgrid operation cost minimization objective function satisfies:
[0068]
[0069] In the formula, γ and τ are the discount factor and the soft update rate, C1, C2, and C3 are the cost coefficients of the electrolyzer, the fuel cell, and the energy storage, respectively;
[0070] The hydrogen production maximization objective function satisfies:
[0071]
[0072] The objective function is defined as:
[0073] R = a1r1 + a2r2 (27)
[0074] where a1 and a2 are weight coefficients set by the two optimization objectives;
[0075] The microgrid tracks the wind and solar power fluctuations through the electrolyzer, the fuel cell assists in regulation, and the remaining power error is balanced by the energy storage, and the power balance constraint is satisfied:
[0076] P ES = P AEL -P new -P FC (28)
[0077] In the real-time control phase, the electrolyzer-related state has working power constraints, working temperature constraints, and
[0078] The ramping power constraint is satisfied:
[0079]
[0080] where and represent the minimum and maximum values of the normal working power of electrolyzer i, respectively; and represent the minimum and maximum values of the internal temperature of the electrolyzer, respectively; R AEL,i represents the upper limit of the ramping rate of electrolyzer i; the fuel cell-related state has an output power constraint and a ramping power constraint, which is satisfied:
[0081] where and represent the minimum and maximum values of the normal working power of the fuel cell, respectively; R FC represents the upper limit of the ramping rate of electrolyzer i;
[0082] The hydrogen storage tank contains a hydrogen storage state constraint, which is satisfied:
[0083] SOH min <SOH i <SOH max (34)
[0084] where SOH min and SOH max represent the minimum and maximum values of the safe hydrogen storage state of the hydrogen storage tank, respectively.
[0085] Further, in step two, the controller of the optimization regulation model of the wind-solar-hydrogen storage micro-grid is designed as a learning agent, which responds to the random environment by sequentially selecting actions in discrete time steps, and is constructed as a Markov decision process with state space S and action space U. At each time step t, the deep reinforcement learning agent observes the current state s t , selects a control action u t , and after the action is executed in the system, the system state will transfer to s t+1 ∈S, where the action u t corresponds to the control input vector u of the control system at time t. To solve this optimization problem, a deterministic policy is used, denoted as μ θ :S→U. For the deterministic policy, the expression of the total discounted reward from time t is:
[0086]
[0087] The following is the discrete form of the objective function value function in equation (35), whose state value function V μ is equal to its action value function Q μ :
[0088] V μ (s t )=π θ (u t |s t )Q μ (s t ,μ θ (s t ))=Q μ (s t ,μ θ (s t )) (36)
[0089] Where the policy π satisfies π θ (u t |s t )=1.
[0090] For the regulation of the wind-solar-hydrogen storage micro-grid, the goal is to solve an optimal control law u(t)=μ θ (s t ) to maximize the cumulative discounted reward.
[0091] Since the electrolyzer allows short-time operation in a certain overrun interval, the constraint of is a soft constraint, which is embedded in the reward function. When the working power of the electrolyzer exceeds , the expression is:
[0092]
[0093] The hydrogen storage tank still has a buffer margin outside the safety hydrogen storage range of the calibration. The constraint is embedded in the reward function. When the hydrogen storage state is less than SOH min or greater than SOH max , the corresponding penalty is triggered, respectively, and the expression is as follows:
[0094]
[0095] Finally, the expression of the hybrid reward function with multiple constraint embeddings is constructed as follows:
[0096] V μ (s t )=α1r1+α2r2+α3r3+α4r4+α5r5 (40)
[0097] Further, in step three, let x i (k), x i (k+1) be the state variables of electrolytic cell i at the kth and k+1th moments in the iteration process, respectively, and x i ∈R represents a state variable of electrolytic cell i. The exchange process of state information between adjacent stacks is regarded as a linear system. By observing and calculating the difference between the hydrogen storage states of adjacent hydrogen storage tanks, the output value of the coordinated electrolytic cell is calculated to finally realize the state balance of the hydrogen storage tank,
[0098]
[0099] where h i (k) and h j (k) are the differences between hydrogen storage tanks i and j at the kth moment; a ij is the correlation coefficient between hydrogen storage tanks i and j. If they are related, 0 < a ij ≤ 1, otherwise a ij = 0; γ is the convergence coefficient;
[0100] The expression of the improved consistency algorithm based on gradient compensation is as follows:
[0101]
[0102] where ζ i (k) is an intermediate variable in the iteration, and λ i is the gradient compensation factor;
[0103] The convergence criterion of the improved consistency algorithm is as follows, where ε is a real number less than 0.01,
[0104] |h i (k+1)-h i (k)|<ε (43)
[0105] The algorithm implementation process is as follows: first, initialize the network structure information, construct the system communication topology adjacency matrix A=[a ij ], set the gradient compensation factor and the convergence threshold, and let the iteration number k=0; then, information interaction is carried out, the Agent i samples the local state information and exchanges state information with the adjacent Agent j ; then, the gradient value of the variable in the iteration process is determined; the improved consensus algorithm is executed, and the result is output; the distance of the gradient descent of the Agent i is determined, and the convergence is judged; when the difference between the current iteration result and the last iteration result is within the specified value ε, the variable reaches convergence, and the calculation is stopped.
[0106] Further, in step four, the improved TD3 deep reinforcement learning adopts a stacked data sampling strategy, and the system state of the experience pool is composed of continuous M intermediate states, which are voltage and current information at each sampling step.
[0107] Further, in step four, the improved TD3 deep reinforcement learning includes two evaluation Q value networks, two target Q value networks, one evaluation Actor network and one target Actor network, and the average value of the output values of the two Q value networks is used as the target value approximation, and its expression is:
[0108]
[0109] In the formula, s k+1 is the state in the k+1 sampling transition; d k is the termination signal; r k is the immediate reward; and u′ k+1 is the action output after adding Gaussian noise to the target Ator, and its expression is:
[0110] u′ k =clip((μ′ θ (s k )+ε′),u min ,u max ) (45)
[0111] In the formula, ε'~clip(N(0,σ),-l,l) is Gaussian noise; u min and u max represent the lower limit and upper limit of the action output, respectively;
[0112] The training of the double-evaluation Critic network is realized by minimizing the loss function, and its expression is:
[0113]
[0114] In the formula, B is the batch data of empirical playback buffer samples, the kth transition is denoted as B k = (s k , u k , r k , d k , s k+1 ) ; L (φ i ) is a loss function representing the difference between the predicted value of the Critic network on the sampling batch and the target value; the Critic network parameter φ i is updated by calculating the gradient of L (φ i ) with respect to φ i ;
[0115] The gradient descent method is used to update the parameters of the Actor to maximize the performance index, and the expression for estimating the gradient of the Actor parameters using samples is:
[0116]
[0117] For the parameter update of the target network, the expression of the TD3 algorithm using the soft update strategy is:
[0118] φ' i ← (1-τ) φ i + τ φ' i , i = 1, 2 (48)
[0119] θ' <- (1-τ) θ + τ θ' (49).
[0120] The application also provides a wind-solar-hydrogen storage micro-grid regulation system combined with reinforcement learning and consistency theory, which realizes the process of the method according to any one of the above through a modular program, and regulates the wind-solar-hydrogen storage micro-grid, and the system comprises:
[0121] A model construction module: a multi-physical field coupling model of a wind-solar-hydrogen storage micro-grid is established, which covers: an electrochemical process model: including a voltage model and a power model of an alkaline electrolytic cell; a thermodynamic process model: involving electrolytic cell heat balance, equipment working temperature dynamic characteristics; a mass transfer process model: including hydrogen generation, transmission and hydrogen storage tank hydrogen storage pressure model; the micro-grid comprises wind power generation equipment, photovoltaic power generation equipment, an electrolytic cell, a hydrogen storage tank, a fuel cell and an energy storage device;
[0122] An optimization target and constraint module: an optimization regulation framework of a wind-solar-hydrogen storage micro-grid is designed, including: a double-objective function: taking the minimization of micro-grid operation cost, including equipment energy consumption, maintenance cost and hydrogen production maximization, as the core optimization target; a safety constraint system: covering power balance constraints, electrolytic cell working power / temperature / climbing constraints, fuel cell output power / climbing constraints, hydrogen storage tank hydrogen storage state upper and lower limit constraints;
[0123] The lower regulation module is used for power distribution and state balance: an improved consistency algorithm based on gradient compensation is adopted to realize dynamic power distribution of multiple groups of heterogeneous electrolytic cells and balance regulation of multiple hydrogen storage tank inventory states, and the convergence speed is accelerated by introducing a gradient compensation factor, and the iteration is terminated when the convergence criterion is met.
[0124] The upper regulation module is used for multi-device collaborative optimization: based on the improved TD3 deep reinforcement learning algorithm, a stacked data sampling strategy is constructed, the experience pool contains voltage and current information of continuous M intermediate states, a double Q value network and a target network soft update mechanism are used to dynamically cooperate electrolytic cells, hydrogen storage tanks, fuel cells and energy storage devices, and a globally optimized operation strategy is output through learning iteration.
[0125] The hierarchical regulation architecture module integrates the above modules to form a hierarchical regulation system, wherein the upper regulation module is responsible for multi-device collaborative optimization decision, the lower regulation module is based on the upper decision to execute power distribution and hydrogen storage state balance control, and the global optimization goal is realized.
[0126] The application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method according to any one of the above methods when executing the computer program.
[0127] The application also provides a computer readable storage medium for storing computer instructions, wherein the computer instructions are executed by a processor to implement the steps of the method according to any one of the above methods.
[0128] The application has the following advantages and beneficial effects:
[0129] 1. Accurate simulation of system dynamic characteristics
[0130] The application constructs a wind-solar-hydrogen storage micro-grid coupling model covering electrochemistry, thermodynamics and mass transfer process, comprehensively considers the dynamic characteristics and multiple safety operation boundary constraints (such as power, temperature, hydrogen storage state, etc.) of electrolytic cells, hydrogen storage tanks, fuel cells and other devices, and can more accurately simulate the dynamic response of the real system, providing a reliable model basis for the regulation strategy.
[0131] 2. Efficiently cope with system complexity and uncertainty
[0132] The improved TD3 deep reinforcement learning algorithm is adopted to effectively solve the nonlinearity, non-convexity and uncertainty problems (such as wind and light output fluctuation, device heterogeneity) of the wind-solar-hydrogen storage micro-grid. Compared with the traditional optimization method which is easy to fall into local optimum, the algorithm improves the global optimization ability and stability of the regulation strategy through double Critic network mean estimation, soft update strategy and other improvements, and guarantees the efficient operation of the system in the dynamic environment.
[0133] 3. Improve training efficiency and optimize accuracy
[0134] An innovative hierarchical control architecture combining reinforcement learning and consensus theory is adopted: the upper layer achieves multi-device collaborative optimization through deep reinforcement learning, while the lower layer dynamically allocates power in heterogeneous electrolyzers and balances the state of hydrogen storage tanks based on an improved consensus algorithm, compressing the high-dimensional action space and significantly improving training efficiency. This architecture makes the control results closer to the theoretical optimal solution, balancing the multiple objectives of minimizing microgrid operating costs and maximizing hydrogen production.
[0135] 4. Enhance system flexibility and economy
[0136] By fully exploring the potential of the green hydrogen system in terms of "wide operating conditions and continuous and rapid adjustment", and through dynamic coordination of hydrogen production, hydrogen storage, fuel cells and energy storage equipment, the system can improve the new energy consumption rate, ensure the demand for green hydrogen production, reduce operating costs, enhance the reliability and economy of electrolytic hydrogen production projects, and provide intelligent control technology support for large-scale new energy green electricity hydrogen production projects. Attached Figure Description
[0137] Figure 1 This is a flowchart of the method of the present invention;
[0138] Figure 2 To improve the consensus algorithm flowchart;
[0139] Figure 3 To improve the TD3 algorithm structure diagram. Detailed Implementation
[0140] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0141] Example 1
[0142] like Figure 1 As shown, a combined reinforcement learning and consensus theory approach for the control of wind-solar-hydrogen-storage microgrids is presented, with the following specific implementation steps:
[0143] Step 1: Establish a detailed model of the wind-solar-hydrogen-storage microgrid, covering electrochemical, thermodynamic and mass transfer processes;
[0144] Step 2: With the optimization objectives of minimizing microgrid operating costs and maximizing hydrogen production, and considering multiple constraints such as equipment safety operation boundaries, design an optimization control model for the wind-solar-hydrogen-storage microgrid.
[0145] Step 3: Design a theoretical algorithm. Through an improved consistency algorithm based on gradient compensation, dynamically allocate the power of multiple heterogeneous electrolyzers to achieve balanced control of the storage status of multiple hydrogen storage tanks.
[0146] Step four, using deep reinforcement learning based on improved TD3, dynamically coordinating hydrogen production and hydrogen storage and other types of equipment, through learning iteration, finally obtaining the optimization result of microgrid operation.
[0147] Specifically:
[0148] In step one, a detailed model of the wind-solar-hydrogen storage microgrid is built, considering its electrochemical, thermodynamic and mass transfer processes. The wind power generation, photovoltaic power generation and electrochemical energy storage are not described in detail in this embodiment, and only the models of hydrogen production and hydrogen storage are listed.
[0149] The following is the mathematical model of the electrochemical process.
[0150] The voltage of the alkaline electrolyzer stack is the sum of the reversible voltage, ohmic loss, activation loss and diffusion loss, and its expression is:
[0151] V el =N el V cell =N el (V rev +V ohm +V act +V diff ) (1)
[0152] In the formula, N el is the number of series of alkaline electrolyzers in the stack; V cell is the voltage of a single electrolyzer; V rev is the reversible voltage; V ohm is the ohmic loss; V act is the activation loss; and V diff is the diffusion loss.
[0153] The reversible voltage can be modeled as a voltage source that depends on the operating temperature and pressure. The reversible voltage V rev is equal to 1.229V under standard conditions, but this value decreases with increasing temperature and increases slightly with increasing pressure. From the Nernst equation, we have
[0154]
[0155] In the formula, z is the number of electrons transferred by each hydrogen molecule, which is 2; F is the Faraday constant; R is the universal gas constant; P is the working pressure; T is the working temperature, ℃; is the reversible voltage as a function of temperature under standard pressure; P v,KOH is the vapor pressure of the electrolytic solution; is the water activity in the electrolyte;
[0156] The ohmic loss V ohmis the ohmic loss related to the resistance of the charge (electrons and ions) flow through the cell. This value increases linearly with the current I through the cell, following Ohm's law as shown in the following equation
[0157]
[0158] where A is the electrode surface area of the cell, r1 and r2 are empirical parameters, r2 simulates the increase of the resistance value with temperature, and I is the cell current.
[0159] Activation loss V act is caused by the redox reactions occurring at the anode and cathode. The formation of hydrogen and oxygen requires additional energy, thus creating an activation loss. The relationship between the current in the electrode and the activation overvoltage is approximated by the following equation for the entire current process in the case of high activation voltage:
[0160] V act = V act,a + V act,c (4)
[0161]
[0162] where s1, s2, s3, t1, t2, t3, v1, v2, v3, w1, w2, w3 are empirical parameters that represent the dependence of the current of the anode and cathode current sources on the activation overvoltage, indicating that the activation loss is related to temperature.
[0163] Diffusion loss V diff Since the working current density of the alkaline electrolyzer is relatively low, its impact can be ignored. Therefore, V diff is not included in the present model.
[0164] The power consumption of the alkaline electrolyzer is therefore:
[0165]
[0166] where I i and V el,i are the working current and working voltage of the i-th electrolyzer stack, respectively;
[0167] The following is a mathematical model of the thermodynamic process.
[0168] The temperature difference between the main components of the electrolyzer is usually less than 10°C, so the entire electrolyzer is considered as a single heat body, and the overall heat balance can be expressed as:
[0169]
[0170] where the left term describes the change in the temperature of the electrolyzer over time, which depends on the total heat capacity Ct .
[0171] is the internally generated heat, which is the portion of the electrical power input to the cell that is lost as heat. When the electrical power input to the cell exceeds the thermodynamic energy requirement, heat is generated, which can be expressed as:
[0172]
[0173] V tn = V Δh + V irrev (10)
[0174]
[0175] where V tn is the heat-neutral voltage, V Δh is the enthalpy voltage calculated from the Faraday's law of electrolysis, and in the ideal case, these two values are equal. However, in the actual electrolysis process, due to the irreversibility of the thermodynamic process, there is a V irrev , where y is the number of moles of water vapor produced per 1 mole of hydrogen gas electrolyzed; is the molar enthalpy of water vapor produced by the electrolyte at the working temperature and pressure; is the molar enthalpy of liquid water under standard conditions. V Δh , is an empirical parameter that can be obtained from experiments.
[0176] Considering only the heat transfer process between the overall surface of the electrolysis cell and the environment, it can be expressed as:
[0177]
[0178] where h is the convective-radiative heat transfer coefficient, A stack is the outer surface of the electrolysis cell chamber stack, A sep is the total outer surface of the two gas separators, T amb is the ambient temperature.
[0179] is the heat removed by the cooling water, which can be expressed as:
[0180]
[0181] where C cw is the heat capacity of the cooling water; T cw,i and T cw,o are the inlet and outlet temperatures of the cooling water, respectively,
[0182] Generally, the inlet temperature of the cooling water is known and constant, and the outlet temperature can be obtained from the following equation:
[0183]
[0184] UA HX = h cond + h conv I (16)
[0185] where UA HX is the heat transfer coefficient of the heat exchanger obtained by an empirical formula; h cond and h conv represent the conduction heat transfer coefficient and the convection heat transfer coefficient, respectively; and I is the current of the electrolyzer.
[0186] This value is much smaller than the other values and is ignored because it is related to the loss of substances (hydrogen and oxygen) leaving the system and water added, etc.
[0187] The following is a mathematical model of the mass transfer process, with hydrogen as the main substance.
[0188] According to Faraday's law, the molar flow rate of hydrogen generated at the electrode is related to the activation current of the cathode:
[0189]
[0190] where η F represents the Faraday efficiency, defined as the ratio of the actual gas flow rate to the theoretical flow rate, caused by parasitic current loss; I act,c and I act,a are the activation currents of the cathode and anode, respectively.
[0191] The separator of the electrolyzer is usually impermeable to gas, but due to the concentration gradient between the cathode and the anode, the dissolved substances will still diffuse into each other, and the molar flow rate through the separator is:
[0192]
[0193] where D is the molecular diffusion coefficient of hydrogen; d, ε, and τ are the thickness, porosity, and tortuosity of the separator, respectively; A cell is the surface area of the electrolyzer cell; and C are the molar concentrations of hydrogen at the cathode and anode, respectively.
[0194] Generally, the electrolyzer will install a device between the separators to balance the OH - charge consumed / generated during the electrochemical reaction process; the flow rate of hydrogen transferred to the opposite separator through the mixing tube can be represented as a fraction of the net flow rate into its original separator:
[0195]
[0196] where the empirical expression of λ and μ can be defined according to current, temperature and pressure as:
[0197]
[0198] Finally, the available hydrogen flow rate at the electrolyzer outlet is:
[0199]
[0200] After the hydrogen is produced by the electrolyzer, it enters the hydrogen storage tank for storage. The mathematical model of the hydrogen storage pressure of the hydrogen storage tank is:
[0201]
[0202] where n sto is the hydrogen storage amount of the hydrogen storage tank; n sto (t0) is the hydrogen storage amount of the hydrogen storage tank at t0; is the hydrogen flow rate from the hydrogen storage tank; p sto is the pressure of the hydrogen storage tank; T sto is the temperature of the hydrogen storage tank, in K; and V sto is the volume of the hydrogen storage tank.
[0203] In step two, the control variable vector u and the state variable vector s of the wind-solar-hydrogen storage micro-grid model are represented as:
[0204]
[0205] where I AEL,m T is the current reference value of the m alkaline electrolyzers, P FC is the fuel cell power reference value; P ES , P new T , P AEL,m T are the current power values of the battery energy storage, the wind-solar field and the conventional electrical load, and the m alkaline electrolyzers, respectively; T AEL,m T are the internal temperatures of the m electrolyzers; SOH m T are the internal hydrogen storage states of the m hydrogen storage tanks; are the hydrogen production rates of the electrolyzers and the hydrogen flow rates from the hydrogen storage tanks for fuel cell power generation, respectively.
[0206] The regulation of the microgrid system is established as a multi-objective optimization problem, which has two main objectives: one is to minimize the operation cost of the microgrid, and the other is to maximize the hydrogen production.
[0207] The first optimization objective is achieved by minimizing the following function:
[0208]
[0209] where γ and τ are the discount factor and the soft update rate; C1, C2 and C3 are the cost coefficients of the electrolyzer, the fuel cell and the energy storage, respectively.
[0210] The second optimization objective is achieved by maximizing the following function:
[0211]
[0212] Accordingly, the corresponding objective function is defined as:
[0213] R = α1r1 + α2r2 (27)
[0214] where α1 and α2 are the weight coefficients set for the two optimization objectives.
[0215] The regulation problem of the wind-solar-hydrogen-storage microgrid system includes the following device safety operation boundary constraints: the power balance constraint of the microgrid, the working power constraint, the working temperature constraint and the climbing power constraint of the electrolyzer, the output power constraint and the climbing power constraint of the fuel cell, the state of charge constraint, the charging and discharging power constraint and the climbing constraint of the battery energy storage, and the hydrogen storage state constraint of the hydrogen storage tank.
[0216] The wind-solar-hydrogen-storage microgrid tracks the wind-solar power fluctuations in the microgrid through the electrolyzer, the fuel cell assists, and the other power errors are finally balanced by the energy storage, so there is a power balance constraint:
[0217] P ES = P AEL -P new -P FC (28)
[0218] In the real-time control stage, considering that the device is always in working condition and the safety production requirement, the electrolyzer related states have working power constraint, working temperature constraint and climbing power constraint:
[0219]
[0220] where and respectively represent the minimum and maximum values of the normal working power of the electrolyzer i; and respectively represent the minimum and maximum values of the temperature inside the electrolyzer; R AEL,i represents the upper limit of the ramp rate of electrolyzer i.
[0221] The fuel cell is similar to the electrolyzer, considering that the equipment is always in operation, the fuel cell related state has output power constraint and ramp power constraint to meet:
[0222]
[0223]
[0224] where, and respectively represent the minimum and maximum values of the normal working power of the fuel cell; R FC represents the upper limit of the ramp rate of electrolyzer i.
[0225] In order to ensure the safety of production, the hydrogen storage tank mainly contains the hydrogen storage state constraint to meet:
[0226] SOH min <SOH i <SOH max (34)
[0227] where, SOH min and SOH max respectively represent the minimum and maximum values of the safe hydrogen storage state of the hydrogen storage tank.
[0228] In order to deal with the non-convexity and uncertainty of the wind-solar-hydrogen storage microgrid regulation process, the controller of the system in this embodiment is designed as a deep reinforcement learning agent. The agent responds to the random environment by sequentially selecting actions in discrete time steps. Accordingly, the optimization problem is constructed as a Markov decision process with state space S and action space U. At each time step t, the deep reinforcement learning agent observes the current state s t , selects a control action u t , and after the action is executed in the system, the system state will transfer to s t+1 ∈S. Here, the action u t corresponds to the control input vector u of the control system at time t.
[0229] In order to solve this optimization problem, a deterministic policy is needed, denoted as μ θ : S→U, for the deterministic policy, the total discounted reward from time t is:
[0230]
[0231] The following formula is the discrete form of the objective function value function value in formula (35), and the state value function V μequal to its action-value function Q μ :
[0232] V μ (s t ) = π θ (u t | s t ) Q μ (s t , μ θ (s t )) = Q μ (s t , μ θ (s t )) (36)
[0233] where the policy π satisfies π θ (u t | s t ) = 1.
[0234] For the regulation of the wind-solar-hydrogen microgrid, the objective is to solve an optimal control law u(t) = μ θ (s t ) to maximize the cumulative discounted reward.
[0235] Since the electrolyzer allows short-time operation in a certain overrun interval, the constraint of is a soft constraint, which can be embedded in the reward function. When the electrolyzer operating power exceeds , there is
[0236]
[0237] Similarly, the hydrogen storage tank still has a buffer margin outside the calibrated safe hydrogen storage range. In order to enhance the flexibility of system control, the constraint is embedded in the reward function. When the hydrogen storage state is less than SOH min or greater than SOH max , the corresponding penalties are triggered respectively:
[0238]
[0239] Finally, the hybrid reward function with multiple constraint embeddings is constructed as:
[0240] V μ (s t ) = α1r1+ α2r2+ α3r3+ α4r4+ α5r5 (40)
[0241] The key advantage of this function design is to convert the discrete constraint conditions in traditional optimization problems into continuous differentiable reward signals. This method not only effectively alleviates the discontinuity problem of the feasible region and reduces the complexity of potential constraint conflicts, but more importantly, it can guide the deep reinforcement learning agent to explore the policy space more efficiently in the early stage of training, reduce the probability of falling into invalid regions or causing safety hazards, and improve the system's regulation margin to adapt to dynamic changes and uncertainties in operation.
[0242] In step three, consistency refers to the process of gradually converging the specific states or behaviors of intelligent agents in complex systems over time through coordinated control, which is one of the core fundamental problems in multi-agent system research. Using this mechanism to make key state variables globally consistent has become an effective method to coordinate the operation of distributed systems.
[0243] For the hydrogen storage state balancing problem of m heterogeneous hydrogen storage tanks in a wind-solar-hydrogen storage microgrid, let x i (k) and x i (k+1) be the state variables of electrolyzer i at the kth and k+1th iterations, respectively, and x i ∈R represents a state variable of electrolyzer i. The exchange process of state information between adjacent stacks can be regarded as a linear system, which can be calculated by traditional consistency algorithms such as equation (41). By observing and calculating the difference in hydrogen storage state between adjacent hydrogen storage tanks, the output value of the coordinating electrolyzer is calculated, and finally the state balancing of the hydrogen storage tank is achieved.
[0244]
[0245] where h i (k) and h j (k) are the differences between hydrogen storage tanks i and j at the kth iteration; a ij is the correlation coefficient between hydrogen storage tanks i and j, if they are related, then 0 < a ij ≤ 1, otherwise a ij = 0; γ is the convergence coefficient.
[0246] To consider the communication problem and speed up the average consistency convergence, the gradient of the function at a certain point is the direction in which the function value changes the fastest, and the function value changes the fastest when moving in the negative gradient direction of the function at that point. Therefore, the gradient term of the state variable is introduced in the fixed consistency algorithm to obtain the improved consistency algorithm based on gradient compensation, as follows:
[0247]
[0248] where ζ i (k) is an intermediate variable in the iteration, and λi is the gradient compensation factor.
[0249] The convergence criterion of the consensus algorithm is as follows: where ε is a real number less than 0.01.
[0250] |h i (k+1)-h i (k)|<ε (43)
[0251] The specific implementation process of the algorithm is shown in Figure 2 First, the network structure information is initialized, the system communication topology adjacency matrix A = [a ij ] is constructed, the gradient compensation factor and the convergence threshold are set, and the iteration number k = 0 is set; then the information interaction is carried out, the Agent i samples the local state information and exchanges state information with the adjacent Agent j ; then the gradient value of the variable in this iteration process is determined; the improved consensus algorithm is executed, and the result is output; the distance of the Agent i gradient descent is determined, and the convergence is judged. When the difference between the current iteration result and the last iteration result is within the specified value ε, the variable reaches convergence, and the calculation is stopped.
[0252] In step four, as mentioned earlier, it is quite difficult to design a control law for such a nonlinear complex system relying on traditional optimization methods. Although deep reinforcement learning algorithms such as DQN, DDPG, and TD3 can be used for such problems, there are essential differences in their applicability: DQN, as a value-based algorithm, is suitable for low-dimensional discrete action spaces and is not suitable for the control requirements of the continuous action space of the electrolytic cell; DDPG, as an Actor-Critic type algorithm, can handle continuous action spaces, and its model-free nature has advantages for nonlinear uncertain systems. However, in practical applications, it is found that in fast dynamic control systems with input constraints, convergence problems easily occur: the frequent touching of the action boundary by the control command leads to the attenuation of the neural network gradient, and the overestimation bias of the Q function further limits its performance.
[0253] As an optimization extension of DDPG, the TD3 algorithm can effectively overcome the above defects.
[0254] Therefore, an optimal control method based on an improved TD3 (Double Delay Deep Deterministic Policy Gradient) algorithm is proposed, as shown in Figure 3 This algorithm has significant advantages in dealing with systems with uncertainty and nonlinearity.
[0255] In view of the dynamic characteristics of the system, the algorithm uses a stacked data sampling strategy. By collecting the voltage and current information at each sampling step, intermediate states are formed, which are finally stored in the system state s t composed of consecutive M intermediate states: In order to guarantee the dynamic feature capture, the memory depth M needs to be set reasonably by balancing data storage and processing capacity.
[0256] TD3 uses double Critic neural networks to approximate Q function values and applies the Actor-Critic learning framework. To stabilize the training process, both the evaluation Actor network and the evaluation Critic network are equipped with corresponding copy networks, i.e., target networks. In summary, TD3 utilizes six neural networks: two evaluation Q value networks (parameters φ1, φ2); two target Q value networks (parameters φ'1, φ'2); an evaluation Actor network (parameters θ); a target Actor
[0257] network (parameters θ').
[0258] The traditional TD3 method uses the minimum value in the output value of the double target Critic network as the approximation of the Critic network target value. This method is prone to cause large underestimation bias and affect the approximation accuracy. Therefore, the average value of the output values of the two Q value networks is used as the approximation to achieve more accurate and stable estimation:
[0259]
[0260] In the formula, s k+1 is the state in the k+1 sampling transition; d k is the termination signal; r k is the immediate reward; u′ k+1 is the action output after adding Gaussian noise to the target Ator:
[0261] u′ k = clip((μ′ θ (s k )+ε′),u min ,u max ) (45)
[0262] In the formula, ε' ~ clip(N(0, σ), -1, 1) is Gaussian noise; u min and u max represent the lower limit and upper limit of the action output, respectively.
[0263] The training of the double evaluation Critic network is realized by minimizing the loss function:
[0264]
[0265] In the formula, B is the batch data sampled from the experience replay buffer, and the kth transition is denoted as B k =(s k ,u k ,rk ,d k ,s k+1 );L(φ i The loss function characterizes the difference between the predicted and target values of the Critic network on a sampling batch; the Critic network parameter φ i i = 1, 2 By calculating L(φ) i For φ i The gradient is updated.
[0266] The parameters of the Actor can be updated using gradient descent to maximize performance metrics. The gradient for estimating the Actor's parameters using samples is as follows:
[0267]
[0268] For parameter updates of the target network, the TD3 algorithm adopts a soft update strategy, as shown below:
[0269] φ′ i ←(1-τ)φ i +τφ′ i ,i=1,2 (48)
[0270] θ′←(1-τ)θ+τθ′ (49)
[0271] This update strategy significantly slows down the update rate of the target network parameters, effectively ensuring the stability of the learning process.
Claims
1. A method for regulating a wind-solar-hydrogen-storage microgrid using a combination of reinforcement learning and consistency theory, characterized in that, Includes the following steps: Step 1: Establish a model of the wind-solar-hydrogen-storage microgrid. The model includes an electrochemical process model, a thermodynamic process model, and a mass transfer process model. The wind-solar-hydrogen-storage microgrid includes wind power generation equipment, photovoltaic power generation equipment, electrolyzers, hydrogen storage tanks, fuel cells, and energy storage equipment. The electrochemical process model includes the following: the voltage of the alkaline electrolytic cell stack is equal to the sum of the reversible voltage, ohmic loss, activation loss, and diffusion loss, and its expression is: (1) In the formula, This represents the number of alkaline electrolytic cells connected in series within the stack. The voltage of a single electrolytic cell; It is a reversible voltage; For ohmic loss; This is due to activation loss; For diffusion loss; Reversible voltage From the Nernst equation: (2) In the formula, The number of electrons transferred for each hydrogen molecule is set to 2; It is Faraday's constant; It is the universal gas constant; Due to work pressure; Operating temperature, unit: °C; This is the reversible voltage that varies with temperature under standard pressure; This is the vapor pressure of the electrolyte solution; The water activity in the electrolyte; Ohmic loss As shown in the following formula: (3) In the formula, This refers to the surface area of the electrodes in the electrolytic cell. and These are empirical parameters. The resistance value increased with temperature was simulated. This refers to the current in the electrolytic cell; Activation loss This is caused by the redox reactions occurring at the anode and cathode. The approximate formula for the relationship between the current in the electrode and the activation overvoltage, covering the entire current process under high activation voltage conditions, is as follows: (4) (5) (6) In the formula, , , , , , , , , , , , These are all empirical parameters, reflecting the dependence between the current of the anode and cathode current sources and the activation overvoltage; diffusion loss Since the operating current density of the alkaline electrolyzer is relatively low, its influence is negligible and therefore it is not included in this model. Therefore, the total power consumption of the alkaline electrolytic cell is: (7) In the formula, and The first Operating current and operating voltage of the electrolytic cell stack; The aforementioned mathematical model of the thermodynamic process includes: The overall heat balance expression for the electrolytic cell is: (8) In the formula, the term on the left describes the change in electrolytic cell temperature over time, which depends on the total heat capacity. ; The heat generated internally is the portion of the electrical power input to the electrolytic cell that is lost as heat. Heat is generated when the electrical energy input to the electrolytic cell exceeds the thermodynamic energy requirement. Its expression is: (9) (10) (11) (12) In the formula, It is the thermal neutral voltage; This refers to the enthalpy voltage calculated using Faraday's law of electrolysis; however, in actual electrolysis processes, there exists... ,in, This refers to the number of moles of water vapor produced per 1 mol of hydrogen electrolysis. The molar enthalpy of water vapor generated by the electrolyte at the operating temperature and pressure; The molar enthalpy of liquid water under standard conditions; , , Empirical parameters are obtained through experiments; Considering only the heat transfer process between the overall surface of the electrolytic cell and the environment, its expression is: (13) In the formula, The convective-radiative heat transfer coefficient; It is the outer surface of the stacked electrolytic cell chambers; It is the total outer surface of the two gas separators; It is the ambient temperature; The heat removed by the cooling water is expressed as: (14) In the formula, The heat capacity of cooling water; and Here, the inlet and outlet temperatures of the cooling water are respectively: The inlet temperature of the cooling water is known and remains constant, while the outlet temperature is calculated using the following formula: (15) (16) In the formula, The heat transfer coefficient of the heat exchanger is obtained through an empirical formula; and These represent the conductive heat transfer coefficient and the convective heat transfer coefficient, respectively. This value is related to the loss of substances leaving the system and water addition, and is much smaller than other values, so it is ignored. The mathematical models of the mass transfer process include: According to Faraday's law, the molar flow rate of hydrogen gas generated at the electrode is related to the activation current of the cathode: (17) In the formula, The Faraday efficiency is defined as the ratio of the actual gas flow rate to the theoretical flow rate, caused by parasitic current losses. and These are the activation currents for the cathode and anode, respectively; The dissolved substances in the separator of the electrolytic cell undergo interdiffusion. The expression for the molar flow rate through the separator is: (18) In the formula, is the molecular diffusion coefficient of hydrogen; , and These are the separator's thickness, porosity, and tortuosity, respectively. This refers to the surface area of the electrolytic cell chamber. , These are the molar concentrations of hydrogen at the cathode and anode, respectively. Equipment is installed between the separators within the electrolytic cell to balance the OH groups consumed / generated during the electrochemical reaction. - The charge, transferred through the mixing tube to the hydrogen gas flowing into the opposite separator, is expressed as a portion of the net flow rate entering its original separator, and its expression is: (19) Among them, it is defined according to current, temperature and pressure. and The empirical expression is: (20) Finally, the expression for the available hydrogen flow rate at the electrolyzer outlet is obtained as follows: (21) After hydrogen is produced by the electrolyzer, it enters a hydrogen storage tank for storage. The expression for the hydrogen storage pressure in the storage tank is: (22) (23) In the formula, This refers to the hydrogen storage capacity of the hydrogen storage tank. Hydrogen storage tank Hydrogen storage capacity at any given time; The flow rate used for hydrogen gas flowing out of the hydrogen storage tank; This refers to the pressure of the hydrogen storage tank. Temperature of the hydrogen storage tank, unit: K; This refers to the volume of the hydrogen storage tank; Step 2: With the optimization objectives of minimizing microgrid operating costs and maximizing hydrogen production, and considering the boundary constraints of equipment safe operation, design an optimization control model for the wind-solar-hydrogen-storage microgrid. The constraints include power balance constraints, electrolyzer operating power constraints, operating temperature constraints, ramp-up power constraints, fuel cell output power constraints, ramp-up power constraints, and hydrogen storage state constraints of the hydrogen storage tank. Among them, the control variable vector of the wind-solar-hydrogen-storage microgrid model and state variable vector They are represented as follows: (24) in, Here are the reference current values for m alkaline electrolytic cells. This is a reference value for fuel cell power. , , These represent the current power values of battery energy storage, several wind and solar power stations and conventional electrical loads, and m alkaline electrolyzers, respectively. Let m be the internal temperatures of the electrolytic cells; The hydrogen storage status inside m hydrogen storage tanks; , These represent the hydrogen production rate of the electrolyzer and the rate at which hydrogen flows out of the hydrogen storage tank for fuel cell power generation, respectively. The objective function for minimizing the operating cost of the microgrid satisfies: (25) in, and It is the discount factor and the soft update rate. , and These are the cost coefficients for electrolyzers, fuel cells, and energy storage, respectively. The objective function for maximizing hydrogen production satisfies: (26) The defined objective function is obtained as follows: (27) In the formula, and These are the weighting coefficients that are set by the user for the two optimization objectives; The microgrid tracks wind and solar power fluctuations via electrolyzers, with fuel cells providing auxiliary regulation. Remaining power errors are balanced by energy storage, and the power balance constraints are satisfied as follows: (28) During the real-time control phase, the electrolytic cell's relevant states are subject to constraints including operating power, operating temperature, and ramp-up power. (29) (30) (31) In the formula, and These represent electrolytic cells. Minimum and maximum normal operating power; and These represent the minimum and maximum temperatures inside the electrolytic cell, respectively. Indicates electrolytic cell The upper limit of the ramp rate; the fuel cell related states have output power constraints and ramp power constraints that must be satisfied: (32) (33) In the formula, and These represent the minimum and maximum values of the fuel cell's normal operating power, respectively. Indicates electrolytic cell The upper limit of the climbing rate; The hydrogen storage tank contains hydrogen storage state constraints that satisfy: (34) In the formula, and These represent the minimum and maximum values of the safe hydrogen storage state of the hydrogen storage tank, respectively; In this model, the controller of the optimized control model for the wind-solar-hydrogen-storage microgrid is designed as a learning agent. It responds to the stochastic environment by sequentially selecting actions in discrete time steps, constructing a Markov decision process with a state space and an action space . At each time step... The deep reinforcement learning agent observes the current state. Select a control action After this action is executed in the system, the system state will transition to... Here, action Corresponding to the control system at time control input vector To solve this optimization problem, a deterministic strategy is required, denoted as . For deterministic strategies, from a time perspective The expression for the total discount reward is: (35) The following equation is the discrete form of the objective function value in equation (35), and its state value function is... Equal to its action value function : (36) Among them, strategy satisfy ; For the regulation of wind-solar-hydrogen-storage microgrids, the goal is to find an optimal control law. To maximize cumulative discount rewards; Since the electrolyzer is allowed to operate for short periods within a specific out-of-limit range, the constraint is a soft constraint, which is embedded in the reward function when the electrolyzer's operating power exceeds... The expression for time is: (37) The hydrogen storage tank still has a buffer margin outside the calibrated safe hydrogen storage range. This constraint is embedded in the reward function; when the hydrogen storage state is less than... or greater than At that time, the corresponding penalty is triggered, and its expression is as follows: (38) (39) Finally, the expression for constructing the hybrid reward function with multiple constraint embeddings is: (40) Step 3: By using an improved consensus algorithm based on gradient compensation, the power of multiple heterogeneous electrolyzers is dynamically allocated to achieve balanced control of the stock status of multiple hydrogen storage tanks. Step 4: Employ a deep reinforcement learning algorithm based on an improved TD3 to dynamically coordinate the electrolyzer, hydrogen storage tank, fuel cell, and energy storage equipment, and obtain the optimized operation results of the microgrid through learning iterations; Steps three and four constitute a hierarchical control architecture. Step four is the upper-level control, which realizes multi-device collaborative optimization, while step three is the lower-level control, which realizes power distribution and hydrogen storage state balance.
2. In the wind-solar-hydrogen-storage microgrid control method based on reinforcement learning and consistency theory as described in claim 1, in step three, let... , Electrolytic cells In the iteration process Time and the The state variable at time t, Indicates electrolytic cell The state variable is considered as a linear system, and the exchange of state information between adjacent stacks is treated as a linear system. The consensus algorithm shown in equation (41) calculates the output value of the coordinated electrolyzer by observing and calculating the difference in hydrogen storage state between adjacent hydrogen storage tanks, and finally achieves the state equilibrium of the hydrogen storage tanks. (41) In the formula, and Hydrogen storage tank and In the The difference in time; Hydrogen storage tank and The correlation coefficient between them; if there is a correlation between them, then ,otherwise ; The convergence coefficient; The expression for the improved consensus algorithm based on gradient compensation is as follows: (42) In the formula, For intermediate variables in the iteration, This is the gradient compensation factor; The convergence criterion of the improved consensus algorithm is as follows: For real numbers less than 0.01, (43) The specific implementation process of the algorithm is as follows: First, initialize the network structure information and construct the system communication topology adjacency matrix. Set the gradient compensation factor and convergence threshold, and let the number of iterations... Next, information exchange takes place. Sample local state information and compare it with neighboring Exchange state information; then determine the gradient values of the variables in this iteration; execute the improved consensus algorithm and output the results; determine... The gradient descent distance is used to determine convergence; if the difference between the current iteration result and the previous iteration result is within a specified value... When the time is up, the variables converge and the calculation stops.
3. The method for regulating a wind-solar-hydrogen-storage microgrid based on reinforcement learning and consistency theory as described in claim 2, characterized in that, In step four, the improved TD3 deep reinforcement learning adopts a stacked data sampling strategy. The system state of the experience pool consists of M consecutive intermediate states, and the intermediate states are voltage and current information of each sampling step.
4. The wind-solar-hydrogen-storage microgrid control method based on reinforcement learning and consistency theory as described in claim 3, characterized in that, In step four, the improved TD3 deep reinforcement learning includes two evaluation Q-value networks, two target Q-value networks, one evaluation Actor network, and one target Actor network. The average of the output values of the two Q-value networks is used as an approximation of the target value, and its expression is: (44) In the formula, This refers to the state during the (k+1)th sample transition; This is a termination signal; For instant rewards; The expression for the action output of the target Ator after adding Gaussian noise is: (45) In the formula, It is Gaussian noise; and These represent the lower and upper limits of the action output, respectively; The training of the dual-evaluation Critic network is achieved by minimizing the loss function, which is expressed as: (46) In the formula, The batch data sampled for the experience playback buffer is denoted as k-th transition. ; The loss function characterizes the difference between the predicted and target values of the Critic network on a sampling batch; Critic network parameters Through calculation for Update the gradient; The parameters of the Actor are updated using gradient descent to maximize the performance metric. The expression for estimating the gradient of the Actor's parameters using samples is as follows: (47) For parameter updates of the target network, the expression for the soft update strategy used by the TD3 algorithm is as follows: (48) (49)。 5. A wind-solar-hydrogen-storage microgrid control system combining reinforcement learning and consistency theory, characterized in that, The system implements the process of the method described in any one of claims 1-4 through a modular program to regulate a wind-solar-hydrogen-storage microgrid, the system comprising: Model building module: Establishes a multi-physics coupling model for the wind-solar-hydrogen-storage microgrid, covering: electrochemical process model: including voltage and power models of the alkaline electrolyzer; thermodynamic process model: involving the thermal balance of the electrolyzer and the dynamic characteristics of the equipment operating temperature; mass transfer process model: including hydrogen generation, transport, and hydrogen storage pressure models of the hydrogen storage tank; the microgrid includes wind power generation equipment, photovoltaic power generation equipment, electrolyzer, hydrogen storage tank, fuel cell, and energy storage equipment; Optimization Objectives and Constraints Module: Design an optimization and control framework for a wind-solar-hydrogen-storage microgrid, including: Dual objective functions: Minimizing the operating cost of the microgrid, including equipment energy consumption, maintenance costs, and maximizing hydrogen production as the core optimization objectives; Safety constraint system: Covering power balance constraints, electrolyzer operating power / temperature / ramp constraints, fuel cell output power / ramp constraints, and upper and lower limits of hydrogen storage state constraints for hydrogen storage tanks; The lower-level control module is used for power allocation and state equilibrium: it adopts an improved consensus algorithm based on gradient compensation to achieve: dynamic power allocation of multiple heterogeneous electrolyzers; balanced control of the stock state of multiple hydrogen storage tanks, and accelerates the convergence speed by introducing a gradient compensation factor, and terminates the iteration when the convergence criterion is met. The upper-level control module optimizes multi-device collaboration by constructing a stacked data sampling strategy based on the improved TD3 deep reinforcement learning algorithm. The experience pool contains voltage and current information from M consecutive intermediate states. A soft update mechanism for the dual-Q value network and the target network dynamically coordinates the electrolyzer, hydrogen storage tank, fuel cell, and energy storage equipment, and outputs a globally optimized operating strategy through learning iteration. Hierarchical control architecture module: The above modules are integrated to form a hierarchical control system. The upper-level control module is responsible for multi-device collaborative optimization decision-making, and the lower-level control module executes power allocation and hydrogen storage state balance control based on the upper-level decision to achieve the global optimization goal.
6. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1-4.
7. A computer-readable storage medium for storing computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1-4.
Citation Information
Patent Citations
Virtual power plant double-layer optimization method considering refined demand response and electrolytic hydrogen production
CN116307193A
Optimized dispatching method for electricity-hydrogen-heat comprehensive energy microgrid based on improved TD3 algorithm
CN119787352A