A robust and secure deep reinforcement learning-based risk scheduling method for electricity-hydrogen coupled systems
By constructing a robust and secure deep reinforcement learning framework, the modeling problems of Faraday efficiency variation and renewable energy uncertainty in the electric-hydrogen coupling system are solved, realizing efficient and safe scheduling of the electric-hydrogen coupling system and improving the system's reliability and solution efficiency.
Patent Information
- Application Number
- CN202411636803.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2044-11-15
AI Technical Summary
Existing risk scheduling models for electric-hydrogen coupling systems suffer from problems such as coarse modeling, difficulty in solving problems, and insufficient uncertainty handling capabilities when dealing with Faraday efficiency variations, complex energy balances, and uncertainties in renewable energy. This leads to deviations between scheduling results and actual operation, making it difficult to meet high safety and complexity requirements.
A robust safety deep reinforcement learning framework is constructed, including a dynamic efficiency model for electrolyzers and a risk scheduling model considering wind power uncertainties. Combined with a flexible actor-evaluator baseline algorithm, the risk scheduling strategy of the electro-hydrogen coupling system is optimized through robustly constrained Markov decision processes.
It improves the convergence efficiency and constraint satisfaction of the scheduling model for the electro-hydrogen coupling system, enhances the ability to handle uncertainties, ensures the reliability and safety of system operation, and improves the solution efficiency and the rationality of the strategy.
Smart Images

Figure CN119863051B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of risk scheduling for electro-hydrogen coupling systems, specifically a risk scheduling method for electro-hydrogen coupling systems based on robust and secure deep reinforcement learning. Background Technology
[0002] To alleviate environmental pollution and energy depletion, transitioning to a power energy structure dominated by high-penetration clean energy is particularly urgent. This also exacerbates the risks to grid operation posed by the highly random nature of new energy sources such as wind and solar power. As an emerging energy carrier, hydrogen possesses advantages such as being green and carbon-free, having high energy density, and being highly flexible, which can mitigate the randomness of new energy sources and reduce operational risks. Among these, risk scheduling of the electro-hydrogen coupling system, as a key technology supporting the reliable development and overall integration and synergy of electro-hydrogen energy, delves into the two dimensions of operational limit exceedance probability and severity from the perspective of a comprehensive management system, aiming to fully explore the application potential of hydrogen energy in ensuring the sustainable and safe advancement of renewable energy plans.
[0003] In electro-hydrogen hybrid systems, the electrolyzer is one of the core devices for achieving electro-hydrogen coupling. From a risk-based dispatching perspective, reasonable modeling of the electrolyzer is the cornerstone and key to accurate quantification and rigorous management of electro-hydrogen hybrid energy systems. Existing research establishes physical models of the electrolyzer based on the electrochemical reaction process of water electrolysis. For example, some literature considers the fixed efficiency operation of the electrolyzer in the day-ahead economic dispatching model of the integrated energy system, neglecting the changes in electrolyzer efficiency during actual operation, leading to significant deviations in practical applications. Other literature establishes operating models based on the polarization relationship within the electrolyzer, considering the nonlinear relationship between hydrogen production and power consumption, to analyze the impact of electrolyzer characteristics on system operating economy. However, the establishment of the above electrolyzer models all ignore the impact of Faraday efficiency changes on electrolyzer operating efficiency, resulting in a relatively coarse modeling process that is difficult to accurately describe the dynamic efficiency operating characteristics of the electrolyzer. Therefore, in actual power grids, dispatching behavior may deviate significantly from actual operating conditions, leading to unreasonable results.
[0004] A safe and reliable clean energy solution based on risk-based scheduling of electro-hydrogen coupling systems faces the following significant challenges: a) The complex heterogeneous energy balance relationships and dynamic physical characteristics of the electrolyzers in electro-hydrogen coupling systems mathematically classify the operation model of these systems as a nonlinear, non-convex, high-dimensional mixed-integer programming problem, especially when considering Faraday efficiency variations. b) The risk of branch overloads is typically modeled as an integral term of the probability and severity of branch overloads, stemming from the uncertainty of renewable energy distributed through the affine characteristics of power flow. This presents extremely difficult challenges to existing solution methods. c) Furthermore, handling the volatility of renewable energy is a challenging task. The prediction error of new energy sources has been shown to follow a non-Gaussian distribution, indicating that forward-looking scheduling, from an optimization perspective, places high demands on the perception and adaptability to complex uncertain environments.
[0005] Existing work primarily focuses on scenario-based methods to address the risk scheduling problem in complex electro-hydrogen coupling systems. Scenario-based methods attempt to simplify the model through linearization and convexity assumptions, while utilizing techniques such as scenario discretization to handle risk terms arising from complex uncertainties. However, in practical applications, the effectiveness of balancing the number of scenarios with uncertainty perception capabilities under non-Gaussian distributions of uncertainty has not been verified, suggesting that increasing the number of scenarios may be more urgent. Therefore, although this approach sacrifices model accuracy to some extent, it does not significantly improve solution efficiency.
[0006] Recent exploratory research has focused on the deployment and application of deep reinforcement learning in the operation of electro-hydrogen coupling systems, due to its superior performance in handling uncertainty, nonlinear dynamics optimization problems, and online decision-making. However, electro-hydrogen coupling systems are constrained by the physical limitations of energy equipment and the availability of resources, and their high safety and complexity requirements significantly increase the demands and difficulties for deep reinforcement learning in meeting these constraints. Therefore, designing a deep reinforcement learning safety paradigm that can satisfy these constraints is crucial to ensuring the reliable operation of electro-hydrogen coupling systems. Existing deep reinforcement learning safety paradigms mainly fall into three categories:
[0007] 1) Incorporating constraint violation signals as penalty augmentations into the reward function. For example, existing literature has proposed energy optimization management methods for electric-hydrogen coupling systems based on DQN. Other literature has proposed scheduling strategies for hybrid energy storage electric-hydrogen coupling systems based on SAC and DDPG, respectively. However, designing the penalty factor is quite challenging, especially when dealing with the risk scheduling problem of electric-hydrogen coupling systems, as the penalty factor involves a large number of constraint terms with different units. In this case, an extremely complex and time-consuming combinatorial parameter tuning process is required, leading to serious challenges to training performance and stability.
[0008] 2) Physical safety layer correction. Safety layer correction methods, through the design of optimization models or projection algorithms, make minimal adjustments to the output actions to ensure the system operates within the safety domain. Some literature has proposed a deep reinforcement learning algorithm for optimizing the operation of an electro-hydrogen coupling system by integrating an adaptive safety module, effectively ensuring the satisfaction of constraints on hydrogen storage state and voltage amplitude. However, the optimization correction process of the safety layer significantly increases the computational burden of training. Furthermore, this correction process, independent of the training environment, hinders the agent's learning of the problem itself and may lead to the generation of suboptimal strategies.
[0009] 3) Introducing additional constraint training mechanisms. Pioneering work involves guiding agents to conduct safety surveys through the establishment of Constrained Markov Processes (C-SAC), PD-DDPG, and IPO. C-SAC treats constraint breaches as the safety survey cost for the agent, uses the expected value of this cost to characterize constraint satisfaction, and performs additional training with the goal of limiting the constraint to a cost threshold. This provides a suitable approach for solving the operation of electro-hydrogen coupling systems. However, C-SAC relies solely on the expected cost as a safety indicator, meaning that this method may have limitations in the presence of unexpected operational risks, especially when uncertainties are introduced and the state space is excessively large.
[0010] To address these challenges, it is necessary to explore alternative deep reinforcement learning methods that incorporate more robust constraint learning mechanisms that can tailor strategies to the different risk requirements of high-safety-demand electro-hydrogen coupling systems. Summary of the Invention
[0011] The purpose of this invention is to provide a risk scheduling method for an electro-hydrogen coupling system based on robust and secure deep reinforcement learning, comprising the following steps:
[0012] 1) Considering the Faraday efficiency loss effect of the electrolyzer, a dynamic operation model for hybrid electro-hydrogen energy storage is constructed;
[0013] 2) Taking into account the uncertainties of wind power, a risk scheduling model for the electric-hydrogen coupling system that takes into account the dynamic operation of electric-hydrogen hybrid energy storage is established;
[0014] 3) Construct a robust constrained Markov decision-making architecture;
[0015] 4) The robust constrained Markov decision architecture is parameterized using the flexible actor-evaluator baseline algorithm;
[0016] 5) Solve the risk scheduling model of the electric-hydrogen coupling system using the parameterized robust constrained Markov decision process architecture to obtain the risk scheduling scheme of the electric-hydrogen coupling system.
[0017] Furthermore, the dynamic operation model of the hybrid electric-hydrogen energy storage is shown below:
[0018]
[0019] In the formula, HHV H The highest calorific value of hydrogen under standard conditions; η e,t To improve the operating efficiency of hybrid electric-hydrogen energy storage; The power consumption and hydrogen production of the hybrid energy storage electrolyzer for hydrogen;
[0020] Among them, the power consumption of the electrolytic cell and hydrogen production As shown below:
[0021] P t e =N E V cell,t I cell,t (2)
[0022]
[0023] In the formula, I represents the hydrogen production rate of the electrolyzer at time t. cell,t =A cell i cell,t N is the operating current of the electrolytic cell. E The number of electrolytic cells in series is related to the maximum capacity of the electrolytic cells; k1 represents the conversion factor; η f,t For Faraday efficiency;
[0024] Operating voltage V of hybrid electric-hydrogen energy storage cell,t As shown below:
[0025]
[0026] In the formula, t is the scheduling interval; V rev The voltage is the reversible voltage; T' is the operating temperature; R and F are the rational gas constant and Faraday constant, respectively. and These are the partial pressures of hydrogen and oxygen, respectively. The activity of H2O; R cell For internal equivalent ohmic resistance; α an and α cat These are the charge transfer coefficients for the anode and cathode, respectively; i an and i cat These are the anode and cathode exchange current densities, respectively; i cell,t Current density;
[0027] Faraday efficiency η f,t As shown below:
[0028]
[0029] In the formula, Oxygen permeability in the diffusion mechanism; Hydrogen permeability in the diffusion mechanism; These represent the partial pressures of hydrogen at the cathode and oxygen at the anode in the catalyst layer, respectively; A H2 A O2 These are the fitting parameters for the increase in hydrogen and oxygen partial pressure in the cathode catalyst layer, respectively; Hydrogen permeability caused by pressure difference;
[0030] Furthermore, the risk scheduling model for the electro-hydrogen coupling system is shown below:
[0031]
[0032] sth(x,y,ζ)=0 (9)
[0033] Ax + Dy + Mζ + e ≤ 0 (10)
[0034] g(x,y,ζ)≤0 (11)
[0035] u(x,y,ζ)≤0 (12)
[0036]
[0037]
[0038] In the formula, the superscript t represents the scheduling interval; T is the scheduling period; vector x represents the decision variables including the scheduling strategy; vector y represents other state variables in HEES; vector ζ represents the uncertain prediction error; A, D and M are coefficient matrices, and e is the coefficient vector; subscript l represents the l-th branch; l represents the branch set; the symbol '-' represents the upper limit of each variable; (·)+ represents max{·,0}; Represents the degree of branch overload; d B Let Pr{x} = 1 - F(x) represent the probability of x, and F() be the cumulative distribution function; l,t Indicates branch current flow; d B Indicates the risk threshold; f t (x) represents the operating cost function; h(x,y,ζ)=0 represents the AC power flow equation and energy balance equation.
[0039] Furthermore, robust constrained Markov decision-making framework This includes action networks, reward evaluation networks, safety evaluation networks, Lagrange multipliers, and goal networks; Represents the state space. Represents the action space. Let r be the state transition function, representing the state transition in state s. t Next, execute action a t The reward obtained, c represents the reward in state s. t Next, execute action a t The cost of safe exploration obtained, where d is the threshold for exceeding the limit and γ is the discount factor; γ∈[0,1];
[0040] The action network encapsulates the action policy distribution π with neural network parameters θ, establishing a mapping from the state space to the action space.
[0041] The reward evaluation network is used to evaluate the operational status of HEES. t Lowering the control decision a t Reward soft Q function (s t ,a t );
[0042] The security assessment network is used to evaluate the operational status of HEES. t Lowering the control decision a t Security assessment function Γ π (s t ,a t ,β);
[0043] Lagrange multipliers are used to adjust the dual multipliers related to entropy and excess cost weights during the dual update process;
[0044] The target network is used to enhance the training stability of the reward value discrimination network and the cost value discrimination network modules.
[0045] Furthermore, the state space As shown below:
[0046]
[0047] In the formula, P L,t m H2,t P W,t , P G,t-1 , E B,t-1 These represent the electric load prediction vector, hydrogen load prediction vector, new energy prediction value, the first derivative of the new energy prediction value with respect to time, the thermal power output vector at the previous moment, the hydrogen storage capacity of the hydrogen storage tank at the previous moment, and the state of charge of the electrical energy storage at the previous moment, respectively; the subscripts L, W, G, H, and B represent the load, wind farm, thermal power unit, hydrogen storage tank, and battery set, respectively.
[0048] The first derivative of the predicted new energy value with respect to time is shown below:
[0049]
[0050] In the formula, Δt is the time difference; P W,t-1 This represents the predicted value of new energy sources at time t-1;
[0051] The action space As shown below:
[0052]
[0053] ΔP G,t =a G,t Θ[-T 60 D G ,T 60 U G (18)
[0054]
[0055] rd G,t =a D,t Θ[0,min(P G,t - P G ,T 10 D G (20)
[0056]
[0057] In the formula, the subscripts E and FC represent the sets of electrolyzers, fuel cells, and energy storage devices, respectively. The subscript '_' indicates the lower limit of the corresponding variable; These represent the charging / discharging power of the energy storage, which are combined into a single decision P. B,t To reduce the action space; U G With D G These represent the ramp rate and descent rate of the generator unit, respectively. T 60 With T 10 These refer to time intervals of one hour and ten minutes, respectively; a G,t a U,t a D,t a E,t a F,t a B,t a H,t These respectively represent the power output increment action of thermal power units, the upper rotation standby action of thermal power units, the lower rotation standby action, the operating current density action of electrolyzers, the power generation action of fuel cells, the charging / discharging power action of energy storage, and the hydrogen output action of hydrogen storage tanks; G,t rd G,t These represent the reserve capacity for upward and downward rotation, respectively; i cell,E,tP represents the current density limiting the operation of the electrolyzer; B,t This indicates the power of the electrical energy storage.
[0058] Furthermore, the state transition function of the robust constrained Markov decision architecture is shown below:
[0059] P G,t =P G,t-1 +ΔP G,t (25)
[0060]
[0061] In the formula, η represents the hydrogen inlet vector of the hydrogen storage tank; H,in η H,out Efficiency vectors for hydrogen input and output from the hydrogen storage tank, respectively; η cha,B η dis,B Efficiency vectors for charging and discharging of electrical energy storage, respectively.
[0062] Furthermore, the cost function c π,ρ (s t ,a t ,ζ) are shown below:
[0063] c π,ρ (s t ,a t ,ζ)=[ε e c e,t +ε b c b,t +ε pf c pf,t +ε pb c pb,t +ε ph c ph,t ]|π,ρ (31)
[0064]
[0065]
[0066] In the formula, π represents the distribution of the agent's regulatory decision-making strategy; ε e ε b ε pf ε pb With ε ph These represent scaling respectively; c e,t This indicates that the capacity safety constraint of the hydrogen storage tank exceeds the limit; c b,t This indicates that the safety constraint on energy storage capacity has exceeded the limit; c pf,t This indicates that the branch overload risk constraint exceeds the limit; c pb,t Indicates the power imbalance in HEES; c ph,tP represents the hydrogen energy imbalance quantity; e,t Let t represent the power consumption of the e-th PEM electrolytic cell at time t.
[0067] The reward function is as follows:
[0068] r t =-(f t +f C,t )=-(f G,t +f Q,t +f L,t +f HESS,t +f C,t (37)
[0069]
[0070] f HESS,t =λ h (f HS,om,t +f HS,loss,t +f B,om,t +f B,loss,t (41)
[0071] f C,t =λ c (c π,ρ (s t ,a t ,ζ)-d) + (42)
[0072]
[0073] In the formula, f G,t Indicates the operating cost of thermal power plants; f Q,t f L,t These represent the penalties for abandoning wind and solar power and for load shedding, respectively; f HESS,t Indicates the operating cost of an electric-hydrogen hybrid energy storage system; f C,t Indicates the penalty cost for exceeding the limit; a g b g c g u g d g λ represents the unit cost of thermal power unit constants, primary, secondary, upward rotational reserve, and downward rotational reserve; g , λ q , λ l , λ h , λ c ΔP represents the cost coefficient. i,t (ζ) represents the power imbalance caused by uncertainty at the i-th node; f HS,om,t f BS,om,t These represent the operation and maintenance costs of hydrogen energy storage and electric energy storage, respectively; f HS,loss,t fBS,loss,t The energy consumption costs for hydrogen energy storage and electric energy storage are respectively; C e,inv C fc,inv These represent the investment and construction costs of the electrolyzer e and the fuel cell fc, respectively; T e,life T fc,life These represent the maximum operating times of the electrolyzer e and the fuel cell fc, respectively; λ BS,om λ represents the unit operation and maintenance cost coefficient for energy storage. ep This refers to the transmission and distribution price.
[0074] Furthermore, the security assessment function Γ π (s t ,a t ,β) are shown below:
[0075]
[0076] In the formula, φ(·) and Φ(·) represent the probability density function and cumulative probability distribution function of the standard normal distribution, respectively; F π,ρ Indicate C π,ρ probability density function
[0077] Among them, C π Second central moments As shown below:
[0078]
[0079] Furthermore, when parameterizing the robust constrained Markov decision architecture using the flexible actor-judge baseline algorithm, the objective function and constraints are as follows:
[0080]
[0081] In the formula, This represents the entropy threshold.
[0082] Furthermore, when parameterizing the robust constrained Markov decision architecture, the loss function of the security evaluation network... As shown below:
[0083]
[0084] In the formula, This indicates a replay of an experience.
[0085] The loss function of the reward evaluation network is shown below:
[0086]
[0087] The loss function J of the action network π (θ) is shown below:
[0088]
[0089] The update equations for the Lagrange multipliers are shown below:
[0090]
[0091] In the formula, α and κ are Lagrange multipliers; ω is the learning rate.
[0092] The technical effects of this invention are undeniable. Addressing the critical needs of low-carbon and green energy transformation and energy security, this invention proposes a risk scheduling method for electro-hydrogen coupling systems based on a robust flexible actuator-evaluator, achieving efficient and strict control of risks in electro-hydrogen coupling systems and ensuring the reliability and safety of system operation.
[0093] This invention proposes a risk scheduling model for an electro-hydrogen coupling system that considers the dynamic efficiency characteristics of the electrolyzer. This model takes into account Faraday efficiency losses, effectively reducing deviations in practical applications and managing the risk of branch overload caused by the propagation of renewable energy uncertainties through AC power flow, thereby ensuring the safe operation of the electro-hydrogen coupling system.
[0094] This invention addresses the risk scheduling problem in electro-hydrogen coupling systems by developing a robust constrained Markov decision process framework. In this framework, the risk of constraint violation is considered the safe exploration cost of the agent. By using the conditional risk value of this cost as a safety evaluation criterion, replacing the expected value used in traditional constrained Markov decision processes, this invention establishes a risk-avoidance-oriented constraint learning mechanism to improve the constraint satisfaction of risk scheduling in electro-hydrogen coupling systems. Furthermore, a safety evaluation function based on a second-order distributed Bellman operator is designed for efficient estimation of the conditional risk value.
[0095] This invention proposes a more robust and secure deep reinforcement learning algorithm, the Robust Flexible Actor-Evaluator. This algorithm uses a deep neural network to parameterize a robust constrained Markov decision process framework, enabling the agent to express decisions and learn policies in risk scheduling of an electro-hydrogen coupling system. This process incorporates entropy and cost constraints into the actor-evaluator framework. A primal-dual optimization method is introduced to address constrained optimization training, driving the robust constrained Markov decision process framework to adaptively learn towards maximum entropy. The proposed robust flexible actor-evaluator algorithm ensures global optimality and constraint satisfaction in the formulation of risk scheduling strategies for electro-hydrogen coupling systems and possesses online decision-making capabilities.
[0096] In terms of practical application, the robust flexible actor-evaluator algorithm framework proposed in this invention demonstrates significantly improved convergence performance and safe exploration capabilities. Compared to algorithms employing traditional constrained Markov decision process (CDM) frameworks and those based on CDM frameworks, the proposed robust constrained Markov decision process framework improves convergence efficiency by at least 45.03%, while rigorously ensuring the satisfaction of operational constraints. Compared to existing optimization-based methods, the proposed algorithm exhibits superior uncertainty handling capabilities. Compared to scenario-based methods, the robust flexible actor-evaluator algorithm possesses more comprehensive uncertainty perception and enhanced risk resistance, effectively addressing extreme operating conditions of the electro-hydrogen coupling system. Compared to scenario-based methods, the scheduling strategy for the electro-hydrogen coupling system formulated by the method proposed in this invention demonstrates a more reasonable trade-off between safety and economic efficiency. It not only reliably meets load demands but also makes fuller use of hybrid electric-hydrogen energy storage to effectively manage operational risks. Regarding the deployment time of the electric-hydrogen coupling system, the robust flexible actor-evaluator algorithm achieves a solution time of only 0.412 seconds in the IEEE-118 node system, improving solution efficiency by more than 15,907 times compared to other risk-based scheduling optimization methods. The algorithm proposed in this invention demonstrates the enormous potential for advancing the application and deployment of next-generation artificial intelligence technologies in promoting the safe operation of green energy systems. Attached Figure Description
[0097] Figure 1 It is a graph showing the operating efficiency of the electrolyzer and the hydrogen production curve;
[0098] Figure 2 This is a diagram of the electro-hydrogen coupling system architecture;
[0099] Figure 3 This is a framework diagram of the proposed robust constrained Markov decision process;
[0100] Figure 4 This is a schematic diagram of the proposed robust constrained Markov decision process and the traditional constrained Markov decision process during the constraint learning phase.
[0101] Figure 5 This is a schematic diagram of the proposed robust flexible actuator-evaluator algorithm;
[0102] Figure 6 This describes the training process of a robust, flexible actor-evaluator. It includes reward, cost, cost rate, action network loss function, reward evaluation network loss function, cost evaluation network loss function, and policy entropy.
[0103] Figure 7 It is a comparison chart of rewards and costs for various deep reinforcement learning training processes;
[0104] Figure 8This represents the load factor distribution of heavily loaded branches under uncertain conditions. Here, Q0 represents the lower critical value, Q1 represents the first quartile, Q2 is the median, Q3 is the third quartile, and Q4 is the upper critical value.
[0105] Figure 9 It is a scheduling strategy for the electro-hydrogen coupling system based on the proposed robust flexible actor-evaluator method and optimization method;
[0106] Figure 10 The figures show the changes in hydrogen storage capacity of the hydrogen storage tank, state of charge of the battery, and efficiency of the electrolyzer under the proposed robust flexible actuator-evaluator method and the scheduling strategy of the electro-hydrogen coupling system formulated based on the optimization method. Detailed Implementation
[0107] The present invention will be further described below with reference to embodiments, but it should not be construed that the scope of the present invention is limited to the following embodiments. Various substitutions and modifications made based on ordinary technical knowledge and common practices in the art without departing from the above-described technical concept of the present invention should be included within the scope of protection of the present invention.
[0108] Example 1:
[0109] A risk scheduling method for an electro-hydrogen coupling system based on robust safety deep reinforcement learning includes the following steps:
[0110] 1) Considering the Faraday efficiency loss effect of the electrolyzer, a dynamic operation model for hybrid electro-hydrogen energy storage is constructed;
[0111] 2) Taking into account the uncertainties of wind power, a risk scheduling model for the electric-hydrogen coupling system that takes into account the dynamic operation of electric-hydrogen hybrid energy storage is established;
[0112] 3) Construct a robust constrained Markov decision-making architecture;
[0113] 4) The robust constrained Markov decision architecture is parameterized using the flexible actor-evaluator baseline algorithm;
[0114] 5) Solve the risk scheduling model of the electric-hydrogen coupling system using the parameterized robust constrained Markov decision process architecture to obtain the risk scheduling scheme of the electric-hydrogen coupling system.
[0115] Example 2:
[0116] A risk scheduling method for an electro-hydrogen coupling system based on robust safety deep reinforcement learning is presented, with the same technical content as in Example 1. Further, the dynamic operation model for electro-hydrogen hybrid energy storage is shown below:
[0117]
[0118] In the formula, HHV HThe highest calorific value of hydrogen under standard conditions; η e,t To improve the operating efficiency of hybrid electric-hydrogen energy storage; The power consumption and hydrogen production of the hybrid energy storage electrolyzer for hydrogen;
[0119] Among them, the power consumption of the electrolytic cell and hydrogen production As shown below:
[0120]
[0121] In the formula, I represents the hydrogen production rate of the electrolyzer at time t. cell,t =A cell i cell,t N is the operating current of the electrolytic cell. E The number of electrolytic cells in series is related to the maximum capacity of the electrolytic cells; k1 represents the conversion factor; η f,t For Faraday efficiency;
[0122] Operating voltage V of hybrid electric-hydrogen energy storage cell,t As shown below:
[0123]
[0124] In the formula, t is the scheduling interval; V rev The voltage is the reversible voltage; T' is the operating temperature; R and F are the rational gas constant and Faraday constant, respectively. and These are the partial pressures of hydrogen and oxygen, respectively. The activity of H2O; R cell For internal equivalent ohmic resistance; α an and α cat These are the charge transfer coefficients for the anode and cathode, respectively; i an and i cat These are the anode and cathode exchange current densities, respectively; i cell,t Current density;
[0125] Faraday efficiency η f,t As shown below:
[0126]
[0127] In the formula, Oxygen permeability in the diffusion mechanism; Hydrogen permeability in the diffusion mechanism; These represent the partial pressures of hydrogen at the cathode and oxygen at the anode in the catalyst layer, respectively; A H2 A O2 These are the fitting parameters for the increase in hydrogen and oxygen partial pressure in the cathode catalyst layer, respectively; Hydrogen permeability caused by pressure difference;
[0128] Example 3:
[0129] A risk scheduling method for an electro-hydrogen coupling system based on robust safety deep reinforcement learning is provided. The technical content is the same as any one of Examples 1-2. Furthermore, the risk scheduling model for the electro-hydrogen coupling system is as follows:
[0130]
[0131] sth(x,y,ζ)=0 (9)
[0132] Ax + Dy + Mζ + e ≤ 0 (10)
[0133] g(x,y,ζ)≤0 (11)
[0134] u(x,y,ζ)≤0 (12)
[0135]
[0136] In the formula, the superscript t represents the scheduling interval; T is the scheduling period; vector x represents the decision variables including the scheduling strategy; vector y represents other state variables in HEES; vector ζ represents the uncertain prediction error; A, D and M are coefficient matrices, and e is the coefficient vector; subscript l represents the l-th branch; l represents the branch set; the symbol '-' represents the upper limit of each variable; (·)+ represents max{·,0}; Represents the degree of branch overload; d B Let Pr{x} = 1 - F(x) represent the probability of x, and F() be the cumulative distribution function; l,t Indicates branch current flow; d B Indicates the risk threshold; f t (x) represents the operating cost function; h(x,y,ζ)=0 represents the AC power flow equation and energy balance equation.
[0137] Example 4:
[0138] A risk scheduling method for an electro-hydrogen coupling system based on robust safety deep reinforcement learning is proposed, with the technical content being the same as any one of Examples 1-3, further comprising a robust constrained Markov decision architecture. This includes action networks, reward evaluation networks, safety evaluation networks, Lagrange multipliers, and goal networks; Represents the state space. Represents the action space. Let r be the state transition function, representing the state transition in state s. t Next, execute action a tThe reward obtained, c represents the reward in state s. t Next, execute action a t The cost of safe exploration obtained, where d is the threshold for exceeding the limit and γ is the discount factor; γ∈[0,1];
[0139] The action network encapsulates the action policy distribution π with neural network parameters θ, establishing a mapping from the state space to the action space.
[0140] The reward evaluation network is used to evaluate the operational status of HEES. t Lowering the control decision a t Reward soft Q function (s t ,a t );
[0141] The security assessment network is used to evaluate the operational status of HEES. t Lowering the control decision a t Security assessment function Γ π (s t ,a t ,β);
[0142] Lagrange multipliers are used to adjust the dual multipliers related to entropy and excess cost weights during the dual update process;
[0143] The target network is used to enhance the training stability of the reward value discrimination network and the cost value discrimination network modules.
[0144] Example 5:
[0145] A risk scheduling method for an electro-hydrogen coupling system based on robust and secure deep reinforcement learning, with the same technical content as any one of embodiments 1-4, further wherein the state space... As shown below:
[0146]
[0147] In the formula, P L,t m H2,t P W,t , P G,t-1 , E B,t-1 These represent the electric load prediction vector, hydrogen load prediction vector, new energy prediction value, the first derivative of the new energy prediction value with respect to time, the thermal power output vector at the previous moment, the hydrogen storage capacity of the hydrogen storage tank at the previous moment, and the state of charge of the electrical energy storage at the previous moment, respectively.
[0148] The first derivative of the predicted new energy value with respect to time is shown below:
[0149]
[0150] In the formula, the subscripts L, W, G, H, and B represent the load, wind farm, thermal power unit, hydrogen storage tank, and battery pack, respectively. Δt is the time difference; P W,t-1 This represents the predicted value of new energy sources at time t-1;
[0151] The action space As shown below:
[0152]
[0153] ΔP G,t =a G,t Θ[-T 60 D G ,T 60 U G (18)
[0154]
[0155] rd G,t =a D,t Θ[0,min(P G,t - P G ,T 10 D G (20)
[0156]
[0157] In the formula, the subscripts E and FC represent the sets of electrolyzers, fuel cells, and energy storage devices, respectively. The subscript '_' indicates the lower limit of the corresponding variable; These represent the charging / discharging power of the energy storage, which are combined into a single decision P. B,t To reduce the action space; U G With D G These represent the ramp rate and descent rate of the generator unit, respectively. T 60 With T 10 These refer to time intervals of one hour and ten minutes, respectively; a G,t a U,t a D,t a E,t a F,t a B,t a H,t These respectively represent the power output increment action of thermal power units, the upper rotation standby action of thermal power units, the lower rotation standby action, the operating current density action of electrolyzers, the power generation action of fuel cells, the charging / discharging power action of energy storage, and the hydrogen output action of hydrogen storage tanks; G,t rd G,tThese represent the reserve capacity for upward and downward rotation, respectively; i cell,E,t P represents the current density limiting the operation of the electrolyzer; B,t This indicates the power of the electrical energy storage.
[0158] Example 6:
[0159] A risk scheduling method for an electro-hydrogen coupling system based on robust safety deep reinforcement learning is provided. The technical content is the same as any one of Examples 1-5. Furthermore, the state transition function of the robust constrained Markov decision architecture is as follows:
[0160] P G,t =P G,t-1 +ΔP G,t (25)
[0161]
[0162] In the formula, η represents the hydrogen inlet vector of the hydrogen storage tank; H,in η H,out Efficiency vectors for hydrogen input and output from the hydrogen storage tank, respectively; η cha,B η dis,B Efficiency vectors for charging and discharging of electrical energy storage, respectively.
[0163] Example 7:
[0164] A risk scheduling method for an electro-hydrogen coupling system based on robust safety deep reinforcement learning, with the same technical content as any one of embodiments 1-6, further wherein the cost function c π,ρ (s t ,a t ,ζ) are shown below:
[0165] c π,ρ (s t ,a t ,ζ)=[ε e c e,t +ε b c b,t +ε pf c pf,t +ε pb c pb,t +ε ph c ph,t ]|π,ρ (31)
[0166]
[0167] In the formula, π represents the distribution of the agent's regulatory decision-making strategy; ε e ε b ε pf ε pb With εph These represent scaling respectively; c e,t This indicates that the capacity safety constraint of the hydrogen storage tank exceeds the limit; c b,t This indicates that the safety constraint on energy storage capacity has exceeded the limit; c pf,t This indicates that the branch overload risk constraint exceeds the limit; c pb,t Indicates the power imbalance in HEES; c ph,t P represents the hydrogen energy imbalance quantity; e,t This represents the power consumption of the ethPEM electrolyzer at time t.
[0168] The reward function is as follows:
[0169] r t =-(f t (+f C,t )=-(f G,t +f Q,t (+f L,t +f HESS,t +f C,t (37)
[0170]
[0171]
[0172] f HESS,t =λ h (f HS,om,t +f HS,loss,t +f B,om,t +f B,loss,t (41)
[0173] f C,t =λ c (c π,ρ (s t ,a t ,ζ)-d) + (42)
[0174]
[0175] In the formula, f G,t Indicates the operating cost of thermal power plants; f Q,t f L,t These represent the penalties for abandoning wind and solar power and for load shedding, respectively; f HESS,t Indicates the operating cost of an electric-hydrogen hybrid energy storage system; f C,t Indicates the penalty cost for exceeding the limit; a g b g c g u g d gλ represents the unit cost of thermal power unit constants, primary, secondary, upward rotational reserve, and downward rotational reserve; g , λ q , λ l , λ h , λ c ΔP represents the cost coefficient. i,t (ζ) represents the power imbalance caused by uncertainty at the i-th node; f HS,om,t f BS,om,t These represent the operation and maintenance costs of hydrogen energy storage and electric energy storage, respectively; f HS,loss,t f BS,loss,t The energy consumption costs for hydrogen energy storage and electric energy storage are respectively; C e,inv C fc,inv These represent the investment and construction costs of the electrolyzer e and the fuel cell fc, respectively; T e,life T fc,life These represent the maximum operating times of the electrolyzer e and the fuel cell fc, respectively; λ BS,om λ represents the unit operation and maintenance cost coefficient for energy storage. ep This refers to the transmission and distribution price.
[0176] Example 8:
[0177] A risk scheduling method for an electro-hydrogen coupling system based on robust safety deep reinforcement learning is proposed, with the technical content being the same as any one of Examples 1-7. Further, the safety evaluation function Γ... π (s t ,a t ,β) are shown below:
[0178]
[0179] In the formula, φ(·) and Φ(·) represent the probability density function and cumulative probability distribution function of the standard normal distribution, respectively; F π,ρ Indicate C π,ρ probability density function
[0180] Among them, C π Second central moments As shown below:
[0181]
[0182] Example 9:
[0183] A risk scheduling method for an electro-hydrogen coupling system based on robust safety deep reinforcement learning is provided. The technical content is the same as any one of Examples 1-8. Further, when parameterizing the robust constrained Markov decision architecture using the flexible actor-evaluator baseline algorithm, the objective function and constraints are as follows:
[0184]
[0185] In the formula, This represents the entropy threshold.
[0186] Example 10:
[0187] A risk scheduling method for an electro-hydrogen coupling system based on robust safety deep reinforcement learning is provided. The technical content is the same as any one of Examples 1-9. Furthermore, when parameterizing the robust constrained Markov decision architecture, the loss function of the safety evaluation network is as follows:
[0188]
[0189] In the formula, This indicates a replay of an experience.
[0190] The loss function of the reward evaluation network is shown below:
[0191]
[0192] The loss function for the action network is shown below:
[0193]
[0194] The update equations for the Lagrange multipliers are shown below:
[0195]
[0196] In the formula, α and κ are Lagrange multipliers; ω is the learning rate.
[0197] Example 11:
[0198] The verification of a risk scheduling method for an electro-hydrogen coupling system based on robust safety deep reinforcement learning is as follows:
[0199] This embodiment was executed on a modified IEEE-118 node test system, with four wind farms connected to nodes 4, 8, 15, and 40, respectively. The total installed wind power capacity is 6400MW, with a penetration rate of approximately 39%. A hybrid electric-hydrogen energy storage system operates in conjunction with the wind farms on the same nodes, with the hydrogen storage capacity set at 5-10% of the total installed capacity. The maximum electrical load and hydrogen load are 8860W and 11t, respectively. Historical renewable energy and load data are derived from actual grid data from a city in southern China. The permissible branch overload range is 0.1 times the rated capacity, and the confidence level β is set to 0.05. Other parameter details are shown in Table 1. Technical analysis and performance evaluation were performed using the TensorFlow framework in Python, with backend support including an Intel Core i9-10900K GPU, an NVIDIA GeForce RTX 3090 GPU, and a lightweight 8-core 16GB 18M server.
[0200] Table 1 Simulation Technical Parameters
[0201]
[0202] The specific embodiments of the present invention are as follows:
[0203] 1) Establish an electrolyzer operation model that considers Faraday efficiency loss.
[0204] The relationship between PEM operating voltage and current density is shown below:
[0205]
[0206] In the formula, t is the scheduling interval; V rev P is the reversible voltage; T' is the operating temperature; R and F are the rational gas constant and Faraday constant, respectively; H2 and P O2 These are the partial pressures of hydrogen and oxygen, respectively; a H2O The activity of H2O; R cell For internal equivalent ohmic resistance; α an and α cat These are the charge transfer coefficients for the anode and cathode, respectively; i an and i cat These are the anode and cathode exchange current densities, respectively; i cell,t denoted as current density.
[0207] Electrolytic cell power consumption and hydrogen production As shown below:
[0208] P t e =N E V cell,t Icell,t (2)
[0209]
[0210] In the formula, I represents the hydrogen production rate of the electrolyzer at time t. cell,t =A cell i cell,t N is the operating current of the electrolytic cell. E The number of electrolytic cells in series is related to the maximum capacity of the electrolytic cells; k1 represents the conversion factor; η f The Faraday efficiency is related to the generation of hydrogen and oxygen in the electrolyzer. Due to the permeability of the PME, hydrogen and oxygen diffusion occurs during mass transfer, especially at low current densities, where gas cross-contamination is more pronounced, leading to Faraday losses. The dynamic efficiency operating model of the PEM considering Faraday losses can be expressed as a function of membrane flux density and hydrogen production density, as shown below:
[0211]
[0212] In the formula, Γ H2 Hydrogen production density; The hydrogen permeation flux density; This represents the oxygen permeation flux density.
[0213] According to Faraday's law, the hydrogen production density is as follows:
[0214]
[0215] During PEM electrolysis, hydrogen diffusion across the membrane is caused by the thermal motion of molecules and the concentration difference between the catalyst layers at the anode and cathode. Therefore, the hydrogen permeation flux density is the sum of the diffusion flux density and the permeation flux density caused by the pressure difference. Based on Fick's law, the hydrogen permeation flux density function is as follows:
[0216]
[0217] In the formula, Hydrogen permeability in the diffusion mechanism; Hydrogen permeability caused by pressure difference; , respectively, are the cathode hydrogen partial pressure and the anode oxygen partial pressure in the catalyst layer, defined in (7)-(8); d is the film thickness.
[0218]
[0219] In the formula, A H2 A O2 These are fitting parameters for the increase in hydrogen and oxygen partial pressures in the cathode catalyst layer, respectively. These fitting parameters depend on the structure, thickness, and permeability of the catalyst layer.
[0220] Oxygen permeation in PEM mainly occurs through diffusion. According to Fick's law, the oxygen permeation flux density is as follows:
[0221]
[0222] In the formula, Oxygen permeability in the diffusion mechanism.
[0223] Based on equations (5)-(7), further derivation of (4) yields the dynamic Faraday efficiency as:
[0224]
[0225] To compare the effects of dynamic Faraday efficiency and fixed Faraday efficiency (0.9) on the operating characteristics of PEM electrolyzers. Specifically, considering the input, output, and energy conversion relationships of the PEM, the PEM operating efficiency η... e,t As shown below:
[0226]
[0227] In the formula, HHV H This represents the high calorific value of hydrogen under standard conditions.
[0228] 2) Establish a risk scheduling model for the electric-hydrogen hybrid energy storage system that takes into account the dynamic operation of the system.
[0229] The established HEES risk dispatch model (Risk-based dispatch for HEES, HEES-RD) is described in the compact form of formula (12).
[0230]
[0231] sth(x,y,ζ)=0 (12b)
[0232] Ax + Dy + Mζ + e ≤ 0 (12c)
[0233] g(x,y,ζ)≤0 (12d)
[0234] u(x,y,ζ)≤0 (12e)
[0235] In the formula, the superscript t represents the scheduling interval; T is the scheduling period; vector x represents the decision variables including the scheduling strategy; vector y represents other state variables in HEES; vector ζ represents the uncertain prediction error, which follows the distribution ρ; A, D and M are the coefficient matrices, and e is the coefficient vector.
[0236]
[0237] In the formula, the subscript l represents the l-th branch; l represents the set of branches; the symbol '-' represents the upper limit of each variable; (·)+ represents max{·,0}; d represents the degree of branch overload, which is a function of ζ; B Let Pr{x} = 1 - F(x) represent the probability of x, where F() is the cumulative distribution function; l,t This represents the branch power flow. Formula (13) defines the branch overload rate, which measures the severity of the overload; constraint (14) limits the expected value of the overload risk to a given risk threshold d. B the following.
[0238] 3) Establish a robust constrained Markov decision process for the risk scheduling model of the electro-hydrogen coupling system.
[0239] This invention establishes a Robust Constrained Markov Decision Process (R-CMDP) that includes the core elements of R-CMDP and a risk-averse training mechanism. For the HEES-RD problem, the core elements of R-CMDP are defined as follows: in Represents the state space. Represents the action space. Let r be the state transition function, where r represents the state transition function. Next action The reward obtained, c represents the reward in state s. t Next, execute action a t The obtained safe exploration cost is denoted by d, where d is the threshold for exceeding the limit, and γ (γ∈[0,1]) is the discount factor. The proposed R-CMDP adopts an iterative interactive mode of "state perception-decision making-posterior simulation-feedback learning" and combines it with a risk avoidance training mechanism to explore the optimal action strategy.
[0240] 3.1) The core elements of establishing R-CMDP are as follows:
[0241] state space
[0242] The state space specifically includes: the electrical load prediction vector P L,t Hydrogen load prediction vector m H2,t New energy forecast value P W,t The first derivative of the new energy forecast value with respect to time, P (1)W,t Used to capture the temporal changes of new energy sources; the output vector P of the thermal power unit at the previous moment. G,t-1 The amount of hydrogen stored in the hydrogen storage tank at the previous moment. The state of charge E of the stored energy at the previous moment B,t-1 .
[0243]
[0244] The first derivative of the predicted new energy value with respect to time is calculated as follows:
[0245]
[0246] In the formula, the subscripts L, W, G, H, and B represent the set of load, wind farms, units, hydrogen storage tanks, and batteries, respectively.
[0247] Action space
[0248] The action space includes: thermal power unit output increment action a G,t ; Thermal power unit up / down rotation standby operation a U,t a D,t Electrolytic cell operating current density action a E,t Fuel cell power generation operation a F,t ; Energy storage charging / discharging power operation a B,t Hydrogen storage tank hydrogen output action a H,t To project the action space onto the actual physical space that takes into account the operating limitations of the hydrogen power plant, a linear scaling operator Θ is introduced, as shown in equations (18)-(24). Equation (18) maps the increment ΔP of the thermal power unit. H,t This is to ensure that the ramp constraint is met. Equations (19) and (20) respectively map the up and down rotating reserve capacity ru G,t with rd G,t This is to meet the actual unit capacity limitations. Equation (21) reflects the current density i required to meet the operating limits of the electrolyzer. cell,E,t Equation (22) calculates the power generation P of the fuel cell. FC,t Equation (23) maps the power P of the electrical energy storage. B,t This is to meet the charging and discharging power constraints. Equation (24) maps the hydrogen output of the hydrogen storage tank. Hydrogen output limits.
[0249]
[0250] ΔP G,t =a G,t Θ[-T 60 D G ,T 60 U G (18)
[0251]
[0252] rd G,t =a D,t Θ[0,min(P G,t - P G ,T 10 D G (20)
[0253]
[0254] In the formula, the subscripts E and FC represent the sets of electrolyzers, fuel cells, and energy storage devices, respectively. The subscript '_' indicates the lower limit of the corresponding variable; These represent the charging / discharging power of the energy storage, which are combined into a single decision P. B,t To reduce the action space; U G With D G These represent the ramp rate and descent rate of the generator unit, respectively. T 60 With T 10 These refer to time intervals of one hour and ten minutes, respectively.
[0255] State transition function
[0256] The state transition functions are shown in equations (25)-(29). Equation (25) calculates the unit output; equation (26) represents the energy conversion relationship of the gas storage tank at adjacent times; equation (27) represents the charge state relationship of the electrical storage at adjacent times; and equations (28)-(29) represent the start-up and shutdown state relationships of the electrolyzer and fuel cell equipment, respectively.
[0257] P G,t =P G,t-1 +ΔP G,t (25)
[0258]
[0259] In the formula, The hydrogen inlet vector of the hydrogen storage tank, in hybrid electro-hydrogen energy storage, is equal to the hydrogen outlet of the electrolyzer, and therefore can be expressed by the electrolyzer current density i. cell,E,t η is calculated using equations (3) and (10). H,in η H,out Efficiency vectors for hydrogen input and output from the hydrogen storage tank, respectively; η cha,B η dis,B The efficiency vectors for charging and discharging of the energy storage are respectively; in addition, in order to ensure that the thermal power unit meets the capacity constraint, the thermal power output decision space is limited to Equation (30).
[0260]
[0261] Cost function
[0262] The cost function is defined as shown in equation (31).
[0263] c π,ρ (s t ,a t ,ζ)=[ε e c e,t +ε b c b,t +ε pf c pf,t +ε pb c pb,t +ε ph c ph,t ]|π,ρ (31)
[0264] In the formula, π represents the distribution of the agent's regulatory decision-making strategy; ε e ε b ε pf ε pb With ε ph Each represents scaling, therefore, it normalizes the cost item to the same dimension; c e,t The capacity safety constraint of the hydrogen storage tank is defined in equation (32); c b,t The limit of the safety constraint on energy storage capacity is defined in equation (33); c pf,t The overload risk constraint of the branch circuit exceeds the limit and is defined in equation (34); c pb,t The unbalanced electrical power of HEES is represented by equation (35); c ph,t The unbalanced amount of hydrogen energy is represented by equation (36).
[0265]
[0266] In the formula, P e,t This represents the power consumption of the PEM electrolyzer at time t, expressed as i ∈ [the current density of the electrolyzer]. cell,E,t Calculated using equations (1) and (2).
[0267] reward function
[0268] The reward function is defined as shown in equation (37).
[0269] r t =-(f t +f C,t )=-(f G,t +f Q,t +f L,t +f HESS,t +f C,t (37)
[0270] In the formula, f G,tThe operating costs of thermal power plants are detailed in (38); f Q,t f L,t These represent the penalties for curtailing wind and solar power and for load shedding, respectively. They are calculated under cybersecurity constraints, taking into account the response capability of thermal power units' spinning reserve to uncertain wind power, and are defined in (39)-(40); f HESS,t The operating cost of the electric-hydrogen hybrid energy storage is shown in equation (41); f C,t This represents the penalty cost for exceeding the constraint limit, used to improve the training efficiency of the agent, see equation (42) for details.
[0271]
[0272] f HESS,t =λ h (f HS,om,t +f HS,loss,t +f B,om,t +f B,loss,t (41)
[0273] f C,t =λ c (c π,ρ (s t ,a t ,ζ)-d) + (42)
[0274] In the formula, a g b g c g u g d g λ represents the unit cost of thermal power unit constants, primary, secondary, upward rotational reserve, and downward rotational reserve; g , λ q , λ l , λ h , λ c ΔP represents the cost coefficient. i,t (ζ) represents the power imbalance at the i-th node caused by uncertainty, as shown in equation (43). HS,om,t f BS,om,t These represent the operation and maintenance costs of hydrogen energy storage and electric energy storage, respectively; f HS,loss,t f BS,loss,t The energy consumption costs for hydrogen energy storage and electric energy storage are defined as shown in equation (44).
[0275]
[0276] In the formula, C e,inv C fc,inv These represent the investment and construction costs of the electrolyzer e and the fuel cell fc, respectively; T e,life T fc,lifeThese represent the maximum operating times of the electrolyzer e and the fuel cell fc, respectively; λ B,om λ represents the unit operation and maintenance cost coefficient for energy storage. ep This refers to the transmission and distribution price.
[0277] 3.2) The training mechanism of the proposed R-CMDP is as follows:
[0278] The R-CMDP proposed in this invention introduces a risk aversion constraint learning mechanism, no longer limited by constraint satisfaction and low safety exploration efficiency. R-CMDP replaces the expectation-based safety exploration index of CMDP with a risk control safety exploration index based on CVaR (Conditional Value at Risk), as defined in formula (45). Figure 4 As shown in (b). By adjusting the confidence level β, R-CMDP introduces an infeasible region within a controllable range, thereby achieving constraint overrun identification under risk control. R-CMDP aims to learn more robust cost targets to ensure stricter constraint satisfaction and to cope with unexpected operating conditions in the HEES-RD system. This mechanism, from a safety perspective, quantifies and limits the degree of risk avoidance for constraint overruns.
[0279]
[0280] 4) Establish a security evaluation function based on the second-order distributed Bellman operator.
[0281] The established security evaluation function based on the second-order distributed Bellman operator is shown below:
[0282]
[0283] In the formula, F π,ρ Indicate C π,ρ The probability density function of is often considered to follow a normal distribution in the exploration space of deep reinforcement learning, i.e. here, Indicate C π The second central moments are as follows:
[0284]
[0285] Distributed Bellman Operator Defined in (49). Based on this, Second-order distributed Bellman operator The result is derived in equation (50), which implements a temporal difference estimate, avoiding the burden of full trajectory sampling.
[0286]
[0287] Based on the above second-order distributed Bell operator, C is established. π,ρ (s t ,a t ,ζ) CVaR safety assessment function Γ with confidence level α π (s t ,a t ,β) are shown below:
[0288]
[0289] In the formula, φ(·) and Φ(·) represent the probability density function and cumulative probability distribution function of the standard normal distribution, respectively.
[0290] 5) A risk scheduling method for the electro-hydrogen coupling system of robust flexible actuator-evaluator is proposed.
[0291] This invention proposes a Robust Soft Actor-Critic (R-SAC) algorithm, which is built upon the SAC baseline neural network and parameterizes the HEES-RD R-CMDP architecture as described in Equation (51). This process integrates entropy constraints and over-limit cost constraints. To effectively solve this constrained optimization problem, a policy search method based on Primal-Dual Optimization (PDO) is introduced. R-SAC demonstrates efficient adaptive learning and powerful global exploration capabilities, ensuring global optimality and constraint satisfaction in the HEES-RD problem.
[0292]
[0293] In the formula, This represents the entropy threshold.
[0294] 5.1) Evaluate the training of the network
[0295] For the loss function of RCN, the reward soft Q function is first established. Distributed Bellman Operator As shown below:
[0296]
[0297] Furthermore, based on The loss function for establishing the reward value discrimination network is as follows:
[0298]
[0299] In the formula, This refers to the experience replay pool, which utilizes an asynchronous training strategy to ensure the full utilization of safe exploration experience, thereby improving training efficiency.
[0300] For the loss function of SCN, based on the distributed Bellman operator and The training loss function is established as shown in equations (54)-(55).
[0301]
[0302] 5.2) Training of Action Networks and Lagrange Multipliers
[0303] To find the optimal action policy, a policy exploration method (PDO) is introduced. This method transforms the constrained maximization problem (51) into a dual problem by constructing a Lagrangian function (as shown in Equation (56)). By iterating alternately between the original policy update and the dual Lagrangian multiplier update, this method drives the action network to perform maximum entropy adaptive learning.
[0304]
[0305] Based on the Lagrangian function above, the loss function of the action network is defined as follows:
[0306]
[0307] Furthermore, the original dual update process of the k-th generation is derived as follows:
[0308] For the update of the original strategy, the Lagrange multiplier α (k) and κ (k) The learning rate is fixed, and gradient descent is performed using the policy with a learning rate of ω, as shown below:
[0309]
[0310] For the improvement of dual Lagrange multipliers, the strategy The dual gradient ascent is fixed and performed as follows:
[0311]
[0312] 7) Verification of the effectiveness of the method proposed in this invention
[0313] The performance of the proposed R-SAC algorithm in solving the HEES-RD problem is demonstrated by comparing it with the following five advanced deep reinforcement learning algorithms and optimization methods, in terms of training performance, economy, security, and solution efficiency.
[0314] The network structure and parameters of the R-SAC algorithm proposed in this invention are detailed in Table 2;
[0315] The C-SAC algorithm based on the CMDP framework uses the cost soft Q function based solely on expectation as the safety criterion for learning constraints. Apart from this, the network structure and parameter settings of C-SAC are completely consistent with those of R-SAC.
[0316] The SAC algorithm based on the MDP framework optimizes the network structure and parameters by incorporating the penalty function into the reward function. It removes the SCN and LM modules, and its network structure and parameters are consistent with the R-SAC algorithm.
[0317] The scenario-based approach (SA) handles uncertainty by sampling scenarios and uses piecewise linearization and the Big M method to linearize the HEES-RD model.
[0318] Chance-constrained programming (CCP) models risk constraints from a probabilistic perspective, uses a sample average approximation method to address uncertainty, and sets the confidence level for constraint satisfaction to 0.95.
[0319] Table 2. R-SAC neural network architecture and parameters of the IEEE-118 test system
[0320]
[0321]
[0322] 7.1) Training Performance Analysis
[0323] To verify the convergence and secure exploration capability of the proposed R-SAC algorithm, key steps in the training process of this invention are demonstrated in... Figure 6 . Figure 7 The convergence performance of this algorithm in terms of reward and cost was compared with that of the advanced deep reinforcement learning algorithms C-SAC and SAC. The detailed training performance metrics are shown in Table 3.
[0324] Figure 6 The study presents the average reward and cost changes per iteration during training, clearly demonstrating the excellent stable convergence characteristics of the R-SAC algorithm. Figure 6 The graph depicts the cost rate metric throughout the training process. This metric, representing the average cost at each stage, visually demonstrates the efficiency of safety exploration. The rapid decrease in the cost rate indicates excellent performance in HEES-RD problem constraint learning. Figure 6 The loss function of the actor network is shown to change with the number of training epochs. Figure 6 The loss scenarios for the two reward evaluation networks were plotted separately. Figure 6 The loss of Q-value is demonstrated by evaluating the network around the cost value. Figure 6The focus is then on the loss due to the second-order center distance at the cost of exceeding the limit. It can be seen that the RCN and SCN networks converge around the 800th round, exhibiting extremely high convergence efficiency. Figure 6 The changes in entropy in R-SAC are shown, and the convergence of the entropy value indicates that the agent is exploring the globally optimal policy and has safety.
[0325] Table 3. Metrics of various deep reinforcement learning algorithms during the training process.
[0326]
[0327]
[0328] Figure 4 As shown in (a)-(b) and Table 3, compared to C-SAC and SAC, the R-SAC algorithm achieves the most rapid reward improvement, converging in only 3694 rounds, with a convergence efficiency improvement of over 45.03%. Furthermore, R-SAC outperforms other deep reinforcement learning algorithms in both final reward and cost convergence, demonstrating the optimal overall operating cost and lowest constraint violation risk in the HEES-RD problem. This superior performance is attributed to the proposed risk-averse training mechanism based on the CVaR safety criterion, which imposes stricter requirements on constraint learning, thereby achieving more effective safe exploration and risk avoidance performance during training. In contrast, C-SAC's expectation-based safety criterion leads to a conservative and insufficiently robust constraint learning approach. SAC, on the other hand, lacks an effective learning operation constraint mechanism, relying solely on avoiding high penalties, resulting in lower training efficiency and a tendency to get trapped in local optima. These results demonstrate that the proposed R-SAC algorithm exhibits superior convergence performance and safe exploration capabilities in the HEES-RD problem. This conclusion provides strong technical support for the optimized scheduling of electro-hydrogen coupling systems and broadens the scope for the application of deep reinforcement learning in complex systems.
[0329] 7.2) Economic Analysis
[0330] This section evaluates the economic efficiency of HEES-RD strategies developed using different methods by comparing their overall operating costs; detailed data is shown in Table 4. The proposed R-SAC outperforms C-SAC and SAC in terms of operating cost, primarily due to its superior constraint satisfaction capability, avoiding the substantial penalties for constraint violations. Although R-SAC's operating cost is 0.79% higher than SA, this is mainly due to its more efficient utilization of hybrid electric-hydrogen storage and reserve capacity to address uncertainty. Subsequent analysis will show that this slight increase in cost has a positive impact on the rationality of safety and dispatch strategies. Meanwhile, CCP incurs higher costs compared to R-SAC; in fact, this depends on the tail distribution of uncertainty. This indicates that, in this specific case, the tail probability of control constraint violations is higher than the risk of management violations.
[0331] Table 4. Running costs of various algorithms
[0332]
[0333] 7.3) Security Analysis
[0334] This section compares the resilience of various scheduling strategies under uncertainty. To quantify the operational risk of HEES, we conduct posterior tests on the scheduling strategies of each method under uncertainty scenarios. Furthermore, we introduce the overload probability δ. P Expected overload δ E and maximum overload level δ M The three risk assessment indicators are shown in formulas (61)-(63). Figure 8 The results of post-tests using strategies developed through different methods are presented. Details of the risk assessment metrics are shown in Table 5.
[0335] δ P =n v / (N s ×T)×100%(61)
[0336]
[0337] In the formula, n v N represents the number of times a branch exceeds its limit; s The number of random scenes.
[0338] Table 5 Risk Measurement for Each Algorithm
[0339]
[0340]
[0341] Depend on Figure 8It is evident that all methods exhibit load factors exceeding 1, indicating that high-penetration renewable energy sources increase the risk of branch overload exceeding limits. Compared to C-SAC and SAC, the overall load factor distribution range of R-SAC (Q0–Q4) is [0.28–1.16], significantly closer to the safety boundary. As shown in Table 5, R-SAC's risk assessment index is significantly lower, indicating that its scheduling strategy has higher security. This advantage stems from the proposed R-CMDP safety paradigm, which employs a more robust risk-averse constraint learning mechanism, enhancing constraint satisfaction and ensuring the safety and feasibility of the risk-based scheduling strategy.
[0342] The load factor distribution of SA is relatively compact, with an overall range of [0.48–1.23]. Although SA effectively handles uncertainty through multi-scenario expectations, its overload probability δ P The expected overload δ is 6.95%. E The maximum overload level δ is 0.082. M The values were 0.281, all higher than the corresponding indicators of R-SAC (5.08%, 0.064, and 0.158, respectively). This indicates that the proposed R-SAC method formulates a safer and more reliable scheduling strategy to cope with the uncertainties brought about by high-penetration renewable energy. This is because the limited number of scenarios in SA fails to fully capture the distribution information needed to hedge against uncertainties. In contrast, R-SAC interacts with the uncertain environment through a process of "state awareness-decision-posterior simulation-feedback learning," thereby comprehensively perceiving uncertainties and developing strong resilience.
[0343] CCP focuses on the probabilistic satisfaction of operational constraints, with its main load factor distribution ranging from [0.25 to 0.97]. Compared to R-SAC, its overload probability is lower, at only 3.42%. However, CCP exhibits a large number of high-severity outliers (exceeding the upper neighbor value Q4), with an expected overload δ. E The maximum overload level is δ, which is 0.181. M The probability of exceeding limits is 0.513, more than three times that of R-SAC (0.158). This indicates that the CCP may pose a potentially serious safety risk when HEES encounters sudden extreme operating conditions. Although R-SAC has a higher probability of exceeding limits than the high-risk CCP, it demonstrates significantly superior risk resistance and can reliably handle extreme scenarios.
[0344] 7.4) Rationality Analysis of Scheduling Strategy
[0345] This section evaluates the rationality of the scheduling strategy for the HEES-RD problem by comparing the commonly used risk-based scheduling method SA with the proposed R-SAC algorithm. Figure 9 The scheduling strategies developed by R-SAC and SA are demonstrated. Figure 10This displays the state changes of the hydrogen storage tank and battery, as well as the average operating efficiency of the electrolyzer under different scheduling strategies. For ease of visualization, the hydrogen mass has been converted into electrical power using the higher calorific value (HHVH).
[0346] Figure 9 This indicates that the overall strategies adopted by R-SAC and SA are similar. During periods of renewable energy surplus, excess electricity is absorbed through hydrogen production via electrolyzers and battery charging; while during peak load periods, such as midday and evening, power shortages are alleviated through fuel cell and battery discharge. This synergy between electricity and hydrogen storage ensures the reliability of power supply and improves the absorption rate of renewable energy, demonstrating the feasibility of R-SAC in peak-shaving and valley-filling hybrid energy storage. Compared to SA, R-SAC initiates charging earlier and for a longer duration before peak load periods, such as... Figure 10 The state of charge (SOC) of the medium-capacity battery varies during the periods of 3:00-9:00 and 17:00. During peak load periods of 10:00-12:00 and 19:00-20:00, the R-SAC achieves a higher net discharge capacity from the hybrid electric-hydrogen energy storage system compared to the SA, thereby reducing the required output of the thermal power unit and freeing up additional spinning reserve to cope with uncertainties. This is from... Figure 9 The lower power consumption of the electrolytic cell during these periods and Figure 10 The rapid depletion of hydrogen reserves is evident. As a trade-off, electrolyzers operating under R-SAC conditions have lower efficiency during these peak periods, such as... Figure 10 The time slots are shown as 10:00-12:00 and 19:00-20:00.
[0347] Based on the above economic and safety analyses, R-SAC prioritizes safety over economic efficiency, leveraging its strong uncertainty hedging capabilities to formulate a reasonable and safe dispatch strategy. These findings demonstrate that R-SAC can develop dispatch strategies that meet load demands, and compared to SA, it makes more efficient use of the hybrid electric-hydrogen energy storage system to address uncertainties.
[0348] 7.5) Computational Efficiency Analysis
[0349] This section evaluates the computational efficiency of the proposed R-SAC algorithm by comparing solution times, with the relevant results shown in Table 6. The results indicate that for complex HEES-RD problems, optimization-based methods such as SA and CCP require several hours to complete the solution. In contrast, R-SAC completes the solution in just 0.412 seconds, achieving a computational efficiency improvement of over 99.994% (i.e., 15907 times). This significant difference stems from the high dimensionality, non-convexity, and nonlinear probabilistic optimization characteristics of the HEES-RD problem itself. Existing commercial optimization solvers, constrained by the requirements of linearization, convexity, and scenario discretization, introduce a large number of auxiliary variables, leading to long solution times. The proposed R-SAC method, however, uses forward propagation of a neural network to formulate a scheduling strategy, replacing the traditional modeling and iterative optimization process. Therefore, R-SAC significantly improves risk management efficiency while ensuring the feasibility and safety of the HEES scheduling strategy.
[0350] Table 6 shows the solution time for each algorithm.
[0351]
Claims
1. A robust security deep reinforcement learning-based risk scheduling method for an electricity-hydrogen coupled system, characterized in that, Includes the following steps: 1) Considering the Faraday efficiency loss effect of the electrolyzer, a dynamic operation model for hybrid electro-hydrogen energy storage is constructed; 2) Taking into account the uncertainties of wind power, a risk scheduling model for the electric-hydrogen coupling system that takes into account the dynamic operation of electric-hydrogen hybrid energy storage is established; 3) Construct a robust constrained Markov decision-making architecture; 4) The robust constrained Markov decision architecture is parameterized using the flexible actor-evaluator baseline algorithm; 5) Solve the risk scheduling model of the electric-hydrogen coupling system using the parameterized robust constrained Markov decision process architecture to obtain the risk scheduling scheme of the electric-hydrogen coupling system; The risk scheduling model for the electro-hydrogen coupling system is shown below: sth(x,y,ζ)=0 (9) Ax + Dy + Mζ + e ≤ 0 (10) g(x,y,ζ)≤0 (11) u(x,y,ζ)≤0 (12) In the formula, t represents the scheduling interval; T is the scheduling period; vector x represents the decision variables including the scheduling strategy; vector y represents other state variables in the HEES (hydrogen-electric coupling system); vector ζ represents the uncertain prediction error; A, D, and M are coefficient matrices, and e is the coefficient vector; the subscript l indicates the l-th branch; l represents the branch set; the symbol '-' represents the upper limit of each variable; (·) + It represents max{·,0}; Represents the degree of branch overload; d B Let Pr{x} = 1 - F(x) represent the probability of x, and F() be the cumulative distribution function; l,t Indicates branch current flow; d B Indicates the risk threshold; f t (x) represents the operating cost function; h(x,y,ζ)=0 represents the AC power flow equation and energy balance equation; ρ is the uncertain prediction error distribution; Robust Constrained Markov Decision Architecture This includes action networks, reward evaluation networks, safety evaluation networks, Lagrange multipliers, and goal networks; Represents the state space. Represents the action space. Let r be the state transition function, representing the state transition in state s. t Next, execute action a t The reward obtained, c represents the reward in state s. t Next, execute action a t The cost of safe exploration obtained, where d is the threshold for exceeding the limit and γ is the discount factor; γ∈[0,1]; The action network encapsulates the action policy distribution π with neural network parameters θ, establishing a mapping from the state space to the action space. Reward-judgmenting networks are used to evaluate the operational status of the hydrogen-electro-hydrogen coupling system (HEES). t Lowering the control decision a t Reward soft Q function A safety assessment network for assessing the operating state s of an electrical hydrogen coupling system HEES t Down-regulation decision a t The safety assessment function Γ π (s t , a t , β); Lagrange multipliers are used to adjust the dual multipliers related to entropy and excess cost weights during the dual update process; The target network is used to enhance the training stability of the reward value discrimination network and the cost value discrimination network modules; When parameterizing the robust constrained Markov decision architecture using the flexible actor-evaluator baseline algorithm, the objective function and constraints are as follows: In the formula, Indicates the entropy threshold; β represents the confidence level.
2. The risk scheduling method for an electro-hydrogen coupling system based on robust and secure deep reinforcement learning according to claim 1, characterized in that, The dynamic operation model of the hybrid electric-hydrogen energy storage is shown below: In the formula, HHV H The highest calorific value of hydrogen under standard conditions; η e,t To improve the operating efficiency of hybrid electric-hydrogen energy storage; The power consumption and hydrogen production of the hybrid energy storage electrolyzer for hydrogen; Among them, the power consumption of the electrolytic cell and hydrogen production As shown below: P t e = N E V cell,t I cell,t (2) In the formula, I cell,t is the operating current of the electrolytic cell; N E is the number of series electrolytic cells, which is related to the maximum capacity of the electrolytic cell; k1 represents the conversion coefficient; η f,t is the Faraday efficiency; F is the Faraday constant; Electro-hydrogen hybrid energy storage operating voltage V cell,t As shown below: In the formula, t is the scheduling interval; V rev The voltage is the reversible voltage; T' is the operating temperature; R and F are the rational gas constant and Faraday constant, respectively. and These are the partial pressures of hydrogen and oxygen, respectively. The activity of H2O; R cell For internal equivalent ohmic resistance; α an and α cat These are the charge transfer coefficients for the anode and cathode, respectively; i an and i cat These are the anode and cathode exchange current densities, respectively; i cell,t Current density; Faraday efficiency η f,t As shown below: In the formula, Oxygen permeability in the diffusion mechanism; Hydrogen permeability in the diffusion mechanism; These represent the partial pressures of hydrogen at the cathode and oxygen at the anode in the catalyst layer, respectively. These are the fitting parameters for the increase in hydrogen and oxygen partial pressure in the cathode catalyst layer, respectively; This represents the hydrogen permeability caused by the pressure difference.
3. The risk scheduling method for an electro-hydrogen coupling system based on robust and secure deep reinforcement learning according to claim 1, characterized in that, The state space As shown below: In the formula, P L,t , P W,t , P G,t-1 , E B,t-1 These represent the electric load prediction vector, hydrogen load prediction vector, new energy prediction value, the first derivative of the new energy prediction value with respect to time, the thermal power output vector at the previous moment, the hydrogen storage capacity of the hydrogen storage tank at the previous moment, and the state of charge of the electrical energy storage at the previous moment, respectively; the subscripts L, W, G, H, and B represent the load, wind farm, thermal power unit, hydrogen storage tank, and electrical energy storage set, respectively. The first derivative of the predicted new energy value with respect to time is shown below: In the formula, Δt is a time difference; P W,t-1 represents a new energy prediction value at t-1 The action space As shown below: ΔP G,t = a G,t Θ[-T 60 D G ,T 60 U G ] (18) rd G,t = a D,t Θ [0, min(P G,t - P G , T 10 D G )] (20) In the formula, the subscripts E, FC, and B represent the sets of electrolyzers, fuel cells, and electrical energy storage, respectively; the subscript '_' indicates the lower limit of the corresponding variable. These represent the charging / discharging power of the energy storage, which are combined into a single decision P. B,t To reduce the action space; U G With D G These represent the ramp rate and descent rate of the generator set, respectively; T 60 With T 10 These refer to time intervals of one hour and ten minutes, respectively; a G,t a U,t a D,t a E,t a F,t a B,t a H,t These respectively represent the power output increment action of thermal power units, the upper rotation standby action of thermal power units, the lower rotation standby action, the operating current density action of electrolyzers, the power generation action of fuel cells, the charging / discharging power action of energy storage, and the hydrogen output action of hydrogen storage tanks; G,t rd G,t These represent the reserve capacity for upward and downward rotation, respectively; i cell,E,t P represents the current density limiting the operation of the electrolyzer; B,t This indicates the power of the electrical energy storage.
4. The risk scheduling method for an electro-hydrogen coupling system based on robust and secure deep reinforcement learning according to claim 1, characterized in that, The state transition function of the robust constrained Markov decision architecture is shown below: P G,t = P G,t-1 + ΔP G,t (25) In the formula, η represents the hydrogen inlet vector of the hydrogen storage tank; H,in η H,out Efficiency vectors for hydrogen input and output from the hydrogen storage tank, respectively; η cha,B η dis,B These represent the efficiency vectors for charging and discharging of electrical energy storage, respectively. This represents the amount of hydrogen stored in the hydrogen storage tank at the previous moment; Δt is the time difference; P B,t E represents the power of electrical energy storage. B,t-1 This indicates the state of charge of the stored electrical energy at the previous moment.
5. The risk scheduling method for an electro-hydrogen coupling system based on robust and secure deep reinforcement learning according to claim 1, characterized in that, Cost function c π,ρ (s t ,a t ,ζ) is given by: c π,ρ (s t ,a t ,ζ)=[ε e c e,t +ε b c b,t +ε pf c pf,t +ε pb c pb,t +ε ph c ph,t ]|π,ρ (31) In the formula, π represents the distribution of the agent's regulatory decision-making strategy; ε e ε b ε pf ε pb With ε ph These represent scaling respectively; c e,t This indicates that the capacity safety constraint of the hydrogen storage tank exceeds the limit; c b,t This indicates that the safety constraint on energy storage capacity has exceeded the limit; c pf,t This indicates that the branch overload risk constraint exceeds the limit; c pb,t c represents the power imbalance in the HEES (hydrogen-electro-hydrogen coupling) system. ph,t P represents the hydrogen energy imbalance quantity; e,t Let represent the power consumption of the e-th PEM electrolyzer at time t; ρ is the distribution of uncertain prediction errors. Represents the hydrogen load prediction vector; The reward function is as follows: r t = -(f t + f C,t ) = -(f G,t + f Q,t + f L,t + f HESS,t + f C,t ) (37) f HESS,t = λ h (f HS,om,t +f HS,loss,t +f BS,om,t +f BS,loss,t ) (41) f C,t = λ c (c π,ρ (s t ,a t ,ζ)-d) + (42) In the formula, f G,t represents the operation cost of thermal power; f Q,t , f L,t respectively represent the penalty terms of abandoned wind power and cut load; f HESS,t represents the operation cost of electric-hydrogen hybrid energy storage; f C,t represents the constraint overrun penalty cost; a g , b g , c g , u g , d g respectively represent the unit cost of constant, first-order, second-order, upward rotation reserve, and downward rotation reserve of the thermal power unit; λ q 、λ l 、λ h 、λ c represents the cost coefficient; ΔP i,t (ζ) represents the power imbalance of the i-th node caused by uncertainty; f HS,om,t , f BS,om,t respectively are the hydrogen storage energy, the operation and maintenance cost of the electric storage energy; f HS,loss,t , f BS,loss,t respectively are the hydrogen storage energy, the energy consumption cost of the electric storage energy; C e,inv , C fc,inv respectively are the investment and construction costs of the electrolyzer e and the fuel cell fc; T e,life , T fc,life respectively are the maximum running time of the electrolyzer e and the fuel cell fc; λ BS,om is the unit operation and maintenance cost coefficient of the electric storage energy; λ ep P for power distribution price; c π,ρ (s t , a t , ζ) is the cost function; η e,t is the electricity-hydrogen hybrid energy storage operation efficiency.
6. The risk scheduling method for an electro-hydrogen coupling system based on robust and secure deep reinforcement learning according to claim 1, characterized in that, Security evaluation function Γ π (s t ,a t ,β) is as follows: where φ(·), Φ(·) represent the probability density distribution function and the cumulative probability distribution function of the standard normal distribution, respectively; F π,ρ denotes the probability density function of C π,ρ ; c π,ρ (s t , a t , ζ) is the cost function; Among them, C π Second central moments As shown below:
7. The risk scheduling method for an electro-hydrogen coupling system based on robust and secure deep reinforcement learning according to claim 1, characterized in that, When parameterizing a robust constrained Markov decision architecture, the loss function of the security evaluation network As shown below: In the formula, This represents the replay of experience; α is a Lagrange multiplier; The loss function of the reward evaluation network is shown below: The loss function J of the action network π (θ) is as follows: In the formula, κ is a Lagrange multiplier; The update equations for the Lagrange multipliers are shown below: In the formula, ω is the learning rate.