Green hydrogen comprehensive energy park cooperative regulation and control method based on federal safety reinforcement learning

By constructing a multiphysics model and a robust prediction mechanism for multiple scenarios through federated security reinforcement learning, the security and privacy issues of heterogeneous clusters in green hydrogen production were solved, and the collaborative optimization of ALK and PEM electrolyzers was achieved, improving the system's operational resilience and economy.

CN121684529APending Publication Date: 2026-03-17HOHAI UNIV +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202610173033.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-06
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively address the issues of multi-physics coupling security, data privacy protection, and response speed in large-scale green hydrogen production, especially in distributed hydrogen production systems. Traditional scheduling schemes suffer from high communication burdens, privacy risks, and difficulties in quickly solving high-dimensional nonlinear differential equations using mathematical programming.

Method used

A high-fidelity multiphysics dynamic model is constructed using a federated security reinforcement learning approach. Combined with a multi-scenario robust prediction mechanism, the collaborative optimization of ALK and PEM electrolyzers is achieved through a federated multi-agent reinforcement learning framework, ensuring data privacy and improving response speed.

Benefits of technology

This approach effectively optimizes the operation of ALK and PEM electrolyzers while protecting data privacy, significantly improving the system's operational resilience and economy, reducing the risk of hydrogen-oxygen penetration, and increasing energy utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121684529A_ABST
    Figure CN121684529A_ABST
Patent Text Reader

Abstract

The invention provides a green hydrogen comprehensive energy park cooperative regulation and control method based on federal safety reinforcement learning. The method comprises the following steps: 1, obtaining operation parameters of a heterogeneous hydrogen production cluster; step 2, according to the acquired heterogeneous hydrogen production cluster operation parameters, aiming at different response characteristics of an alkaline electrolytic cell ALK and a proton exchange membrane electrolytic cell PEM, respectively establishing nonlinear multi-physical field models covering electrochemical reaction, transient thermodynamic evolution and gas purity evolution processes; step 3, according to the model established in the step 2, aiming at uncertainty fluctuation of renewable energy sources, based on multi-scene robust prediction, constructing a robust safe operation constraint of the system; 4, constructing a multi-agent reinforcement learning reward function integrated with a security defense mechanism according to the established robust security operation constraint, and generating a local distributed operation strategy; and 5, according to the local distributed operation strategy generated in the step 4, realizing global cooperative control of the heterogeneous hydrogen production cluster by using a federated learning mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of energy dispatch and distributed control technology, and in particular to a collaborative control method for green hydrogen integrated energy parks based on federated security reinforcement learning. Background Technology

[0002] Renewable energy sources such as wind and solar power are inherently random and volatile, and large-scale grid connection poses a severe challenge to the frequency stability and energy balance of the power system. Hydrogen energy, as a high-energy-density, zero-carbon-emission secondary energy carrier, can achieve on-site consumption and long-term storage of fluctuating electricity through water electrolysis technology.

[0003] Currently, industrial green hydrogen production is developing towards large-scale, distributed clustering, typically employing heterogeneous cluster systems composed of alkaline electrolyzers (ALK) and proton exchange membrane electrolyzers (PEM). ALK technology is mature and low-cost, but its dynamic response is slow; PEM, on the other hand, has an extremely fast response rate and can better track high-frequency fluctuations. However, in actual operation, heterogeneous clusters face multiple challenges: First, the electrolysis process involves complex nonlinear multiphysics coupling, including thermodynamic hysteresis effects and the risk of hydrogen-oxygen penetration (HTO) under low-load conditions; improper scheduling can easily lead to equipment lifespan damage or safety accidents. Second, traditional centralized scheduling schemes face significant communication burdens and the risk of privacy leaks of underlying physical parameters when dealing with geographically dispersed hydrogen production stations. Finally, existing mathematical programming methods struggle to solve high-dimensional nonlinear differential equations within milliseconds. Therefore, there is an urgent need for an intelligent collaborative control method that can ensure the safe operation of multiphysics while also balancing data privacy and response speed. Summary of the Invention

[0004] The purpose of this invention is to provide a collaborative control method for green hydrogen integrated energy parks based on federated security reinforcement learning. This method aims to construct a high-fidelity multi-physics dynamic model, utilize a multi-scenario robust prediction mechanism to proactively assess the potential risks brought about by renewable energy fluctuations, and combine a federated multi-agent reinforcement learning framework to achieve collaborative optimization of ALK and PEM heterogeneous units while protecting the data privacy of each hydrogen production unit.

[0005] Technical Solution: To achieve the aforementioned objectives, this invention proposes a collaborative regulation method for green hydrogen integrated energy parks based on federated security reinforcement learning. This method includes the following steps:

[0006] Step 1: Obtain the operating parameters of the heterogeneous hydrogen production cluster. The operating parameters include the real-time predicted power of renewable energy, the electrochemical reaction coefficient, thermodynamic coefficient and gas purity evolution coefficient of the alkaline electrolyzer (ALK) and proton exchange membrane electrolyzer (PEM) in the heterogeneous hybrid electrolyzer cluster.

[0007] Step 2: Based on the operating parameters of the heterogeneous hydrogen production cluster obtained in Step 1, nonlinear multiphysics models covering electrochemical reactions, transient thermodynamic evolution, and gas purity evolution processes are established for the different response characteristics of ALK and PEM electrolyzers.

[0008] Step 3: Based on the model established in Step 2, construct robust safety operation constraints for the system based on multi-scenario robust prediction to address the uncertain fluctuations of renewable energy.

[0009] Step 4: Based on the robust and safe operation constraints established in Step 3, construct a multi-agent reinforcement learning reward function with an integrated security defense mechanism, and generate a local distributed operation strategy;

[0010] Step 5: Based on the local distributed operation strategy generated in Step 4, use the federated learning mechanism to achieve global collaborative control of the heterogeneous hydrogen production cluster.

[0011] Furthermore, in step 2, based on the operating parameters of the heterogeneous hydrogen production cluster obtained in step 1, nonlinear multiphysics models covering electrochemical reactions, transient thermodynamic evolution, and gas purity evolution processes are established for the different response characteristics of ALK and PEM electrolyzers, as follows;

[0012] 1) Electrochemical model

[0013] (A-1)

[0014] (A-2)

[0015] (A-3)

[0016] (A-4)

[0017] (A-5)

[0018] (A-6)

[0019] (A-7)

[0020] (A-8)

[0021] (A-9)

[0022] (A-10)

[0023] (A-11)

[0024] (A-12)

[0025] (A-13)

[0026] in, To be at current density Operating temperature Operating pressure The operating voltage refers to the single-cell terminal voltage during the operation of the electrolytic cell. To operate at temperature Operating pressure The reversible voltage refers to the minimum voltage required to initiate an electrolytic reaction under thermodynamic equilibrium. To be at current density Operating temperature The activation overpotential refers to the loss generated by overcoming the resistance of electrochemical reactions on the electrode surface. To be at current density Operating temperature The ohmic overpotential refers to the loss generated by overcoming the resistance to ion and electron conduction. To be at current density Operating temperature Concentration overpotential refers to the voltage loss caused by the reactant concentration gradient. For the real-time operating temperature of the electrode, To withstand operating pressure, Standard atmospheric pressure This represents the change in Gibbs free energy. For the number of transferred electrons, It is Faraday's constant. It is an inverse hyperbolic sine function. This is the universal gas constant. The charge transfer coefficient, For current density, Wherein is the electrode active area, This represents the total current flowing through the electrolytic cell. Operating temperature The exchange current under, For reference exchange current, For activation energy, For reference temperature, This is the limiting current. Operating temperature The resistance per unit area of ​​ohms below The thickness of the electrode or film. For reference conductivity, The temperature coefficient of electrical conductivity For contact resistance, The hydrogen molar flow rate refers to the amount of hydrogen produced per unit time. Faraday efficiency refers to the ratio of actual hydrogen production to theoretical hydrogen production. The number of individual cells refers to the number of cells connected in series in the electrolytic cell. and These are the empirical fitting coefficients used to calculate Faraday efficiency. Electrolysis power refers to the electrical energy consumed by the electrolytic cell during operation. For system energy efficiency, The mass flow rate of hydrogen gas. The molar mass of hydrogen gas is... Due to the high calorific value of hydrogen, Compression power consumption refers to the auxiliary power consumption required to pressurize hydrogen to the hydrogen storage pressure. Pressure to meet export targets, The adiabatic index of the gas. For compressor efficiency;

[0027] 2) Thermodynamic model

[0028] (A-14) (A-15)

[0029] (A-16)

[0030] (A-17)

[0031] (A-18)

[0032] in, For the heat capacity of the ALK electrolytic reactor, For time variables, This refers to the core temperature of the ALK electrolytic reactor. for The heat production power of the reaction at any given moment for The natural heat dissipation power at any given time refers to the heat lost from the electrolytic cell to the environment. The heat transfer coefficient between the ALK electrolyzer core and the electrolyte. This refers to the real-time temperature of the electrolyte. This represents the total heat capacity of the electrolyte in the ALK electrolyzer. The density of the electrolyte. This refers to the electrolyte circulation volumetric flow rate. The specific heat capacity of the electrolyte. For the cyclic delay time, For the electrolyte to undergo time Temperature after cycle delay for The cooling system removes heat at all times. The equivalent heat capacity of the PEM electrolytic cell system, The temperature of the electrolytic cell. The heat conversion coefficient reflects the proportion of electrical energy loss that is converted into heat energy. The natural convection heat transfer coefficient is... Ambient temperature;

[0033] 3) Gas purity evolution model

[0034] (A-19)

[0035] (A-20)

[0036] (A-21) (A-22)

[0037] (A-23)

[0038] (A-24)

[0039] (A-25)

[0040] (A-26)

[0041] (A-27)

[0042] (A-28)

[0043] (A-29)

[0044] in, For diffusion and permeation rate, The diffusion coefficient is... The solubility coefficient of the gas. For the diaphragm area, For the diaphragm thickness, For convection permeation rate, Where is the permeability constant. For pressure difference, These are membrane structure-related constants. For the mixed dissolution and permeation rate, This refers to the electrolyte circulation volumetric flow rate. The total impurity permeation rate, This is the solubility fitting constant. The temperature of the electrolytic cell. This represents the amount of impurities on the anode side. This represents the volume of the anode gas-liquid mixture. This refers to the amount of impurities in the liquid phase of the separator. The separator time constant, This refers to the amount of gaseous impurities in the separator. The total impurity permeation rate at the outlet. The total outlet gas velocity, This refers to the total amount of gaseous substances within the separator. For separator pressure, The gas phase volume of the separator. To achieve static equilibrium hydrogen and oxygen penetration concentrations, This represents the real-time hydrogen-oxygen penetration concentration, which is the hydrogen content in oxygen.

[0045] Furthermore, in step 3, based on the model established in step 2, robust safety operation constraints for the system are constructed based on multi-scenario robust prediction to address the uncertain fluctuations of renewable energy, as detailed below:

[0046] (A-30)

[0047] (A-31) (A-32)

[0048] (A-33)

[0049] (A-34)

[0050] (A-35)

[0051] (A-36)

[0052] (A-37)

[0053] (A-38)

[0054] (A-39)

[0055] in, for Power deviation rate at time t, For the unit number, For the predicted number of future time steps, The predicted random scenario number, For the continuous operating time of the unit, To predict from From the moment to The offset of the prediction process between steps, for The operating power of the electrolytic cell at any given time. Rated operating power, This represents the power deviation rate after filtering and smoothing. As a smoothing factor, Penalty for high-load operation, i.e., power deviation rate Greater than the high load deviation threshold The load operation risk value triggered at any time. For high load penalty coefficient, This is the high-load safety assessment value, i.e. Exceeding Total runtime, For high-load operation time threshold, For high load deviation threshold, Penalty for low-load operation, i.e., power deviation rate Less than the low load deviation threshold The load operation risk value triggered at any time. For low load penalty coefficient, For low load deviation threshold, This is the low-load safety state value, i.e. Below Total runtime, For reference safety thresholds, The scaling factor for the penalty function. For the first In this scenario Forecast of renewable energy power at any time for The prediction baseline value at time, To predict random errors, For the error variance, To obey arrive The standard normal distribution The decision variable refers to the currently issued electrolytic cell scheduling power command. For decision making The next The device in the In this scenario The predicted state sequence at time 10:00. In the first In this scenario Real-time forecast of new energy output For Consider a predicted scenario and decision-making. and the The state transition function for predicting the output of new energy sources under this scenario. For decision making The risk loss function is as follows: and For loss weighting coefficients, To The highest risk item after sorting by risk value For the first The rated operating range of each device For low-load safety boundary, This is a symbolic function; its value is 1 when the internal variables are greater than 0. To The maximum loss for each scenario after ranking by risk value. Conditional Value at Risk (VaR) refers to the average risk loss exceeding a certain threshold. For risk confidence level, This represents the total number of scenarios. To provide a comprehensive risk probability score, and For risk assessment weighting coefficients, As the normalization factor, To robustly predict security penalties, To predict the system scaling factor, This is the global penalty scaling factor.

[0056] Furthermore, in step 4, based on the robust and safe operation constraints established in step 3, a multi-agent reinforcement learning reward function integrating a security defense mechanism is constructed, and a local distributed operation strategy is generated. The specific method is as follows:

[0057] (A-40)

[0058] (A-41)

[0059] (A-42)

[0060] in, for The state-space vector at time t, For the total power of renewable energy, For the first The core temperature of the ALK electrolyzer. For the first The liquid temperature of the alkaline electrolytic cell, For the first Operating temperature of the PEM electrolyzer, for The action space vector at time t, To be assigned to the The power of the ALK unit To be assigned to the The power of the PEM unit For ALK operating pressure, For PEM operating pressure, This refers to the ALK electrolyte circulation flow rate. This refers to the PEM electrolyte circulation flow rate. This is a comprehensive reward function used to guide policy optimization. To reduce the penalty weight for power reduction, As a weighting of hydrogen production revenue, As a penalty weight for gas purity, This represents the maximum safe limit for the permissible hydrogen and oxygen penetration concentration.

[0061] Furthermore, in step 5, based on the local distributed operation strategy generated in step 4, a federated learning mechanism is used to achieve global collaborative control of the heterogeneous hydrogen production cluster. The specific method is as follows:

[0062] (A-43)

[0063] (A-44)

[0064] (A-45)

[0065] (A-46)

[0066] in, For local agent network parameters, As a global smoothing factor, These are the network parameters obtained through global aggregation. This serves as the reference target value for the algorithm when updating the evaluator network. As a termination signal, As a discount factor, This is the current state. For the currently executing action, Current state Execute action After that, the system transitions to the state of the next moment. To determine the state at the next time step based on the current strategy. The optimal action to take as expected. To evaluate the network's estimate for the next time step, This is the current evaluation network's estimate for the current moment. For the centered loss function, For sample batch size, This is the consistency weighting coefficient. To predict state values, These are actual observations. Let be the loss function for the action network. To take the average of random variables, For natural logarithm operations, Here, is the policy improvement term in the policy gradient algorithm, where... For the evaluator network to evaluate the current state Execution strategy The evaluation value, For the policy network in state The action vectors generated directly below To replay experience pool A batch of states were randomly sampled from the pool. And average the evaluation results under these states. These are the current operating parameters of the action network. This is the proximal regularization coefficient. Here is the entropy regularization coefficient. This represents the action probability distribution output by the policy network.

[0067] Beneficial effects: Compared with the prior art, the technical solution of the present invention has the following beneficial effects:

[0068] This invention focuses on optimizing the operational safety and distributed privacy collaboration of heterogeneous hydrogen production clusters. By constructing a high-fidelity multiphysics model and a robust multi-scenario prediction mechanism, combined with a federated security reinforcement learning framework, it achieves collaborative proactive defense against core temperature, gas purity, and power allocation issues in both ALK and PEM electrolyzers. Compared to traditional centralized or penalty-based scheduling methods, the proposed method can achieve risk prediction and security closure while protecting data privacy, effectively suppressing physical boundary violations and improving energy utilization efficiency, significantly enhancing the operational resilience and economy of large-scale green hydrogen systems. Attached Figure Description

[0069] Figure 1 This is a flowchart of the method of the present invention;

[0070] Figure 2 This is a comparison chart of unit output under scenarios of fluctuating renewable energy output. Figure 2 Figure a in the diagram shows the power output of the ALK generator unit under the baseline strategy for fluctuating wind and solar power output scenarios. Figure 2 Figure b in the figure is the power output diagram of the ALK unit under the improvement strategy for fluctuating wind and solar power output scenarios;

[0071] Figure 3 This is a comparison chart of unit output under a stable renewable energy output scenario. Figure 3 Figure a in the diagram shows the power output of the ALK generator unit under the baseline strategy for a stable wind and solar power output scenario. Figure 3Figure b in the figure is the power output diagram of the ALK unit under the improved strategy for stable wind and solar power output scenarios;

[0072] Figure 4 This is a comparison of the results of gaseous impurity concentration control. Detailed Implementation

[0073] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. After reading this invention, any modifications of the invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.

[0074] like Figure 1 As shown, this invention proposes a collaborative regulation method for green hydrogen integrated energy parks based on federated security reinforcement learning. This method includes the following steps:

[0075] Step 1: Obtain the operating parameters of the heterogeneous hydrogen production cluster. The operating parameters include the real-time predicted power of renewable energy, the electrochemical reaction coefficient, thermodynamic coefficient and gas purity evolution coefficient of the alkaline electrolyzer (ALK) and proton exchange membrane electrolyzer (PEM) in the heterogeneous hybrid electrolyzer cluster.

[0076] Step 2: Based on the operating parameters of the heterogeneous hydrogen production cluster obtained in Step 1, nonlinear multiphysics models covering electrochemical reactions, transient thermodynamic evolution, and gas purity evolution processes are established for the different response characteristics of ALK and PEM electrolyzers.

[0077] Step 3: Based on the model established in Step 2, construct robust safety operation constraints for the system based on multi-scenario robust prediction to address the uncertain fluctuations of renewable energy.

[0078] Step 4: Based on the robust and safe operation constraints established in Step 3, construct a multi-agent reinforcement learning reward function with an integrated security defense mechanism, and generate a local distributed operation strategy;

[0079] Step 5: Based on the local distributed operation strategy generated in Step 4, use the federated learning mechanism to achieve global collaborative control of the heterogeneous hydrogen production cluster.

[0080] Case Analysis

[0081] The following example illustrates the superiority of the collaborative control method for green hydrogen integrated energy parks based on federated security reinforcement learning described in this invention. Two typical 120-minute operating scenarios—one with stable new energy output and the other with fluctuating new energy output—are used to compare the performance of this scheduling strategy in terms of operational safety and hydrogen production efficiency.

[0082] Table 1 summarizes the energy efficiency statistics of the proposed algorithm and the MATD3 algorithm under two operating scenarios: stable and fluctuating. All other model parameters remain the same. Under stable conditions, the proposed algorithm produces 503.12 kg of hydrogen with an energy utilization rate of 87.89%. In contrast, the MATD3 algorithm produces only 490.07 kg of hydrogen with an energy utilization rate of 85.06%. In this scenario, the proposed algorithm improves hydrogen production by approximately 2.66%. Under fluctuating conditions, the proposed algorithm exhibits stronger robustness: its hydrogen production is 485.70 kg, and its energy utilization rate remains at a high level of 87.81%. The MATD3 algorithm's energy utilization rate drops to 83.34% under fluctuating conditions, with an optimization rate of 4.47%. Furthermore, the proposed algorithm maintains a stable hydrogen conversion efficiency of around 52.2% in both scenarios, significantly outperforming MATD3. This performance advantage is mainly due to the superior decision-making accuracy of the proposed algorithm when dealing with multi-agent cooperation and operational constraints: through precise perception of the electric-hydrogen coupling system, the algorithm can achieve more flexible energy management in power fluctuation environments, effectively avoiding the inefficient operating range of energy during the conversion process, thereby maximizing hydrogen production and comprehensive energy utilization efficiency while ensuring system operation safety.

[0083] Table 1. Comparison of energy efficiency statistics of the two algorithms under two types of operating scenarios.

[0084]

[0085] Compare the proposed algorithm with the comparative algorithm MATD3 under different operating conditions, such as Figure 2 As shown. It can be concluded that, compared to the algorithms, Figure 2 Figure a in the middle and Figure 3 Figure a shows significant power fluctuations during operation, exhibiting numerous power peaks and deep troughs. This indicates that the unit undergoes frequent and drastic adjustments in response to environmental changes. In contrast, the proposed algorithm... Figure 2 Figure b in the middle and Figure 3Figure b shows that after introducing robust prediction across multiple scenarios, power allocation is more stable under different scenarios. Observations indicate that the proposed algorithm significantly reduces the occurrence of extreme high-power output by the units, while also effectively preventing units from operating at low power for extended periods. This stable output curve demonstrates that, under the scheduling of the proposed algorithm, units ALK1-ALK4 can more easily share load demands, avoiding abrupt load switching for individual units. By reducing drastic power fluctuations, the system can effectively reduce mechanical and thermal stresses within the electrolyzer units, thereby extending the service life of core equipment. Furthermore, avoiding low-power operating ranges also helps improve the overall energy efficiency of the hydrogen production system, achieving a synergistic improvement in system flexibility and economy while ensuring operational safety.

[0086] Compare the HTO hydrogen-oxygen interpenetration concentration evolution characteristics of the proposed method and the comparative method during operation, such as Figure 4 As shown, in the initial stage, the HTO concentration of both methods was at an initial level of 1.50%. With increasing running time, the proposed method demonstrated superior concentration suppression and optimization capabilities: its HTO concentration curve showed a clear and steady downward trend, successfully reducing the concentration to approximately 1.13% by the end of the 120-minute running cycle. In contrast, the HTO concentration of the comparative method fluctuated only slightly throughout the entire run, remaining consistently at a high level of around 1.48%, failing to achieve effective concentration reduction.

[0087] This significant difference demonstrates that the proposed method, through more precise power allocation and operational control, can more effectively suppress hydrogen-oxygen interpenetration within the electrolyzer. During long-term operation, this continuous optimization of HTO concentration not only significantly improves hydrogen production purity but, more importantly, maintains the system within a safer operating range, substantially reducing safety risks. This reflects the significant advantages of the proposed algorithm in ensuring the inherent safety of the system.

[0088] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A green hydrogen comprehensive energy park collaborative regulation method based on federal security reinforcement learning, characterized in that, The method comprises the following steps: Step 1, obtaining the operation parameters of the heterogeneous hydrogen production cluster, the operation parameters comprising real-time renewable energy predicted power, electrochemical reaction coefficients, thermal dynamics coefficients and gas purity evolution coefficients of the alkaline electrolyzer ALK and the proton exchange membrane electrolyzer PEM in the heterogeneous mixed electrolyzer cluster; Step 2, according to the operation parameters of the heterogeneous hydrogen production cluster obtained in step 1, different response characteristics of the ALK and the PEM electrolyzer are established respectively, and a nonlinear multi-physical field model covering electrochemical reaction, transient thermal dynamics evolution and gas purity evolution process is established; Step 3, according to the model established in step 2, the robust safety operation constraint of the system is constructed based on multi-scenario robust prediction for the uncertainty fluctuation of renewable energy; Step 4, according to the robust safety operation constraint established in step 3, a multi-agent reinforcement learning reward function integrated with a safety defense mechanism is constructed, and a local distributed operation strategy is generated; Step 5, according to the local distributed operation strategy generated in step 4, the global collaborative control of the heterogeneous hydrogen production cluster is realized by using the federated learning mechanism.

2. The method of claim 1, wherein the method is based on federated security reinforcement learning for coordinated control of a green hydrogen integrated energy park. In step 2, according to the operation parameters of the heterogeneous hydrogen production cluster obtained in step 1, different response characteristics of the ALK and the PEM electrolyzer are established respectively, and a nonlinear multi-physical field model covering electrochemical reaction, transient thermal dynamics evolution and gas purity evolution process is established, as follows: 1) Electrochemical model (A-1) (A-2) (A-3) (A-4) (A-5) (A-6) (A-7) (A-8) (A-9) (A-10) (A-11) (A-12) (A-13) wherein, is the operating voltage at current density , operating temperature , operating pressure , is the single cell terminal voltage when the electrolyzer is running, is the reversible voltage at operating temperature , operating pressure , is the minimum voltage required to start the electrolysis reaction in the thermodynamic equilibrium state, is the activation overpotential at current density , operating temperature , is the loss generated by overcoming the resistance of the electrochemical reaction on the electrode surface, is the ohmic overpotential at current density , operating temperature , is the loss generated by overcoming the resistance of ion and electron conduction, is the concentration overpotential at current density , operating temperature , is the voltage loss caused by the concentration gradient of the reactants; is the real-time operating temperature of the electrode, is the operating pressure, is the standard atmospheric pressure, is the Gibbs free energy change, is the number of transferred electrons, is the Faraday constant, is the inverse hyperbolic sine function, is the universal gas constant, is the charge transfer coefficient, is the current density, is the active area of the electrode, wherein, represents the total current flowing through the electrolyzer, is the exchange current at operating temperature , is the reference exchange current, is the activation energy, is the reference temperature, is the limiting current, is the ohmic resistance per unit area at operating temperature , is the thickness of the electrode or membrane, is the reference conductivity, is the temperature coefficient of conductivity, is the contact resistance, is the hydrogen molar flow rate, indicating the amount of hydrogen substance produced per unit time, is the Faraday efficiency, indicating the ratio of the actual hydrogen production to the theoretical hydrogen production, Ncell is the number of cells, which refers to the number of cells connected in series in the electrolysis stack, Ncell is the number of cells, which refers to the number of cells connected in series in the electrolysis stack, Ncell is the number of cells, which refers to the number of cells connected in series in the electrolysis stack, Ncell is the number of cells, which refers to the number of cells connected in series in the electrolysis stack, Ncell is the number of cells, which refers to the number of cells connected in series in the electrolysis stack, Ncell is the number of cells, which refers to the number of cells connected in series in the electrolysis stack, Ncell is the number of cells, which refers to the number of cells connected in series in the electrolysis stack, Ncell is the number of cells, which refers to the number of cells connected in series in the electrolysis stack, Ncell is the number of cells, which refers to the number of cells connected in series in the electrolysis stack, Ncell is the number of cells, which refers to the number of cells connected in series in the electrolysis stack, Ncell is the number of cells, which refers to the number of cells connected in series in the electrolysis stack, Ncell is the number of cells, which refers to the number of cells connected in series in the electrolysis stack, 2) Thermal dynamics model (A-14) (A-15) (A-16) (A-17) (A-18) wherein, is the heat capacity of the ALK electrolysis stack, is the time variable, is the ALK electrolysis stack core temperature, is the reaction heat generation power at time t, is the natural heat dissipation power at time t, which refers to the heat dissipated by the electrolyzer to the environment, is the heat transfer coefficient between the ALK electrolyzer core and the electrolyte, is the real-time temperature of the electrolyte, is the total heat capacity of the electrolyte of the ALK electrolyzer, is the density of the electrolyte, is the circulating volume flow rate of the electrolyte, is the specific heat capacity of the electrolyte, is the circulation delay time, is the temperature of the electrolyte after experiencing the time circulation delay, is the heat removed by the cooling system at time t, is the equivalent heat capacity of the PEM electrolyzer system, is the electrolyzer temperature, is the heat conversion coefficient, which reflects the proportion of electrical energy loss converted into heat energy, is the natural convection heat transfer coefficient, is the ambient temperature; 3) Gas purity evolution model (A-19) (A-20) (A-21) (A-22) (A-23) (A-24) (A-25) (A-26) (A-27) (A-28) (A-29) wherein, is the diffusion permeation rate, is the diffusion coefficient, is the gas solubility coefficient, is the membrane area, is the membrane thickness, is the convective permeation rate, is the permeability constant, is the pressure differential, is the membrane structure related constant, is the mixed dissolution permeation rate, is the electrolyte circulation volume flow rate, is the total impurity permeation rate, is the solubility fitting constant, is the electrolyte cell temperature, is the amount of impurity substance on the anode side, is the anode gas-liquid mixed volume, is the amount of impurity substance in the separator liquid phase, is the separator time constant, is the amount of impurity substance in the separator gas phase, is the outlet total impurity permeation rate, is the outlet total gas flow rate, is the amount of total gas substance in the separator, is the separator pressure, is the separator gas phase volume, is the static equilibrium hydrogen-oxygen breakthrough concentration, is the real-time hydrogen-oxygen breakthrough concentration, i.e., the hydrogen content in oxygen.

3. The method of claim 2, wherein the method is based on federated security reinforcement learning for coordinated control of a green hydrogen integrated energy park. In step 3, according to the model established in step 2, the robust safety operation constraint of the system is constructed based on multi-scenario robust prediction for the uncertainty fluctuation of renewable energy, as follows: (A-30) (A-31) (A-32) (A-33) (A-34) (A-35) (A-36) (A-37) (A-38) (A-39) wherein, is the power deviation rate at time is the unit number, is the predicted future time step, is the predicted random scenario number, is the unit on-duration time, is the predicted process offset between time and time , is the electrolyzer operating power at time , is the rated operating power, is the filtered and smoothed power deviation rate, is the smoothing factor, is the high load operating penalty, i.e. the power deviation rate greater than the high load deviation threshold , is the high load penalty coefficient, is the high load safety assessment value, i.e. the total operating time exceeding , is the high load operating time threshold, is the high load deviation threshold, is the low load operating penalty, i.e. the power deviation rate less than the low load deviation threshold , is the low load penalty coefficient, is the low load deviation threshold, is the low load safety state value, i.e. the total operating time below , is the reference safety threshold, is the penalty function scaling factor, is the predicted renewable energy power at time under scenario , is the predicted reference value at time , is the predicted random error, is the error variance, is the standard normal distribution subject to to , is the decision variable, referring to the currently issued electrolyzer dispatching power instruction, is the decision under scenario , In this scenario The predicted state sequence at time 10:

00. In the first In this scenario Real-time forecast of new energy output For Consider a predicted scenario and decision-making. and the The state transition function for predicting the output of new energy sources under this scenario. For decision making The risk loss function is as follows: and For loss weighting coefficients, To The highest risk item after sorting by risk value For the first The rated operating range of each device For low-load safety boundary, This is a symbolic function; its value is 1 when the internal variables are greater than 0. To The maximum loss for each scenario after ranking by risk value. Conditional Value at Risk (VaR) refers to the average risk loss exceeding a certain threshold. For risk confidence level, This represents the total number of scenarios. To provide a comprehensive risk probability score, and For risk assessment weighting coefficients, As the normalization factor, To robustly predict security penalties, To predict the system scaling factor, This is the global penalty scaling factor.

4. The method of claim 3, wherein, In step 4, according to the robust safety operation constraint established in step 3, a multi-agent reinforcement learning reward function integrated with a safety defense mechanism is constructed, and a local distributed operation strategy is generated, as follows: (A-40) (A-41) (A-42) wherein, is the total power of renewable energy sources, is the state space vector at time instant, is the total power of renewable energy sources, is the core temperature of the nth alkali electrolyzer, is the liquid temperature of the nth alkali electrolyzer, is the core temperature of the nth alkali electrolyzer, is the liquid temperature of the nth alkali electrolyzer, is the operating temperature of the nth PEM electrolyzer, is the operating temperature of the nth PEM electrolyzer, is the action space vector at time instant, is the power allocated to the nth alkali unit, is the power allocated to the nth alkali unit, is the power allocated to the nth PEM unit, is the power allocated to the nth PEM unit, is the alkali operating pressure, is the PEM operating pressure, is the alkali electrolyte circulation flow rate, is the PEM electrolyte circulation flow rate, is the PEM electrolyte circulation flow rate, is the integrated reward function for guiding policy optimization, is the power reduction penalty weight, is the hydrogen production revenue weight, is the gas purity penalty weight, is the maximum safety limit of the allowed hydrogen-oxygen penetration concentration.

5. The method of claim 4, wherein, In step 5, according to the local distributed operation strategy generated in step 4, the global collaborative control of the heterogeneous hydrogen production cluster is realized by using the federated learning mechanism, as follows: (A-43) (A-44) (A-45) (A-46) in, For local agent network parameters, As a global smoothing factor, These are the network parameters obtained through global aggregation. This serves as the reference target value for the algorithm when updating the evaluator network. As a termination signal, As a discount factor, This is the current state. For the currently executing action, Current state Execute action After that, the system transitions to the state of the next moment. To determine the state at the next time step based on the current strategy. The optimal action to take as expected. To evaluate the network's estimate for the next time step, This is the current evaluation network's estimate for the current moment. For the centered loss function, For sample batch size, This is the consistency weighting coefficient. To predict state values, These are actual observations. Let be the loss function for the action network. To take the average of random variables, For natural logarithm operations, Here, is the policy improvement term in the policy gradient algorithm, where... For the evaluator network to evaluate the current state Execution strategy The evaluation value, For the policy network in state The action vectors generated directly below To replay from the experience pool A batch of states were randomly sampled from the pool. And average the evaluation results under these states. These are the current operating parameters of the action network. This is the proximal regularization coefficient. Here is the entropy regularization coefficient. This represents the action probability distribution output by the policy network.

Citation Information

Patent Citations

  • Electricity-hydrogen coupling system risk scheduling method based on robust safety deep reinforcement learning

    CN119863051A

  • Green hydrogen comprehensive energy park cooperative scheduling method considering thermal inertia of electrolytic cell

    CN120511718A

  • Multi-park integrated energy system scheduling method based on personalized federal reinforcement learning

    CN120688813A

  • Multi-energy micro-grid cooperative regulation and control system and method based on cross-layer knowledge injection and federated distillation

    CN121436443A