Electrolytic water hydrogen production control method and system based on reinforcement learning

CN122833653APending Publication Date: 2026-09-29SHENYANG INST OF ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611006055.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

该方法依赖于预先设定的控制逻辑和固定参数,针对额定工况或特定稳态工作点进行设计,难以快速准确地响应可再生能源发电功率的剧烈波动

Benefits of technology

本发明通过融合电化学、热力学、传质及性能衰减物理场模型构建电解水制氢动态仿真环境,并基于深度强化学习设计多目标融合的奖励函数,实现了对电解槽电流和冷却水流量的协同优化控制,显著提升了系统在可再生能源波动供电条件下的综合能效和可再生能源消纳率,有效抑制了温度越限和功率剧烈波动,保障了运行安全并延长了设备使用寿命,整体上大幅提升了电解水制氢系统的动态响应能力和智能化控制水平。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122833653A_ABST
    Figure CN122833653A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of hydrogen production by water electrolysis, and proposes a hydrogen production by water electrolysis control method and system based on reinforcement learning, which comprises the following steps: constructing a dynamic simulation environment of a hydrogen production by water electrolysis system including multiple physical models; designing a reinforcement learning agent, wherein the state space contains the current system operating state, the power generation power prediction value and the hydrogen demand prediction value in the future prediction time domain, the action space contains the electrolytic cell input current adjustment increment and the cooling water flow adjustment increment, and the reward function is a multi-objective fusion function considering system energy efficiency, renewable energy consumption, thermal safety and power fluctuation suppression; training the reinforcement learning agent offline in the simulation environment to obtain a control strategy model; and according to the trained control strategy model, outputting an action instruction according to real-time observation state in each control cycle and executing the action instruction to control the electrolytic cell operation. The present application realizes adaptive control of the hydrogen production by water electrolysis system under dynamic renewable energy power supply.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of water electrolysis for hydrogen production technology, and more specifically, to a control method and system for water electrolysis for hydrogen production based on reinforcement learning. Background Technology

[0002] Electrolysis of water to produce hydrogen is one of the main methods for obtaining high-purity hydrogen in industry. Its basic principle is that under the action of a DC electric field, water molecules undergo oxidation-reduction reactions at the anode and cathode of the electrolyzer, respectively. Oxygen is produced at the anode and hydrogen is produced at the cathode, thereby achieving the decomposition of water.

[0003] However, current operation control of water electrolysis hydrogen production systems mostly employs rule-based control strategies or traditional proportional-integral-derivative (PID) control methods. These methods rely on pre-set control logic and fixed parameters, designed for rated operating conditions or specific steady-state operating points, making it difficult to respond quickly and accurately to drastic fluctuations in renewable energy power generation. When input power changes frequently, traditional control methods often fail to optimize multiple variables such as electrolyzer current and temperature, leading to decreased system efficiency, frequent equipment start-ups and shutdowns, or disruptions to operating conditions, ultimately impacting equipment lifespan and operational economics.

[0004] While existing research includes methods based on Model Predictive Control (MPC) or dynamic programming that attempt to optimize system operation by predicting future power and demand, these methods typically rely on precise mechanistic models and linearization assumptions. They have limited ability to characterize the complex characteristics of electrolyzer operation, such as strong nonlinearity, multi-physics coupling, and performance degradation over long-term operation, making it difficult to achieve globally optimal control and adaptive adjustment. Furthermore, traditional optimization methods often rely on expert experience when facing multi-objective conflicts and lack consideration for long-term cumulative returns.

[0005] Therefore, there is an urgent need for an intelligent control method for hydrogen production by water electrolysis that can adapt to the fluctuations in renewable energy, take into account the long-term performance degradation of equipment, and make forward-looking decisions, so as to overcome the above-mentioned defects of existing technologies and improve the overall operating performance and long-term economic efficiency of the system. Summary of the Invention

[0006] In view of this, the present invention proposes a method and system for controlling hydrogen production by water electrolysis based on reinforcement learning, in order to solve the problems existing in the prior art.

[0007] To achieve the above objectives, the present invention provides the following technical solution: A reinforcement learning-based method for controlling hydrogen production via water electrolysis includes the following steps: A dynamic simulation environment for a water electrolysis hydrogen production system is constructed. The dynamic simulation environment includes electrochemical, thermodynamic and mass transfer dynamic models of the electrolyzer, as well as renewable energy power generation prediction models and hydrogen demand models. The design includes a state space, action space, and reward function for the reinforcement learning agent. The state space includes the current system operating state and the predicted power generation and hydrogen demand in the future prediction time domain. The action space includes the incremental adjustment of the electrolyzer input current and the incremental adjustment of the cooling water flow rate. The reward function is a multi-objective fusion function that takes into account system energy efficiency, renewable energy consumption, thermal safety, and power fluctuation suppression. The reinforcement learning agent is trained offline in the simulation environment to obtain a control policy model. The trained control strategy model is deployed to the controller of a real water electrolysis hydrogen production system. In each control cycle, action commands are output based on the real-time observed status, and the action commands are executed to control the operation of the electrolyzer.

[0008] Furthermore, the electrochemical, thermodynamic, and mass transfer dynamics models of the electrolyzer include: An electrolytic cell voltage-current characteristic model is established based on the equivalent circuit and Nernst equation, wherein the voltage includes reversible voltage, activation overpotential, ohmic overpotential and concentration overpotential. A dynamic temperature model for the electrolytic cell is established based on the lumped heat capacity method. The dynamic temperature model calculates temperature changes based on heat generation power, heat dissipation power, and cooling power. A membrane conductivity model is dynamically established based on the water content inside the membrane. The membrane conductivity model calculates the membrane conductivity based on the water content and temperature of the membrane and affects the ohmic overpotential. A cumulative degradation factor is introduced, which increases cumulatively over time based on the operating current and temperature, and is superimposed on the voltage of a single cell in the electrolyzer to simulate long-term performance degradation.

[0009] Furthermore, the renewable energy power generation prediction model includes: A preliminary power prediction sequence was generated based on the photovoltaic physical model and the wind power generation power curve model; A long short-term memory neural network is used as the residual correction module. The input is the preliminary power sequence output by the physical model, the numerical weather prediction feature vector, and the historical measured power sequence. The output is the corrected residual. The corrected residual is superimposed on the preliminary power prediction sequence to obtain the final predicted power.

[0010] Furthermore, the hydrogen demand model is generated by superimposing low-frequency fluctuations and high-frequency noise onto a baseline load curve. The low-frequency fluctuations adopt a first-order autoregressive model, and the high-frequency noise is independent Gaussian white noise, which is used to simulate the random fluctuations and short-term uncertainties of hydrogen demand.

[0011] Furthermore, the state space also includes the current electrolyzer power, membrane electrode temperature, current hydrogen production rate, system efficiency, current electricity price information, and cumulative system operating time; the prediction time domain is multiple discrete time steps in the future, with each time step spaced 15 minutes apart.

[0012] Furthermore, the reward function is shown in the following equation:

[0013] in: For system energy efficiency, For the curtailment of wind and solar power, This is the temperature over-limit penalty function. This is the power fluctuation penalty function.

[0014] Furthermore, the offline training employs a proximal policy optimization algorithm to construct a deep neural network containing an Actor network and a Critic network. It uses generalized advantage estimation to calculate the advantage function and updates the network parameters by pruning the objective function until the cumulative reward converges.

[0015] Furthermore, the step of deploying the trained control strategy model into the controller of a real water electrolysis hydrogen production system includes: The trained model is lightweighted and deployed to a programmable logic controller or industrial control computer via industrial Ethernet or fieldbus protocol. A safety review module based on hard constraints is set before the command output to forcibly override action commands that cause the current or temperature to exceed the hard safety limit.

[0016] Furthermore, it also includes continuous learning steps: During the actual operation of the system, real state transition data is collected and stored in the experience playback buffer. When the amount of data in the buffer reaches a set threshold, the deployed model is fine-tuned online or periodically using a learning rate lower than that used for offline training, combined with historical offline training data, to adapt to equipment aging and changes in the operating environment.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention constructs a dynamic simulation environment for hydrogen production via water electrolysis by integrating electrochemical, thermodynamic, mass transfer, and performance degradation physical field models. Based on deep reinforcement learning, a multi-objective fusion reward function is designed to achieve coordinated optimization control of electrolyzer current and cooling water flow. This significantly improves the overall energy efficiency and renewable energy absorption rate of the system under conditions of fluctuating renewable energy power supply, effectively suppresses temperature exceedances and drastic power fluctuations, ensures operational safety, and extends equipment lifespan. Overall, it greatly enhances the dynamic response capability and intelligent control level of the water electrolysis hydrogen production system. Attached Figure Description

[0018] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. In the drawings: Figure 1 This is a schematic diagram of the overall process of the reinforcement learning-based water electrolysis hydrogen production control method in an embodiment of the present invention. Detailed Implementation

[0019] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the disclosure to those skilled in the art. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0020] This embodiment provides a reinforcement learning-based control method for hydrogen production via water electrolysis, applied to a proton exchange membrane (PEM) water electrolysis system. The system includes a PEM electrolyzer, power supply, water pump, gas-liquid separator, various sensors, and a programmable logic controller (PLC) or industrial control computer. The PLC or industrial control computer serves as the hardware platform for the intelligent agent, and contains a pre-trained reinforcement learning model.

[0021] The power source is connected to a renewable energy generation device, and the sensors include current, voltage, temperature, pressure, and hydrogen flow sensors.

[0022] like Figure 1 As shown, the method in this embodiment is mainly implemented through the following stages and steps: (I) System Modeling and Simulation Environment Construction Step S110: Establish electrochemical, thermodynamic, and mass transfer dynamic models of the PEM electrolyzer. Based on the equivalent circuit, energy conservation, and material conservation principles of the electrolyzer, calculate the dynamic response of hydrogen production rate, system efficiency, and temperature state variables of key components according to the input current, water supply temperature, and operating pressure. Simultaneously, consider performance degradation factors and simulate the irreversible voltage increase through the degradation factor accumulated over operating time and conditions.

[0023] Among them, the electrochemical sub-model is used to describe the voltage-current characteristics and hydrogen production rate of the electrolyzer, and the electrolyzer stack consists of... The cells are connected in series, and the operating voltage of each cell is... It is composed of the superposition of reversible voltage and various irreversible overpotentials, and is expressed as follows: Specifically, it is calculated using the following formula:

[0024] in, It is a reversible voltage, calculated according to the Nernst equation. and These are the activation overpotentials for the oxygen evolution at the anodic and hydrogen evolution at the cathode, respectively, simplified using the Tafel form. For Ohm overpotential, This represents the concentration overpotential.

[0025] The total voltage of the electrolytic cell stack is: .

[0026] The thermodynamic sub-model describes the core temperature of the electrolyzer based on the lumped heat capacity method. The dynamic changes are governed by the energy conservation equation:

[0027] in, For heat generation power, For environmental heat dissipation power, This is the cooling heat power.

[0028] The proton transfer model mainly describes the dynamic effect of water content in the membrane on membrane conductivity and the change in current efficiency, coupling electrochemistry with thermal state.

[0029] Membrane water content dynamics: The water content within the membrane is represented by λ, and its variation is affected by the difference in water activity across the membrane, electroosmotic drag, and reverse diffusion. To balance model accuracy and computational efficiency, this embodiment employs a first-order dynamic approximation:

[0030] in, The hydration time constant, To match the water activity on the anode side Temperature-related equilibrium moisture content It is approximately the ratio of the partial pressure of water vapor on the anode side to the saturated vapor pressure at the same temperature.

[0031] membrane conductivity It is λ and The strong function, commonly referred to by empirical formulas, is:

[0032] When insufficient water supply or excessively high temperature causes membrane dehydration decline, Increased voltage leads to increased heat generation, creating a vicious cycle.

[0033] Current efficiency The main influence is the cross-permeation loss of hydrogen through the membrane, which is related to the current density and membrane water content. The following empirical model is adopted:

[0034] in, The equivalent hydrogen permeation current density varies with λ and It increases with the rise in [something].

[0035] To simulate the irreversible voltage increase caused by membrane chemical degradation and catalyst aggregation during long-term operation of an electrolyzer, a cumulative decay factor D, in volts, is introduced and directly superimposed on the single-cell voltage to form a cell voltage with decay:

[0036] The evolution of the attenuation factor D is related to the cumulative effect of operating conditions, and is specifically described by the following differential equation or discrete iterative form:

[0037]

[0038] in, This represents the accumulated voltage decay up to time t. The reference decay rate constant is Here, γ is the reference current, and γ is the current acceleration exponent. The activation energy of the decay process. This refers to the membrane electrode temperature, which is output in real time using a thermodynamic model.

[0039] High temperature and high current density accelerate the chemical degradation of the membrane and the loss of catalyst active sites, leading to a slow increase in ohmic resistance and activation overpotential, macroscopically manifested as an irreversible increase in electrolysis voltage at the same current. In the simulation, ΔD is calculated based on the current and temperature in each control cycle, accumulated into D, and then applied to the electrochemical model, thus affecting subsequent calculations of voltage, heat generation, and efficiency. Since increased voltage further increases heat generation... This creates positive feedback, accurately simulating the phenomenon of accelerated performance degradation in the later stages of a real system. The decay factor D is further used as a health indicator in the state space, helping the reinforcement learning agent perceive the degree of device aging and tend to execute control strategies to mitigate the degradation.

[0040] Using the modeling methods described above, the dynamic model can simulate the dynamic behavior of PEM electrolyzers under variable operating conditions, including electrical, thermal, and mass multi-physics fields, as well as the long-term performance degradation process, with high fidelity. This provides a sufficiently realistic and generalizable training environment for reinforcement learning agents.

[0041] Step S120: Establish a renewable energy power generation forecasting model and a hydrogen demand model. The power generation forecasting model generates photovoltaic and wind power generation curves at 15-minute intervals for the next 24 hours based on historical meteorological data and weather forecast information. The hydrogen demand model generates a baseline demand curve based on downstream hydrogen consumption plans. These two models constitute the external environment input for the reinforcement learning agent to interact with.

[0042] Specifically, for photovoltaic power generation, the physical model is based on the single-diode equivalent circuit of a photovoltaic cell, and the relationship between output power and irradiance and temperature is expressed as follows:

[0043] in, For component efficiency under standard test conditions, For the array area, The irradiance of the inclined plane, The power temperature coefficient is estimated empirically from ambient temperature and wind speed. The irradiance and meteorological parameters required by the model are derived from numerical weather prediction data and converted into an initial power prediction sequence through a physical model.

[0044] For wind power generation, the turbine power curve method is used. The standard power curve describes the nonlinear relationship between wind speed and output power, covering the power-up region and the constant power region between the cut-in wind speed, rated wind speed, and cut-out wind speed. First, the wind speed forecast value provided by numerical weather prediction is corrected to the hub height. Then, the power of a single turbine is obtained by interpolating the power curve, and multiplying it by the number of turbines gives the total power of the wind farm. .

[0045] Due to simplification errors in physical models and inherent uncertainties in weather forecasts, a Long Short-Term Memory (LSTM) neural network is introduced as a residual correction module based on the aforementioned physical predictions. The inputs include the preliminary power sequence output by the physical model, the numerical weather prediction feature vector for the corresponding time period, and the historical measured sequence of recent actual power generation. The network's training objective is to learn the residual relationship between the physical model's predicted values ​​and the measured values, and output the corrected residual sequence. Final predicted power:

[0046] in, Output for the physical model. This is the correction value for the LSTM output.

[0047] Step S130: Integrate the above electrochemical model, power generation prediction model, and hydrogen demand model into a virtual simulation environment. The state transition function of this environment simulates how the system state evolves to the next moment after receiving a control action; the reward function is the total operating benefit of the system minus the penalty term, and the specific formula will be detailed below.

[0048] This simulation environment is used for offline pre-training of reinforcement learning models, avoiding the safety risks and efficiency losses caused by trial and error on real devices.

[0049] The hydrogen demand model employs a baseline-plus-perturbation approach to generate a demand curve that reflects downstream hydrogen consumption plans and includes stochastic factors. First, a daily baseline load curve is determined. Two layers of perturbations are then superimposed on this baseline curve to simulate the random fluctuations and short-term uncertainties of real demand. One of these perturbations is a low-frequency fluctuation, which is represented by a first-order autoregressive model.

[0050] The first is the generation of slowly changing trend bias; the second is high-frequency noise, which is directly superimposed with small-variance independent Gaussian white noise, so that the synthesized demand sequence follows the hydrogen use plan, while reflecting the characteristics of short-term small adjustments.

[0051] The demand sequence output by the model is input into the state space of reinforcement learning, enabling the agent to perceive the future trend of hydrogen demand changes in advance.

[0052] (II) Reinforcement Learning Framework Design Step S210: Design of the state space, the states observed by the agent at time step t include, but are not limited to: Current electrolytic cell power

[0053] membrane electrode temperature ; Current hydrogen production rate ; System efficiency ; Predicted values ​​of available renewable energy power generation within a future time domain [t+1, t+N]. ; Future hydrogen demand forecasts within the same forecast time domain ; Current electricity price information ; The system's cumulative uptime is used to reflect the health status of the equipment.

[0054] Step S220: Design of the action space, actions output by the agent. These are continuous value vectors, mapped to control commands that can be directly applied to the physical system, including: Adjustment increment of electrolytic cell input current Adjusting the cooling water flow rate by increment .

[0055] By sending incremental commands, the control process is naturally smoothed, reducing equipment impact. The operating range is within the safe operating range of the electrolyzer, ensuring that the current does not fall below the minimum sustaining current, nor does it exceed the maximum current determined by the power supply and membrane electrode tolerance, while ensuring that the temperature is always below the safe threshold.

[0056] Step S230: Design of the reward function. To guide the agent to learn the optimal strategy that balances efficiency, safety, and device lifespan, this embodiment designs a multi-objective fusion reward function:

[0057] in: System energy efficiency is expressed as hydrogen production per unit of energy consumption. For the curtailment of wind and solar power, Here is the temperature over-limit penalty function, when When the temperature is below the lower limit of the optimal temperature range or above the upper limit, this function generates a sharply increased penalty value to ensure thermal safety. This is a power fluctuation penalty function, proportional to the absolute value of the current increment |ΔI|, used to suppress drastic and frequent changes in current and extend the service life of the electrolytic cell.

[0058] (III) Offline Training Step S310: In this embodiment, the Proximal Policy Optimization (PPO) algorithm is selected as the reinforcement learning algorithm. A deep neural network containing an actor network and a critic network is constructed. The input to the actor network is the state. The output is the mean of the Gaussian policy. and logarithmic standard deviation Used for sampling continuous actions The input to the Critic network is the state. The output is a value estimate for that state. .

[0059] Step S320: In the simulation environment constructed in step S130, the PPO agent is iteratively trained using a historical annual dataset of renewable energy. In each training round, the agent starts from a random initial state, interacts with the simulation environment according to the current policy, collects trajectories containing state, action, and reward, and calculates the advantage function using the Generalized Advantage Estimation (GAE) algorithm. The specific steps are as follows: The agent runs the current policy in a simulation environment and collects trajectories of length T:

[0060] Simultaneously, the value predicted by the Critic network for each state is recorded. .

[0061] Traverse the entire trajectory and calculate for each time step t:

[0062] in, It's an instant reward. It is a discount factor. It is the current state value estimate output by the Critic network; Recursively calculate backwards from the last time step T-1 of the trajectory:

[0063] For t=T-2,T-3,...,0, recursively calculate:

[0064] resulting sequence Estimating the advantage for each transfer.

[0065] To subsequently update the Critic network, the objective value of the value function needs to be constructed. :

[0066] Based on the PPO pruning objective function and mean squared error value loss function, the network parameters are updated using mini-batch stochastic gradient descent:

[0067]

[0068] Continue training until the cumulative reward curve converges.

[0069] (iv) Online Deployment and Continuous Learning Phase Step S410: The PPO model that has converged through offline training in step S320 is lightweighted and deployed to the PLC or industrial control computer in the actual hydrogen production system via Ethernet or fieldbus protocol, receiving real-time sensor data and prediction information from the data acquisition and monitoring control system as status input.

[0070] Step S420: In each control cycle of model inference, the agent determines the current observed state. Output Action The controller calculates the absolute control command for the next moment based on the current increment ΔI and the cooling water flow increment ΔF. This command, along with the cooling water flow command, is sent to the actuator. To ensure absolute safety, a hard-constraint-based safety review module is implemented before the final command output. This module enforces the coverage of any events that could cause I or ΔF to fail. Instructions that exceed the hard safety limits ensure inherent safety.

[0071] Step S4410: Continuous Learning During the actual operation of the system, the real state transition data collected after the agent performs an action will be used. The model continuously stores data in an experience replay buffer. When the buffer reaches a set threshold, fine-tuning training can be triggered online or periodically. During fine-tuning, a smaller learning rate can be used, and the model can be trained together with historical datasets from offline training to prevent it from forgetting learned core strategies and to adapt to the actual characteristics of the field, automatically adapting to equipment aging and changes in the operating environment.

[0072] In summary, this invention constructs a highly realistic electrolyzer model, designs a multi-objective reward function oriented towards energy efficiency, absorption, and durability, and employs a deep reinforcement learning algorithm for training and deployment. Ultimately, it achieves adaptive control of the water electrolysis hydrogen production system under dynamic renewable energy power supply, breaking the limitations of traditional fixed-rule or single-objective optimization models and significantly improving the overall operating performance and long-term economic efficiency of the system.

[0073] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A reinforcement learning-based method for controlling hydrogen production via water electrolysis, characterized in that, include: A dynamic simulation environment for a water electrolysis hydrogen production system is constructed. The dynamic simulation environment includes electrochemical, thermodynamic and mass transfer dynamic models of the electrolyzer, as well as renewable energy power generation prediction models and hydrogen demand models. The design includes a state space, action space, and reward function for the reinforcement learning agent. The state space includes the current system operating state and the predicted power generation and hydrogen demand in the future prediction time domain. The action space includes the incremental adjustment of the electrolyzer input current and the incremental adjustment of the cooling water flow rate. The reward function is a multi-objective fusion function that takes into account system energy efficiency, renewable energy consumption, thermal safety, and power fluctuation suppression. The reinforcement learning agent is trained offline in the simulation environment to obtain a control policy model. The trained control strategy model is deployed to the controller of a real water electrolysis hydrogen production system. In each control cycle, action commands are output based on the real-time observed status, and the action commands are executed to control the operation of the electrolyzer.

2. The method according to claim 1, characterized in that, The electrochemical, thermodynamic, and mass transfer dynamics models of the electrolyzer include: An electrolytic cell voltage-current characteristic model is established based on the equivalent circuit and Nernst equation, wherein the voltage includes reversible voltage, activation overpotential, ohmic overpotential and concentration overpotential. A dynamic temperature model for the electrolytic cell is established based on the lumped heat capacity method. The dynamic temperature model calculates temperature changes based on heat generation power, heat dissipation power, and cooling power. A membrane conductivity model is dynamically established based on the water content inside the membrane. The membrane conductivity model calculates the membrane conductivity based on the water content and temperature of the membrane and affects the ohmic overpotential. A cumulative degradation factor is introduced, which increases cumulatively over time based on the operating current and temperature, and is superimposed on the single cell voltage of the electrolyzer to simulate long-term performance degradation.

3. The method according to claim 1, characterized in that, The renewable energy power generation prediction model includes: A preliminary power prediction sequence was generated based on the photovoltaic physical model and the wind power generation power curve model; A long short-term memory neural network is used as the residual correction module. The input is the preliminary power sequence output by the physical model, the numerical weather prediction feature vector, and the historical measured power sequence. The output is the corrected residual. The corrected residual is superimposed on the preliminary power prediction sequence to obtain the final predicted power.

4. The method according to claim 1, characterized in that, The hydrogen demand model is generated by superimposing low-frequency fluctuations and high-frequency noise onto a baseline load curve. The low-frequency fluctuations adopt a first-order autoregressive model, and the high-frequency noise is independent Gaussian white noise, which is used to simulate the random fluctuations and short-term uncertainties of hydrogen demand.

5. The method according to claim 1, characterized in that, The state space also includes the current electrolyzer power, membrane electrode temperature, current hydrogen production rate, system efficiency, current electricity price information, and cumulative system operating time; the prediction time domain is multiple discrete time steps in the future, with each time step spaced 15 minutes apart.

6. The method according to claim 1, characterized in that, The reward function is shown in the following formula: , in: For system energy efficiency, For the curtailment of wind and solar power, This is the penalty function for exceeding the temperature limit. This is the power fluctuation penalty function.

7. The method according to claim 1, characterized in that, The offline training employs a proximal policy optimization algorithm to construct a deep neural network containing an Actor network and a Critic network. It uses generalized advantage estimation to calculate the advantage function and updates the network parameters by pruning the objective function until the cumulative reward converges.

8. The method according to claim 1, characterized in that, The step of deploying the trained control strategy model to the controller of a real water electrolysis hydrogen production system includes: The trained model is lightweighted and deployed to a programmable logic controller or industrial control computer via industrial Ethernet or fieldbus protocol. A safety review module based on hard constraints is set before the command output to forcibly override action commands that cause the current or temperature to exceed the hard safety limit.

9. The method according to claim 1, characterized in that, It also includes continuous learning steps: During the actual operation of the system, real state transition data is collected and stored in the experience playback buffer. When the amount of data in the buffer reaches a set threshold, the deployed model is fine-tuned online or periodically using a learning rate lower than that used for offline training, combined with historical offline training data, to adapt to equipment aging and changes in the operating environment.

10. A water electrolysis hydrogen production control system based on reinforcement learning, characterized in that, The system executes the reinforcement learning-based water electrolysis hydrogen production control method according to any one of claims 1-9.