Start-stop planning and scheduling method and system for electrolytic cell

By optimizing the start-stop planning and scheduling of the electrolyzer through deep reinforcement learning intelligent agents, the problems of power regulation rate limitation and long start-stop time of traditional electrolyzers are solved, and the rapid response and efficient coordinated control of the water electrolysis hydrogen production system are achieved, thereby improving the stability and economic benefits of the system.

CN120683561APending Publication Date: 2025-09-23DATANG (INNER MONGOLIA) ENERGY DEV CO LTD +4
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510782806.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The power regulation rate limitations, long start-up and stop times, and coordinated control issues of multiple groups of electrolyzers in traditional electrolyzers result in slow response speed and low economic benefits of water electrolysis hydrogen production systems.

Method used

A deep reinforcement learning agent is used to design an electrolyzer start-stop planning and scheduling method based on the deep Q-network algorithm. By constructing an environmental model, defining the state and action space, and designing a reward function, the start-stop operation and power regulation strategy of the electrolyzer are optimized, and the start-stop priority and time are dynamically adjusted to achieve coordinated scheduling of multiple groups of electrolyzers.

Benefits of technology

It improves the dynamic response capability of the electrolytic cell system, reduces power waste, shortens start-up and stop time, realizes efficient coordinated control of multiple groups of electrolytic cells, and improves the system's response speed and economic benefits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120683561A_ABST
    Figure CN120683561A_ABST
Patent Text Reader

Abstract

The invention discloses an electrolytic bath start-stop planning and scheduling method and system, and the method comprises the steps: constructing an electrolytic bath environment model, and obtaining the operation state, power input, hydrogen output and future power demand prediction of an electrolytic bath; based on the deep Q network, the intelligent agent decides start-stop operation of the electrolytic cell and optimizes the power regulation rate according to the current state and future power demand prediction; defining a state space and an action space, designing a reward function, updating a Q value through Q learning, and optimizing an agent decision strategy; the start-stop priority is dynamically adjusted according to the running state, historical data and power adjusting rate of the electrolytic cell; calculating the starting and stopping time of the electrolytic cell according to the starting and stopping decision and the power regulation rate; outputting the state plan of each electrolytic cell, the total load adjustable upper and lower limits and the power adjusting plan; the problems of power regulation rate limitation, long start-stop time and coordinated control of multiple groups of electrolytic cells are solved, and the start-stop response speed, the economic benefit and the stability of the electrolytic cells are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hydrogen production by electrolysis of water, and more particularly to a method and system for planning and scheduling the start and stop of an electrolyzer. Background Art

[0002] Electrolyzers, key equipment for hydrogen production, split water into oxygen and hydrogen through an electrochemical reaction. In recent years, with the growing global demand for clean energy, water electrolysis hydrogen production technology has attracted widespread attention due to its environmentally friendly and sustainable features.

[0003] However, the operation and management of traditional electrolyzers face many challenges, including but not limited to power regulation rate limitations, long start-up and stop times, and complex coordinated control issues for multiple groups of electrolyzers.

[0004] First, the electrolyzer's power regulation rate directly affects the system's response speed. When grid load or renewable energy generation fluctuates dramatically, if the electrolyzer fails to quickly adjust its input power, it may waste power or cause system instability. Specifically, the electrolyzer's maximum power increase and decrease rates limit the maximum power that can be increased or decreased per unit time. This constraint requires us to fully consider the dynamic characteristics of power changes when designing start-up and shutdown plans to avoid operations that exceed the allowable range.

[0005] Secondly, the time required for the electrolyzer to start up and fully enter the working state, as well as the time from the issuance of the stop command to complete shutdown, both involve multiple stages of physical processes, such as heating and pressurization. These processes are not only time-consuming but also consume additional energy. Therefore, when formulating the start-stop plan, the specific start-up and stop characteristics of each electrolyzer must be taken into account, and the start-up and stop sequence must be reasonably arranged to minimize energy loss and waiting time.

[0006] Furthermore, a large-scale hydrogen production system may include multiple types of electrolyzer combinations, such as a one-to-one electrolyzer and a one-to-four electrolyzer. Each type of electrolyzer has different power requirements and regulation characteristics. Effectively managing and coordinating these different types of electrolyzers becomes a complex but necessary task. For example, in a one-to-four electrolyzer group, not only the power regulation rate of each individual electrolyzer unit must be considered, but also the maximum increase and decrease power regulation rate of the entire group. This requires comprehensive consideration of the status of each electrolyzer unit and their mutual influence when formulating a scheduling strategy to ensure the smooth operation of the entire system.

[0007] Therefore, how to solve the problems existing in the existing electrolyzer start-stop planning and scheduling, such as power regulation rate limitation, long start-up and stop time, and coordinated control of multiple groups of electrolyzers, so as to improve the response speed and economic benefits of the hydrogen production system, is an issue that technical personnel in this field urgently need to solve. Summary of the Invention

[0008] In view of this, the present invention provides an electrolytic cell start-stop planning and scheduling method and system to solve some of the technical problems mentioned in the background technology.

[0009] In order to achieve the above object, the present invention adopts the following technical solutions:

[0010] A method for planning and scheduling the start and stop of an electrolytic cell, comprising the following steps:

[0011] S1. Build an electrolyzer environmental model to obtain the electrolyzer's operating status, power input, hydrogen output, and power demand forecast for the next 24 hours;

[0012] S2. Design a deep reinforcement learning agent based on a deep Q-network algorithm. This agent will make decisions about starting and stopping the electrolyzer and optimize the power regulation rate based on the current state and future power demand forecasts.

[0013] S3. Define the state space and action space of the electrolyzer, design a reward function, and use the Q-learning algorithm to update the Q value and optimize the agent's decision-making strategy.

[0014] S4. Dynamically adjust the start and stop priority of the electrolyzer according to the operating status, historical data and power regulation rate of the electrolyzer;

[0015] S5. Calculate the start and stop times of the electrolyzer based on the start and stop decisions and power regulation rate to ensure the timing of the start and stop operations is reasonable;

[0016] S6. Output the scheduling results, including the status planning of each electrolytic cell in the next 24 hours, the upper and lower limits of the total load adjustment, and the power adjustment plan.

[0017] Preferably, the state space of the electrolytic cell includes the current total power of the electrolytic cell, the state of each electrolytic cell, the future power demand forecast, the historical operation data of each electrolytic cell, the power regulation rate and the start and stop time characteristics.

[0018] Preferably, the action space of the electrolytic cell includes starting the electrolytic cell, stopping the electrolytic cell, maintaining the current state and adjusting the power.

[0019] Preferably, the reward function calculates the immediate reward based on whether the total power of the electrolyzer meets the power demand, whether the power regulation rate is within the allowable range, the frequency of start-stop operations, and the electrolyzer fault conditions.

[0020] Preferably, the deep reinforcement learning agent adopts the deep Q network DQN algorithm, and the Q value update formula is:

[0021]

[0022] Among them, α is the learning rate, γ is the discount factor, and r t+1 For instant rewards, is the maximum Q value of the next state, Q(S t+1 ,a) is the current state Q value.

[0023] Preferably, the state space of the electrolyzer is:

[0024] S={P t ,S state ,D t+1 ,D t+2 ,...,D t+24 ,T run , N start / stop ,T standby ,R adj ,T start ,T stop}

[0025] Among them, P t is the total power of the electrolytic cell at the current moment, S state is the state of each electrolytic cell, D t+i is the electricity demand forecast for the next hour i, T run is the total operating time of the electrolyzer, N start / stop is the number of starts and stops of the electrolytic cell, T standby is the standby time of the electrolytic cell, R adj is the power regulation rate, T start is the start-up time characteristic of the electrolytic cell, T stop is the stopping time characteristic of the electrolytic cell;

[0026] The action space of the electrolytic cell is:

[0027] A={a start , a stop , a hold , a adjust}

[0028] Among them, a start Indicates starting the electrolytic cell, a stop Indicates stopping the electrolytic cell, a hold Indicates maintaining the current state, a adjust Indicates power adjustment.

[0029] Preferably, the reward function is designed as follows:

[0030]

[0031] Preferably, the specific method for dynamically adjusting the start and stop priority of the electrolytic cell is:

[0032] F1=w1·Trun +w2·N start / stop +w3·T standby +w6·R adj

[0033] F2=w4·P fluctuation +w5·T group_run +w7·R adj

[0034] Among them, F1 is the start-stop priority, F2 is the load adjustment priority; ω1, ω2, ω3, ω4, ω5, ω6, ω7 are weight coefficients, and the value range is [0,1]. fluctuation T is the operating time of the electrolyzer under fluctuating power, group_run is the operating time of the electrolytic cell group, R adj is the power regulation rate.

[0035] Preferably, the start-up time of the electrolyzer is:

[0036]

[0037] The stop time of the electrolyzer is:

[0038]

[0039] Among them, P target is the target power, P current is the current power, R adj is the power regulation rate.

[0040] An electrolytic cell start-stop planning and scheduling system, based on the electrolytic cell start-stop planning and scheduling method, comprises: an electrolytic cell environment acquisition module, an electrolytic cell start-stop control intelligent agent, an electrolytic cell priority module, an electrolytic cell countdown module, and a scheduling result output module;

[0041] The electrolyzer environment acquisition module is used to build an electrolyzer environment model to obtain the electrolyzer's operating status, power input, hydrogen output, and power demand forecast for the next 24 hours;

[0042] The electrolyzer start-stop control agent is used to design a deep reinforcement learning agent. Based on the deep Q-network algorithm, the agent decides on the start and stop operation of the electrolyzer based on the current state and future power demand forecast, and optimizes the power regulation rate. It also defines the state space and action space of the electrolyzer, designs a reward function, updates the Q value through the Q-learning algorithm, and optimizes the agent's decision-making strategy.

[0043] The electrolyzer priority module is used to dynamically adjust the start and stop priorities of the electrolyzer according to the operating status, historical data and power regulation rate of the electrolyzer;

[0044] The electrolyzer countdown module is used to calculate the start and stop time of the electrolyzer based on the start and stop decision and power regulation rate to ensure the timing rationality of the start and stop operations;

[0045] The scheduling result output module is used to output the scheduling results, including the status planning of each electrolytic cell in the next 24 hours, the upper and lower limits of the total load adjustment, and the power adjustment plan.

[0046] As can be seen from the above technical solutions, compared with the prior art, the present invention provides a method and system for planning and scheduling the start and stop of electrolytic cells. Through the interaction between the intelligent agent and the environment, the start and stop strategies and power regulation plans of the electrolytic cells are dynamically optimized. This solves the problems of traditional methods in terms of power regulation rate limitation, long start and stop times, and coordinated control of multiple groups of electrolytic cells, significantly improving the system's response speed, economic benefits, and stability. Specifically:

[0047] By dynamically optimizing the electrolyzer's power regulation strategy through a deep reinforcement learning agent and combining it with constraints on the power regulation rate, the system ensures that the electrolyzer can quickly respond to changes in power demand within the permitted power regulation rate. This significantly improves the electrolyzer system's dynamic response capability and reduces power waste. When grid load or renewable energy generation fluctuates dramatically, the system can quickly adjust the electrolyzer's input power to ensure stable system operation.

[0048] The deep reinforcement learning agent comprehensively considers the start-up and stop-time characteristics of the electrolyzer, dynamically optimizes the start-stop sequence and timing, and reduces energy loss during the start-up and stop process of the electrolyzer. By rationally arranging the start-stop sequence, waiting time is minimized and the system operation efficiency is improved.

[0049] Through deep reinforcement learning agents, the power regulation rate, start-stop time characteristics and operating status of different types of electrolytic cells are comprehensively considered to achieve coordinated scheduling of multiple groups of electrolytic cells, and efficient coordinated control of multiple groups of electrolytic cells is achieved to ensure the smooth operation of the entire system. By dynamically adjusting the power distribution and start-stop sequence of each electrolytic cell, the overall operating efficiency of the system is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0051] Figure 1 A schematic diagram of a method for planning and scheduling the start and stop of an electrolytic cell provided by the present invention;

[0052] Figure 2 Schematic diagram of the Deep Q Network algorithm agent provided by the present invention;

[0053] Figure 3 This is a schematic diagram of an electrolytic cell start-stop planning and scheduling system provided by the present invention. DETAILED DESCRIPTION

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0055] An embodiment of the present invention discloses a method for planning and scheduling the start and stop of an electrolytic cell, comprising the following steps:

[0056] S1. Build an electrolyzer environmental model to obtain the electrolyzer's operating status, power input, hydrogen output, and power demand forecast for the next 24 hours;

[0057] S2. Design a deep reinforcement learning agent based on a deep Q-network algorithm. This agent will make decisions about starting and stopping the electrolyzer and optimize the power regulation rate based on the current state and future power demand forecasts.

[0058] S3. Define the state space and action space of the electrolyzer, design a reward function, and use the Q-learning algorithm to update the Q value and optimize the agent's decision-making strategy.

[0059] S4. Dynamically adjust the start and stop priority of the electrolyzer according to the operating status, historical data and power regulation rate of the electrolyzer;

[0060] S5. Calculate the start and stop times of the electrolyzer based on the start and stop decisions and power regulation rate to ensure the timing of the start and stop operations is reasonable;

[0061] S6. Output the scheduling results, including the status planning of each electrolytic cell in the next 24 hours, the upper and lower limits of the total load adjustment, and the power adjustment plan.

[0062] In step S1, the electrolytic cell environment model includes the following parameters: the total power of the electrolytic cells at the current time t, the status of each electrolytic cell, including the operating status, hot standby status, cold standby status and fault status, the power demand forecast for the next 24 hours, the historical operating data of each electrolytic cell, including the total operating time, the number of starts and stops and the standby time, the power regulation rate, including the maximum rising power regulation rate and the maximum falling power regulation rate, the start time characteristics and the stop time characteristics.

[0063] In order to further implement the above technical solution, the state space of the electrolytic cell includes the current total power of the electrolytic cell, the status of each electrolytic cell, the future power demand forecast, the historical operation data of each electrolytic cell, the power regulation rate and the start and stop time characteristics.

[0064] In order to further implement the above technical solution, the action space of the electrolytic cell includes starting the electrolytic cell, stopping the electrolytic cell, maintaining the current state and adjusting the power.

[0065] In order to further implement the above technical solution, the reward function is to calculate the immediate reward based on whether the total power of the electrolyzer meets the power demand, whether the power regulation rate is within the allowable range, the frequency of start and stop operations, and the electrolyzer fault conditions.

[0066] In order to further implement the above technical solutions, the deep reinforcement learning agent adopts the deep Q network DQN algorithm, and the Q value update formula is:

[0067]

[0068] Among them, α is the learning rate, the range is (0,1], γ is the discount factor, the range is (0,1], r t+1 For instant rewards, is the maximum Q value of the next state, Q(S t+1 ,a) is the current state Q value.

[0069] In order to further implement the above technical solution, the state space of the electrolytic cell is:

[0070] S={P t ,S state ,D t+1 ,D t+2 ,...,D t+24 , T run ,N start / stop ,T standby ,R adj ,T start ,T stop}

[0071] Among them, P t is the total power of the electrolytic cell at the current moment, S state is the state of each electrolytic cell, D t+i is the electricity demand forecast for the next hour i, T run is the total operating time of the electrolyzer, N start / stop is the number of starts and stops of the electrolytic cell, T standby is the standby time of the electrolytic cell, R adj is the power regulation rate, T start is the start-up time characteristic of the electrolytic cell, T stop is the stopping time characteristic of the electrolytic cell;

[0072] The action space of the electrolytic cell is:

[0073] A={a start , a stop , a hold , a adjust}

[0074] Among them, a start Indicates starting the electrolytic cell, switching the electrolytic cell in cold standby or hot standby state to running state, a stop Indicates stopping the electrolytic cell and switching the running electrolytic cell to cold standby or hot standby state. hold Indicates maintaining the current state, a adjust Indicates power adjustment.

[0075] In order to further implement the above technical solution, the reward function designed is as follows:

[0076]

[0077] In this implementation, the network structure of the DQN agent includes an input layer, a hidden layer, and an output layer;

[0078] The input layer receives information from the state space S, including: the total power P of the electrolyzer at the current moment t , the state S of each electrolytic cell state , the power demand forecast D for the next hour i t+i , the total operating time of the electrolyzer T run , the number of starts and stops of the electrolytic cell N start / stop , the standby time of the electrolyzer is T standby , power regulation rate R adj , the start-up time characteristic T of the electrolytic cell start , the stopping time characteristic T of the electrolytic cell stop ;

[0079] The hidden layer consists of multiple fully connected layers, which are used to extract the features of the state space. The activation function of the hidden layer is ReLU.

[0080] ReLU(x)=max(0,x);

[0081] The number of hidden layers and the number of neurons in each layer can be adjusted according to the specific problem;

[0082] The output layer outputs the Q value Q(s,a) of each action a, whose dimension is the dimension of the action space A, and the activation function of the output layer is a linear function.

[0083] The training process of the DQN agent is as follows:

[0084] DQN updates the Q value through the Q-learning algorithm Q value update formula;

[0085] DQN trains the neural network by minimizing the loss function, which is defined as:

[0086]

[0087] Among them, θ is the parameter of the current neural network, θ - is the parameter of the target network, and E represents the expected value;

[0088] The parameters θ of the neural network are updated by the gradient descent method, and the update formula is:

[0089]

[0090] Where η is the learning rate, is the gradient of the loss function with respect to the parameter θ;

[0091] Experience replay mechanism:

[0092] The experience of each time step is stored in the replay buffer D. During training, a batch of experience is randomly sampled from the replay buffer D, and the sampled experience is used to batch update the parameters of the neural network;

[0093] Target network update strategy:

[0094] The parameters θ of the target network - The update formula is:

[0095] θ-θ

[0096] C is the target network update frequency, which is usually set to a large value (such as 1000 steps);

[0097] The DQN algorithm agent approximates the Q-value function through a deep neural network, combines the experience replay mechanism and the target network update strategy, and realizes dynamic decision-making for the start-stop planning and scheduling of the electrolyzer. The method of the present invention can effectively cope with fluctuations in electricity demand and improve the response speed and economic benefits of the hydrogen production system.

[0098] In order to further implement the above technical solution, the specific method of dynamically adjusting the start and stop priority of the electrolyzer is as follows:

[0099] F1=w1·T run +w2·N start / stop +w3·T standby +w6·R adj

[0100] F2=w4·P fluctuation +w5·T group_run +w7·Radj

[0101] Among them, F1 is the start-stop priority, F2 is the load adjustment priority; ω1, ω2, ω3, ω4, ω5, ω6, ω7 are weight coefficients, and the value range is [0,1]. fluctuation T is the operating time of the electrolyzer under fluctuating power, group_run is the operating time of the electrolytic cell group, R adj is the power regulation rate.

[0102] In actual applications, the startup priority sequence is determined based on the value of F1. When determining the sequence, the electrolytic cell grouping situation needs to be taken into consideration, and the priority of the largest group is followed. For example, when the number of cells added or reduced is greater than or equal to 2, two cells in the same group cannot be started or stopped at the same time. When the first cell and the second cell are started or stopped successively, the time interval between them can be configured through the interface. Based on the value of F2, the overall operating status of the largest group is first evaluated. According to the total operating time of the group and the operating time under fluctuating power, the group is determined to have priority for power adjustment. After the group is determined, the specific operating parameters of the individual electrolytic cells are further refined to finally determine which electrolytic cell should have priority for load adjustment.

[0103] During operation, the F1 priority is used to determine the start and stop sequence of the electrolyzers, and the F2 priority is used to adjust the load distribution. The system can flexibly adjust the start and stop strategy according to load fluctuations under different operating conditions, ensuring that the system can respond quickly when the total load fluctuates significantly. By balancing the load among different electrolyzers, stable and efficient operation is achieved.

[0104] In order to further implement the above technical solution, the start-up time of the electrolyzer is:

[0105]

[0106] The stop time of the electrolyzer is:

[0107]

[0108] Among them, P target is the target power, P current is the current power, R adj is the power regulation rate.

[0109] In practical applications, the start and stop time of the electrolytic cell includes the countdown for starting the electrolytic cell with added cells and the countdown for stopping the electrolytic cell with reduced cells;

[0110] Electrolytic cell filling start countdown: predict the time T at a certain moment in the future through the electrolytic cell input power prediction curve i The number of slots to be increased is determined firstly whether it is in the prohibited start cycle. If so, Ti The start and stop planning of the electrolytic cell is not done at this moment. If not, there are two situations. The first one is that if one electrolytic cell needs to be added at time Ti, the priority module is used to determine which one needs to be added. Then, the start time △T is obtained based on the start characteristics of the electrolytic cell that needs to be started. It can be obtained that the time when the electrolytic cell start instruction is issued is T s =T i -△T; The second method is that if i (i≥2) electrolytic cells need to be added at time Ti, the priority module is used to determine which cells should be added in order, and the start-up time △T of the electrolytic cells that need to be started is obtained respectively. i , it is necessary to calculate the time when the start command of different slots is T s , compare any two starting times T si and T sm If the two deviations are within Tm, the second slot start time is postponed to the time after the first slot starts plus T m (T m configurable via the interface).

[0111] Electrolytic cell reduction stop countdown: predict the time T at a certain moment in the future through the electrolytic cell input power prediction curve i The number of slots to be increased, first determine T i Is it in the prohibited start cycle? If yes, T i Do not make the start and stop plan of the electrolyzer at all times. If not, there are two situations to analyze. The first one is if T i At this moment, one electrolytic cell needs to be reduced. The priority module is used to determine which cell needs to be reduced. The stop time △T′ is obtained according to the stop characteristics, and the time when the electrolytic cell stop instruction is issued is T s =T i ; The second type if T i At the moment, i (i≥2) electrolytic cells need to be reduced. The priority module is used to determine which cells need to be reduced in order, and the start-up time △T of the electrolytic cells that need to be stopped is obtained respectively. i , it is necessary to calculate the time when the stop instruction of different slots is T s , compare any two starting times T si and T sm , if the two deviations are in T m If the second stop slot time is within, the second stop slot time is postponed to the time after the first stop slot plus T m (T m configurable via the interface).

[0112] An electrolytic cell start-stop planning and scheduling system, based on an electrolytic cell start-stop planning and scheduling method, includes: an electrolytic cell environment acquisition module, an electrolytic cell start-stop control intelligent agent, an electrolytic cell priority module, an electrolytic cell countdown module, and a scheduling result output module;

[0113] The electrolyzer environment acquisition module is used to build an electrolyzer environment model to obtain the electrolyzer's operating status, power input, hydrogen output, and power demand forecast for the next 24 hours;

[0114] The electrolyzer start-stop control agent is used to design a deep reinforcement learning agent. Based on the deep Q-network algorithm, the agent decides on the start and stop operation of the electrolyzer based on the current state and future power demand forecast, and optimizes the power regulation rate. It also defines the state space and action space of the electrolyzer, designs a reward function, updates the Q value through the Q-learning algorithm, and optimizes the agent's decision-making strategy.

[0115] The electrolyzer priority module is used to dynamically adjust the start and stop priorities of the electrolyzer according to the operating status, historical data and power regulation rate of the electrolyzer;

[0116] The electrolyzer countdown module is used to calculate the start and stop time of the electrolyzer based on the start and stop decision and power regulation rate to ensure the timing rationality of the start and stop operations;

[0117] The scheduling result output module is used to output the scheduling results, including the status planning of each electrolytic cell in the next 24 hours, the upper and lower limits of the total load adjustment, and the power adjustment plan.

[0118] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0119] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for scheduling the start and stop of an electrolytic cell, characterized in that: The following steps are involved: S1. Build an electrolyzer environmental model to obtain the electrolyzer's operating status, power input, hydrogen output, and power demand forecast for the next 24 hours; S2. Design a deep reinforcement learning agent based on a deep Q-network algorithm. This agent will make decisions about starting and stopping the electrolyzer and optimize the power regulation rate based on the current state and future power demand forecasts. S3. Define the state space and action space of the electrolyzer, design a reward function, and use the Q-learning algorithm to update the Q value and optimize the agent's decision-making strategy. S4. Dynamically adjust the start and stop priority of the electrolyzer according to the operating status, historical data and power regulation rate of the electrolyzer; S5. Calculate the start and stop times of the electrolyzer based on the start and stop decisions and power regulation rate to ensure the timing of the start and stop operations is reasonable; S6. Output the scheduling results, including the status planning of each electrolytic cell in the next 24 hours, the upper and lower limits of the total load adjustment, and the power adjustment plan.

2. The electrolytic cell start-stop planning and scheduling method according to claim 1, characterized in that: The state space of the electrolyzer includes the current total power of the electrolyzer, the status of each electrolyzer, the future power demand forecast, the historical operation data of each electrolyzer, the power regulation rate, and the start and stop time characteristics.

3. The electrolytic cell start-stop planning and scheduling method according to claim 1, characterized in that: The action space of the electrolyzer includes starting the electrolyzer, stopping the electrolyzer, maintaining the current state, and adjusting the power.

4. The electrolytic cell start-stop planning and scheduling method according to claim 1, characterized in that: The reward function calculates the immediate reward based on whether the total power of the electrolyzer meets the power demand, whether the power regulation rate is within the allowable range, the frequency of start and stop operations, and the electrolyzer fault conditions.

5. The method for scheduling the start and stop of an electrolytic cell according to claim 1, wherein: The deep reinforcement learning agent uses the deep Q network DQN algorithm, and the Q value update formula is: Among them, α is the learning rate, γ is the discount factor, and r t+1 For instant rewards, is the maximum Q value of the next state, Q(S t+1 ,a) is the current state Q value.

6. The electrolytic cell start-stop planning and scheduling method according to claims 2 and 3, characterized in that: The state space of the electrolytic cell is: S={P t ,S state ,D t+1 ,D t+2 ,...D t+24 ,T run ,N start / stop ,T standby ,R adj ,T start ,T stop } Among them, P t is the total power of the electrolytic cell at the current moment, S state is the state of each electrolytic cell, D t+i is the electricity demand forecast for the next hour i, T run is the total operating time of the electrolyzer, N start / stop is the number of starts and stops of the electrolytic cell, T standby is the standby time of the electrolytic cell, R adj is the power regulation rate, T start is the start-up time characteristic of the electrolytic cell, T stop is the stopping time characteristic of the electrolytic cell; The action space of the electrolytic cell is: A={a start ,a stop ,a hold ,a adjust } Among them, a start Indicates starting the electrolytic cell, a stop Indicates stopping the electrolytic cell, a hold Indicates maintaining the current state, a adjust Indicates power adjustment.

7. The method for scheduling the start and stop of an electrolytic cell according to claim 4, wherein: The designed reward function is specifically:

8. The method for scheduling the start and stop of an electrolytic cell according to claim 6, characterized in that: The specific method for dynamically adjusting the start and stop priority of the electrolyzer is: F1=w1·T run +w2·N start / stop +w3·T stabdby +w6·R adj F2=w4·P fluctuation +w5·T griyo_run +w7·R adj Among them, F1 is the start-stop priority, F2 is the load adjustment priority; ω1, ω2, ω3, ω4, ω5, ω6, ω7 are weight coefficients, and the value range is [0,1]. fluctuation T is the operating time of the electrolyzer under fluctuating power, group_run is the operating time of the electrolytic cell group, R adj is the power regulation rate.

9. The method for planning and scheduling the start and stop of an electrolytic cell according to claim 1, characterized in that: The start-up time of the electrolyzer is: The stop time of the electrolyzer is: Among them, P target is the target power, P current is the current power, R adj is the power regulation rate.

10. An electrolytic cell start-stop planning and scheduling system, characterized in that: An electrolytic cell start-stop planning and scheduling method according to any one of claims 1 to 9, comprising: an electrolytic cell environment acquisition module, an electrolytic cell start-stop control intelligent agent, an electrolytic cell priority module, an electrolytic cell countdown module, and a scheduling result output module; The electrolyzer environment acquisition module is used to build an electrolyzer environment model to obtain the electrolyzer's operating status, power input, hydrogen output, and power demand forecast for the next 24 hours; The electrolyzer start-stop control agent is used to design a deep reinforcement learning agent. Based on the deep Q-network algorithm, the agent decides on the start and stop operation of the electrolyzer based on the current state and future power demand forecast, and optimizes the power regulation rate. It also defines the state space and action space of the electrolyzer, designs a reward function, updates the Q value through the Q-learning algorithm, and optimizes the agent's decision-making strategy. The electrolyzer priority module is used to dynamically adjust the start and stop priorities of the electrolyzer according to the operating status, historical data and power regulation rate of the electrolyzer; The electrolyzer countdown module is used to calculate the start and stop time of the electrolyzer based on the start and stop decision and power regulation rate to ensure the timing rationality of the start and stop operations; The scheduling result output module is used to output the scheduling results, including the status planning of each electrolytic cell in the next 24 hours, the upper and lower limits of the total load adjustment, and the power adjustment plan.