A Multi-Time-Scale Scheduling Method for Electro-Hydrogen Coupled Systems Based on Deep Reinforcement Learning

CN122137007APending Publication Date: 2026-06-02POWER CHINA KUNMING ENG CORP LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
POWER CHINA KUNMING ENG CORP LTD
Filing Date
2026-02-05
Publication Date
2026-06-02

Smart Images

  • Figure CN122137007A_ABST
    Figure CN122137007A_ABST
Patent Text Reader

Abstract

This invention provides a multi-timescale scheduling method for an electro-hydrogen coupling system based on deep reinforcement learning. The method includes building a hierarchical optimization model encompassing wind and solar power generation, electrochemical energy storage, electrolytic hydrogen production, fuel cells, and hydrogen storage; collecting real-time and predicted wind and solar power output data; using a trained deep reinforcement learning agent to make decisions on power commands for electrolytic hydrogen production and fuel cells at the long-term optimization layer; calculating real-time power deviations after execution; and adaptively allocating power for electrochemical energy storage charging and discharging and grid interaction at the short-term optimization layer using signal decomposition, thus achieving multi-timescale collaborative optimization. This invention enables efficient scheduling of the electro-hydrogen coupling system, improves the utilization rate of renewable energy, enhances the system's power balance capability and economy, and ensures the safe operation of energy storage and hydrogen storage devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of energy system optimization and scheduling technology, and more specifically, to a multi-timescale scheduling method for an electric-hydrogen coupled system based on deep reinforcement learning. Background Technology

[0002] With the rapid development of renewable energy, the proportion of intermittent energy sources such as wind and solar power integrated into the power system is constantly increasing. The randomness and uncertainty of their output pose significant challenges to the stable operation of the power grid. Traditional power dispatching methods mainly rely on empirical rules or optimization models based on a single time scale, which are insufficient to effectively address the volatility of renewable energy. Meanwhile, the electro-hydrogen coupling system, as an emerging integrated energy utilization model, achieves the mutual conversion of electrical and hydrogen energy through electrolysis and fuel cells, providing a new approach to solving the problems of renewable energy consumption and storage. However, the complexity of the electro-hydrogen coupling system lies in its involvement of multiple energy forms and equipment, requiring coordinated and optimized dispatching across different time scales to achieve efficient system operation and economic goals.

[0003] In implementing the embodiments of the present invention, the prior art has at least the following problems or defects: traditional scheduling methods cannot fully consider the uncertainty of renewable energy and the dynamic characteristics of each device in the electric-hydrogen coupling system, resulting in insufficient system power balance capability, poor economy, and difficulty in ensuring the safe operation of energy storage and hydrogen storage devices. Summary of the Invention

[0004] This invention provides a multi-timescale scheduling method for an electro-hydrogen coupling system based on deep reinforcement learning, comprising the following steps: S1: Construct a multi-timescale optimization scheduling model for the electro-hydrogen coupling system. The model adopts a hierarchical optimization framework, including a long-timescale optimization layer and a short-timescale optimization layer. The electro-hydrogen coupling system includes a wind turbine, a photovoltaic system, an electrochemical energy storage unit, a fuel cell, an electrolysis hydrogen production system, and a hydrogen storage system. S2: Collect real-time power output data of the wind turbine and the photovoltaic system, and obtain power prediction data for multiple future time periods; S3: Based on the real-time output data, the power prediction data, and the system load data, the trained deep reinforcement learning agent makes optimization decisions in the long-term optimization layer to generate a power scheduling instruction sequence for the electrolytic hydrogen production system and the fuel cell within the scheduling cycle. S4: Execute the power scheduling command sequence, calculate the real-time power deviation of the system, and use the adaptive allocation method based on signal decomposition to generate the charging and discharging power commands of the electrochemical energy storage unit and the power interaction commands with the grid in the short time scale optimization layer.

[0005] Further, in step S1, the optimization objective of the long-term optimization layer is to minimize the net power imbalance of the system within the scheduling cycle. The net power imbalance of the system is the algebraic sum of the actual output of renewable energy, load demand, power consumption of the electrolysis hydrogen production system, and output power of the fuel cell. The optimization objective of the short-term optimization layer is to quickly compensate for and smooth the remaining power deviation after the long-term optimization.

[0006] Further, in step S1, a scheduling model for the electrochemical energy storage unit is established, the model satisfying the following capacity state update equation and operating constraints: The capacity state update equation is:

[0007] in, for The state of charge of the electrochemical energy storage unit at that time. for The state of charge of the electrochemical energy storage unit at that time; for The charging power of the electrochemical energy storage unit at that time. for The discharge power of the electrochemical energy storage unit at that time; The charging efficiency of the electrochemical energy storage unit is... The discharge efficiency of the electrochemical energy storage unit; The rated capacity of the electrochemical energy storage unit; The operational constraints include: Charge and discharge power constraints: ; Charge state range constraints: ; mutual exclusion constraint between charge and discharge states: ; in, This refers to the maximum permissible charging power of the electrochemical energy storage unit. This refers to the maximum permissible discharge power of the electrochemical energy storage unit; This is the lower limit of the permissible state of charge of the electrochemical energy storage unit. This represents the upper limit of the allowed state of charge of the electrochemical energy storage unit; To indicate A binary variable representing the start / stop state of the charging action at all times. To indicate A binary variable representing the start and stop states of the discharge action at any given time.

[0008] Further, in step S1, a scheduling model for the hydrogen storage system is established, the model satisfying the following hydrogen storage state update equation and operational constraints: The hydrogen storage status update equation is as follows:

[0009] in, for The amount of hydrogen stored in the hydrogen storage system at any given time. for The amount of hydrogen stored in the hydrogen storage system at any given time; for The hydrogen storage capacity of the hydrogen storage system at that time for The hydrogen release power of the hydrogen storage system at that time; The hydrogen storage efficiency of the hydrogen storage system is given. The hydrogen release efficiency of the hydrogen storage system; The scheduling time interval; The operational constraints include: Hydrogen storage power constraints: ; Hydrogen storage capacity range constraints: ; Mutual exclusion constraints on hydrogen storage and release states: ; in, This represents the maximum permissible hydrogen storage capacity of the hydrogen storage system. This represents the maximum permissible hydrogen release power of the hydrogen storage system. This is the lower limit of the allowable hydrogen storage capacity of the hydrogen storage system. This represents the maximum allowable hydrogen storage capacity of the hydrogen storage system. To indicate A binary variable representing the start-stop state of hydrogen storage operation at any given moment. To indicate A binary variable representing the start and stop status of the hydrogen release action at all times.

[0010] Further, in step S2, the obtained power prediction data for multiple future time periods specifically includes a sequence of predicted wind and solar power output values ​​and a corresponding quantified sequence of prediction uncertainty; for the future... k The power over a given time period, at the current decision moment. The prediction information is represented as: point prediction value and its confidence interval half-width ,in This is a baseline forecast value obtained based on historical data. In order to be in Constantly assessing, for the future The relative error coefficient of the prediction at time t. The prediction relative error coefficient is calculated based on the historical statistical characteristics of the predicted relative error coefficient.

[0011] Furthermore, in step S3, the deep reinforcement learning agent is trained and makes decisions using a deep deterministic policy gradient algorithm framework. The state space of the intelligent agent Defined as:

[0012] in They represent from From this moment to the future H A sequence of predicted values ​​for wind power, solar power and load power points for each time period; These are the uncertainty quantification sequences for the corresponding wind power and photovoltaic power point prediction value sequences, respectively; for The state of charge of the electrochemical energy storage unit at that time; for The amount of hydrogen stored in the hydrogen storage system at any given time; for Real-time power imbalance of the system at any given moment; for Real-time electricity market price signals; The action space of the intelligent agent Defined as: ,in To t The scheduling instruction for the power consumption of the electrolytic hydrogen production system at time +1. To The scheduling command for the output power of the fuel cell at any given time.

[0013] Furthermore, in step S3, the reward function of the deep reinforcement learning agent... The multi-objective weighted summation method is adopted, specifically as follows: ; in, These are preset weighting coefficients; As a power balance reward, for Real-time power imbalance of the system at any given moment; As an economic reward, , for The power of constant interaction with the power grid. The preset penalty coefficient for wind and solar power curtailment. for The amount of wind and solar power curtailed at any given moment; For safety constraint penalties, , and This is the preset penalty coefficient.

[0014] Furthermore, the implementation of the deep deterministic policy gradient algorithm is based on the Actor-Critic framework, including an online policy network. Online value network Target-Policy Network and target value network ; The parameters of the online value network By minimizing the time difference error loss function Update, among which For the goal Q value, Discount factor; The parameters of the online policy network By outputting the action gradient along the online value network The direction is updated, and the update formula is: ; The parameters of the target network and Synchronize parameters with the online network via soft update: ,in This is the soft update coefficient.

[0015] Further, in step S4, the adaptive allocation method based on signal decomposition includes the following steps: S41: Calculate the real-time power deviation; S42: The real-time power deviation is decomposed into multiple frequency components using a signal decomposition algorithm; S43: Based on the real-time operating status of the electrochemical energy storage unit, adaptively determine the frequency division threshold; S44: High-frequency components with frequencies higher than the threshold are allocated to the electrochemical energy storage unit for smoothing, and low-frequency components with frequencies lower than the threshold are allocated to the power grid for regulation.

[0016] Furthermore, in step S42, the signal decomposition algorithm is an empirical mode decomposition algorithm, which decomposes the real-time power deviation into a set of intrinsic mode function components. and a residual component ,in Highest frequency Lowest frequency; In step S43, the adaptive determination of the frequency segmentation threshold is performed based on the current available adjustable power of the electrochemical energy storage unit. and its state of charge Determine a frequency index The rule is determined as follows: From Start accumulating each amplitude until the cumulative amplitude is reached. First time exceeding ,or Approaching the preset safety boundary limits the range of usable frequencies. The value is Then with The corresponding average frequency is used as the frequency segmentation threshold.

[0017] The embodiments of the present invention have at least the following beneficial effects: 1. By employing deep reinforcement learning agents for long-term optimization decision-making, it can comprehensively consider the real-time output of renewable energy, power prediction data, and system load demand, and generate a precise power scheduling command sequence for the electrolysis hydrogen production system and fuel cells. This effectively improves the power balance capability of the system and solves the problem of insufficient power balance capability of traditional scheduling methods when dealing with the uncertainty of renewable energy.

[0018] 2. The application of the adaptive allocation method based on signal decomposition in the short-timescale optimization layer enables rapid decomposition and precise allocation of real-time power deviations. High-frequency components are allocated to electrochemical energy storage units for smoothing, while low-frequency components are allocated to grid regulation, thereby improving the system's dynamic response speed and regulation accuracy. This solves the problem of difficulty in quickly compensating for and smoothing power deviations in existing technologies and enhances system stability.

[0019] 3. This method utilizes a multi-objective weighted reward function for deep reinforcement learning agents, balancing multiple objectives such as power balance, economy, and safe equipment operation. During the optimization scheduling process, it considers both system economy and ensures the safe operation of energy storage and hydrogen storage devices through penalty terms. This solves the problem of balancing economy and safety caused by a single optimization objective in existing technologies, thus improving the overall performance of the system. Attached Figure Description

[0020] The above and other objects, features, and advantages of exemplary embodiments of the present invention will become readily apparent from the following detailed description taken in conjunction with the accompanying drawings. Several embodiments of the invention are illustrated in the drawings by way of example and not limitation, wherein: Figure 1This is a flowchart illustrating a multi-timescale scheduling method for an electro-hydrogen coupling system based on deep reinforcement learning, provided in an embodiment of the present invention. Detailed Implementation

[0021] The principles and spirit of the invention will now be described with reference to several exemplary embodiments. It should be understood that these embodiments are provided merely to enable those skilled in the art to better understand and implement the invention, and are not intended to limit the scope of the invention in any way. Rather, these embodiments are provided to make the invention more thorough and complete, and to fully convey the scope of the invention to those skilled in the art.

[0022] The following is for reference. Figure 1 , Figure 1 This is a flowchart illustrating a multi-timescale scheduling method for an electro-hydrogen coupling system based on deep reinforcement learning, provided in an embodiment of the present invention. Figure 1 As shown, a multi-timescale scheduling method for an electro-hydrogen coupling system based on deep reinforcement learning includes: S1: Construct a multi-timescale optimization scheduling model for the electro-hydrogen coupling system. The model adopts a hierarchical optimization framework, including a long-timescale optimization layer and a short-timescale optimization layer. The electro-hydrogen coupling system includes a wind turbine, a photovoltaic system, an electrochemical energy storage unit, a fuel cell, an electrolysis hydrogen production system, and a hydrogen storage system. S2: Collect real-time power output data of the wind turbine and the photovoltaic system, and obtain power prediction data for multiple future time periods; S3: Based on the real-time output data, the power prediction data, and the system load data, the trained deep reinforcement learning agent makes optimization decisions in the long-term optimization layer to generate a power scheduling instruction sequence for the electrolytic hydrogen production system and the fuel cell within the scheduling cycle. S4: Execute the power scheduling command sequence, calculate the real-time power deviation of the system, and use the adaptive allocation method based on signal decomposition to generate the charging and discharging power commands of the electrochemical energy storage unit and the power interaction commands with the grid in the short time scale optimization layer.

[0023] The optimization objective of the long-timescale optimization layer is to minimize the net power imbalance of the system within the scheduling cycle. This is achieved by rationally scheduling the power of the electrolysis hydrogen production system and the fuel cell to minimize the algebraic sum between the actual output of renewable energy, load demand, power consumption of the electrolysis hydrogen production system, and output power of the fuel cell. This helps improve the overall power balance capability of the system. The optimization objective of the short-timescale optimization layer is to quickly compensate for and smooth the remaining power deviation after long-timescale optimization to ensure the real-time power balance of the system. The scheduling model of the electrochemical energy storage unit includes a capacity state update equation and operational constraints. The capacity state update equation describes the change in the state of charge of the electrochemical energy storage unit at different time points, involving parameters such as charging power, discharging power, charging efficiency, discharging efficiency, and rated capacity. Operational constraints include charging and discharging power constraints, state of charge range constraints, and charging and discharging state mutual exclusion constraints. These constraints ensure that the electrochemical energy storage unit operates within a safe range. The scheduling model of the hydrogen storage system also includes a hydrogen storage capacity state update equation and operational constraints. These equations and constraints ensure that the hydrogen storage system performs hydrogen storage and discharging operations within a safe range.

[0024] In some embodiments, in step S1, the optimization objective of the long-time-scale optimization layer is to minimize the net power imbalance of the system within the scheduling cycle. The net power imbalance of the system is the algebraic sum of the actual output of renewable energy, load demand, power consumption of the electrolytic hydrogen production system, and output power of the fuel cell. The optimization objective of the short-time-scale optimization layer is to quickly compensate for and smooth the remaining power deviation after the long-time-scale optimization.

[0025] The calculation of the system's net power imbalance in the long-term optimization layer involves several key parameters. The actual output of renewable energy refers to the electrical power generated by wind turbines and photovoltaic systems during actual operation; load demand refers to the electrical load power the system needs to meet; the power consumed by the hydrogen electrolysis system refers to the electrical power consumed during the water electrolysis process; and the fuel cell output power refers to the power output after the fuel cell converts hydrogen energy into electrical energy. The algebraic sum of these parameters reflects the system's power balance state over a long timescale. The short-term optimization layer focuses on the rapid compensation of real-time power deviations, which refer to the difference between the system power and the long-term optimized scheduling results during actual operation. By rapidly compensating for and mitigating these deviations, the short-term optimization layer can effectively cope with the volatility of renewable energy and ensure the real-time stable operation of the system.

[0026] When constructing the long-term optimization layer, historical data and predictive models can be used to obtain forecasts of renewable energy output and load demand. These forecasts can serve as input parameters for the optimization model. Combined with the operating characteristics of the electrolytic hydrogen production system and fuel cells, mathematical optimization algorithms, such as linear or nonlinear programming, can be used to generate the optimal power dispatch command sequence. For the short-term optimization layer, an adaptive allocation method based on signal decomposition can be used to handle real-time power deviations. Specifically, the power deviation is decomposed into high-frequency and low-frequency components. The high-frequency component is allocated to electrochemical energy storage units for rapid mitigation, while the low-frequency component is regulated through interaction with the grid. This hierarchical optimization and adaptive allocation strategy can effectively improve the system's power balance capability and dynamic response performance.

[0027] In some embodiments, in step S1, a scheduling model for the electrochemical energy storage unit is established, the model satisfying the following capacity state update equation and operating constraints: The capacity state update equation is:

[0028] in, for The state of charge of the electrochemical energy storage unit at that time. for The state of charge of the electrochemical energy storage unit at that time; for The charging power of the electrochemical energy storage unit at that time. for The discharge power of the electrochemical energy storage unit at that time; The charging efficiency of the electrochemical energy storage unit is... The discharge efficiency of the electrochemical energy storage unit; The rated capacity of the electrochemical energy storage unit; The operational constraints include: Charge and discharge power constraints: ; Charge state range constraints: ; mutual exclusion constraint between charge and discharge states: ; in, This refers to the maximum permissible charging power of the electrochemical energy storage unit. This refers to the maximum permissible discharge power of the electrochemical energy storage unit; This is the lower limit of the permissible state of charge of the electrochemical energy storage unit. This represents the upper limit of the allowed state of charge of the electrochemical energy storage unit; To indicate A binary variable representing the start / stop state of the charging action at all times. To indicate A binary variable representing the start and stop states of the discharge action at any given time.

[0029] The capacity state update equation of an electrochemical energy storage unit involves several key parameters. Among them, SOC State of charge (SOC) represents the remaining electrical charge of an electrochemical energy storage unit at a certain moment and is an important indicator for measuring its charge and discharge state. and These represent charging power and discharging power, respectively, reflecting the power input and output of the energy storage unit under different operating modes; and These are charging efficiency and discharging efficiency, used to correct for energy loss during actual charging and discharging processes; This is the rated capacity of the electrochemical energy storage unit, representing its maximum stored capacity. Operating constraints include charge / discharge power constraints, state of charge (SOC) range constraints, and mutually exclusive charge / discharge constraints. Charge / discharge power constraints ensure that the energy storage unit's charge / discharge power does not exceed its maximum allowable value; SOC range constraints limit the SOC to between the allowable upper and lower limits, preventing overcharging or over-discharging; and mutually exclusive charge / discharge constraints ensure that the energy storage unit will not simultaneously perform charging and discharging operations.

[0030] More specifically, the scheduling model for electrochemical energy storage units can be constructed and applied through the following steps. First, based on the characteristic parameters of the electrochemical energy storage unit, such as rated capacity, maximum charge / discharge power, and charge / discharge efficiency, a capacity state update equation is established. These parameters can be obtained through experimental testing or equipment manuals. Second, considering the actual operating requirements of the system, reasonable state of charge (SOC) ranges and charge / discharge power limits are set. For example, based on the health status of the energy storage unit and the application scenario, the allowable upper and lower limits of SOC are set to 20% and 80%, respectively, to ensure its safety during long-term operation. During operation, the SOC and charge / discharge power of the energy storage unit are monitored in real time, and its charge / discharge strategy is dynamically adjusted according to the capacity state update equation. When the system needs to respond quickly to power fluctuations, high-frequency power components are allocated to the electrochemical energy storage unit through an adaptive allocation method to smooth them out, thereby achieving rapid adjustment of the system's power balance.

[0031] In some embodiments, in step S1, a scheduling model for the hydrogen storage system is established, the model satisfying the following hydrogen storage state update equation and operational constraints: The hydrogen storage status update equation is as follows:

[0032] in, The amount of hydrogen stored in the hydrogen storage system at any given time. The amount of hydrogen stored in the hydrogen storage system at any given time; The hydrogen storage capacity of the hydrogen storage system at that time The hydrogen release power of the hydrogen storage system at that time; The hydrogen storage efficiency of the hydrogen storage system is given. The hydrogen release efficiency of the hydrogen storage system; The scheduling time interval; The operational constraints include: Hydrogen storage power constraints: ; Hydrogen storage capacity range constraints: ; Mutual exclusion constraints on hydrogen storage and release states: ; in, This represents the maximum permissible hydrogen storage capacity of the hydrogen storage system. This represents the maximum permissible hydrogen release power of the hydrogen storage system. This is the lower limit of the allowable hydrogen storage capacity of the hydrogen storage system. This represents the maximum allowable hydrogen storage capacity of the hydrogen storage system. for A binary variable representing the start-stop state of hydrogen storage operation at any given moment. for t A binary variable representing the start and stop status of the hydrogen release action at all times.

[0033] The state update equation for the hydrogen storage capacity of a hydrogen storage system involves several key parameters. Among them, It represents the amount of hydrogen stored in a hydrogen storage system at a certain moment and is an important indicator for measuring its hydrogen storage status; and These represent hydrogen storage power and hydrogen release power, respectively, reflecting the power input and output of the hydrogen storage system under different operating modes; and These are hydrogen storage efficiency and hydrogen release efficiency, used to correct for energy losses during actual hydrogen storage and release processes; This refers to the scheduling time interval, representing the time span of each scheduling operation. Operational constraints include hydrogen storage / discharge power constraints, hydrogen storage capacity range constraints, and hydrogen storage / discharge state mutual exclusion constraints. The hydrogen storage / discharge power constraints ensure that the hydrogen storage / discharge power of the hydrogen storage system does not exceed its maximum allowable value; the hydrogen storage capacity range constraints limit the hydrogen storage capacity to between the allowable upper and lower limits, preventing overcharging or over-discharging of the equipment; and the hydrogen storage / discharge state mutual exclusion constraints ensure that the hydrogen storage system will not simultaneously perform hydrogen storage and discharging operations at the same time.

[0034] Based on the characteristic parameters of the hydrogen storage system, such as maximum hydrogen storage capacity, maximum hydrogen storage / discharge power, and hydrogen storage / discharge efficiency, a state update equation for the hydrogen storage capacity is established. These parameters can be obtained through experimental testing or equipment manuals. Secondly, considering the actual operational requirements of the system, reasonable ranges for hydrogen storage capacity and limits for hydrogen storage / discharge power are set. For example, based on the health status and application scenario of the hydrogen storage system, allowable upper and lower limits for hydrogen storage capacity are set at 10% and 90%, respectively, to ensure its safety during long-term operation. During operation, the hydrogen storage capacity and hydrogen storage / discharge power of the hydrogen storage system are monitored in real time, and its hydrogen storage / discharge strategy is dynamically adjusted according to the state update equation. When the system needs to respond quickly to power fluctuations, an adaptive allocation method is used to distribute low-frequency power components to the hydrogen storage system for adjustment, thereby achieving rapid adjustment of the system's power balance.

[0035] In some embodiments, the power prediction data for the multiple future time periods obtained in step S2 specifically includes a sequence of predicted wind power and photovoltaic output values ​​and a corresponding quantified sequence of prediction uncertainty; for the future... k The power over a given time period, at the current decision moment. The prediction information is represented as: point prediction value and its confidence interval half-width , This is a baseline forecast value obtained based on historical data. In order to be in Constantly assessing, for the future The relative error coefficient of the prediction at time t. The prediction relative error coefficient is calculated based on the historical statistical characteristics of the predicted relative error coefficient.

[0036] The power forecast data for multiple future time periods includes a series of point forecast values ​​for wind and solar power output, along with corresponding quantified uncertainty sequences. Point forecast value series It is derived from historical data and predictive models, representing the current decision-making moment. For the future Expected wind or solar power output for a given period. Quantification of forecast uncertainty sequence. This is obtained by analyzing the historical statistical characteristics of the prediction error, and is used to quantify the uncertainty of the predicted value. Among them, It is the relative error coefficient of the forecast, used to adjust the baseline forecast value. This allows for more accurate point predictions. This prediction method provides richer information to scheduling models, enabling them to fully consider the volatility of renewable energy sources when making optimization decisions.

[0037] Historical data is used to generate baseline forecasts for wind and solar power output using predictive algorithms, such as machine learning models or physical models. The relative prediction error coefficient is calculated based on historical prediction error data. And through statistical analysis, a quantified sequence of predictive uncertainty is obtained. .

[0038] In practical applications, these predicted data can be used as input parameters into the scheduling model. For example, the scheduling model can dynamically adjust the power scheduling commands of the electrolysis hydrogen production system and fuel cell based on point predicted values ​​and uncertainty quantification sequences through optimization algorithms, thereby reducing the scheduling risk caused by prediction errors while satisfying system power balance.

[0039] In some embodiments, in step S3, the deep reinforcement learning agent is trained and makes decisions using a deep deterministic policy gradient algorithm framework. The state space of the intelligent agent Defined as:

[0040] in, They represent from From this moment to the future H A sequence of predicted values ​​for wind power, solar power and load power points for each time period; These are the uncertainty quantification sequences for the corresponding wind power and photovoltaic power point prediction value sequences, respectively; for The state of charge of the electrochemical energy storage unit at that time; for The amount of hydrogen stored in the hydrogen storage system at any given time; for Real-time power imbalance of the system at any given moment; for Real-time electricity market price signals; Action space of an intelligent agent Defined as: ,in To The scheduling instructions for the power consumption of the electrolysis hydrogen production system at the specified time. To The scheduling command for the output power of the fuel cell at any given time.

[0041] The state space of a deep reinforcement learning agent This includes point-predicted value sequences for several key parameters, such as wind power, solar power, and load power. , and These parameters reflect the system's power demand and expected renewable energy output over a future period. Uncertainty Quantification Sequence and The uncertainties in wind and solar power forecasting are described, providing a reference for prediction errors for intelligent agents. The state of charge of electrochemical energy storage units is also discussed. Hydrogen storage capacity of hydrogen storage system It reflects the current status of the energy storage device, while the real-time power imbalance... and electricity market price signals This provides the intelligent agent with real-time operational status and economic information of the system. Action Space Including the power consumption of the hydrogen electrolysis system and fuel cell output power The scheduling instructions directly affect the power balance and economy of the system.

[0042] The training of a deep reinforcement learning agent is achieved through the following steps. First, when constructing the state space and action space, relevant parameters need to be obtained from the system's real-time monitoring data and prediction model. For example, power prediction values ​​can be generated by combining historical data with machine learning or physical models, while uncertainty quantification sequences can be obtained by analyzing the historical statistical characteristics of prediction errors. During training, the agent continuously learns through interaction with the environment, optimizing its policy network and value network. The policy network generates actions based on the current state, while the value network evaluates the quality of the actions. The agent's goal is to maximize cumulative rewards, and the reward function comprehensively considers multiple objectives such as power balance, economy, and safe equipment operation. In actual operation, the agent generates scheduling instructions based on the current state and sends these instructions to the electrolysis hydrogen production system and fuel cells, thereby achieving optimized system scheduling.

[0043] In some embodiments, in step S3, the reward function of the deep reinforcement learning agent A multi-objective weighted summation method is adopted, specifically as follows: ; in, These are preset weighting coefficients; As a power balance reward, for Real-time power imbalance of the system at any given moment; As an economic reward, , for The power of constant interaction with the power grid. The preset penalty coefficient for wind and solar power curtailment. for The amount of wind and solar power curtailed at any given moment; For safety constraint penalties, , and This is the preset penalty coefficient.

[0044] The reward function consists of three parts: power balance reward. Economic rewards and safety constraint penalties Power balancing reward This represents a negative value indicating the real-time power imbalance of the system, with the aim of improving system stability by minimizing the power imbalance. Economic incentives. The power costs associated with grid interaction and the penalties for wind and solar power curtailment are considered, aiming to reduce system operating costs. Security constraint penalties are also included. This ensures the safe operation of energy storage and hydrogen storage devices by penalizing situations where the state of charge and hydrogen storage capacity exceed safe limits. (Weighting coefficient) and This is used to adjust the proportion of each part in the total reward, thereby achieving a balance between different objectives.

[0045] The construction of the reward function includes setting weight coefficients based on the actual needs and priorities of the system. and For example, if the system prioritizes power balance, it can increase... The weighting. Secondly, the real-time power imbalance. Electricity market price signals can be obtained through system power monitoring equipment. This can be obtained from electricity market data. Wind and solar curtailment penalty coefficient. This can be set according to the system's economic strategy, typically reflecting the negative impact of wind and solar power curtailment on the system's economics. The penalty coefficient in the safety constraint penalty item... and Adjustments can be made based on the equipment's safe operation requirements. During the agent's training process, the shape of the reward function is optimized by continuously adjusting the weight coefficients and penalty coefficients, thereby guiding the agent to better balance goals such as power balance, economy, and equipment safe operation during the decision-making process.

[0046] In some embodiments, the deep deterministic policy gradient algorithm is implemented based on the Actor-Critic framework, including an online policy network. Online value network Target-Policy Network and target value network ; The parameters of the online value network By minimizing the time difference error loss function Update, among which For the goal Q value, Discount factor; The parameters of the online policy network By outputting the action gradient along the online value network The direction is updated, and the update formula is: ; The parameters of the target network and Synchronize parameters with the online network via soft update: ,in This is the soft update coefficient.

[0047] Online policy network in deep deterministic policy gradient algorithm It is a neural network that determines the current state. Output optimal action Its parameters Updated using gradient ascent. Online value network. The parameters used to evaluate the value of state-action pairs Updates are performed by minimizing the temporal difference error loss function. Target Policy Network and target value network The parameters are synchronized with the parameters of the online network via soft updates, updating the coefficients. It is usually set to a small positive number to smooth the update process of the target network. Discount factor The temporal difference error loss function measures the current value of future rewards, typically ranging from 0 to 1. It compares the target... Q Value and actual Q The differences in values ​​are used to update the parameters of the online value network, thereby optimizing the agent's decision-making process.

[0048] During implementation, the parameters of the online policy network and the online value network are first initialized. and And set the parameters of the target network. and Similar to online networks, during training, the agent adjusts its behavior based on the current state. Actions are generated via an online policy network. And receive rewards for interacting with the environment. and the next state .Target Q The values ​​are calculated through the target network and used to update the parameters of the online value network. Parameters of the online policy network The policy is then optimized by updating along the action gradient direction output by the value network. The parameters of the target network are synchronized with the online network through soft updates, using the following formula: and ,in These are soft update coefficients. In this way, the agent can continuously learn and optimize during training, ultimately achieving efficient scheduling of the electro-hydrogen coupling system.

[0049] In some embodiments, step S4 of the adaptive allocation method based on signal decomposition includes the following steps: S41: Calculate the real-time power deviation; S42: The real-time power deviation is decomposed into multiple frequency components using a signal decomposition algorithm; S43: Based on the real-time operating status of the electrochemical energy storage unit, adaptively determine the frequency division threshold; S44: High-frequency components with frequencies higher than the threshold are allocated to the electrochemical energy storage unit for smoothing, and low-frequency components with frequencies lower than the threshold are allocated to the power grid for regulation.

[0050] Real-time power deviation refers to the difference between the actual power of the system and the dispatch command, reflecting the immediate imbalance of the system. The signal decomposition algorithm decomposes the power deviation into multiple frequency components, the frequency of which reflects the rate of power fluctuation. The adaptive frequency threshold is determined based on the current available adjustable power and state of charge of the electrochemical energy storage unit, used to distinguish between high-frequency and low-frequency components. High-frequency components typically correspond to rapid fluctuations and are suitable for smoothing by the electrochemical energy storage unit; low-frequency components correspond to slower power changes and are more suitable for regulation through interaction with the grid. This adaptive allocation method can fully utilize the rapid response capability of the electrochemical energy storage unit and the regulation capability of the grid.

[0051] Real-time power deviation is acquired through system power monitoring equipment and calculated as the difference between the actual power and the dispatch commanded power. An empirical mode decomposition algorithm is used to decompose the power deviation into a set of intrinsic mode functions (IMF) components and a residual component. Each IMF component has different frequency characteristics, arranged sequentially from high to low frequency. The frequency segmentation threshold is adaptively determined based on the current available adjustable power and state of charge of the electrochemical energy storage unit. For example, the amplitude of the first IMF component is accumulated until the accumulated amplitude exceeds the available adjustable power of the energy storage unit, or the state of charge approaches the safety boundary; the frequency of the corresponding IMF component at this point is the threshold. High-frequency components above the threshold are allocated to the electrochemical energy storage unit for smoothing, while low-frequency components are allocated to grid regulation. In this way, the system can flexibly respond to power fluctuations at different time scales, ensuring stable system operation.

[0052] In some embodiments, in step S42, the signal decomposition algorithm is an empirical mode decomposition algorithm, which decomposes the real-time power deviation into a set of intrinsic mode function components. and a residual component ,in Highest frequency Lowest frequency; In step S43, the adaptive determination of the frequency segmentation threshold is performed based on the current available adjustable power of the electrochemical energy storage unit. and its state of charge Determine a frequency index The rule is determined as follows: From Start accumulating each amplitude until the cumulative amplitude is reached. First time exceeding ,or Approaching the preset safety boundary limits the range of usable frequencies. The value is Then with The corresponding average frequency is used as the frequency segmentation threshold.

[0053] The Empirical Mode Decomposition (EMD) algorithm decomposes real-time power deviation into a set of Intrinsic Mode Function (IMF) components, each representing fluctuations in the power deviation across different frequency ranges. The IMF components are arranged in descending order of frequency, with high-frequency components reflecting rapid power fluctuations, suitable for rapid smoothing by electrochemical energy storage units; low-frequency components reflect slow power changes, more suitable for regulation through interaction with the grid. The adaptive frequency threshold is determined based on the currently available adjustable power of the electrochemical energy storage unit. and state of charge The frequency division threshold is determined by accumulating the amplitude of the IMF components until the accumulated amplitude exceeds the available adjustable power of the energy storage unit or the state of charge approaches a preset safety boundary. This adaptive method can dynamically adjust the frequency threshold according to the actual state of the energy storage unit, ensuring stable operation of the system under different operating conditions.

[0054] The real-time power deviation is decomposed using an empirical mode decomposition algorithm, yielding a set of IMF components and a residual component. Starting with the first IMF component, the amplitudes of each IMF component are sequentially accumulated until the accumulated amplitude first exceeds the currently available adjustable power of the electrochemical energy storage unit. or state of charge Approaching the preset safety boundary, the corresponding IMF component frequency at this point becomes the frequency segmentation threshold. High-frequency components exceeding the threshold are allocated to the electrochemical energy storage unit for smoothing, while low-frequency components are allocated to grid regulation. In this way, the system can flexibly respond to power fluctuations at different time scales, ensuring stable system operation while optimizing the utilization efficiency of the energy storage unit and the overall economic efficiency of the system.

[0055] The above description is merely an explanation of some preferred embodiments of the present invention and the technical principles employed. Those skilled in the art should understand that the scope of the invention as described in the embodiments of the present invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of the present invention.

Claims

1. A multi-timescale scheduling method for an electro-hydrogen coupling system based on deep reinforcement learning, characterized in that, Includes the following steps: S1: Construct a multi-timescale optimization scheduling model for the electro-hydrogen coupling system. The model adopts a hierarchical optimization framework, including a long-timescale optimization layer and a short-timescale optimization layer. The electro-hydrogen coupling system includes a wind turbine, a photovoltaic system, an electrochemical energy storage unit, a fuel cell, an electrolysis hydrogen production system, and a hydrogen storage system. S2: Collect real-time power output data of the wind turbine and the photovoltaic system, and obtain power prediction data for multiple future time periods; S3: Based on the real-time output data, the power prediction data, and the system load data, the trained deep reinforcement learning agent makes optimization decisions in the long-term optimization layer to generate a power scheduling instruction sequence for the electrolytic hydrogen production system and the fuel cell within the scheduling cycle. S4: Execute the power scheduling command sequence, calculate the real-time power deviation of the system, and use the adaptive allocation method based on signal decomposition to generate the charging and discharging power commands of the electrochemical energy storage unit and the power interaction commands with the grid in the short time scale optimization layer.

2. The method according to claim 1, characterized in that, In step S1, the optimization objective of the long-term optimization layer is to minimize the net power imbalance of the system within the scheduling cycle. The net power imbalance of the system is the algebraic sum of the actual output of renewable energy, load demand, power consumption of the electrolysis hydrogen production system, and output power of the fuel cell. The optimization objective of the short-term optimization layer is to quickly compensate for and smooth the remaining power deviation after the long-term optimization.

3. The method according to claim 2, characterized in that, In step S1, a scheduling model for the electrochemical energy storage unit is established, and the model satisfies the following capacity state update equation and operating constraints: The capacity state update equation is: in, for The state of charge of the electrochemical energy storage unit at that time. for The state of charge of the electrochemical energy storage unit at that time; for The charging power of the electrochemical energy storage unit at that time. for The discharge power of the electrochemical energy storage unit at that time; The charging efficiency of the electrochemical energy storage unit is... The discharge efficiency of the electrochemical energy storage unit; The rated capacity of the electrochemical energy storage unit; The operational constraints include: Charge and discharge power constraints: ; Charge state range constraints: ; mutual exclusion constraint between charge and discharge states: ; in, This refers to the maximum permissible charging power of the electrochemical energy storage unit. This refers to the maximum permissible discharge power of the electrochemical energy storage unit; This is the lower limit of the permissible state of charge of the electrochemical energy storage unit. This represents the upper limit of the allowed state of charge of the electrochemical energy storage unit; To indicate A binary variable representing the start / stop state of the charging action at all times. To indicate A binary variable representing the start and stop states of the discharge action at any given time.

4. The method according to claim 2, characterized in that, In step S1, a scheduling model for the hydrogen storage system is established, which satisfies the following hydrogen storage state update equation and operational constraints: The hydrogen storage status update equation is as follows: in, for The amount of hydrogen stored in the hydrogen storage system at any given time. for The amount of hydrogen stored in the hydrogen storage system at any given time; for The hydrogen storage capacity of the hydrogen storage system at that time for The hydrogen release power of the hydrogen storage system at that time; The hydrogen storage efficiency of the hydrogen storage system is given. The hydrogen release efficiency of the hydrogen storage system; The scheduling time interval; The operational constraints include: Hydrogen storage power constraints: ; Hydrogen storage capacity range constraints: ; Mutual exclusion constraints on hydrogen storage and release states: ; in, This represents the maximum permissible hydrogen storage capacity of the hydrogen storage system. This represents the maximum permissible hydrogen release power of the hydrogen storage system. This is the lower limit of the allowable hydrogen storage capacity of the hydrogen storage system. This represents the maximum allowable hydrogen storage capacity of the hydrogen storage system. for A binary variable representing the start-stop state of hydrogen storage operation at any given moment. for A binary variable representing the start and stop status of the hydrogen release action at all times.

5. The method according to claim 1, characterized in that, In step S2, the obtained power prediction data for multiple future time periods specifically includes a sequence of predicted wind and solar power output values ​​and a corresponding quantified sequence of prediction uncertainties; for the future... k The power over a given time period, at the current decision moment. The prediction information is represented as: point prediction value and its confidence interval half-width ,in, This is a baseline forecast value obtained based on historical data. In order to be in Constantly assessing, for the future k The relative error coefficient of the prediction at time t. The prediction relative error coefficient is calculated based on the historical statistical characteristics of the predicted relative error coefficient.

6. The method according to claim 1, characterized in that, In step S3, the deep reinforcement learning agent is trained and makes decisions using a deep deterministic policy gradient algorithm framework. The state space of the intelligent agent Defined as: in, They represent from From this moment to the future H A sequence of predicted values ​​for wind power, solar power and load power points for each time period; These are the uncertainty quantification sequences for the corresponding wind power and photovoltaic power point prediction value sequences, respectively; for The state of charge of the electrochemical energy storage unit at that time; for The amount of hydrogen stored in the hydrogen storage system at any given time; for Real-time power imbalance of the system at any given moment; for Real-time electricity market price signals; The action space of the intelligent agent Defined as: ,in To The scheduling instructions for the power consumption of the electrolysis hydrogen production system at the specified time. To The scheduling command for the output power of the fuel cell at any given time.

7. The method according to claim 6, characterized in that, In step S3, the reward function of the deep reinforcement learning agent... The multi-objective weighted summation method is adopted, specifically as follows: ; in, The preset weighting coefficients, As a power balance reward, As an economic reward, This is a safety constraint penalty item.

8. The method according to claim 6, characterized in that, The implementation of the deep deterministic policy gradient algorithm is based on the Actor-Critic framework, including an online policy network. Online value network Target-Policy Network and target value network ; The parameters of the online value network Updates are performed by minimizing the time-series difference error loss function. The parameters of the online policy network The direction of the action gradient is updated along the output of the online value network; The parameters of the target network and The parameters are synchronized with the online network via a soft update method.

9. The method according to claim 1, characterized in that, In step S4, the adaptive allocation method based on signal decomposition includes the following steps: S41: Calculate the real-time power deviation; S42: The real-time power deviation is decomposed into multiple frequency components using a signal decomposition algorithm; S43: Based on the real-time operating status of the electrochemical energy storage unit, adaptively determine the frequency division threshold; S44: High-frequency components with frequencies higher than the threshold are allocated to the electrochemical energy storage unit for smoothing, and low-frequency components with frequencies lower than the threshold are allocated to the power grid for regulation.

10. The method according to claim 9, characterized in that, In step S42, the signal decomposition algorithm is an empirical mode decomposition algorithm, which decomposes the real-time power deviation into a set of intrinsic mode function components. and a residual component ,in Highest frequency Lowest frequency; In step S43, the adaptive determination of the frequency segmentation threshold is performed based on the current available adjustable power of the electrochemical energy storage unit. and its state of charge Determine the frequency index The rule is determined as follows: From Start accumulating each amplitude Until the cumulative amplitude first exceeds ,or Approaching the preset safety boundary limits the range of usable frequencies. The value is Then with The corresponding average frequency is used as the frequency segmentation threshold.