Hydrogen energy micro-grid energy management method based on Lyapunov safety reinforcement learning

By employing the Lyapunov secure reinforcement learning method, combined with an improved Actor-Critic algorithm and a mixed integer programming model, the nonlinear characteristics and safety hazards of hydrogen microgrid systems were addressed, achieving safe and stable energy management and cost optimization.

CN121660490APending Publication Date: 2026-03-13SOUTHWEST PETROLEUM UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively address the non-convex and nonlinear characteristics of hydrogen microgrid systems and the safety hazards inherent in traditional reinforcement learning, leading to operational instability and high operating costs.

Method used

We employ a Lyapunov-based safety reinforcement learning approach, combined with an improved Actor-Critic algorithm, to construct a safety assessment mechanism and a mixed-integer programming model. We then define a safety-inducing policy set using Lyapunov functions to optimize energy management strategies.

Benefits of technology

While ensuring system security, it optimizes operating costs and enhances adaptability to dynamic scenarios and generalization capabilities for complex systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121660490A_ABST
    Figure CN121660490A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of power systems, and particularly discloses a hydrogen energy micro-grid energy management method based on Lyapunov safety reinforcement learning, and the method comprises the steps: constructing a system operation cost model; defining an energy scheduling strategy as a reinforcement learning action, and performing real-time security assessment and correction by using a mixed integer programming model; introducing a Lyapunov function to construct a security induction strategy set; the economy of the reinforcement learning strategy and the operation safety of the system are balanced in combination with a safety evaluation mechanism; and constructing a Markov decision process of safety evaluation, and adopting an improved Actor-Critic algorithm to realize strategy optimization to obtain a management scheme. The method has the advantages that the problems that an existing convex optimization algorithm is difficult to process non-convex and non-linear characteristics of the system, traditional reinforcement learning depends on soft constraint, potential safety hazards exist and the like are solved, and finally the operation cost is reduced while the operation safety of the system is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system technology, and in particular to an energy management method for hydrogen microgrids based on Lyapunov secure reinforcement learning. Background Technology

[0002] Currently, the global energy system is undergoing profound transformation, and microgrids and hydrogen energy technology, as key supporting means to promote the optimization of the energy structure, are gradually receiving widespread attention. The "hydrogen microgrid," formed by the combination of these two technologies, is considered an important development direction for future energy systems. Hydrogen microgrids integrate renewable energy sources such as photovoltaics and wind power with processes like water electrolysis for hydrogen production, hydrogen storage, and fuel cell power generation, constructing a compact, flexible, and independently autonomous clean energy system. Compared to traditional electrified microgrids, hydrogen energy systems not only possess medium- and long-term energy storage capabilities, effectively alleviating the problem of unstable renewable energy output, but also achieve efficient peak shaving and rapid response through fuel cells, improving the overall operational resilience and energy utilization efficiency of the system. As of the end of 2023, more than 500 hydrogen-related microgrids or demonstration projects were in operation or under construction globally, widely distributed in countries such as Germany, Japan, the United States, China, and South Africa. Europe accounts for approximately 37% of the total number of hydrogen microgrid projects globally, Asia 34%, and North America approximately 21%. China accounts for approximately 28% of the global new capacity in green hydrogen energy project deployments, becoming a significant force driving the development of global hydrogen microgrids. Representative projects include Germany's "Energy Park" project, which integrates wind power, water electrolysis for hydrogen production, and fuel cell load regulation, reducing carbon emissions by 40,000 tons annually; and California's Calistoga resilient microgrid, which provides up to 48 hours of off-grid operation capability, significantly enhancing regional energy security. According to market research firm DataIntelo, the global green hydrogen microgrid market is projected to grow from $1.2 billion in 2023 to $6.8 billion in 2030, with a compound annual growth rate (CAGR) exceeding 28%, demonstrating strong technological maturity and investment attractiveness.

[0003] With the widespread application of hydrogen-electric hybrid energy systems, the development of an efficient energy management framework based on safety reinforcement learning will become a key driver for the low-carbon transformation of energy systems. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide an energy management method for hydrogen microgrids based on Lyapunov secure reinforcement learning.

[0005] The objective of this invention is achieved through the following technical solution: an energy management method for hydrogen microgrids based on Lyapunov secure reinforcement learning, the method comprising,

[0006] Construct a system operating cost model;

[0007] The energy scheduling strategy is defined as a reinforcement learning action, and a mixed-integer programming model is used for real-time safety assessment and correction.

[0008] A Lyapunov function is introduced to construct a set of security-inducible strategies; a security assessment mechanism is combined to balance the economy of reinforcement learning strategies with the security of system operation.

[0009] A Markov decision process for security assessment is constructed, and an improved Actor-Critic algorithm is used to optimize the strategy and obtain a management plan.

[0010] Specifically, the system operating cost model is as follows:

[0011] ;

[0012] In the formula, T represents the total number of time periods within the optimization time domain; , , These are the operating costs of gas turbines, battery energy storage systems, and hydrogen energy systems, respectively. The cost of purchasing electricity from the main power grid;

[0013] ;

[0014] ;

[0015] ;

[0016] ;

[0017] ;

[0018] ;

[0019] ;

[0020] ;

[0021] In the formula, This refers to the output power of the gas turbine. , and For coefficients; and These are the start-up and shutdown costs of the gas turbine, respectively. and These are binary variables indicating the start / stop status of the gas turbine; This refers to the volume of electricity imports or exports. For electricity price; Indicates the number of batteries; The absolute value of the energy change for each battery; The degradation cost per unit energy change; Costs related to battery degradation; For time The input power of the battery at that time; For time The battery's output power at that time; This is the degradation cost coefficient; For time The power of the electrolytic cell; This refers to the unit operating cost of the electrolytic cell; It serves as a start-up indicator for the electrolytic cell; The startup cost of the electrolytic cell; For time The power of the hydrogen fuel cell; The unit operating cost of hydrogen fuel cells; For the start-up indicator of hydrogen fuel cells; For the start-up cost of hydrogen fuel cells, It is the unit operating cost of a hydrogen energy system. This indicates the energy change in a hydrogen energy system; It is the degradation cost coefficient of the hydrogen energy system.

[0022] Specifically, a photovoltaic equipment operation model, a gas turbine model, a battery energy storage system model, and a hydrogen energy system model are constructed, and these models are embedded as constraints into the safety assessment mechanism.

[0023] Specifically, the reinforcement learning action 'a' is a vector that includes the outputs of photovoltaic, gas turbine, battery energy storage system, electrolyzer, hydrogen fuel cell, and thermal energy storage system.

[0024] Specifically, the safety of reinforcement learning action 'a' is evaluated using a mixed-integer programming model, which is represented as:

[0025] ;

[0026] In the formula, For safety reasons, This is the original action;

[0027] Constraint violation degree (cv) and corrective action (a) m Determined based on the following two situations:

[0028] If the energy management scheme is considered safe, then the constraint violation degree (cv) and the corrected action are... The following formula is derived:

[0029] ;

[0030] ;

[0031] If the energy management plan is unsafe, use the following formula to determine c. v and a m :

[0032] ;

[0033] In the formula, This is the proportionality coefficient; It is a constant; c v and a m The strategy used to update the agent.

[0034] Specifically, the introduction of Lyapunov functions to construct a set of security-inducing strategies includes:

[0035] Introducing a security baseline strategy constructed from a security assessment optimization model Its cumulative constraint violations satisfy:

[0036]

[0037] In the formula, To constrain the discount factor; This represents the cumulative constraint cost of the baseline policy under the initial state; This represents the maximum allowed cumulative constraint violation. Define a set of Lyapunov functions:

[0038] ;

[0039] In the formula, Denotes the set of Lyapunov functions. As a strategy, For strategy parameters; It is a Lyapunov function; it represents the state space of the microgrid. Mapping to the set of real numbers ; In strategy The Bellman constraint operator is given below, where h is the immediate safety index function; It is the initial state. This allows for the cumulative constraint to violate the maximum value;

[0040] ;

[0041] In the formula, The mathematical expectation is represented by action a and the next state. produce; For the constraint cost function; To constrain the discount factor;

[0042] Based on any Define a set of security induction strategies:

[0043] ;

[0044] In the formula, For a set of Markov-stationary policies;

[0045] like Its cumulative constraint cost satisfies:

[0046] ;

[0047] In the formula, D represents the expected value of the cumulative discount constraint violation; This represents the parameterization strategy to be optimized. This indicates the security baseline policy.

[0048] Specifically, the Markov decision process for constructing security assessments includes:

[0049] State vector :

[0050] ;

[0051] In the formula, The charging status of the battery energy storage system; This represents the operating status of the first gas turbine. This refers to the operating status of the second gas turbine; The power of the proton exchange membrane electrolyzer; The power of the hydrogen fuel cell; This refers to the hydrogen storage level; The predicted photovoltaic power; This represents the actual photovoltaic power. For real-time electricity prices;

[0052] action :

[0053] ;

[0054] In the formula, Adjusting the charge and discharge cycle of BESS; Adjustments were made for the power generation of the first gas turbine. Adjustments were made for the power generation of the second gas turbine; For adjusting the power of the hydrogen electrolyzer; Adjustments for power generation from fuel cells; Adjustments for hydrogen energy storage; For adjusting photovoltaic power;

[0055] award:

[0056] ;

[0057] In the formula, Scaling of rewards; Let HHEES be the operating cost at time t; γ is a predefined constant and is a scaling factor;

[0058] ;

[0059] In the formula, A collection of devices; For the secondary operating costs of each piece of equipment, Let be the power of each device at time t. , , Cost parameters for each piece of equipment; For electricity purchase costs, For power purchase status indication, For the purchased power output, The electricity price at the time of purchase; For electricity sales revenue, This is an indicator of electricity sales status. The amount of electricity sold. Let be the electricity price at time t; and These are penalties for over-generation and under-generation, respectively.

[0060] Constraint violation:

[0061] .

[0062] In the formula, Let be the degree of constraint violation at time t; Let v be the specific value of the constraint violation at time t, where v is the abbreviation for violation.

[0063] Specifically, the strategy optimization is achieved using an improved Actor-Critic algorithm as follows:

[0064] At time t, the environment first provides the state. The agent acts according to the policy. Will Mapped to action probability distribution, output action Next, an assessment. Security, constrained violation degree Returning the modified action to the agent. The reward is passed to the environment. Then, the environment determines the reward based on the reward function R and the state transition probability P. and the next state Repeating this process will generate a trajectory. The goal of the intelligent agent is to implement a trajectory optimization strategy. To maximize the cumulative discount reward J while ensuring the degree of violation of the cumulative discount constraint. If the limit is not exceeded, it can be represented as a constrained optimization problem:

[0065]

[0066]

[0067] Given a safe baseline policy The Lyapunov family of functions is defined as follows:

[0068] ;

[0069] In the formula, It is the general Bellman operator, represented as:

[0070] ;

[0071] In the formula, For state Any function; For state and actions Any function; To start from the current state Take action The next state reached afterward; if ,but ;

[0072] The update process must be located in the induced set. Therefore, the constraint in the above equation can be expressed in one step as:

[0073] ;

[0074] ;

[0075] in, For enhanced security Q-function, To integrate optimal auxiliary constraint costs Lyapunov function, As a safety baseline strategy; The optimal auxiliary constraint cost.

[0076] Using the deterministic policy gradient theorem, the gradient of the accumulated discount reward with respect to the policy parameters is expressed as:

[0077] ;

[0078] In the formula, For state-based The state expectation operator; For strategy The action value function is given by R, which is short for reward.

[0079] Subsequently, using a first-order Taylor expansion, the objective function is... The change at point can be approximated as follows:

[0080]

[0081] In the formula, g represents the objective function with respect to the network parameters. The gradient of the policy update, where k represents the number of iterations for the policy update; Represented as Therefore, the above formula can be further expressed as:

[0082] ;

[0083] In the formula, when When fixed, Become a constant; for small trust regions, maximize Updates can be obtained within this area. :

[0084] ;

[0085] ;

[0086] In the formula, This is the parameter vector for the (k+1)th iteration; To find the function that maximizes the subsequent function g is the objective function in gradient vector at To optimize variables; This is the current parameter vector for the k-th iteration. This is a first-order Taylor approximation of the objective function; It is a quadratic expression; It is the Hessian matrix related to the Kolb-Leibler divergence; This is the trust region radius parameter. The present invention has the following advantages:

[0087] The proposed safety reinforcement learning framework based on Lyapunov constraints, combined with an improved actor-critic algorithm, constructs a Markov decision process that includes safety assessment, defines a safety-induced policy set using Lyapunov functions, and evaluates and corrects actions in real time using a mixed-integer programming model. This effectively solves the problems of existing convex optimization algorithms being unable to handle the non-convex and nonlinear characteristics of systems and the safety hazards caused by traditional reinforcement learning relying on soft constraints. While ensuring the safe operation of hydrogen-electric hybrid energy systems, it also optimizes operating costs and improves adaptability to dynamic scenarios and generalization to complex systems. Attached Figure Description

[0088] Figure 1 This is a topology diagram of HHEES;

[0089] Figure 2 This is a diagram illustrating the changing trend of reward value.

[0090] Figure 3 This is a diagram illustrating the total training cost.

[0091] Figure 4 This is a schematic diagram illustrating the overall energy fitting effect;

[0092] Figure 5 This is a schematic diagram of the comprehensive energy flow analysis over a 24-hour period.

[0093] Figure 6 This diagram illustrates the changes in imbalance during the training process of the SAC system.

[0094] Figure 7 A schematic diagram illustrating how the imbalance changes over time.

[0095] Figure 8 This is a diagram illustrating the system's 24-hour cost records. Detailed Implementation

[0096] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the invention and are not intended to limit the invention; that is, the described embodiments are merely some embodiments of the invention, and not all embodiments. The components of the embodiments of the invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0097] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0098] It should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0099] The present invention will be further described below with reference to the accompanying drawings, but the scope of protection of the present invention is not limited to the following description.

[0100] like Figures 1 to 8 As shown, an energy management method for hydrogen microgrids based on Lyapunov secure reinforcement learning is proposed. This method includes:

[0101] Construct a system operating cost model; the system operating cost model is as follows:

[0102] ;

[0103] In the formula, T represents the total number of time periods within the optimization time domain; , , These are the operating costs of gas turbines, battery energy storage systems, and hydrogen energy systems, respectively. The cost of purchasing electricity from the main power grid;

[0104] ;

[0105] ;

[0106] ;

[0107] ;

[0108] ;

[0109] ;

[0110] ;

[0111] ;

[0112] In the formula, This refers to the output power of the gas turbine. , and For coefficients; and These are the start-up and shutdown costs of the gas turbine, respectively. and These are binary variables indicating the start / stop status of the gas turbine; This refers to the volume of electricity imports or exports. For electricity price; Indicates the number of batteries; The absolute value of the energy change for each battery; The degradation cost per unit energy change; Costs related to battery degradation; For time The input power of the battery at that time; For time The battery's output power at that time; This is the degradation cost coefficient; For time The power of the electrolytic cell; This refers to the unit operating cost of the electrolytic cell; It serves as a start-up indicator for the electrolytic cell; The startup cost of the electrolytic cell; For time The power of the hydrogen fuel cell; The unit operating cost of hydrogen fuel cells; For the start-up indicator of hydrogen fuel cells; For the start-up cost of hydrogen fuel cells, It is the unit operating cost of a hydrogen energy system. This indicates the energy change in a hydrogen energy system; It is the degradation cost coefficient of the hydrogen energy system.

[0113] The energy scheduling strategy is defined as a reinforcement learning action, and a mixed-integer programming model is used for real-time safety assessment and correction.

[0114] The action 'a' derived from reinforcement learning is defined as the energy management strategy of the Hybrid Hydrogen Energy Storage System (HHEES). The reinforcement learning action 'a' is a vector that includes the outputs of the photovoltaic, gas turbine, battery storage system, electrolyzer, hydrogen fuel cell, and thermal energy storage system. The safety of the reinforcement learning action 'a' is evaluated using a mixed-integer programming model, expressed as:

[0115] ;

[0116] In the formula, For safety reasons, For the original action; to make (i.e., safe actions) Minimize the difference from the original action a. The system must meet the operational constraints of core equipment such as photovoltaics, gas turbines, battery energy storage systems, electrolyzers, hydrogen fuel cells, and hydrogen storage systems. This includes constraints on the power output range, energy conversion efficiency relationships, state change rate limits, and energy storage boundaries of each device. Specifically, this includes: matching constraints between actual and predicted power output of photovoltaics and the reduction ratio; constraints on the start-up and shutdown states, ramp rate, and minimum continuous operation / downtime of gas turbines; constraints on the charging and discharging power limits, state of charge boundaries, and charging and discharging efficiency of battery energy storage; constraints on the power-hydrogen production conversion of electrolyzers; constraints on the power-hydrogen consumption matching of hydrogen fuel cells; and constraints on the hydrogen storage capacity range, charging and discharging rates, and losses of hydrogen storage systems.

[0117] Constraint violation degree (cv) and corrective action (a) m Determined based on the following two situations:

[0118] Case 1: If the energy management scheme is considered safe, then the constraint violation degree (cv) and the corrected action are... The following formula is derived:

[0119] ;

[0120] ;

[0121] Scenario 2: If the energy management plan is unsafe, use the following formula to determine c. v and a m :

[0122] ;

[0123] In the formula, This is the proportionality coefficient; It is a constant; c v and a m The strategy used to update the agent.

[0124] A photovoltaic equipment operation model, a gas turbine model, a battery energy storage system model, and a hydrogen energy system model are constructed, and these models are embedded as constraints in a safety assessment mechanism.

[0125] The photovoltaic (PV) equipment operation model primarily reflects the relationship between actual output power, predicted power, and output adjustment. Actual output power is influenced by predicted power, PV reduction ratio, and inverter conversion efficiency. It is necessary to ensure that the actual output power does not exceed the upper limit of the predicted power, and that the reduction ratio is within a reasonable range, in order to balance energy utilization efficiency and system stability. This can be expressed as follows:

[0126] ;

[0127] ;

[0128] ;

[0129] In the formula, This represents the actual output power of the photovoltaic system at that moment. Indicates the predicted photovoltaic power; Indicates the percentage reduction in photovoltaic power generation; This indicates the inverter efficiency.

[0130] The gas turbine model focuses on power output range and operating state constraints: its output power must be between the minimum and maximum allowable values ​​and is directly affected by start-up and shutdown states (operation or shutdown); simultaneously, considering the physical characteristics of the equipment, the power change rate (ramp rate) must be controlled within a safe range to avoid equipment damage or system fluctuations caused by sudden increases or decreases; furthermore, it must meet minimum continuous operating time and downtime constraints to prevent frequent start-ups and shutdowns from shortening equipment lifespan. This is expressed as follows:

[0131] ;

[0132] ;

[0133] ;

[0134] ;

[0135] In the formula, This represents the minimum power output of the gas turbine. This represents the maximum operating power of the gas turbine. Indicates the on / off status of the gas turbine; and It is a slope rate limitation; as well as It is the minimum on / off state time.

[0136] The battery energy storage system model revolves around charging and discharging behavior and energy state: Both charging and discharging power have their own maximum limits, and charging and discharging states are mutually exclusive (only one state can be charging, discharging, or idle at a time); the energy storage capacity needs to be maintained between the minimum and maximum capacity to prevent overcharging and over-discharging from damaging the battery; energy loss occurs during charging and discharging, and the difference between actual usable energy and theoretical energy needs to be reflected through an efficiency coefficient to ensure the accuracy of energy storage scheduling. This is expressed as follows:

[0137] ;

[0138] ;

[0139] ;

[0140] ;

[0141] In the formula, and These represent charging power and discharging power, respectively. This refers to the battery's maximum charge and discharge power. Let be the discharge power of the battery at time t; For battery charging efficiency; Indicates the charging / discharging state; This indicates the energy stored in a battery energy storage system (BESS). and These represent the minimum and maximum allowable energy storage capacity of the battery, respectively.

[0142] The hydrogen energy system model encompasses the synergistic constraints of the electrolyzer, hydrogen fuel cell (HFC), and hydrogen storage system (HES). The limiting factors of the hydrogen energy system are as follows:

[0143] The operation of an electrolyzer must be matched to its power input range (minimum and maximum power limits). The amount of hydrogen produced and the electrical energy consumed are related through electrolysis efficiency and energy conversion coefficient, reflecting the conversion law of electrical energy to hydrogen energy. The model is as follows:

[0144] ;

[0145] In the formula, The amount of hydrogen produced in the electrolyzer at time t; Let t be the amount of hydrogen stored in the hydrogen storage tank at time t; The amount of hydrogen consumed by the hydrogen fuel cell at time t; is the electrolysis efficiency; k is the conversion coefficient from electricity to hydrogen. The power consumption of the electrolyzer is [value missing]; the power of the hydrogen fuel cell is directly related to the amount of hydrogen consumed and the energy conversion efficiency, subject to the following constraints:

[0146] ;

[0147] ;

[0148] ;

[0149] ;

[0150] In the formula, The output power of a hydrogen fuel cell (HFC); The amount of hydrogen available for a hydrogen fuel cell; Energy conversion efficiency; Power consumed by hydrogen; This represents the hydrogen consumption coefficient of a hydrogen fuel cell. and These are the minimum and maximum hydrogen consumption power, respectively. and These represent the maximum permissible rate of decrease and the maximum permissible rate of increase of hydrogen consumption power, respectively; when the output power... When the value is not zero, constraints ensure the safe operation of the hydrogen fuel cell and limit its ramp rate.

[0151] Hydrogen storage systems must control the amount of hydrogen stored within a safe capacity range, while limiting the hydrogen storage and release rates (including energy losses during storage and release) to ensure the physical feasibility of hydrogen energy storage and retrieval. The constraints are as follows:

[0152] ;

[0153] ;

[0154] ;

[0155] ;

[0156] In the formula, Let t be the amount of hydrogen stored in the hydrogen storage device. / These represent the hydrogen storage capacity and hydrogen release capacity of the hydrogen storage device at time t, respectively. This is the upper limit for hydrogen storage; and These are the upper limits of the hydrogen storage rate and the hydrogen release rate per unit time for the hydrogen storage device, respectively. This is the storage loss coefficient; and These are the storage loss coefficient and release loss coefficient of the hydrogen storage device, respectively.

[0157] For electrolyzers, hydrogen fuel cells (HFCs), and hydrogen storage systems (HS), the following relationships exist regarding hydrogen energy storage:

[0158] .

[0159] The aforementioned equipment model defines the safe operating boundaries of each component from dimensions such as power output, energy conversion, and state changes. This forms the basis for subsequent evaluation of action safety using a mixed-integer programming model. These boundaries will be embedded as constraints into the safety assessment mechanism to ensure that actions generated by reinforcement learning strategies (such as photovoltaic power output adjustment, energy storage charging and discharging control, and hydrogen energy equipment scheduling) do not exceed the physical limits of the equipment, ultimately achieving synergistic optimization of safety and economy.

[0160] A Lyapunov function is introduced to construct a set of security-inducible strategies; a security assessment mechanism is combined to balance the economy of reinforcement learning strategies with the security of system operation.

[0161] The Lyapunov function is introduced to construct a set of security-inducing strategies, including:

[0162] In the security reinforcement learning framework, the key to policy optimization lies in gradually approaching the optimal solution while ensuring system security. To this end, a security baseline policy constructed by a security assessment and optimization model is introduced. Its cumulative constraint violations satisfy:

[0163] ;

[0164] In the formula, To constrain the discount factor; This represents the cumulative constraint cost of the baseline policy under the initial state; Let Lyapunov functions be the maximum allowed threshold. Define a set of such functions:

[0165] ;

[0166] In the formula, Denotes the set of Lyapunov functions. As a strategy, For strategy parameters; It is a Lyapunov function; it represents the state space of the microgrid. Mapping to the set of real numbers ; In strategy The Bellman constraint operator is given below, where h is the immediate safety index function; It is the initial state; This allows for the cumulative constraint to violate the maximum value;

[0167] ;

[0168] In the formula, The mathematical expectation is represented by action a and the next state. produce; For the constraint cost function; This is a constraint discount factor; this condition is equivalent to the Lyapunov descent condition, used to ensure that the system constraint level does not exceed the safety limit during dynamic evolution.

[0169] Based on any Define a set of security induction strategies:

[0170] ;

[0171] In the formula, For a set of Markov-stationary policies;

[0172] like Its cumulative constraint cost satisfies:

[0173] .

[0174] In the formula, D represents the expected value of the cumulative discount constraint violation; This represents the parameterization strategy to be optimized. This indicates the security baseline policy.

[0175] This ensures that the strategy remains within the safe and feasible domain during both the training and execution phases.

[0176] A Markov decision process for security assessment is constructed, and an improved Actor-Critic algorithm is used to optimize the strategy, enabling efficient solution of the security and economic synergy optimization problem and obtaining management solutions.

[0177] HHEES first provides status information. Subsequently, the agent, based on the status... and its strategy Output the corresponding action .here, Indicates the state Select action The probability distribution is determined by the parameters. Decision (i.e.) If a constraint violation occurs, it will return a violation message. Give the agent the correct action to generate a new action. The modified action is then output to HHEES. Next, using the reward function and state transition probabilities, the hydrogen microgrid returns the reward corresponding to the action. And the state in the next moment. This iterative loop generates states, actions, rewards, and constraint violation information, ultimately forming a trajectory: .

[0178] The agent aims to enhance its control policy by maximizing the cumulative discount reward J, while ensuring that the cumulative discount constraint violation Dis remains within acceptable limits. This objective can be formulated as the following constraint optimization problem:

[0179] ;

[0180] ;

[0181] In the formula, and Discount factor; For the trajectory length; and It is the maximum allowable value for the expected cumulative constraint violation strategy.

[0182] The Markov decision process for constructing security assessments includes:

[0183] State vector :

[0184] ;

[0185] In the formula, The charging status of the battery energy storage system; This represents the operating status of the first gas turbine. This refers to the operating status of the second gas turbine; The power of the proton exchange membrane electrolyzer; The power of the hydrogen fuel cell; This refers to the hydrogen storage level; The predicted photovoltaic power; This represents the actual photovoltaic power. This represents the real-time electricity price; this comprehensive status representation captures the operational status of HHEES at time t.

[0186] action : Represents an energy management strategy, defined by the following equation:

[0187] ;

[0188] In the formula, Adjusting the charge and discharge cycle of BESS; Adjustments were made for the power generation of the first gas turbine. Adjustments were made for the power generation of the second gas turbine; For adjusting the power of the hydrogen electrolyzer; Adjustments for power generation from fuel cells; Adjustments for hydrogen energy storage; For adjusting photovoltaic power;

[0189] Incentive: The incentive is defined as a negative deviation from the operating cost of HHEES, designed to optimize economic efficiency.

[0190] ;

[0191] In the formula, Scaling of rewards; Let HHEES be the operating cost at time t; A predefined constant γ is a scaling factor that ensures that rewards increase as the strategy is optimized, while remaining within a reasonable range;

[0192] ;

[0193] In the formula, A collection of devices; For the secondary operating costs of each piece of equipment, Let be the power of each device at time t. , , Cost parameters for each piece of equipment; For electricity purchase costs, For power purchase status indication, For the purchased power output, The electricity price at the time of purchase; For electricity sales revenue, This is an indicator of electricity sales status. The amount of electricity sold. Let be the electricity price at time t; and These are penalties for over-generation and under-generation, respectively.

[0194] Constraint violation: Constraint violation is defined as the degree of constraint violation obtained by solving the following mathematical model;

[0195] ;

[0196] In the formula, Let be the degree of constraint violation at time t; Let v be the specific value of the constraint violation at time t, where v is the abbreviation for violation.

[0197] Furthermore, the improved Actor-Critic algorithm is used to optimize the strategy as follows:

[0198] At time t, the environment first provides the state. The agent acts according to the policy. Will Mapped to action probability distribution, output action Next, an assessment. Security, constrained violation degree Returning the modified action to the agent. The reward is passed to the environment. Then, the environment determines the reward based on the reward function R and the state transition probability P. and the next state Repeating this process will generate a trajectory. The goal of the intelligent agent is to implement a trajectory optimization strategy. To maximize the cumulative discount reward J while ensuring the degree of violation of the cumulative discount constraint. If the limit is not exceeded, it can be represented as a constrained optimization problem:

[0199]

[0200]

[0201] This method utilizes Lyapunov functions to construct a set of safety-inducing policies, including the optimal policy. ;

[0202] Given a safe baseline policy The Lyapunov family of functions is defined as follows:

[0203] ;

[0204] In the formula, B is the general Bellman operator, expressed as:

[0205] ;

[0206] In the formula, For state Any function; For state and actions Any function; To start from the current state Take action The next state reached afterward; if ,but ;

[0207] To ensure safety and approximate the optimal policy during policy optimization, the update process must reside within the induced set. Therefore, the constraint in the above equation can be expressed in one step as:

[0208] ;

[0209] ;

[0210] in, For enhanced security Q-function, To integrate optimal auxiliary constraint costs Lyapunov function, As a safety baseline strategy; The optimal auxiliary constraint cost.

[0211] Using the deterministic policy gradient (DPG) theorem, the gradient of the accumulated discount reward with respect to the policy parameters is expressed as:

[0212] ;

[0213] In the formula, For state-based The state expectation operator; For strategy The action value function is given by R, which is short for reward.

[0214] Subsequently, using a first-order Taylor expansion, the objective function is... The change at point can be approximated as follows:

[0215]

[0216] In the formula, g represents the objective function with respect to the network parameters. The gradient of the policy update, where k represents the number of iterations for the policy update; It can be represented as Therefore, the above formula can be further expressed as:

[0217] ;

[0218] In the formula, when When fixed, Become a constant; for small trust regions, maximize Updates can be obtained within this area. :

[0219] ;

[0220] ;

[0221] In the formula, This is the parameter vector for the (k+1)th iteration; To find the function that maximizes the subsequent function g is the objective function in gradient vector at To optimize variables; This is the current parameter vector for the k-th iteration. This is a first-order Taylor approximation of the objective function; It is a quadratic expression; It is the Hessian matrix related to the Kolb-Leibler divergence; This is the trust region radius parameter. Figure 2 The study demonstrates the trend of reward values ​​during reinforcement learning training. In the early stages of training, reward values ​​are low and fluctuate significantly, indicating that the agent's policy is not yet stable and there is considerable exploratory behavior. Subsequently, the reward values ​​gradually increase and stabilize, suggesting that the agent is continuously learning and adjusting its policy, gradually finding better action plans to obtain higher rewards. Finally, the fluctuation range of the reward values ​​decreases and remains at a near-stable level, indicating that the agent's policy is relatively stable, the training process is successful, and the model has effectively optimized the policy to maximize rewards.

[0222] Figure 3 The system demonstrates the power supply and demand situation over a 24-hour period. It performs exceptionally well in handling daily load demand fluctuations. Particularly noteworthy is its effective management through the integration of multiple energy sources, particularly during the peak net load demand of approximately 1250 kW at time step 14. Gas turbine 1 contributes the major share of approximately 750 kW, supplemented by hydrogen fuel cells (approximately 300 kW), solar photovoltaic power generation (approximately 200 kW), and battery discharge (approximately 200 kW). Furthermore, the system significantly improves energy efficiency by charging the battery during off-peak hours (approximately 50-80 kW per charge) and discharging during peak demand periods (e.g., approximately 200 kW at time step 14). Notably, the system exhibits significant grid output during time steps 16-18, demonstrating effective load management.

[0223] Figure 4 The diagram illustrates the varying supply ratios of different energy sources within the SAC system throughout the day: solar power peaks between 3 PM and 6 PM, supplying nearly 100% of the power; batteries primarily provide power between midnight and 4 AM and between 6 PM and 9 PM, reaching a peak of approximately 80%; hydrogen energy contributes significantly in the early morning and at night, exceeding 50% at its peak; gas turbines supply approximately 60% of the power between 4 AM and 10 AM; while the grid only intervenes slightly between 3 AM and 5 AM, accounting for less than 10%. The diverse and complementary energy structure ensures a stable power supply throughout the day.

[0224] Figure 5 The changes in system imbalance on the training set are illustrated. Initially, the system imbalance value fluctuated wildly, with peak values ​​exceeding 200 and even 400 multiple times, indicating a possible significant load imbalance or anomaly in the system. Over time, the imbalance value decreased significantly, remaining near zero for most of the time. This demonstrates that after adjustments or optimizations, the system's operation became more stable, fluctuations were significantly reduced, and overall system performance improved. This proves that the SAOM-Lyapunov strategy can produce positive effects during training.

[0225] Based on broken lines Figure 6 The data shows that the system's load-supply matching exhibits significant synergistic characteristics within 20 hours: the coefficient of determination R² between actual total equipment output (blue) and net load demand (orange) is 0.902, indicating that the model has a strong predictive ability for overall supply and demand trends. Particularly in the 11-15 hour range, the two curves are highly consistent (error < 50 kW). The root mean square error (RMSE) is 80.02 kW (6.4% of maximum power), verifying the system's global stability within a dynamic range of 1250 kW. The green supply surplus curve continuously rises during the 0-5 hour period (from -250 kW to 800 kW), intuitively reflecting the system's flexible response capability in the early low-load phase, providing a clear visual basis for optimizing adjustment strategies.

[0226] Figure 7 The data shows that for 80% of the system's time periods (0-10 hours and 16-20 hours), the operating cost remained consistently below the average of $571.69 per hour, reflecting the intensive nature of basic operations. Although a single cost spike of $3,002 occurred in the 13th hour, the cost rapidly decreased after that period (the cost in the 16-20 hour period decreased by 49.8% compared to the peak), validating the system's resilient response to sudden load spikes. Figure 8 The data shows that for 80% of the system's time periods (0-10 hours and 16-20 hours), the operating cost remained consistently below the average of $571.69 per hour, reflecting the intensive nature of basic operations. Although a single cost spike of $3,002 occurred in the 13th hour, the cost rapidly decreased after that period (the cost in the 16-20 hour period decreased by 49.8% compared to the peak), validating the system's resilient response to sudden load spikes.

[0227] The above description is merely a preferred embodiment of the present invention and does not constitute any limitation on the present invention. Any person skilled in the art can make many possible variations and modifications to the technical solution of the present invention, or modify it into equivalent embodiments, without departing from the scope of the present invention. Therefore, any modifications, equivalent changes, and alterations made to the above embodiments based on the technology of the present invention without departing from the scope of the present invention are within the protection scope of the present invention.

Claims

1. An energy management method for hydrogen microgrids based on Lyapunov secure reinforcement learning, characterized by: The method includes, Construct a system operating cost model; The energy scheduling strategy is defined as a reinforcement learning action, and a mixed-integer programming model is used for real-time safety assessment and correction. A Lyapunov function is introduced to construct a set of security-inducible strategies; a security assessment mechanism is combined to balance the economy of reinforcement learning strategies with the security of system operation. A Markov decision process for security assessment is constructed, and an improved Actor-Critic algorithm is used to optimize the strategy and obtain a management plan.

2. The energy management method for hydrogen microgrids based on Lyapunov secure reinforcement learning according to claim 1, characterized in that: The system operating cost model is as follows: ; In the formula, T represents the total number of time periods within the optimization time domain; , , These are the operating costs of gas turbines, battery energy storage systems, and hydrogen energy systems, respectively. The cost of purchasing electricity from the main power grid; ; ; ; ; ; ; ; ; In the formula, This refers to the output power of the gas turbine. , and For coefficients; and These are the start-up and shutdown costs of the gas turbine, respectively. and These are binary variables indicating the start / stop status of the gas turbine; This refers to the volume of electricity imports or exports. For electricity price; Indicates the number of batteries; The absolute value of the energy change for each battery; The degradation cost per unit energy change; Costs related to battery degradation; For time The input power of the battery at that time; For time The battery's output power at that time; This is the degradation cost coefficient; For time The power of the electrolytic cell; This refers to the unit operating cost of the electrolytic cell; It serves as a start-up indicator for the electrolytic cell; The startup cost of the electrolytic cell; For time The power of the hydrogen fuel cell; The unit operating cost of hydrogen fuel cells; For the start-up indicator of hydrogen fuel cells; For the start-up cost of hydrogen fuel cells, It is the unit operating cost of a hydrogen energy system. This indicates the energy change in a hydrogen energy system; It is the degradation cost coefficient of the hydrogen energy system.

3. The energy management method for hydrogen microgrids based on Lyapunov secure reinforcement learning according to claim 1, characterized in that: A photovoltaic equipment operation model, a gas turbine model, a battery energy storage system model, and a hydrogen energy system model are constructed, and these models are embedded as constraints in a safety assessment mechanism.

4. The energy management method for hydrogen microgrids based on Lyapunov secure reinforcement learning according to claim 1, characterized in that: The reinforcement learning action 'a' is a vector that includes the outputs of photovoltaic, gas turbine, battery energy storage system, electrolyzer, hydrogen fuel cell, and thermal energy storage system.

5. The energy management method for hydrogen microgrids based on Lyapunov secure reinforcement learning according to claim 4, characterized in that: The safety of reinforcement learning action 'a' is evaluated using a mixed-integer programming model, which is represented as: ; In the formula, For safety reasons, This is the original action; Constraint violation degree (cv) and corrective action (a) m Determined based on the following two situations: If the energy management scheme is considered safe, then the constraint violation degree (cv) and the corrected action are... The following formula is derived: ; ; If the energy management plan is unsafe, use the following formula to determine c. v and a m : ; In the formula, This is the proportionality coefficient; It is a constant; c v and a m The strategy used to update the agent.

6. The energy management method for hydrogen microgrids based on Lyapunov secure reinforcement learning according to claim 4, characterized in that: The Lyapunov function is introduced to construct a set of security-inducing strategies, including: Introducing a security baseline strategy constructed from a security assessment optimization model Its cumulative constraint violations satisfy: ; In the formula, To constrain the discount factor; This represents the cumulative constraint cost of the baseline policy under the initial state; This allows for the cumulative constraint to violate the maximum value; Define a set of Lyapunov functions: ; In the formula, Denotes the set of Lyapunov functions. As a strategy, For strategy parameters; It is a Lyapunov function; it represents the state space of the microgrid. Mapping to the set of real numbers ; In strategy The Bellman constraint operator is given below, where h is the immediate safety index function; It is the initial state; This allows for the cumulative constraint to violate the maximum value; ; In the formula, The mathematical expectation is represented by action a and the next state. produce; For the constraint cost function; To constrain the discount factor; Based on any Define a set of security induction strategies: ; In the formula, For a set of Markov-stationary policies; like Its cumulative constraint cost satisfies: ; In the formula, D represents the expected value of the cumulative discount constraint violation; This represents the parameterization strategy to be optimized. This indicates the security baseline policy.

7. The energy management method for hydrogen microgrids based on Lyapunov secure reinforcement learning according to claim 4, characterized in that: The Markov decision process for constructing security assessments includes: State vector : ; In the formula, The charging status of the battery energy storage system; This represents the operating status of the first gas turbine. This refers to the operating status of the second gas turbine; The power of the proton exchange membrane electrolyzer; The power of the hydrogen fuel cell; This refers to the hydrogen storage level; The predicted photovoltaic power; This represents the actual photovoltaic power. For real-time electricity prices; action : ; In the formula, Adjusting the charge and discharge cycle of BESS; Adjustments were made for the power generation of the first gas turbine. Adjustments were made for the power generation of the second gas turbine; For adjusting the power of the hydrogen electrolyzer; Adjustments for power generation from fuel cells; Adjustments for hydrogen energy storage; For adjusting photovoltaic power; award: ; In the formula, Scaling of rewards; Let HHEES be the operating cost at time t; γ is a predefined constant and is a scaling factor; ; In the formula, A collection of devices; The secondary operating cost of each piece of equipment; Let be the power of each device at time t; , , Cost parameters for each piece of equipment; For electricity purchase costs; This indicates the electricity purchase status. The purchased power output; The electricity price at the time of purchase; Revenue from electricity sales; This indicates the electricity sales status. The power output of electricity sold; Let be the electricity price at time t; and These are penalties for over-generation and under-generation, respectively. Constraint violation: ; In the formula, Let be the degree of constraint violation at time t; Let v be the specific value of the constraint violation at time t, where v is the abbreviation for violation.

8. The energy management method for hydrogen microgrids based on Lyapunov secure reinforcement learning according to claim 7, characterized in that: The strategy optimization is achieved using an improved Actor-Critic algorithm: At time t, the environment first provides the state. The agent acts according to the policy Will Mapped to action probability distribution, output action Next, an assessment Security, constrained violation degree Returning the modified action to the agent. The information is passed to the environment, which then determines the reward based on the reward function R and the state transition probability P. and the next state Repeating this process will generate a trajectory. The goal of the intelligent agent is to optimize the trajectory strategy. To maximize the cumulative discount reward J while ensuring the degree of violation of the cumulative discount constraint. If the limit is not exceeded, it can be represented as a constrained optimization problem: ; ; Given security baseline policy The Lyapunov family of functions is defined as follows: ; In the formula, B is the general Bellman operator, expressed as: ; In the formula, For state Any function; For state and actions Any function; To start from the current state Take action The next state reached afterward; if ,but ; The update process must be located in the induced set. Therefore, the constraint in the above equation can be further expressed as: ; ; in, For enhanced security Q-functions; To integrate optimal auxiliary constraint costs Lyapunov functions; As a safety baseline strategy; The optimal auxiliary constraint cost; Using the deterministic policy gradient theorem, the gradient of the accumulated discount reward with respect to the policy parameters is expressed as: ; In the formula, For state-based The state expectation operator; For strategy The action value function is given by R, which is short for reward. Subsequently, using a first-order Taylor expansion, the objective function is... The change at point can be approximated as follows: ; In the formula, g represents the objective function with respect to the network parameters. The gradient of the policy update, where k represents the number of iterations for the policy update; Represented as Therefore, the above formula can be further expressed as: ; In the formula, when When fixed, Become a constant; for small trust regions, maximize Updates can be obtained within this area. : ; ; In the formula, This is the parameter vector for the (k+1)th iteration; To find the function that maximizes the subsequent function g is the objective function in gradient vector at To optimize variables; This is the current parameter vector for the k-th iteration; This is a first-order Taylor approximation of the objective function; It is a quadratic expression; It is the Hessian matrix related to the Kolb-Leibler divergence; This is the trust region radius parameter.