Hydrogen-based community microgrid optimization scheduling method based on clean hydrogen constraint

By constructing a refined electric-carbon coupled electrolyzer model and Lagrange multiplier optimization algorithm, the constraint problem of clean hydrogen standards in hydrogen-based community microgrid systems was solved, and optimal scheduling that simultaneously meets cleanliness and economic benefits was achieved in a dynamic environment.

CN120706641APending Publication Date: 2025-09-26HEBEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510821694.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing technologies fail to integrate clean hydrogen standards as mandatory constraints into the optimized scheduling framework of hydrogen-based community microgrid systems, resulting in carbon emissions from the hydrogen production process unable to meet increasingly stringent cleanliness requirements, and a lack of effective adaptive scheduling strategies to simultaneously meet the dual requirements of economic benefits and clean hydrogen production in a dynamic environment.

Method used

A refined electric-carbon coupled electrolyzer model is constructed, a dynamic hydrogen cleanliness evaluation model and a step-by-step carbon trading model are introduced, and the Lagrange multiplier and the PPO algorithm are combined to form the LC-PPO algorithm. The Lagrange multiplier is adaptively adjusted to meet the clean hydrogen standard constraints and optimize the scheduling strategy.

Benefits of technology

It achieves dynamic adjustment of hydrogen production and use strategies under the premise of meeting clean hydrogen standards, ensures that the hydrogen cleanliness index always remains below the threshold, maximizes economic benefits, accurately characterizes hydrogen production efficiency and carbon emission behavior under different power source and load power combinations, and provides an adaptive optimization scheduling method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706641A_ABST
    Figure CN120706641A_ABST
Patent Text Reader

Abstract

The invention discloses a hydrogen-based community microgrid optimization scheduling method based on clean hydrogen constraint. The method comprises the following steps: firstly, constructing system models including a refined electric carbon coupling electrolytic cell model, a dynamic hydrogen standard evaluation model, a hydrogen fuel cell model, a hydrogen storage tank hydrogen quality updating model, a stepped carbon transaction model, an HCMS scheduling model and optimal scheduling constraints; the refined electric-carbon coupling electrolytic cell model considers multiple operation states, hydrogen production efficiency and power source carbon emission tracking of an electrolytic cell; then, constructing a CMDP model based on the system model; and finally, a Lagrange multiplier is introduced into a PPO algorithm to solve the CMDP model, an optimal strategy is obtained, and the expectation cumulative discount cost constraint is met while the expectation cumulative discount reward is maximized. A clean hydrogen standard is used as a key constraint to be introduced into an optimization process, optimization scheduling considering real-time electricity price and power grid carbon emission intensity at the same time is constructed, and economic benefits are maximized on the premise that the clean hydrogen standard is met by dynamically adjusting a hydrogen production and use strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of microgrid optimization and scheduling, and specifically provides a hydrogen-based community microgrid optimization and scheduling method based on clean hydrogen constraints. Background Art

[0002] As the global energy transition accelerates, hydrogen, as a key clean energy carrier, demonstrates tremendous potential in distributed energy systems such as community microgrids. However, the process of hydrogen production from water electrolysis often relies on grid power, and its inherent carbon emission intensity raises questions about the cleanliness of hydrogen, making it difficult to meet increasingly stringent "clean hydrogen" standards. These standards typically set clear upper limits on greenhouse gas emissions per unit of hydrogen, making carbon footprint management during hydrogen production a key challenge in the planning and operation of hydrogen-based community microgrid systems (HCMSs). Limited by the intermittent nature of renewable energy and installed capacity, HCMSs often rely on grid power to ensure a stable supply of electricity. Grid power consumption directly impacts the carbon emissions of the hydrogen production process. However, existing technologies generally fail to incorporate the dynamic carbon emission intensity of grid power into the accurate calculation of hydrogen's carbon footprint, nor do they integrate clean hydrogen standards as mandatory constraints into system optimization and scheduling frameworks. As a result, operational strategies may fail to ensure that actual hydrogen production meets cleanliness requirements.

[0003] Accurately assessing and meeting clean hydrogen standards first relies on accurate modeling of electrolyzer energy consumption and carbon emissions during the hydrogen production process. In terms of hydrogen production models, traditional technologies usually use fixed efficiency or simplified linear efficiency models, which are difficult to reflect the complexity of actual operations. Although some studies have focused on the nonlinear efficiency characteristics of electrolyzers, revealed the nonlinear relationship between efficiency and input power, and explored the state transition mechanism of electrolyzers at different power levels, there is a lack of systematic analysis of the energy consumption and carbon emission behavior of electrolyzers under multiple operating states, which directly leads to significant deviations in the quantification of hydrogen production energy consumption and corresponding carbon emissions (especially when using grid electricity), hindering the effective implementation of clean hydrogen standards.

[0004] Furthermore, community microgrid operations face multiple time-varying uncertainties and uncertainties in load, renewable energy output, electricity prices, and grid carbon intensity, placing high demands on the adaptive capabilities of dispatch strategies. Reinforcement learning is widely used for its ability to handle complex decision-making problems, but existing reinforcement learning methods often focus on economic optimization objectives and lack effective mechanisms to strictly address operational constraints such as the "Clean Hydrogen Standard." In constrained optimization problems, some existing techniques address constraints by introducing penalty terms into the reward function. However, this approach faces difficulties in selecting penalty coefficients, increasing the complexity of algorithm implementation. Another class of techniques restricts default actions through Lyapunov functions or by integrating safety layers into neural networks. This rigid approach of restricting the action space often results in a limited exploration scope and produces overly conservative strategies.

[0005] In summary, existing technologies lack an effective solution to integrate clean hydrogen standards as explicit constraints into the reinforcement learning framework, making it difficult to simultaneously meet the dual requirements of system economic benefits and clean hydrogen production in a dynamic environment. Summary of the Invention

[0006] In view of the deficiencies of the existing technology, the technical problem to be solved by the present invention is to provide a hydrogen-based community microgrid optimization scheduling method based on clean hydrogen constraints.

[0007] The present invention solves the technical problem by adopting the following technical solutions:

[0008] A method for optimizing and scheduling a hydrogen-based community microgrid based on clean hydrogen constraints, comprising the following steps:

[0009] Step 1: Build a system model, including a refined electro-carbon coupled electrolyzer model, a dynamic hydrogen standard evaluation model, a hydrogen fuel cell model, a hydrogen storage tank hydrogen quality update model, a step-by-step carbon trading model, an HCMS scheduling model, and optimized scheduling constraints;

[0010] The refined electro-carbon coupled electrolytic cell model uses state indicator variables to characterize the state of the electrolytic cell, which are defined as follows:

[0011]

[0012] Where: δ e,s (t) represents the state of the electrolytic cell at time t, where s=0 is the standby state, s=1 is the low-load state, s=2 is the variable-load state, and s=3 is the overload state;

[0013] The carbon emissions of an electrolyzer are:

[0014]

[0015] Where: E e,s(t) represents the carbon emissions of the electrolyzer at time t, is the dynamic carbon emission intensity of the power grid at time t, P grid,ely (t) is the grid power consumed by the electrolyzer at time t, and Δt is the time step;

[0016] The hydrogen production of the electrolyzer is:

[0017]

[0018] Where: represents the hydrogen production of the electrolyzer at time t, η e,s (t) represents the hydrogen production efficiency of the electrolyzer at time t, P ely (t) is the input power of the electrolytic cell at time t, is the energy density of hydrogen;

[0019] The dynamic clean hydrogen standard evaluation model is:

[0020]

[0021] Where, is the hydrogen cleanliness index at time t, They are Carbon emissions and hydrogen production of the electrolyzer at each moment;

[0022] The HCMS scheduling model is:

[0023]

[0024] F(t)=F op (t)+F energy (t)+F carbon (t)+F penalty (t) (15)

[0025] F op (t) = c e P ely (t)Δt+c f P fc (t)Δt (16)

[0026]

[0027] F penalty (t) = τ pv P surplus (t)Δt (18)

[0028] Where: F total represents the total operating cost of the system, T represents the scheduling period, F(t) represents the comprehensive cost of the system at time t, and F op(t) represents the operation and maintenance cost at time t, F energy (t) represents the energy purchase cost at time t, F carbon (t) represents the carbon transaction cost at time t, F penalty (t) represents the penalty cost of abandoning light at time t, c e 、c f is the operation and maintenance cost coefficient of the electrolyzer and hydrogen fuel cell, c ele (t) is the electricity price at time t, P grid (t) is the grid interaction power, is the external hydrogen price at time t, is the external hydrogen purchase amount at time t, τ pv is the light abandonment penalty coefficient, P surplus (t) is the unused photovoltaic power;

[0029] Step 2: Construct a CMDP model based on the system model. The CMDP model is defined by the tuple (S, A, P, R, C, d, γ); the state space S: s t ∈S represents the state at time t, which can be expressed as:

[0030]

[0031] Where: P load (t) represents the building load at time t, P pv (t) represents the photovoltaic power generation at time t, M tank (t) represents the hydrogen mass in the hydrogen storage tank of the community hydrogen station at time t, represents the hydrogen charge of the hydrogen vehicle cluster at time t, Indicates the hydrogen consumption of the hydrogen fuel cell at time t;

[0032] Action space A: a t ∈A represents the agent in state s t the actions taken next;

[0033] a t ={P ely (t),P fc (t)} (26)

[0034] State transition probability P: the probability that the system will transition to the next state after taking an action from the current state;

[0035] Reward function R: r t =R(s t ,a t ) indicates that the agent is in state s t Next, perform action a t Instant rewards received t ;

[0036] r t = -F(t) (27)

[0037] Cost function C: c t =C(s t ,a t ) indicates that the agent is in state s t Next, perform action a t the costs incurred;

[0038]

[0039] Where: represents the upper limit of the clean hydrogen standard; d represents the constraint tolerance, and γ represents the discount factor;

[0040] Step 3: Introduce Lagrange multipliers into the PPO algorithm to solve the CMDP model and obtain the optimal strategy, which maximizes the expected cumulative discounted reward while satisfying the expected cumulative discounted cost constraint.

[0041] Furthermore, in the third step, the objective function to be solved is:

[0042]

[0043] Where: λ is a non-negative Lagrange multiplier, represents the expected cumulative discount reward, represents the expected cumulative discounted cost, and E[·] represents the expectation;

[0044] The objective function L of the policy network of the PPO algorithm policy (θ) is:

[0045] L policy (θ)=L CLIP (θ)-λE[A C,t ] (30)

[0046] Where: L CLIP (θ) is the clipping objective function of the PPO algorithm, A C,t Indicates that based on the cost c t Cost advantages of computing;

[0047] The objective function of the value network of the PPO algorithm is:

[0048]

[0049] Where: is the target state value, Indicates state s t the value of

[0050] The objective function of the Lagrange multiplier is:

[0051]

[0052] Where: κλ is the regularization coefficient, is an estimate of the expected cumulative discounted cost under the current policy.

[0053] Furthermore, the hydrogen tank quality update model includes the hydrogen quality update in the community hydrogen station and the on-board hydrogen tank; the hydrogen quality update formula in the community hydrogen station hydrogen tank is:

[0054]

[0055] Where: M tank (t+1) is the hydrogen mass in the hydrogen storage tank of the community hydrogen station at time t+1, is the change in hydrogen mass in the hydrogen storage tank of the community hydrogen station at time t;

[0056] The formula for updating the hydrogen mass in the on-board hydrogen storage tank is:

[0057]

[0058] Where: M tank,hv (t+1), M tank,hv (t) represents the mass of hydrogen in the onboard hydrogen storage tank at time t+1 and t, M hv,use (t) represents the hydrogen consumption of the hydrogen vehicle cluster at time t.

[0059] Furthermore, the ladder-type carbon trading model includes: in the exempt carbon emission range 0 to E1(t), the carbon trading cost is ξ1(E(t)-E1(t)); in the low carbon emission responsibility range E1(t) to E2(t), the carbon trading cost is ξ2(E(t)-E1(t)); in the medium carbon emission responsibility range E2(t) to E3(t), the carbon trading cost is ξ2(E2(t)-E1(t))+ξ3(E(t) -E2(t)); in the high carbon emission responsibility interval E3(t)~∞, the carbon trading cost is ξ2(E2(t)-E1(t))+ξ3(E3(t)-E2(t))+ξ4(E(t)-E3(t)); where E(t) represents the system carbon emissions, E1(t), E2(t), and E3(t) represent the critical carbon emission quotas of adjacent carbon emission responsibility intervals, and ξ1, ξ2, ξ3, and ξ4 are carbon prices.

[0060] Furthermore, in standby mode, the hydrogen production efficiency of the electrolyzer is 0; in low load, variable load and overload modes, the hydrogen production efficiency of the electrolyzer is expressed as:

[0061]

[0062] Where: Pely,N Rated value for the input power to the electrolyzer.

[0063] Compared with the prior art, the present invention has the following beneficial effects:

[0064] (1) This invention innovatively introduces the clean hydrogen standard as a key constraint into the optimization process, constructs an optimized dispatch that takes into account both the real-time electricity price and the carbon emission intensity of the power grid, and maximizes economic benefits while meeting the clean hydrogen standard by dynamically adjusting the hydrogen production and use strategy. After the clean hydrogen standard constraint is introduced, the system's operating strategy will change. The most obvious change is that the system will actively reduce the use of grid electricity for electrolytic hydrogen production during the nighttime period with low electricity prices. Since grid electricity has carbon emission intensity, large-scale use will cause the hydrogen cleanliness index to exceed the clean hydrogen standard threshold. By dynamically suppressing this behavior, it is ensured that the hydrogen cleanliness index always remains below the clean hydrogen standard threshold.

[0065] (2) The present invention breaks through the traditional fixed efficiency assumption and constructs a precise electric-carbon coupled electrolyzer model that takes into account multiple states such as standby, low load, variable load, and overload. It accurately depicts the hydrogen production efficiency and carbon emission behavior under different power source and load power combinations, providing an accurate basis for the quantification of clean hydrogen standards.

[0066] (3) The present invention proposes a clean hydrogen standard constrained optimization scheduling method that combines the Lagrange multiplier method and the PPO algorithm, transforming the constrained Markov decision process into an equivalent unconstrained maximum-minimum optimization problem, and guiding the control strategy to gradually meet the clean hydrogen standard constraints by adaptively adjusting the Lagrange multiplier. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 This is the framework diagram of the hydrogen-based community microgrid system (HCMS);

[0068] Figure 2 It is the framework diagram of the LC-PPO algorithm;

[0069] Figure 3 It is a real-time electricity price map;

[0070] Figure 4 A real-time carbon emission intensity map;

[0071] Figure 5 Comparison of hydrogen production efficiency and hydrogen production of different models;

[0072] Figure 6 Comparison of energy consumption and carbon emissions of electrolyzers of different models;

[0073] Figure 7 is the average reward curve of different algorithms;

[0074] Figure 8is the cost curve of different algorithms;

[0075] Figure 9 is the electrolyzer power distribution of the PPO algorithm during the scheduling period;

[0076] Figure 10 is the electrolyzer power distribution of the LP-PPO algorithm during the scheduling period;

[0077] Figure 11 is the electrolyzer power distribution of the HP-PPO algorithm during the scheduling period;

[0078] Figure 12 is the electrolyzer power distribution of the LC-PPO algorithm during the scheduling period;

[0079] Figure 13 The hydrogen cleanliness change curves of different algorithms during the scheduling period;

[0080] Figure 14 is the grid power distribution under different clean hydrogen standards during the dispatch period;

[0081] Figure 15 Comparison of hydrogen quality and electrolyzer grid power consumption under different clean hydrogen standards. DETAILED DESCRIPTION

[0082] Specific embodiments are given below in conjunction with the accompanying drawings. The specific embodiments are only used to introduce the technical solutions of the present invention in detail and are not intended to limit the scope of protection of the present application.

[0083] The research object of this invention is HCMS, and its framework is as follows Figure 1 As shown. The system integrates local photovoltaic (PV) power generation systems, hydrogen fuel cells, hydrogen storage tanks, electrolyzers, hydrogen vehicles (HVs), and building loads within the community. Its core lies in the efficient two-way conversion of electricity and hydrogen energy through electrolyzers and hydrogen fuel cells. Photovoltaic power generation can prioritize building loads, and excess electricity can be used to produce hydrogen in electrolyzers; the hydrogen in hydrogen storage tanks can not only generate electricity through hydrogen fuel cells to meet building loads, but also provide refueling services for hydrogen vehicles within the community. The system aims to achieve low-carbon economic operation while meeting the community's diverse energy needs (building loads, hydrogen fuel cells) by optimizing the coordinated operation of each unit. The key challenge is how to ensure that the produced hydrogen meets increasingly stringent clean hydrogen standards when relying on grid-supplemented electricity for electrolysis hydrogen production. This directly drives the need for refined electrolyzer modeling and advanced scheduling strategies.

[0084] This paper proposes a method for optimizing the scheduling of hydrogen-based community microgrids based on clean hydrogen constraints. First, a refined electric-carbon coupled electrolyzer model is established. This model can accurately capture the nonlinear efficiency and carbon emission characteristics of the electrolyzer under multiple operating states, laying the foundation for reliable low-carbon scheduling and avoiding the significant evaluation errors of traditional models. Based on this model, the scheduling problem under clean hydrogen constraints is constructed as a constrained Markov decision process (CMDP), and a Lagrangian constrained proximal policy optimization (LC-PPO) algorithm is designed. By adaptively adjusting the Lagrangian multiplier, the clean hydrogen standard constraint is effectively processed. Specifically, the following steps are included:

[0085] Step 1: Build a system model, including a refined electro-carbon coupled electrolyzer model, a dynamic hydrogen standard evaluation model, a hydrogen fuel cell model, a hydrogen storage tank hydrogen quality update model, a step-by-step carbon trading model, an HCMS scheduling model, and optimized scheduling constraints;

[0086] 1.1 Refined electro-carbon coupled electrolyzer model

[0087] To overcome the limitations of traditional electrolyzer models (which often use fixed efficiency or simplified linear relationships) in accurately evaluating energy consumption and carbon emissions, a refined electric-carbon coupled electrolyzer model with multi-state operation, nonlinear efficiency, and carbon emission tracking from electricity sources is proposed. This provides a solid foundation for the subsequent accurate evaluation of the carbon footprint of hydrogen production and the verification and implementation of clean hydrogen standards. In order to finely characterize the multi-state operating characteristics of the electrolyzer in actual operation, a state indicator variable δ is introduced. e,s (t) represents the state of the electrolytic cell at time t and is defined as follows:

[0088]

[0089] Where: s = 0 is the standby state, s = 1 is the low load state, s = 2 is the variable load state, and s = 3 is the overload state. The electrolytic cell can only be in one operating state at any time, so:

[0090]

[0091] The power constraints corresponding to each state are:

[0092]

[0093] Where: P standby is the standby power of the electrolyzer, P ely (t) is the input power of the electrolytic cell at time t, P ely,N Rated value of power input to the electrolyzer;

[0094] Hydrogen production efficiency is a key indicator for measuring electrolyzer performance. It is not constant but a nonlinear function related to the operating state. When in standby mode (s = 0), the electrolyzer hydrogen production efficiency is 0. When in working mode, its efficiency can be represented by the following polynomial fitting:

[0095]

[0096] Where: η e,s (t) represents the hydrogen production efficiency of the electrolyzer at time t;

[0097] In order to distinguish the source of electricity supplied to the electrolyzer, an explicit electricity-carbon coupling mechanism is established; the electricity source of the electrolyzer is composed of renewable energy generation and the grid, and the carbon emissions of the electrolyzer are only generated when the grid electricity is used. Then the carbon emissions of the electrolyzer at time t E e,s (t) is:

[0098]

[0099] Where: is the dynamic carbon emission intensity of the power grid at time t, which indicates the amount of CO2 released when the power grid provides 1 kWh of electricity; P grid,ely (t) is the grid power consumed by the electrolyzer at time t, and Δt is the time step;

[0100] Based on the above state and efficiency model, the hydrogen production at time t for:

[0101]

[0102] Where: is the energy density of hydrogen.

[0103] 1.2 Dynamic Clean Hydrogen Standard Evaluation Model

[0104] For hydrogen production from water electrolysis in HCMS, its lifecycle carbon emissions are primarily concentrated in the electricity consumed during the hydrogen production phase. If the electricity comes from zero-carbon renewable energy, the emissions are zero; if it comes from a carbon-containing power grid, indirect emissions occur. To dynamically monitor and ensure that the produced hydrogen meets preset standards, the present invention adopts a hydrogen cleanliness index based on time series accumulation. This index is defined as the ratio of cumulative carbon emissions to cumulative hydrogen production from the initial moment to the current moment. This index takes into account the long-term impact of the electrolyzer's operating history and the combination of power sources, providing a clear and quantifiable cleanliness constraint target for optimized scheduling.

[0105]

[0106] Where, is the hydrogen cleanliness index at time t, They are Carbon emissions and hydrogen production of the electrolyzer at each moment.

[0107] 1.3 Hydrogen fuel cell model

[0108] Hydrogen fuel cells consume hydrogen to generate electricity. The relationship between hydrogen consumption and output power is:

[0109]

[0110] Where: P fc (t) is the hydrogen consumption and output power of the hydrogen fuel cell at time t, and ηfc is the efficiency of the hydrogen fuel cell.

[0111] 1.4 Hydrogen storage tank hydrogen quality update model

[0112] The hydrogen storage tank is used to balance the production and consumption of hydrogen. The formula for updating the hydrogen mass in the hydrogen storage tank of the community hydrogen station is:

[0113]

[0114] Where: M tank (t+1), M tank (t) is the hydrogen mass in the hydrogen storage tank of the community hydrogen station at time t+1, t, is the change in hydrogen mass in the hydrogen storage tank of the community hydrogen station at time t, is the hydrogen filling amount of the hydrogen vehicle cluster at time t;

[0115] Taking the hydrogen vehicle cluster within the community into consideration, the hydrogen consumption and charging amount of the hydrogen vehicle cluster are:

[0116]

[0117] Where: M hv,use (t), R is the hydrogen consumption and hydrogen filling amount of the hydrogen vehicle cluster at time t. The hydrogen consumption or hydrogen filling is determined according to the status of the hydrogen vehicle. It is stipulated that: when the vehicle is out and consumes hydrogen, it is considered hydrogen consumption; when the vehicle is in the community and replenishes hydrogen fuel, it is considered hydrogen filling. consume is the fuel consumption rate of hydrogen vehicles, D hv,i (t) is the travel distance of the i-th vehicle group at time t, R hv,i (t) is the hydrogen charging rate of the i-th vehicle group at time t, N hv N is the number of vehicle groups included in the hydrogen vehicle cluster. hv,i is the number of hydrogen vehicles included in the i-th vehicle group; since the change in hydrogen in the hydrogen storage tank of the hydrogen station is calculated in kg, and the hydrogen filling rate is g / s, 3.6 represents the unit conversion;

[0118] The formula for updating the hydrogen mass in the on-board hydrogen storage tank is:

[0119]

[0120] Where: M tank,hv (t+1), M tank,hv (t) represents the mass of hydrogen in the onboard hydrogen storage tank at time t+1 and t, and 1000 represents the unit conversion from kg to g.

[0121] 1.5 Ladder-type carbon trading model

[0122] When the system carbon emissions E(t) deviate from the carbon emission quota, carbon emission quotas can be purchased or sold in the carbon trading market. This invention divides carbon emission responsibilities into four levels and adopts a stepped carbon trading system. In the exempt carbon emission range 0~E1(t), the carbon trading cost F carbon (t) = ξ1(E(t)-E1(t)); in the low carbon emission responsibility range E1(t) ~ E2(t), the carbon trading cost F carbon (t) = ξ2(E(t)-E1(t)); in the carbon emission responsibility range E2(t) ~ E3(t), the carbon transaction cost F carbon (t)=ξ2(E2(t)-E1(t))+ξ3(E(t)-E2(t)); In the high carbon emission responsibility range E3(t)~∞, the carbon trading cost F carbon (t)=ξ2(E2(t)-E1(t))+ξ3(E3(t)-E2(t))+ξ4(E(t)-E3(t)); among them, E1(t), E2(t), E3(t) represent the critical carbon emission quotas of adjacent carbon emission responsibility intervals, ξ1, ξ2, ξ3 and ξ4 are carbon prices, ξ1 and ξ2 take the carbon base price ξ, ξ3=(1+β)ξ, ξ4=(1+2β)ξ, and the price growth rate β is 25%.

[0123] 1.6 HCMS Scheduling Model

[0124] The HCMS scheduling objective is to minimize the total system operating cost F within the scheduling period while satisfying the operating constraints. total , then:

[0125]

[0126] F(t)=F op (t)+F energy (t)+F carbon (t)+F penalty (t) (15)

[0127] F op (t) = c e P ely(t)Δt+c f P fc (t)Δt (16)

[0128]

[0129] F penalty (t) = τ pv P surplus (t)Δt (18)

[0130] Where: T represents the scheduling period, F(t) represents the system comprehensive cost at time t, and F op (t) represents the operation and maintenance cost at time t, F energy (t) represents the energy purchase cost at time t, F penalty (t) represents the penalty cost of abandoning light at time t, c e 、c f is the operation and maintenance cost coefficient of the electrolyzer and hydrogen fuel cell, c ele (t) is the electricity price at time t; P grid (t) is the grid interaction power, positive means power purchase, negative means power sale; is the external hydrogen price at time t, is the external hydrogen purchase amount at time t, τ pv is the light abandonment penalty coefficient, P surplus (t) is the unused photovoltaic power.

[0131] Optimization scheduling constraints include power balance constraints, hydrogen storage tank capacity constraints, electrolyzer operation constraints, hydrogen fuel cell operation constraints, hydrogen vehicle status constraints, and clean hydrogen standard constraints;

[0132] Power balance constraints:

[0133] P grid (t)+P pv (t)+P fc (t) = P load (t)+P ely (t) (19)

[0134] Where, P pv (t) represents the photovoltaic power generation at time t, P load (t) represents the building load at time t;

[0135] Hydrogen storage tank capacity constraints:

[0136]

[0137] Where: The upper and lower limits of hydrogen quality in the hydrogen storage tank of the community hydrogen station, The upper and lower limits of the hydrogen quality in the on-board hydrogen storage tank;

[0138] Electrolyzer operation constraints:

[0139] P standby ≤P ely (t)≤1.3P ely,N (twenty one)

[0140] Hydrogen fuel cell operation constraints:

[0141]

[0142] Where, The upper and lower limits of the hydrogen fuel cell output power;

[0143] Hydrogen vehicle status constraints:

[0144]

[0145] Where: R is the upper and lower limits of hydrogen charging power for hydrogen vehicles. hv,eq (t) is the equivalent hydrogen filling rate used to fill the critical hydrogen that does not overflow the on-board hydrogen storage tank at time t;

[0146] Clean hydrogen standard constraints:

[0147]

[0148] Where: Φ standard is the clean hydrogen standard threshold; this constraint ensures that the hydrogen produced during the entire scheduling period meets the cleanliness requirement.

[0149] Step 2: Build CMDP model based on the system model;

[0150] HCMS optimal scheduling is essentially a sequential decision-making problem. The system needs to make optimal operational decisions in a dynamic environment facing multiple uncertainties such as load, photovoltaic output, real-time electricity prices, and dynamic grid carbon intensity. At the same time, key clean hydrogen standards must be met as constraints. This sequential decision-making problem with explicit operational constraints is modeled using a constrained Markov decision process. A CMDP is typically defined by a tuple (S, A, P, R, C, d, γ), where:

[0151] State space S: s t ∈S represents the state at time t, which can be expressed as:

[0152]

[0153] Action space A: a t ∈A represents the agent in state s tThe actions taken under this rule specifically refer to the electrolyzer input power and hydrogen fuel cell output power;

[0154] a t ={P ely (t),P fc (t)} (26)

[0155] State transition probability P: the system changes from the current state s t Take action a t Then transfer to the next state s t+1 probability;

[0156] Reward function R: r t =R(s t ,a t ) indicates that the agent is performing action a t To achieve the best economic performance, the reward function is designed to be a negative system comprehensive cost:

[0157] r t = -F(t) (27)

[0158] Cost function C: c t =C(s t ,a t ) indicates that the agent is performing action a t The core of the present invention is to convert the clean hydrogen standard constraint into a cost, so the cost is defined as the upper limit of the clean hydrogen standard. Degree of violation:

[0159]

[0160] Constraint tolerance d: the preset upper limit of the expected cumulative discount cost;

[0161] Discount factor γ∈[0,1]: used to balance the importance of immediate rewards / costs and future rewards / costs;

[0162] The goal of CMDP is to find an optimal policy π * :S→A, so that the strategy maximizes the expected cumulative discounted reward At the same time, satisfy the expected cumulative discount cost constraint Where E[·] represents the expectation. The present invention sets the constraint tolerance d to 0, aiming to drive the intelligent agent to learn an operation strategy that strictly complies with the clean hydrogen standard constraints through the CMDP framework.

[0163] Step 3: Solve the CMDP model based on the Lagrange multiplier method and PPO algorithm to obtain the optimal strategy;

[0164] Although standard reinforcement learning algorithms (such as the PPO algorithm) perform well in maximizing cumulative rewards, they are generally difficult to directly handle the explicit constraints in CMDP. C (π)≤d; To solve this problem, the present invention introduces the Lagrangian multiplier method, transforms the constrained optimization problem into an equivalent unconstrained problem, and combines it with the PPO algorithm to obtain the LC-PPO (Lagrangian Constrained PPO) algorithm. The core idea of ​​the LC-PPO algorithm is to solve the following saddle point problem:

[0165]

[0166] Where: λ is a non-negative Lagrange multiplier, which is used as a penalty coefficient to dynamically weigh the relationship between reward maximization and constraints. If the constraints tend to violate J C (π)>d, λ will increase to strengthen the penalty on the strategy; if the constraint satisfies J C (π)≤d, λ will decrease to allow the policy to focus more on reward maximization;

[0167] Construct a policy network πθ, which is represented by a neural network with parameters θ, input state, output policy, and policy objective function L policy (θ) not only includes the original pruning mechanism of the PPO algorithm for stabilizing policy updates, but also introduces additional penalty terms related to the constraint cost, so:

[0168] L policy (θ)=L CLIP (θ)-λE[A C,t ] (30)

[0169] Where: L CLIP (θ)=E[min(ratio t (θ)A R,t ,clip(ratio t (θ),1-ε,1+ε)A R,t )] is the clipping objective function of the PPO algorithm, and the probability ratio t (θ)=πθ(a t |s t ) / π θold (a t |s t ) represents the current strategy π θ (a t |s t ) and the old strategy π θold (a t |s t) is used to measure the degree of change of the current strategy relative to the old strategy after updating the parameter θ; clip(·) represents the clipping operation, ε represents the clipping factor; A R,t is the reward advantage, used to evaluate action a t How good is the action relative to the average; A C,t Based on the cost c t Calculate the cost advantage for evaluating action a t relative risk in causing a constraint violation;

[0170] Reward Advantage A R,t Using the Generalized Advantage Estimation (GAE) calculation, we have A R,t =δ t +(γλ G )δ t+1 +···+(γλ G ) T-t+1 δ T-1 ;in, represents the TD residual, λ G represents the smoothing parameter of the generalized advantage estimate, Represent the value of the current state and the next state respectively, using the value network estimate;

[0171] Value Network By parameters Neural network representation for estimating state value Its purpose is to assist in calculating the bonus advantage A R,t ; The objective function of the value network is:

[0172]

[0173] Where: is the target state value;

[0174] The Lagrange multiplier λ is adaptively adjusted according to the constraint satisfaction during the entire training process, and its goal is to make J C (π) does not exceed the constraint tolerance d, and the regularization term κ is added λ λ 2 to the loss function to stabilize its update process, the objective function L of the Lagrange multiplier λ for:

[0175]

[0176] Where: κλ is the regularization coefficient, is the current strategy for J CThe estimate of (π) is approximated using the average undiscounted cost of a training batch.

[0177] Initialize the policy network, value network and Lagrange multiplier; the agent obtains the state s at time t t , which is input into the policy network to calculate the state s t The action probability distribution logπ(a t |s t ), select action a according to probability sampling t Return to the HCMS environment and update the state; the reward r at time t t and cost c t Feedback to the agent, and calculate the state value through the value network Based on the reward r t and state value Calculate the reward advantage A using GAE R,t and the target state value According to the cost c t To calculate the cost advantage A C,t During each round of training, multiple samples are randomly selected to update the parameters of the policy network and the value network, and the policy target L is maximized by gradient ascent. policy (θ), minimize the value loss by gradient descent According to the average cost of the current strategy Calculate the Lagrangian loss L λ , update the Lagrange multiplier by gradient descent; after multiple iterations, the LC-PPO algorithm will output an optimization strategy π θ , this strategy can maximize the expected cumulative discount reward while making the system operation strictly comply with the clean hydrogen standard constraints.

[0178] Example

[0179] This embodiment takes the hydrogen-based community microgrid system as an example. In order to focus on the scheduling and control of the HCMS core unit while maintaining the feasibility of calculation, a hydrogen vehicle cluster is set up to include three groups of vehicles, each with 5 vehicles, with a fixed average mileage. HCMS related parameters are detailed in Table 1. The system allows the purchase of hydrogen from the external market when hydrogen is insufficient, and the external hydrogen price is set to a fixed value of 14.67$ / kg. Table 2 lists the key hyperparameters of the LC-PPO algorithm. Real-time electricity prices and carbon emission intensity are shown in Table 1. Figure 3 and Figure 4 .

[0180] Table 1 HCMS related parameters

[0181]

[0182] Table 2 Algorithm hyperparameters

[0183]

[0184] In order to quantify the performance difference between the refined electric-carbon coupled electrolyzer model (Model 1) constructed in the present invention and the traditional simplified model (Model 2) and its impact on system evaluation, Model 2 assumes that the electrolyzer has a constant hydrogen production efficiency, and the operating power does not exceed its rated power, and the standby power consumption is ignored. Figure 5 Comparison of hydrogen production efficiency and hydrogen production of different models. Figure 6 Comparison of energy consumption and carbon emissions of electrolyzers of different models. Figure 5 、 6 It can be seen that Model 1 can capture the nonlinear characteristics of hydrogen production efficiency with input power and multiple operating states, including overload state ( Figure 5 The hydrogen production in the midday photovoltaic peak period exceeds that of Model 2) and standby state ( Figure 6 The energy consumption of the electrolyzer is non-zero during low-power periods. However, the hydrogen production efficiency of Model 2 is constant, and its operating power is limited by the rated power of the electrolyzer. This makes it impossible to simulate an overload state to fully absorb excess photovoltaic power, and also ignores energy consumption in standby mode.

[0185] Table 3 Error comparison of different models

[0186]

[0187] Table 3 quantifies the impact of the two models on system performance evaluation. Using the results of Model 1 as a benchmark, Model 2, due to its simplifying assumptions, may overestimate or underestimate efficiency at certain times due to its constant efficiency. Furthermore, Model 2 fails to account for the additional hydrogen generated by Model 1 through the overload state using excess photovoltaic power, resulting in a negative bias of up to 18.36% in the hydrogen production assessment. The electrolyzer power consumption assessment has a negative bias of 18.23%. Model 2 does not consider the standby state and fails to capture power fluctuations during actual operation, which collectively leads to an underestimation of total energy consumption. The associated carbon emissions assessment has a negative bias of 37.78%. This significant bias in carbon emissions stems from the bias in energy consumption assessment. Furthermore, Model 1's electricity-carbon coupling mechanism accurately accounts for carbon emissions associated with grid power consumption (including standby periods), while Model 2 cannot accurately reflect this. It can be seen from this that ignoring the dynamic efficiency of the electrolyzer and multiple operating states (especially overload and standby) will lead to serious evaluation deviations of system performance (hydrogen production, energy consumption) and environmental impact (carbon emissions). The refined electric-carbon coupled electrolyzer model proposed in this invention provides the necessary basis for reliable optimization scheduling and accurate carbon footprint assessment of HCMS.

[0188] In order to evaluate the effectiveness of the LC-PPO algorithm in dealing with the clean hydrogen standard constraint, and to compare it with three benchmark algorithms: PPO, LP-PPO, and HP-PPO; PPO refers to the standard PPO algorithm without considering the clean hydrogen standard constraint; LP-PPO (LowPenalty PPO) adds a fixed penalty term r′t=r in the reward function of PPO. t -κ p c t , penalty coefficient κ p is 5; HP-PPO (High Penalty PPO) adds a fixed penalty term to the reward function of PPO, and the penalty coefficient κ p is 10; LC-PPO (algorithm of the present invention) uses Lagrange multipliers to adaptively adjust the penalty intensity. Figure 7 and Figure 8 The average reward and cost curves of the four algorithms during the training process are shown respectively. For the PPO algorithm, PPO focuses on maximizing the reward and achieves the highest average reward. However, due to the complete neglect of the constraints, its cost remains at a high level, indicating that it continues to violate the clean hydrogen standard. For the LP-PPO algorithm, the smaller fixed penalty coefficient is not enough to effectively guide the strategy to avoid violating the clean hydrogen standard. Although the cost decreases, it does not approach zero, and its reward is slightly lower than that of PPO. For the HP-PPO algorithm, the larger fixed penalty coefficient forces the strategy to strictly adhere to the clean hydrogen standard constraint. The cost eventually converges to zero, but this strong penalty excessively restricts strategy exploration, resulting in its lowest average reward among all algorithms. The LC-PPO proposed in this invention exhibits excellent balancing ability. The cost decreases rapidly as training progresses and stabilizes at zero, indicating that it can effectively meet the clean hydrogen standard constraint. Its average reward is significantly higher than that of HP-PPO and only slightly lower than that of the unconstrained PPO, thanks to the adaptive adjustment mechanism of the Lagrange multiplier. The training curve preliminarily proves that LC-PPO can achieve economic benefits close to the unconstrained case while ensuring that the clean hydrogen constraint is met, which is better than the fixed penalty coefficient method.

[0189] Figures 9-12 The electrolyzer power distribution diagrams of the four algorithms, PPO, LP-PPO, HP-PPO and LC-PPO, within the scheduling period are shown. Figure 13 The hydrogen cleanliness change curves of different algorithms in the scheduling period, the clean hydrogen standard threshold Φ standard Set to 4.9kgCO2e / kg H2, such as Figure 13 Indicated by the middle dotted line; Table 4 shows the compliance with clean hydrogen standards, carbon emissions and total operating costs under different algorithms.

[0190] Table 4 Clean hydrogen standard compliance, carbon emissions and total operating costs under different algorithms

[0191]

[0192] For the PPO algorithm, it shows a strong arbitrage tendency during the nighttime low electricity price period (such as t = 20-30h, t = 45-55h), driving the electrolyzer to operate at high power (mainly consuming grid electricity). This behavior directly leads to a sharp surge in the hydrogen cleanliness index during the corresponding period, far exceeding the clean hydrogen standard threshold of 4.9; the total operating cost of the PPO algorithm is the lowest ($930.9, but it has 20 violations and the clean hydrogen standard compliance rate is only 72.2%. For the LP-PPO algorithm, its penalty is insufficient. Although it reduces the hydrogen production power at night, there is still significant grid electricity consumption in certain periods, resulting in occasional exceeding of the hydrogen cleanliness index. There are a total of 4 violations, the clean hydrogen standard compliance rate is 94.4%, and the total operating cost ($955) is between PPO and HP-PPO. For the HP-PPO algorithm, due to the large fixed penalty coefficient, the activity of using grid electricity for electrolytic hydrogen production at night is almost completely suppressed, which ensures that the hydrogen cleanliness index is always below the clean hydrogen standard threshold. value, 0 violations occurred, and a 100% clean hydrogen standard compliance rate was achieved. However, this conservative strategy sacrificed all low electricity price arbitrage opportunities, resulting in the highest total operating cost ($987.4). For the LC-PPO algorithm, some low electricity price periods were also utilized for hydrogen production at night, but its power level was significantly and dynamically suppressed. This is because when the hydrogen cleanliness index exceeds the clean hydrogen standard threshold, the internal adaptive Lagrange multiplier will increase, thereby increasing the "penalty" for grid-powered hydrogen production, guiding the strategy to reduce power so that the hydrogen cleanliness index can be stably maintained below the clean hydrogen standard threshold, achieving a 100% clean hydrogen standard compliance rate (Table 5: 0 violations); most importantly, LC-PPO can still seize some economic opportunities while strictly meeting the clean hydrogen standard constraints, achieving a total operating cost significantly lower than HP-PPO ($975). Compared with unconstrained PPO, LC-PPO is more efficient in strictly meeting the clean hydrogen standard (100% vs. Under the premise of achieving a 72.2% reduction in system carbon emissions, the system carbon emissions were successfully reduced by 21.2% (from 342.3 kg to 269.7 kg), while the total operating cost only increased slightly by 4.7% (from $930.9 to $975), which fully demonstrates the ability of LC-PPO to achieve an excellent balance between constraint satisfaction and economic dispatch.

[0193] The LC-PPO algorithm of the present invention is used to explore different clean hydrogen standard thresholds Φ standard (From a loose 5.7 to a strict 4.1kgCO2e / kg H2) and its impact on HCMS optimization operation strategy, carbon emissions and economy. Figure 14is the grid power distribution under different clean hydrogen standards during the scheduling period; as the clean hydrogen standard threshold decreases (i.e., the standard becomes stricter), the strategy learned by the LC-PPO algorithm tends to significantly reduce the operation of the electrolyzer during the night time (mainly relying on grid electricity). This is because stricter standards impose stronger inhibitions on the consumption of carbon-containing grid electricity, while the strategy of using photovoltaic hydrogen production during the day is less affected. Figure 15 The comparison of hydrogen quality and electrolyzer grid power consumption under different clean hydrogen standards is shown in Table 5. The reduction in grid-generated hydrogen production at night directly leads to a decrease in the total hydrogen production within the system. In order to meet the community's hydrogen energy needs (such as the daily refueling volume of hydrogen vehicles and the basic power generation needs of fuel cells), the system will increase the amount of hydrogen purchased from the external market. In terms of carbon emissions and total operating costs, as shown in Table 5, the lower the clean hydrogen standard threshold (the more stringent), the lower the total carbon emissions of the system, which verifies the effectiveness of the constraint. However, the total operating cost of the system shows the opposite trend. The stricter the standard, the higher the cost. The main reasons for the cost increase include: (1) the loss of more opportunities for electricity price arbitrage through electrolysis during the low electricity price period at night; (2) the increase in external hydrogen purchases with relatively high costs. It can be seen that the setting of clean hydrogen standard thresholds directly affects the operating strategy, carbon emission level and economic cost of HCMS. There is a clear trade-off between environmental benefits (low carbon emissions) and economic costs. The stricter the standard, the better the emission reduction effect, but the higher the operating cost. This provides a quantitative basis for policymakers or community managers to choose appropriate clean hydrogen standards according to their specific emission reduction targets and economic affordability.

[0194] Table 5 Carbon emissions and total operating costs under different clean hydrogen standards

[0195]

[0196] Any matters not described in the present invention are applicable to the prior art.

Claims

1. A method for optimizing the scheduling of hydrogen-based community microgrids based on clean hydrogen constraints, characterized in that: The following steps are involved: Step 1: Build a system model, including a refined electro-carbon coupled electrolyzer model, a dynamic hydrogen standard evaluation model, a hydrogen fuel cell model, a hydrogen storage tank hydrogen quality update model, a step-by-step carbon trading model, an HCMS scheduling model, and optimized scheduling constraints; The refined electro-carbon coupled electrolytic cell model uses state indicator variables to characterize the state of the electrolytic cell, which are defined as follows: Where: δ e,s (t) represents the state of the electrolytic cell at time t, where s=0 is the standby state, s=1 is the low-load state, s=2 is the variable-load state, and s=3 is the overload state; The carbon emissions of an electrolyzer are: Where: E e,s (t) represents the carbon emissions of the electrolyzer at time t, is the dynamic carbon emission intensity of the power grid at time t, P grid,ely (t) is the grid power consumed by the electrolyzer at time t, and Δt is the time step; The hydrogen production of the electrolyzer is: Where: represents the hydrogen production of the electrolyzer at time t, η e,s (t) represents the hydrogen production efficiency of the electrolyzer at time t, P ely (t) is the input power of the electrolytic cell at time t, D H2 is the energy density of hydrogen; The dynamic clean hydrogen standard evaluation model is: Where, is the hydrogen cleanliness index at time t, They are Carbon emissions and hydrogen production of the electrolyzer at each moment; The HCMS scheduling model is: F(t)=F op (t)+F energy (t)+F carbon (t)+F penalty (t) (7) F op (t)=c e P ely (t)Δt+c f P fc (t)Δt (8) F penalty (t)=τ pv P surplus (t)Δt (10) Where: F total represents the total operating cost of the system, T represents the scheduling period, F(t) represents the comprehensive cost of the system at time t, and F op (t) represents the operation and maintenance cost at time t, F energy (t) represents the energy purchase cost at time t, F carbon (t) represents the carbon transaction cost at time t, F penalty (t) represents the penalty cost of abandoning light at time t, c e 、c f is the operation and maintenance cost coefficient of the electrolyzer and hydrogen fuel cell, c ele (t) is the electricity price at time t, P grid (t) is the grid interaction power, is the external hydrogen price at time t, M H2,buy (t) is the external hydrogen purchase amount at time t, τ pv is the light abandonment penalty coefficient, P surplus (t) is the unused photovoltaic power; Step 2: Construct a CMDP model based on the system model. The CMDP model is defined by the tuple (S, A, P, R, C, d, γ); State space S: s t ∈S represents the state at time t, which can be expressed as: Where: P load (t) represents the building load at time t, P pv (t) represents the photovoltaic power generation at time t, M tank (t) represents the hydrogen mass in the hydrogen storage tank of the community hydrogen station at time t, represents the hydrogen charge of the hydrogen vehicle cluster at time t, Indicates the hydrogen consumption of the hydrogen fuel cell at time t; Action space A: a t ∈A represents the agent in state s t the actions taken next; a t ={P ely (t),P fc (t)} (12) State transition probability P: the probability that the system will transition to the next state after taking an action from the current state; Reward function R: r t =R(s t ,a t ) indicates that the agent is in state s t Next, perform action a t Instant rewards received t ; r t =-F(t) (13) Cost function C: c t =C(s t ,a t ) indicates that the agent is in state s t Next, perform action a t the costs incurred; Where: represents the upper limit of the clean hydrogen standard; d represents the constraint tolerance, and γ represents the discount factor; Step 3: Introduce Lagrange multipliers into the PPO algorithm to solve the CMDP model and obtain the optimal strategy, which maximizes the expected cumulative discounted reward while satisfying the expected cumulative discounted cost constraint.

2. The hydrogen-based community microgrid optimization scheduling method based on clean hydrogen constraints according to claim 1 is characterized in that: In the third step, the objective function to be solved is: Where: λ is a non-negative Lagrange multiplier, represents the expected cumulative discount reward, represents the expected cumulative discounted cost, and E[·] represents the expectation; The objective function L of the policy network of the PPO algorithm policy (θ) is: L policy (θ)=L CLIP (θ)-λE[A C,t ] (16) Where: L CLIP (θ) is the clipping objective function of the PPO algorithm, A C,t Indicates that based on the cost c t Cost advantages of computing; The objective function of the value network of the PPO algorithm is: Where: is the target state value, Indicates state s t the value of The objective function of the Lagrange multiplier is: Where: κλ is the regularization coefficient, is an estimate of the expected cumulative discounted cost under the current policy.

3. The hydrogen-based community microgrid optimization scheduling method based on clean hydrogen constraints according to claim 1 or 2 is characterized in that: The hydrogen storage tank hydrogen quality update model includes hydrogen quality updates in community hydrogen stations and on-board hydrogen storage tanks; The formula for updating the hydrogen mass in the hydrogen storage tank of a community hydrogen station is: Where: M tank (t+1) is the hydrogen mass in the hydrogen storage tank of the community hydrogen station at time t+1, is the change in hydrogen mass in the hydrogen storage tank of the community hydrogen station at time t; The formula for updating the hydrogen mass in the on-board hydrogen storage tank is: Where: M tank,hv (t+1), M tank,hv (t) represents the mass of hydrogen in the onboard hydrogen storage tank at time t+1 and t, M hv,use (t) represents the hydrogen consumption of the hydrogen vehicle cluster at time t.

4. The method for optimizing and scheduling hydrogen-based community microgrids based on clean hydrogen constraints according to claim 3 is characterized in that: The ladder-type carbon trading model includes: in the exempt carbon emission range 0~E1(t), the carbon trading cost F carbon (t) = ξ1(E(t)-E1(t)); in the low carbon emission responsibility range E1(t) ~ E2(t), the carbon trading cost F carbon (t) = ξ2(E(t)-E1(t)); in the carbon emission responsibility range E2(t) ~ E3(t), the carbon transaction cost F carbon (t)=ξ2(E2(t)-E1(t))+ξ3(E(t)-E2(t)); In the high carbon emission responsibility range E3(t)~∞, the carbon trading cost F carbon (t)=ξ2(E2(t)-E1(t))+ξ3(E3(t)-E2(t))+ξ4(E(t)-E3(t)); where E(t) represents the system carbon emissions, E1(t), E2(t), and E3(t) represent the critical carbon emission quotas of adjacent carbon emission responsibility intervals, and ξ1, ξ2, ξ3, and ξ4 are carbon prices.

5. The hydrogen-based community microgrid optimization scheduling method based on clean hydrogen constraints according to claim 1 is characterized in that: In standby mode, the hydrogen production efficiency of the electrolyzer is 0; in low load, variable load and overload modes, the hydrogen production efficiency of the electrolyzer is expressed as: Where: P ely,N Rated value for the input power to the electrolyzer.