Hydrogen energy micro-grid dispatching method and device based on CVaR constraint safety reinforcement learning and medium

By constructing a dynamic equivalent carbon intensity tracking model and a distributed security strategy network, the problems of carbon memory effect and long-tail risk of carbon emissions in microgrids are solved, and the accuracy and security of low-carbon dispatch are achieved.

CN122066232APending Publication Date: 2026-05-19LIAONING UNIVERSITY OF PETROLEUM AND CHEMICAL TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LIAONING UNIVERSITY OF PETROLEUM AND CHEMICAL TECHNOLOGY
Filing Date
2026-02-04
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing low-carbon dispatch methods for microgrids cannot accurately track the carbon memory effect of energy storage systems, and cannot effectively manage the long-tail risk of carbon emissions when renewable energy uncertainties are high, leading to carbon emissions easily exceeding standards.

Method used

A CVaR-constrained security reinforcement learning method is adopted to construct a dynamic equivalent carbon intensity tracking model and a distributed security policy network. The model is trained centrally using a CVaR-constrained Lagrange relaxation algorithm to achieve carbon-energy coordinated scheduling.

Benefits of technology

Accurately track the carbon emission flow of energy storage systems, explicitly optimize the tail risk of carbon emission distribution, reduce the probability of carbon emission violations, and improve system security and the absorption rate of renewable energy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122066232A_ABST
    Figure CN122066232A_ABST
Patent Text Reader

Abstract

The invention discloses a hydrogen energy micro-grid dispatching method based on CVaR constraint safety reinforcement learning, and belongs to the technical field of power system operation and control. The method comprises the following steps: firstly, acquiring heterogeneous multi-microgrid community environment state spatial data; constructing a dynamic equivalent carbon intensity tracking model, and checking node carbon intensity and energy storage dynamic equivalent carbon intensity; then, a multi-agent flexible Actor-Critic distributed security policy network is constructed, the carbon emission long tail risk is captured through a non-cross quantile network, and a CVaR constrained original-dual optimization training model is adopted; and finally, outputting a scheduling instruction, and realizing cooperative scheduling by means of a carbon-energy cooperative time shifting mechanism. The method has the core advantages that the energy storage carbon flow is accurately tracked, the tail risk is explicitly optimized, economic and carbon quota constraints are balanced, the extreme carbon standard exceeding risk can be effectively avoided in an uncertain environment, the violation rate is remarkably reduced, the renewable energy consumption rate is increased, and community low-carbon cooperation and safe operation are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system operation and control technology, specifically to a hydrogen microgrid scheduling method based on CVaR constrained security reinforcement learning. It is particularly suitable for collaborative optimization scheduling of heterogeneous multi-microgrid communities containing hydrogen energy and battery energy storage under conditions of high uncertainty in renewable energy, while meeting strict carbon emission constraints. Background Technology

[0002] With the intensification of global climate change, the deep decarbonization of power systems has become a core task of energy transition. Microgrids, as an effective carrier integrating distributed power sources, energy storage, and flexible loads, play a crucial role in absorbing renewable energy. In recent years, hydrogen energy, with its high energy density and long-cycle storage characteristics, has been widely introduced into microgrids to form power-to-hydrogen-to-power (P2H2P) integrated energy systems. Through hydrogen production via electrolyzers and power generation via hydrogen fuel turbines, flexible energy transfer across time scales is achieved, demonstrating enormous potential for carbon emission reduction.

[0003] However, existing microgrid low-carbon dispatch methods face two major challenges: First, at the carbon emission accounting level, traditional methods are usually based on static average carbon emission factors, measuring only at the generation side. This fails to capture the spatiotemporal dynamics of carbon emission flows during power transmission, especially for energy storage devices such as batteries and hydrogen storage tanks. The "carbon attribute" of their stored energy fluctuates dynamically with changes in grid carbon intensity during charging. Existing static methods ignore this "carbon memory" effect, making it difficult to accurately assess the true emission reduction value of energy storage systems during "low-carbon charging and high-carbon discharging," and thus unable to guide the precise dispatch of heterogeneous energy storage resources. Second, at the level of safe and optimized dispatch, the high uncertainty of renewable energy (wind power, photovoltaic) leads to significant carbon emission violation risks in power dispatch. Traditional multi-agent reinforcement learning... MARL (Marginal Learning) methods are risk-neutral methods that only optimize expected returns or meet expected constraints. Under strict daily carbon emission quota limits, they cannot effectively manage the "long-tail risk" of carbon emission distribution. Even if the average carbon emissions meet the requirements, serious carbon emission exceedances may still occur in extreme scenarios (sudden drop in renewable energy output, soaring electricity prices). They lack a mechanism to explicitly quantify and optimize tail risk, resulting in insufficient system security and robustness.

[0004] In summary, existing technologies for low-carbon dispatch in hydrogen microgrids lack both a dynamic carbon flow tracking mechanism for heterogeneous energy storage (electricity / hydrogen), making it difficult to quantify the low-carbon time-shifting value of energy storage; and a risk constraint strategy based on a distributed perspective, making it impossible to strictly meet rigid carbon emission constraints in highly uncertain environments. Summary of the Invention

[0005] The technical problem to be solved by this invention is the bottleneck of inaccurate carbon flow tracking and easy carbon emission exceedance in low-carbon scheduling of heterogeneous multi-microgrid communities: traditional methods rely on static carbon factors and cannot reflect the carbon memory effect of energy storage systems; moreover, traditional reinforcement learning only constrains the expected value and has weak control over long-tail risks, and is prone to exceeding carbon quota limits when photovoltaic / wind power fluctuates drastically.

[0006] To address the aforementioned technical problems, this invention proposes a hydrogen microgrid scheduling method based on CVaR-constrained security reinforcement learning, comprising the following steps:

[0007] S1: Obtain the environmental state spatial data of the heterogeneous multi-micronet community system;

[0008] S2: Based on the environmental state space data, construct a dynamic equivalent carbon intensity tracking model to calculate the carbon intensity of distribution network nodes and the dynamic equivalent carbon intensity of energy storage devices;

[0009] S3: Construct a multi-agent flexible Actor-Critic distributed security policy network; use the CVaR-constrained Lagrange relaxation algorithm to train it intensively to obtain a low-carbon scheduling policy model;

[0010] S4: Input the real-time status data into the model and output the dispatch scheme of each microgrid. The coordinated dispatch of the heterogeneous multi-microgrid community system under carbon emission constraints is completed through the carbon-energy coordinated time-shifting mechanism.

[0011] Preferably, the environmental state spatial data mentioned in step S1 includes photovoltaic power generation output, wind power generation output, battery state of charge, load power of each microgrid, time-of-use electricity price of the external power grid, and hydrogen level of the hydrogen storage tank. The above data is preprocessed at hourly resolution and used as the state input of the reinforcement learning environment.

[0012] Preferably, the dynamic equivalent carbon intensity tracking model described in step S2 includes a nodal carbon intensity calculation model and a unified carbon flow model for energy storage systems;

[0013] The node carbon intensity calculation model treats carbon emissions as a virtual fluid attached to the active power flow, and determines the real-time carbon intensity of the nodes by calculating the weighted average of the energy carbon intensity of all injected nodes. The unified carbon flow model of the energy storage system adopts a recursive mass-carbon balance equation, and dynamically updates the dynamic equivalent carbon intensity of the energy storage devices at the current moment by combining the carbon intensity of the energy storage devices at the previous moment, the charging / hydrogen production power at the current moment and the corresponding input source carbon intensity.

[0014] Preferably, the distributed security policy network described in step S3 equips each microgrid with an agent, the agent comprising:

[0015] (1) Actor network: used to input local observations and output scheduling actions such as battery charging and discharging, electrolyzer power, and hydrogen combustion turbine power;

[0016] (2) Reward Critic Network: Used to evaluate the economic value of actions;

[0017] (3) Distributed Cost Critic Network: It adopts a non-cross quantile network structure to output multiple discrete quantiles and fit the complete probability distribution of future cumulative carbon emission costs.

[0018] Preferably, in step S3, a "centralized training, distributed execution" architecture is used to train the network, and Conditional Value at Risk (CVaR) is introduced as a risk metric. The CVaR risk value is calculated based on a pre-set confidence level, and the formula is as follows:

[0019] (3)

[0020] in, For confidence level, For system status, For scheduling actions, For distributed CostCritic networks at the quantile level The predicted carbon emission cost output below, , This represents the total number of quantile samples.

[0021] Preferably, the Lagrange loss function constructed in step S3 is:

[0022] (4)

[0023] in, The expected return evaluated by the Reward Critic network. This is the entropy regularization temperature coefficient, used to regulate the balance between exploration and exploitation. Let Shannon entropy be the policy distribution. For Lagrange multipliers, used as adaptive penalty coefficients for violating safety constraints. This is the preset single-round carbon emission budget threshold.

[0024] Preferably, in step S3, the primal-dual method is used to update the parameters. Training convergence is achieved by alternately optimizing the network parameters and Lagrange multipliers. The Lagrange multiplier update formula is as follows:

[0025] (5)

[0026] in, Let be the learning rate of the Lagrange multiplier.

[0027] Preferably, the heterogeneous multi-microgrid community system described in step 4 includes traditional microgrids and hydrogen energy microgrids that integrate electrolyzers, hydrogen storage tanks and hydrogen fuel turbines;

[0028] The carbon-energy coordinated time-shifting mechanism is as follows: when the carbon intensity of the distribution network is lower than the preset threshold or during periods of renewable energy surplus, the hydrogen microgrid is controlled to increase the hydrogen production power of the electrolyzer and the traditional microgrid is used for battery charging; when the carbon intensity of the distribution network is higher than the preset threshold or during peak load periods, the hydrogen fuel turbine is controlled to generate electricity using hydrogen energy and the battery is used to discharge, replacing the power purchase from the distribution network.

[0029] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the steps of the hydrogen microgrid scheduling method based on CVaR constrained security reinforcement learning as described in any of the preceding claims.

[0030] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the hydrogen microgrid scheduling method based on CVaR-constrained security reinforcement learning as described in any of the preceding claims.

[0031] Beneficial effects:

[0032] Compared with existing technologies, this invention innovates in three aspects: system modeling, risk control, and collaborative mechanisms.

[0033] 1. At the modeling level, this invention proposes a unified dynamic equivalent carbon intensity mechanism that can accurately track carbon emission flows in electrochemical energy storage and hydrogen energy storage systems. Unlike traditional static accounting methods, it can reveal the "carbon memory" effect of energy storage devices and effectively quantify their cross-time carbon emission reduction value.

[0034] 2. At the control level, this invention does not merely constrain the expected value of carbon emissions, but introduces distributed reinforcement learning and conditional risk value constraints to explicitly optimize the tail risk of carbon emission distribution. This method can perceive and avoid high carbon emission risks in extreme uncertainty scenarios, significantly reducing the probability of carbon emission violations while ensuring economic efficiency and improving system security.

[0035] 3. At the collaborative level, this invention, through the cooperation of heterogeneous microgrids, utilizes the large-capacity cross-time regulation capability of hydrogen energy systems and the rapid response capability of battery energy storage to achieve carbon emission "peak shaving and valley filling" and carbon-energy time shift at the community level, thereby improving the absorption rate of renewable energy. Attached Figure Description

[0036] Figure 1This is a schematic diagram of the overall process of the method of the present invention;

[0037] Figure 2 This is a schematic diagram of the layered architecture of the heterogeneous multi-micronet community system of the present invention;

[0038] Figure 3 This is a topology diagram of an IEEE 13-node multi-micronet community system in an embodiment of the present invention;

[0039] Figure 4 This is a schematic diagram illustrating the principle of the dynamic equivalent carbon intensity (DECI) tracking mechanism of the present invention;

[0040] Figure 5 This is a schematic diagram of the overall framework of the distributed MASAC-CVaR algorithm of this invention;

[0041] Figure 6 This is a comparison of the reward convergence and carbon emission curves of the present invention and the baseline method during the training process;

[0042] Figure 7 The image shows a heat map of the spatiotemporal distribution of carbon emissions in a microgrid community under the traditional expectation constraint method and the method of this invention. Detailed Implementation

[0043] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0044] In this embodiment, the hydrogen microgrid scheduling method based on Conditional Value at Risk (CVaR) constrained security reinforcement learning is illustrated in the flowchart below. Figure 1 As shown, it includes the following steps:

[0045] A schematic diagram of the layered architecture of a heterogeneous multi-micronet community system is shown below. Figure 2 As shown, this embodiment focuses on a heterogeneous community system comprising six microgrids, employing a modified Institute of Electrical and Electronics Engineers (IEEE) 13-node distribution network topology (structure diagram shown in Figure 3). The system includes three Hydrogen Microgrids (H-MGs) and three Traditional Microgrids (T-MGs). The H-MGs integrate an electrolyzer, hydrogen storage tank, and hydrogen fuel turbine, forming a power-to-hydrogen-to-power (P2H2P) system, leveraging the high-capacity storage capabilities of hydrogen to achieve energy transfer across time scales. The T-MGs integrate battery energy storage to meet baseline load demands and address short-term power fluctuations. The system is set with a daily carbon emission allowance of 5000 kg.

[0046] S1: Obtain the environmental state space data of the heterogeneous multi-microgrid community system from the Supervisory Control and Data Acquisition (SCADA) system and historical database, including photovoltaic power generation output, wind power generation output, battery state of charge (SOC), load power of each microgrid, time-of-use electricity price of the external power grid, and hydrogen level of the hydrogen storage tank. After preprocessing the above data at hourly resolution, use it as the state input of the reinforcement learning environment.

[0047] S2: As Figure 4 As shown, a Dynamic Equivalent Carbon Intensity (DECI) tracking model is constructed based on the aforementioned environmental state space data. This model includes a nodal carbon intensity calculation model and a unified carbon flow model for the energy storage system, used to calculate the carbon intensity of distribution network nodes and the dynamic equivalent carbon intensity of energy storage devices in real time. Specifically:

[0048] (1) Calculate the real-time carbon intensity (NCI) of each node in the distribution network: Based on the carbon emission flow theory, carbon emissions are regarded as a virtual fluid attached to the active power flow. The real-time carbon intensity of each node is determined by calculating the weighted average of the energy carbon intensity injected into all nodes; the formula is:

[0049] (1)

[0050] in, For nodes At any moment carbon strength, For the node Flow to Node active power, The carbon emission factor for the local diesel generator is set at 0.9 kgCO2 / kWh;

[0051] (2) Calculate the dynamic equivalent carbon intensity of energy storage devices: For batteries and hydrogen storage tanks, a recursive mass-carbon balance equation is constructed. Combining the carbon intensity of the energy storage device at the previous moment, the charging / hydrogen production power at the current moment, and the corresponding input source carbon intensity, the dynamic equivalent carbon intensity of the energy stored in the battery or hydrogen storage tank at the current moment is dynamically updated to accurately capture the "carbon memory" effect of carbon emissions. The formula is as follows:

[0052] (2)

[0053] in, For energy storage devices At any moment The dynamic equivalent carbon intensity, This represents the current energy storage level (electricity or hydrogen). For charging / hydrogen production power, The input source carbon intensity at the moment of charging;

[0054] Using this equation, the system can identify whether the stored energy comes from clean periods (such as peak photovoltaic periods) or high-carbon emission periods.

[0055] S3: Construct a distributed security policy network based on Multi-Agent Soft Actor-Critic (MASAC), and use a CVaR-constrained Lagrange relaxation algorithm to train it intensively, finally obtaining a low-carbon scheduling policy model, specifically:

[0056] like Figure 5 As shown, each microgrid is equipped with an agent, and each agent includes:

[0057] (1) Actor network: used to input local observation information and output scheduling actions such as battery charging and discharging, electrolyzer power, and hydrogen combustion turbine power;

[0058] (2) Reward Critic Network: Used to evaluate the economic value of actions ( value);

[0059] (3) Distributed Cost Critic Network: It adopts a non-cross-quantile network structure, abandons single-generation value output, and outputs... discrete quantiles It accurately fits the complete probability distribution of future cumulative carbon emission costs.

[0060] A centralized training and distributed execution architecture is adopted for centralized training and parameter updates of the policy network. During the training phase, to meet strict carbon emission constraints, Conditional Value at Risk (CVaR) is introduced as a risk metric. Environmental state space data and dynamic equivalent carbon intensity information are synchronously input into the distributed security policy network, specifically as follows:

[0061] (1) Calculate the CVaR risk value:

[0062] Based on the discrete quantiles output by the distributed Cost Critic network, all tail quantiles greater than a preset confidence level are selected, and their average value is taken as the CVaR estimate of the current state-action pair (characterizing the expected carbon emission cost in the worst-case scenario, used to quantify long-tail security risk); the tail risk (average worst-case loss) is calculated with a 95% confidence level, using the following formula:

[0063] (3)

[0064] in, For confidence level, For system status, For scheduling actions, For distributed CostCritic networks at the quantile level The predicted carbon emission cost output below, , Let 32 ​​be the total number of quantile samples.

[0065] During network training, the tail risk of carbon emission distribution is accurately estimated by minimizing the quantile loss function and applying non-crossing constraints to ensure quantile monotonicity.

[0066] (2) Construct and update the Lagrange objective function:

[0067] The objective function follows the principle of dynamic adaptation: when the CVaR estimate exceeds the preset carbon emission budget, the Lagrange multipliers... Increase the penalty for policy updates by strengthening the safety constraints; when the CVaR estimate satisfies the constraints, This reduction allows the strategy to explore action spaces with higher economic returns, achieving a dynamic balance between safety and economy. Simultaneously, an augmented Lagrangian loss function is constructed, incorporating a policy entropy regularization term, an economic reward term, and a safety constraint term. Introducing Lagrange multipliers The formula for dynamically penalizing CVaR over-budget is as follows:

[0068] (4)

[0069] in, The expected return evaluated by the Reward Critic network. This is the entropy regularization temperature coefficient, used to regulate the balance between exploration and exploitation. Let Shannon entropy be the policy distribution. For Lagrange multipliers, used as adaptive penalty coefficients for violating safety constraints. The preset single-round carbon emission budget threshold;

[0070] (3) Original-dual parameter update:

[0071] The primal-dual optimization method is adopted to optimize network parameters by alternating strategies. With Lagrange multipliers Achieving training convergence: In the initial update phase, minimize the Lagrange penalty term (Lagrange multiplier). The policy loss (product of CVaR safety margin) is used to complete the Actor network parameters. The update phase; in the dual update phase, the multipliers are dynamically adjusted using the gradient ascent method based on the degree of violation of the CVaR safety margin. The formula is:

[0072] (5)

[0073] in, The learning rate of the Lagrange multiplier;

[0074] Lagrange multipliers Always maintain a non-negative carbon emission risk value Exceeding the budget The penalty weight is automatically increased when it is in a certain condition, and decreased when it is in a certain condition.

[0075] S4: Using the trained low-carbon scheduling strategy model, real-time status data is input into the trained Actor network, and scheduling instructions for each microgrid are output, including: the charging and discharging power of the battery energy storage system, the output power of the diesel generator, the hydrogen production power of the electrolyzer, the power generation power of the hydrogen fuel turbine, and the amount of flexible load reduction participating in demand response.

[0076] The carbon-energy coordinated time-shifting mechanism enables low-carbon coordinated scheduling of heterogeneous multi-microgrid community systems. Specifically, during periods of low carbon intensity in the distribution network or when renewable energy is abundant, the hydrogen microgrid is controlled to increase the hydrogen production power of the electrolyzer and the traditional microgrid is used to charge the battery, thereby reducing the dynamic equivalent carbon intensity of the stored energy. During periods of high carbon intensity in the distribution network or when the load is at its peak, the hydrogen fuel turbine is controlled to generate electricity using low-carbon hydrogen energy and the battery is discharged, replacing the high-carbon emission distribution network's power purchase, thus achieving peak shaving and valley filling of the overall carbon emission curve of the microgrid community.

[0077] In engineering implementation, this embodiment encapsulates DECI calculation, network inference, and control command issuance functions into software modules, which are deployed on the microgrid central controller or cloud server. During online operation, after receiving SCADA data, the system automatically completes DECI updates and network forward inference, and outputs control commands to adjust the operating power of the electrolyzer, battery, and hydrogen fuel turbine.

[0078] Test results for the IEEE 13-node system show that: Figure 6As shown, after training convergence, the cumulative carbon emissions of the MASAC-CVaR method of this invention stabilized at around 3282 kg, significantly lower than the baseline method's 3962 kg based on Lagrangian relaxation (MASAC-Lag), and it met the 5000 kg daily quota constraint throughout the process; Figure 7 As shown in the carbon emission heat map, the method of this invention successfully eliminates high-carbon emission areas during the evening peak hours. This is primarily due to the hydrogen microgrid's ability to produce hydrogen on a large scale using low-carbon electricity during the midday hours (10:00-15:00) when solar power is abundant. The carbon emission risk is extremely low, and during peak evening load periods (17:00-21:00), the stored green hydrogen is used to generate electricity instead of purchasing electricity from the high-carbon grid, achieving a significant "carbon-energy time shift" effect. Compared to traditional expected constraint methods, this invention reduces the risk of carbon emission violations (95% CVaR) by 17.5% by sacrificing only 7% of economic benefits, demonstrating outstanding robustness.

[0079] This application also provides an electronic device that may include a memory and a processor. The memory stores a computer program, and when the processor calls the computer program in the memory, it can implement the steps provided in the above embodiments. Of course, the electronic device may also include various network interfaces, power supplies, and other components.

[0080] This application also provides a readable storage medium storing a computer program thereon, which, when executed, can perform the steps provided in the above embodiments. The storage medium may include various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0081] The method provided by this invention has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are merely for the purpose of helping to understand the core ideas of this invention. It should be noted that those skilled in the art can make various improvements and modifications to this invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this invention.

Claims

1. A hydrogen microgrid scheduling method based on CVaR-constrained security reinforcement learning, characterized in that, Includes the following steps: S1: Obtain the environmental state spatial data of the heterogeneous multi-micronet community system; S2: Based on the environmental state space data, construct a dynamic equivalent carbon intensity tracking model to calculate the carbon intensity of distribution network nodes and the dynamic equivalent carbon intensity of energy storage devices; S3: Construct a multi-agent flexible Actor-Critic distributed security policy network; use the CVaR-constrained Lagrange relaxation algorithm to train it intensively, and obtain a low-carbon scheduling policy model; S4: Input the real-time status data into the model and output the dispatch scheme of each microgrid. The coordinated dispatch of the heterogeneous multi-microgrid community system under carbon emission constraints is completed through the carbon-energy coordinated time-shifting mechanism.

2. The method according to claim 1, characterized in that, The environmental state space data mentioned in step S1 includes photovoltaic power generation output, wind power generation output, battery state of charge, load power of each microgrid, time-of-use electricity price of the external power grid, and hydrogen level of the hydrogen storage tank. After the above data is preprocessed at hourly resolution, it is used as the state input of the reinforcement learning environment.

3. The method according to claim 1, characterized in that, The dynamic equivalent carbon intensity tracking model described in step S2 includes a nodal carbon intensity calculation model and a unified carbon flow model for energy storage systems; The node carbon intensity calculation model treats carbon emissions as a virtual fluid attached to the active power flow, and determines the real-time carbon intensity of the nodes by calculating the weighted average of the energy carbon intensity of all injected nodes. The unified carbon flow model of the energy storage system adopts a recursive mass-carbon balance equation, and dynamically updates the dynamic equivalent carbon intensity of the energy storage devices at the current moment by combining the carbon intensity of the energy storage devices at the previous moment, the charging / hydrogen production power at the current moment and the corresponding input source carbon intensity.

4. The method according to claim 1, characterized in that, The distributed security policy network described in step S3 equips each microgrid with an agent, the agent comprising: (1) Actor network: used to input local observations and output scheduling actions such as battery charging and discharging, electrolyzer power, and hydrogen combustion turbine power; (2) Reward Critic Network: Used to evaluate the economic value of actions; (3) Distributed Cost Critic Network: It adopts a non-cross quantile network structure to output multiple discrete quantiles and fit the complete probability distribution of future cumulative carbon emission costs.

5. The method according to claim 1, characterized in that, Step S3 employs a "centralized training, distributed execution" architecture to train the network, introducing Conditional Value at Risk (CVaR) as a risk metric. The CVaR risk value is calculated based on a pre-set confidence level, using the following formula: (3) in, For confidence level, For system status, For scheduling actions, For distributed Cost Critic networks at the quantile level The predicted carbon emission cost output is as follows. , This represents the total number of quantile samples.

6. The method according to claim 1, characterized in that, The Lagrange loss function constructed in step S3 is: (4) in, The expected return evaluated by the Reward Critic network. This is the entropy regularization temperature coefficient, used to regulate the balance between exploration and exploitation. Let Shannon entropy be the policy distribution. For Lagrange multipliers, used as adaptive penalty coefficients for violating safety constraints. This is the preset single-round carbon emission budget threshold.

7. The method according to claim 1, characterized in that, In step S3, the primal-dual method is used to update the parameters. Training convergence is achieved by alternately optimizing the network parameters and Lagrange multipliers. The Lagrange multiplier update formula is as follows: (5) in, Let be the learning rate of the Lagrange multiplier.

8. The method according to claim 1, characterized in that, The heterogeneous multi-microgrid community system described in step 4 includes traditional microgrids and hydrogen energy microgrids that integrate electrolyzers, hydrogen storage tanks, and hydrogen fuel turbines; The carbon-energy coordinated time-shifting mechanism is as follows: when the carbon intensity of the distribution network is lower than the preset threshold or during periods of renewable energy surplus, the hydrogen microgrid is controlled to increase the hydrogen production power of the electrolyzer and the traditional microgrid is used for battery charging; when the carbon intensity of the distribution network is higher than the preset threshold or during peak load periods, the hydrogen fuel turbine is controlled to generate electricity using hydrogen energy and the battery is used to discharge, replacing the power purchase from the distribution network.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the steps of the hydrogen microgrid scheduling method based on CVaR constrained security reinforcement learning as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the hydrogen microgrid scheduling method based on CVaR constrained security reinforcement learning as described in any one of claims 1 to 8.