Power distribution network operation optimization method and system containing distributed electric-hydrogen coupling system

By constructing an operational model for DPHCS and using deep reinforcement learning algorithms, the coordinated operation of DPHCS and the distribution network was optimized, solving the seasonal energy imbalance problem of renewable energy, improving the stability and energy utilization efficiency of the distribution network, and realizing the green production and efficient utilization of hydrogen energy.

CN120999614BActive Publication Date: 2026-02-06STATE GRID JIANGSU ELECTRIC POWER CO LTD RESEARCH INSTITUTE +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511503499.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2026-02-06
Estimated Expiration
2045-10-21

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively address the seasonal energy imbalance problem of renewable energy, single energy storage systems cannot meet the long-term regulation needs of distribution networks, and there is insufficient research on the optimization of distributed electric-hydrogen coupling systems (DPHCS) in distribution networks, neglecting their supporting role in distribution networks.

Method used

A distributed electric-hydrogen coupled system (DPHCS) operation model is constructed. By combining the Markov decision MDP model and the deep reinforcement learning DRL agent, the DistFlow linearized equation and the near-end policy optimization PPO algorithm (DFPPO) are used to design an expert safety layer and optimize the coordinated operation of DHCS and the distribution network.

Benefits of technology

By leveraging the large capacity and long-term energy storage characteristics of DHCS, the renewable energy carrying capacity and operational stability of the distribution network have been significantly improved, the curtailment rate of solar power has been reduced, medium- and long-term energy allocation and green hydrogen production have been achieved, and the energy transformation on the demand side has been promoted.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120999614B_ABST
    Figure CN120999614B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of new power system, and more particularly to a power distribution network operation optimization method and system containing a distributed power-hydrogen coupling system, the method comprising: constructing operation models of electrolytic cells, hydrogen storage tanks, compressors and fuel cell devices in the DPHCS; constructing a power distribution network operation optimization model considering the DPHCS; converting the optimization task of the operation optimization model into an MDP model; using a DFPPO algorithm based on a DistFlow linearization equation and a PPO algorithm to solve the MDP model, designing a deep reinforcement learning DRL agent based on the PPO algorithm, and using the DistFlow linearization equation to construct an expert safety layer; designing an operation optimization model calculation optimization process based on the DFPPO algorithm, training parameters of the DRL agent, and saving the trained parameters for online execution, which can efficiently handle the differentiated characteristics and complex constraints of multiple devices in the power distribution network.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of new power system technology, and particularly relates to a power distribution network operation optimization method and system containing a distributed electric-hydrogen coupling system. BACKGROUND

[0002] Due to the randomness and volatility of renewable energy, the distribution access brings challenges to the supply and demand balance and voltage stability of the power distribution network. Therefore, in the future, new technologies need to be developed to solve the negative effects brought by distributed grid connection and improve the operation stability of the power distribution network. Energy storage, as a flexible regulation resource that can be used for peak shaving, promoting renewable energy consumption, improving power quality and delaying power distribution network investment, will play a crucial role in the process of building a new type of power system and energy system.

[0003] However, a single energy storage system is difficult to meet the multi-element regulation demand of the modern power system under complex operating conditions due to its own operating principle and medium limitations. Due to the limitations of cost, capacity and user behavior, energy storage resources such as lithium batteries are mainly used for short-term and medium-term power regulation within a day, and still cannot effectively handle the long-time scale energy imbalance problem caused by the seasonal distribution of renewable energy, and still have certain limitations in improving the renewable energy carrying capacity of the power distribution network.

[0004] Hydrogen energy storage system has received widespread attention due to its large capacity and long-time storage advantages. The distributed electric-hydrogen coupling system (DPHCS) composed of electrolytic cell (EL), hydrogen storage tank (HSS), compressor (CO), fuel cell (FC) and hydrogen load can realize the bidirectional energy conversion of "electricity-hydrogen-electricity", improve the distributed photovoltaic (PV) consumption capacity of the power distribution network, realize the medium and long-term energy allocation, and at the same time, realize the green production of hydrogen energy and promote the demand side energy transformation. At present, the research on DPHCS mainly focuses on the park micro-grid level, and uses the multi-energy coupling characteristics of hydrogen energy to improve the energy utilization efficiency and energy supply safety of the park, and few researches are carried out on the coordinated operation of the power distribution network and the DPHCS. The optimization problems of the power distribution network and the DPHCS are often formulated separately, ignoring the supporting role of the DPHCS to the power distribution network.

[0005] The information disclosed in this BACKGROUND section is only for the purpose of enhancing the understanding of the background of the present application and should not be taken as an acknowledgment or any form of suggestion that this information forms prior art with respect to the present application. SUMMARY

[0006] The present application provides a power distribution network operation optimization method and system containing a distributed electric-hydrogen coupling system, thereby effectively solving the problems in the background art.

[0007] In order to achieve the above object, the technical scheme adopted by the present application is: a power distribution network operation optimization method containing a distributed electric-hydrogen coupling system, comprising the following steps:

[0008] An operation model of an electrolyzer, a hydrogen storage tank, a compressor and a fuel cell device in the distributed electric-hydrogen coupling system DPHCS is constructed;

[0009] A power distribution network operation optimization model considering the DPHCS is constructed based on the operation model, and the operation optimization model includes an objective function, a constraint condition and a decision variable;

[0010] The optimization task of the operation optimization model is converted into a Markov decision MDP model, and the MDP model includes a state space, an action space, a reward function, a state transition probability distribution and a power distribution network environment;

[0011] A DFPPO algorithm based on a DistFlow linearization equation and a proximal policy optimization PPO algorithm is used to solve the MDP model, and the DFPPO algorithm is used to design a deep reinforcement learning DRL agent based on the PPO algorithm and to construct an expert safety layer using the DistFlow linearization equation;

[0012] An optimization flow of the operation optimization model based on the DFPPO algorithm is designed, parameters of the DRL agent are trained, and the trained parameters are saved for online execution.

[0013] Further, the operation model of the electrolyzer, the hydrogen storage tank, the compressor and the fuel cell device in the distributed electric-hydrogen coupling system DPHCS includes:

[0014] For a node , representing a node set containing the DPHCS in the power distribution network, the operation of the DPHCS needs to meet the node power balance constraint and the device constraint;

[0015] The node power balance constraint is that the active power of the DPHCS at the node is determined by the active power output of the fuel cell FC and the active power consumption of the electrolyzer EL and the compressor CO , , ,

[0016] ;

[0017] The device constraint includes the electrolyzer EL operation constraint, the fuel cell FC operation constraint, the compressor CO operation constraint and the hydrogen storage tank HSS operation constraint.

[0018] Furthermore, the operating constraints of the electrolytic cell EL include:

[0019] ;

[0020] ;

[0021] ;

[0022] ;

[0023] In the formula, the active power consumption of the electrolytic cell is... Sources include power generated by self-built photovoltaic systems and grid power ; This represents the actual power output of the photovoltaic system. For abandoned light power, This refers to the grid-connected power of photovoltaic power. For the active power consumption of the electrolytic cell, Hydrogen production; The higher heating value of hydrogen. For nodes The efficiency of the electrolytic cell; yes Time Node Minimum / maximum active power consumption allowed in the electrolytic cell;

[0024] The operating constraints of the fuel cell (FC) include:

[0025] ;

[0026] ;

[0027] In the formula, For the active power output of the fuel cell, For hydrogen consumption, This is the lower heating value of hydrogen. For nodes The efficiency of fuel cells; yes Time Node The minimum / maximum active power allowed for the fuel cell;

[0028] The compressor CO operating constraints include:

[0029] ;

[0030] The above formula describes Time Node The active power consumed by the compressor and the hydrogen production of EL ,temperature efficiency of co universal gas constant initial pressure target pressure and isentropic exponent ;

[0031] The hydrogen storage tank HSS operation constraints include:

[0032] ;

[0033] ;

[0034] The above formula defines the storage state of the HSS at the time node , which is determined by the storage state at the previous time, the hydrogen production amount of the electrolytic cell , the hydrogen consumption amount of the fuel cell , the capacity of the hydrogen storage tank , and the hydrogen demand amount ; and are the minimum and maximum limit values of the storage state SOH, respectively. and are the minimum and maximum limit values of the storage state SOH, respectively.

[0035] Further, the power distribution network operation optimization model considering the DPHCS is constructed based on the operation model, including:

[0036] Considering an active power distribution network, including a set of nodes and a set of lines connecting the nodes ; for the nodes and the lines , the operation optimization model of the power distribution network in an optimization period includes:

[0037] Objective function:

[0038] ;

[0039] ;

[0040] ;

[0041] ;

[0042] wherein, is the active power loss, is the operation cost, is the total light rejection cost; , , , , and represent the active power loss, HSS operation, EL, CO, FC and curtailment cost coefficients of the distribution network, respectively;

[0043] Distribution network power flow constraints:

[0044] ;

[0045] ;

[0046] ;

[0047] ;

[0048] ;

[0049] wherein, denotes the node index, is the active power injection at node j, denotes the active power of the DPHCS at node j, is the active power injected by the PV inverter at node j, denotes the load power, and are the active power flow on branch jk or ij, is the resistance of branch ij, is the square of the current of branch ij, is the reactive power injection at node j, is the reactive power injected by the PV inverter at node j, is the reactive power of the load at node j, and are the reactive power flow on branch jk or ij, the direction is defined from i to j, is the reactance of branch ij, and are the square of the voltage magnitude at node i, j, and are the minimum and maximum voltage magnitude allowed at node;

[0050] PV reactive power output constraints:

[0051] ;

[0052] wherein, is the reactive power output of the PV, denotes the maximum capacity of the PV inverter at node ;

[0053] DPHCS power balance constraint:

[0054] ;

[0055] The above formula represents the nodes. The DPHCS power output at the location is generated by FC power generation. and the power consumption of EL and CO , Decide.

[0056] Furthermore, the state space Includes: status This reflects the DRL agent's perception of the distribution network's operating status at time t. , , and These are the active / reactive power vectors and node voltage magnitude vectors for all node loads, respectively. for The vector of actual photovoltaic power generation at any given time. Let be the hydrogen storage state vector of the HSS; This represents the total active power loss of the distribution network at the current moment. This is the hydrogen loading vector;

[0057] The action space Includes: actions It includes the decision variables made by the DRL agent in the power distribution network operating environment; , , The reactive power vector and the amount of light wasted from the PV output; , This is the active power vector output by DPHCS;

[0058] The reward function :

[0059] ;

[0060] In the formula, The instantaneous reward at time t. , Penalty for exceeding node voltage limits. As a penalty for insufficient hydrogen load supply and excessive hydrogen production, and As a penalty weight, Time Node The voltage penalty is defined as:

[0061] ;

[0062] In the formula, Vnom and Vnom represents a positive penalty term when the voltage of the node deviates from the nominal voltage Vnom by more than a certain range, otherwise it is 0;

[0063] The state transition probability distribution P includes: the probability of the state transitioning from s t to s t after the power flow calculation following the agent taking action a t+1 at time t;

[0064] The power distribution network environment includes: constructing an environment for DRL agents to interact according to real node and line parameters, at each decision step , the power distribution network injects variables , , The injection of these random quantities causes changes in the active power loss and voltage distribution of the power distribution network; the DRL agent makes corresponding actions according to the state at time t, inputs the power distribution network simulation environment, and then performs power flow calculation to obtain and the new power flow distribution result.

[0065] Further, the DFPPO algorithm designs a deep reinforcement learning DRL agent based on the PPO algorithm, including the following steps:

[0066] In the PPO algorithm, the DRL agent uses the policy gradient theory to optimize the objective , and the policy gradient estimates the direction of the gradient in the following form:

[0067] ;

[0068] In the formula, is the gradient calculation of the objective J, the policy represents the probability of taking action in state , and the parameter is , which is given by the actor network; represents the expected future cumulative return after taking action , the state value function is introduced, and the advantage function is defined to reflect the good or bad degree of action relative to the average level in this state:

[0069] ;

[0070] wherein, represents the advantage function;

[0071] The objective function of the PPO algorithm is:

[0072]

[0073] wherein, is the probability ratio of the new and old policy is the parameter of the Actor network, the clipping function limits the policy ratio to a range not exceeding , is the clipping parameter to avoid overfitting at each update; The PPO algorithm estimates the value function

[0074] with the help of the Critic network, and calculates the temporal difference error according to the following formula:

[0075]

[0076] wherein, is the temporal difference error at time t, is the immediate reward, is the discount factor, is the value function at time t+1; the objective of the Critic network is to make more accurately approximate the actual value function by minimizing the loss function, and the loss function is defined as:

[0077]

[0078] wherein, is the loss function, is the parameter of the critic function, and N is the number of experience trajectories, is the actual value function at time t, represents the expectation, is the temporal difference error at time t;

[0079] The parameters and of the Actor and Critic of the PPO algorithm are updated by the following formula:

[0080]

[0081]

[0082] wherein,​​​​​​ learning rate for the Actor network, learning rate for the Critic network;

[0083] Further, the constructing the expert safety layer using the DistFlow linearization equation comprises the following steps:

[0084] The topology of the distribution network is defined by a matrix:

[0085] ;

[0086] ;

[0087] ;

[0088] where F and T are incidence matrices representing the start and end nodes of the lines, M0 is the incidence matrix of the distribution network when the slack node is calculated, m0 is the column corresponding to the slack node, M is the incidence matrix of the distribution network when the slack node is not calculated, and is a judgment function, which is 1 if l is i or j, and 0 otherwise,

[0089] The power flow equation of the distribution network is represented as:

[0090] ;

[0091] ;

[0092] ;

[0093] where, and are diagonal matrices, which are the resistance and reactance of the lines, respectively, P and Q are the active and reactive power matrices, is a vector representing the square of the voltage of all nodes except the slack node, is the square of the voltage of the slack node, and its size is , is the number of nodes in the distribution network except the slack node; is the square of the line current, which is used to calculate the power loss, and in the distribution network, is ignored so that the power flow equation becomes linear; therefore, after the active and reactive powers are obtained, the equation is rewritten as:

[0094] ;

[0095] where, , As the identity matrix, the linear expression in the above equation shows the power vector and voltage vector v. 2 The relationship between the PPO strategy and the action space 'a' contains active power commands and reactive power commands.

[0096] Action space a and voltage v 2 The relationship between them is:

[0097] ;

[0098] In the above formula, and The components related to active and reactive power in the action are represented. Based on the above formula, an expert safety layer based on the DistFlow equation is constructed. The purpose of the expert safety layer is to minimize the Euclidean distance between the original unsafe action and the target safe action, and to project the unsafe action of the PPO agent into the safe area.

[0099] Furthermore, the projection of the unsafe actions of the PPO agent onto the safe area includes:

[0100] The projection process is achieved by solving an optimization problem, the optimization objective of which is to find the safest action closest to the original action 'a'. Its mathematical form is:

[0101] ;

[0102] ;

[0103] ;

[0104] In the formula, relaxation parameters It is used to manage the limitations on voltage amplitude.

[0105] Furthermore, the optimization process of the running optimization model based on the DFPPO algorithm includes the following steps:

[0106] During the exploration phase, the agent interacts with the distribution network environment according to the initial strategy, aiming to collect the experience trajectory of the current round; each round selects continuous data of a set value for interaction, including several decision steps; the original action made by the agent in each decision step will be optimized by the expert safety layer, and actions that do not meet the voltage constraints will be projected to the safety domain as much as possible according to the real-time state of the distribution network.

[0107] During the training phase, after each round of exploration, the agent trains the parameters of the Actor and Critic networks based on the collected historical experience trajectories; and updates the parameters using gradients.

[0108] In the execution stage, the agent outputs the scheduling strategy of the power distribution network and the DPHCS through the trained Actor network, and feeds back the regulation and control effect through power flow calculation.

[0109] The application also includes a power distribution network operation optimization system containing a distributed electricity-hydrogen coupling system, which uses the method as described above, and the system comprises:

[0110] An operation model modeling unit is configured to build operation models of electrolytic cells, hydrogen storage tanks, compressors and fuel cell devices in the distributed electricity-hydrogen coupling system DPHCS.

[0111] An optimization model modeling unit is configured to build a power distribution network operation optimization model considering the DPHCS based on the operation models, wherein the operation optimization model comprises an objective function, a constraint condition and a decision variable.

[0112] A conversion unit is configured to convert the optimization task of the operation optimization model into a Markov decision MDP model, wherein the MDP model comprises a state space, an action space, a reward function, a state transition probability distribution and a power distribution network environment.

[0113] A solving unit is configured to solve the MDP model using a DFPPO algorithm based on a DistFlow linearization equation and a proximal policy optimization PPO algorithm, wherein the DFPPO algorithm is used to design a deep reinforcement learning DRL agent based on the PPO algorithm, and an expert safety layer is built using the DistFlow linearization equation.

[0114] An online decision unit is configured to design an optimization flow of the operation optimization model based on the DFPPO algorithm, train parameters of the DRL agent, and save the trained parameters for online execution.

[0115] The application also includes a computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method as described above.

[0116] The application also includes a storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the method as described above.

[0117] The beneficial effects of the present application are that the large capacity and long time energy storage characteristics of DPHCS can solve the cross-season energy imbalance problem, and significantly improve the renewable energy carrying capacity and operation stability of the power distribution network. The proximal strategy optimization (DFPPO) algorithm based on the DistFlow equation ensures that the agent always meets the voltage constraints and other safety constraints of the power distribution network during the exploration process by constructing an expert safety layer. Compared with traditional model-based optimization methods, the DFPPO algorithm does not need to construct an accurate mathematical model, can efficiently handle the differentiated characteristics and complex constraints of multiple devices in the power distribution network, significantly improve the solution efficiency and accuracy of the optimization task, and at the same time ensure the safety of online operation. It can be widely applied to various scenarios such as campus microgrid and urban power distribution network, and provides an innovative solution for building a new power system and energy system. BRIEF DESCRIPTION OF DRAWINGS

[0118] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0119] Figure 1 Flowchart of the method in Example 1;

[0120] Figure 2 Structure schematic diagram of the system in Example 1;

[0121] Figure 3 Flowchart of the method in Example 2;

[0122] Figure 4 DPHCS schematic diagram in the power distribution network in Example 2;

[0123] Figure 5 DPHCS collaborative power distribution network optimization framework based on DFPPO algorithm in Example 2;

[0124] Figure 6 Structure schematic diagram of the computer equipment of the present application. DETAILED DESCRIPTION

[0125] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments.

[0126] Example 1:

[0127] As Figure 1The method comprises the following steps:

[0128] An operation model of an electrolytic cell, a hydrogen storage tank, a compressor and a fuel cell device in the distributed power-hydrogen coupling system DPHCS is constructed.

[0129] A power distribution network operation optimization model considering the DPHCS is constructed based on the operation model, and the operation optimization model comprises an objective function, a constraint condition and a decision variable.

[0130] The optimization task of the operation optimization model is converted into a Markov decision MDP model, and the MDP model comprises a state space, an action space, a reward function, a state transition probability distribution and a power distribution network environment.

[0131] The MDP model is solved by using a DFPPO algorithm based on a DistFlow linearization equation and a proximal policy optimization PPO algorithm, the DFPPO algorithm is used to design a deep reinforcement learning DRL agent based on the PPO algorithm, and an expert safety layer is constructed by using the DistFlow linearization equation.

[0132] An operation optimization model calculation optimization process based on the DFPPO algorithm is designed, parameters of the DRL agent are trained, and the trained parameters are saved for online execution.

[0133] By introducing the DPHCS, bidirectional energy conversion of "electricity-hydrogen-electricity" is realized, and the randomness and volatility of renewable energy (such as photovoltaic) are effectively solved. The large-capacity and long-time energy storage characteristics of the DPHCS can cope with the cross-season energy imbalance problem, significantly improving the renewable energy carrying capacity and operation stability of the power distribution network. At the same time, through the collaborative optimization of the electrolytic cell and the fuel cell, the energy utilization efficiency of the power distribution network is further improved, and the light and wind curtailment rate is reduced.

[0134] The proximal policy optimization (DFPPO) algorithm based on the DistFlow equation constructs an expert safety layer to ensure that the agent always meets the voltage constraints and other safety constraints during the exploration process. Compared with the traditional model-based optimization method, the DFPPO algorithm does not need to construct an accurate mathematical model, can efficiently handle the differentiated characteristics and complex constraints of multiple devices in the power distribution network, significantly improve the solving efficiency and accuracy of the optimization task, and ensure the safety of online operation.

[0135] The embodiment not only optimizes the operation of the power distribution network, but also realizes green production and efficient utilization of hydrogen energy through the DPHCS, and promotes the energy transformation of the demand side. Through the collaborative optimization of the electricity-hydrogen coupling system, the medium and long-term energy allocation capability is provided for the power distribution network, and technical support is provided for the large-scale application of hydrogen energy. In addition, the method of the embodiment can be widely applied to various scenes such as park microgrids and urban power distribution networks, and provides an innovative solution for building a new power system and energy system.

[0136] In the embodiment, the operation models of the electrolyzer, hydrogen storage tank, compressor and fuel cell device in the distributed electricity-hydrogen coupling system DPHCS are constructed, including:

[0137] For node , , the set of nodes containing DPHCS in the power distribution network, the operation of DPHCS needs to meet the node power balance constraint and the device constraint.

[0138] Node power balance constraint: the active power of DPHCS at node is determined by the active power output of the fuel cell FC and the active power consumption of the electrolyzer EL and the compressor CO , ,

[0139] .

[0140] The device constraints include electrolyzer EL operation constraints, fuel cell FC operation constraints, compressor CO operation constraints and hydrogen storage tank HSS operation constraints.

[0141] The electrolyzer EL operation constraints include:

[0142] .

[0143] .

[0144] .

[0145] .

[0146] In the formula, the active power consumption of the electrolyzer includes the power generation of the self-built photovoltaic system and the grid power . is the actual power output of the photovoltaic, is the abandoned light power, is the photovoltaic on-grid power; is the active power consumption of the electrolyzer, ​for the hydrogen production amount; for the higher heating value of hydrogen, for the node at the electrolyzer; is for the node at the electrolyzer;

[0147] The fuel cell FC operating constraints include:

[0148] ;

[0149] ;

[0150] wherein is the active power output of the fuel cell, is the hydrogen consumption amount, is the lower heating value of hydrogen, for the node at the fuel cell; is for the node at the fuel cell;

[0151] The compressor CO operating constraints include:

[0152] ;

[0153] The above equation describes for the node at the compressor, in terms of the hydrogen production amount , the temperature , the efficiency of the CO , the universal gas constant , the initial pressure , the target pressure , and the isentropic exponent ;

[0154] The hydrogen storage tank HSS operating constraints include:

[0155] ;

[0156] ;

[0157] The above equation defines for the node at the HSS , in terms of the storage state at the previous time step , the hydrogen production amount of the electrolyzer , the hydrogen consumption amount of the fuel cell , the capacity of the hydrogen storage tank , and the hydrogen demand determination; and are the minimum and maximum limit values of the storage state SOH, respectively.

[0158] Based on the operation model, the operation optimization model of the distribution network considering the DPHCS is constructed, including:

[0159] Considering an active distribution network, including a set of nodes and a set of lines connecting the nodes ; for the nodes and the lines , the operation optimization model of the distribution network in an optimization period includes:

[0160] Objective function:

[0161] ;

[0162] ;

[0163] ;

[0164] ;

[0165] wherein, is the active power loss, is the operation cost, is the total light rejection cost; , , , , and represent the active loss, HSS operation, EL, CO, FC, and light rejection cost coefficients of the distribution network, respectively;

[0166] Distribution network power flow constraints:

[0167] ;

[0168] ;

[0169] ;

[0170] ;

[0171] ;

[0172] wherein, denotes the node index, is the active injection power of node j, This represents the active power of the DPHCS at node j. The active power output at node j is generated by the grid-connected photovoltaic inverter. Indicates load power. and For the active power flow on branch jk or ij, Let be the resistance of branch ij. Let be the square of the current in branch ij. Inject reactive power into node j. The reactive power output at node j is generated by the grid-connected photovoltaic inverter. Let j be the reactive power of the load at node j. and For reactive power flow on branches jk or ij, the direction is conventionally defined as from i to j. For the reactance of branch ij, and Let be the square of the voltage magnitude at nodes i and j. and These are the minimum and maximum allowable voltage amplitudes for the node;

[0173] Solar reactive power output constraints:

[0174] ;

[0175] In the formula, To provide reactive power output for photovoltaic systems. Represents a node The maximum capacity of the PV inverter;

[0176] DPHCS power balance constraint:

[0177] ;

[0178] The above formula represents the nodes. The DPHCS power output at the location is generated by FC power generation. and the power consumption of EL and CO , Decide.

[0179] Among them, the state space Includes: status This reflects the DRL agent's perception of the distribution network's operating status at time t. , , and These are the active / reactive power vectors and node voltage magnitude vectors for all node loads, respectively. for The vector of actual photovoltaic power generation at any given time. Let be the hydrogen storage state vector of the HSS; This represents the total active power loss of the distribution network at the current moment. This is the hydrogen loading vector;

[0180] Action space Includes: actions It includes the decision variables made by the DRL agent in the power distribution network operating environment; , , The reactive power vector and the amount of light wasted from the PV output; , This is the active power vector output by DPHCS;

[0181] reward function :

[0182] ;

[0183] In the formula, The instantaneous reward at time t. , Penalty for exceeding node voltage limits. As a penalty for insufficient hydrogen load supply and excessive hydrogen production, and As a penalty weight, Time Node The voltage penalty is defined as:

[0184] ;

[0185] In the formula, Rated voltage, and This formula represents the upper and lower limits of the voltage amplitude. Time Node A positive penalty is applied when the voltage deviates from the rated voltage by more than a certain range; otherwise, it is 0.

[0186] State transition probability distribution P include: The agent takes action a at time t. t After power flow calculation, the state changes from s t Transfer to s t+1 The probability of;

[0187] The distribution network environment includes: constructing an environment for DRL agents to interact based on real node and line parameters, and performing actions at each decision step. Variables will be injected into the distribution network. , , The injection of these random variables leads to changes in the active power loss and voltage distribution of the distribution network; the DRL agent adjusts its actions based on the state at time t. Make the corresponding actions Input the data into the power distribution network simulation environment, and then perform power flow calculations to obtain the results. And the new trend distribution results.

[0188] In this embodiment, the DFPPO algorithm designs a deep reinforcement learning DRL agent based on the PPO algorithm, including the following steps:

[0189] In the PPO algorithm, the DRL agent uses policy gradient theory to optimize the objective. The policy gradient is estimated by the following form:

[0190] ;

[0191] In the formula, To calculate the gradient of target J, the strategy is... Indicates the state Take action below The probability, with parameter . This strategy is given through the Actor network; Indicates taking action The expected cumulative return after the current state is introduced into the state value function. And define an advantage function to reflect the action. How good or bad is relative to the average level under this condition:

[0192] ;

[0193] In the formula, Represents the dominant function;

[0194] The objective function of the PPO algorithm is:

[0195] ;

[0196] In the formula, New and old strategies The probability ratio, For the parameters of the Actor network, the shearing function By limiting the strategy ratio The range does not exceed , To truncate parameters and avoid excessive updates each time;

[0197] The PPO algorithm estimates the value function using a critic network. The timing difference error is calculated according to the following formula:

[0198] ;

[0199] In the above formula, is the timing difference error at time t, is the immediate reward, is the discount factor, is the value function at time t+1; the goal of the critic network is to make more accurately approximate the actual value function by minimizing the loss function, which is defined as:

[0200] ;

[0201] In the above formula, loss function, is the parameter of the critic function, and N is the number of experience trajectories, is the actual value function at time t, represents expectation, is the timing difference error at time t;

[0202] PPO algorithm actor and critic parameters and are updated by:

[0203] ;

[0204] ;

[0205] In the formula, is the learning rate of the actor network, is the learning rate of the critic network;

[0206] The expert safety layer is constructed using the DistFlow linearization equation, including the following steps:

[0207] The topology structure of the distribution network is defined by a matrix:

[0208] ;

[0209] ;

[0210] ;

[0211] In the formula, the connection matrix F and T represent the starting node and the terminal node of the line, M0 is the incidence matrix of the distribution network when calculating the balanced node, m0 is the column corresponding to the balanced node, is the incidence matrix of the distribution network when not calculating the balanced node, and is a judgment function, which is 1 if i is j or j, otherwise 0,

[0212] The power flow equation of the distribution network is expressed as:

[0213]

[0214]

[0215]

[0216] wherein, and are diagonal matrices, which are the resistance and reactance of the line respectively, P and Q are active and reactive power matrices, represents the vector of the square value of the voltage of all nodes except the balance node, represents the square value of the voltage of the balance node, and the size is represents the number of nodes in the distribution network except the balance node; represents the square of the line current, which is used to calculate the power loss, and in the distribution network, is ignored so that the power flow equation becomes linear; therefore, after the active power and the reactive power are obtained, the equation is rewritten as:

[0217]

[0218] wherein, , is a unit matrix, and the linear expression in the above formula shows the relationship between the power vector and the voltage vector v 2 The action space a corresponding to the PPO strategy includes the active power instruction and the reactive power instruction;

[0219] The relationship between the action space a and the voltage v 2 is:

[0220]

[0221] In the above formula, and represent the components related to the active and reactive in the action, and according to the above formula, the expert safety layer based on the DistFlow equation is constructed, and the purpose of the expert safety layer is to minimize the Euclidean distance between the original unsafe action and the target safe action, and to project the unsafe action of the PPO agent into the safe region.

[0222] Projecting the unsafe action of the PPO agent into the safe region includes:

[0223] ​​​​​​The projection process is implemented by solving an optimization problem whose objective is to find the closest safe action to the original action a , mathematically,

[0224] ;

[0225] ;

[0226] ;

[0227] wherein, , the relaxation parameter is used to manage the limit condition of the voltage amplitude.

[0228] Wherein, the operation optimization model based on the DFPPO algorithm is designed to calculate the optimization process, including the following steps:

[0229] In the exploration stage, the agent interacts with the power distribution network environment according to the initialized strategy, aiming to collect the experience trajectory of the current round; A set value of continuous data is selected for interaction in each round, including several decision steps; The original action made by the agent at each decision step will be optimized by the expert safety layer, and the action that does not meet the voltage constraint will be projected to the safe domain as much as possible according to the real-time state of the power distribution network;

[0230] In the training stage, after the round exploration is completed, the agent trains the parameters of the actor Actor and critic Critic network according to the collected historical experience trajectory; The parameters are updated by gradient;

[0231] In the execution stage, the agent outputs the scheduling strategy of the power distribution network and the DPHCS through the trained Actor network, and feedbacks the regulation and control effect through power flow calculation.

[0232] As shown in Figure 2 , the embodiment further includes a power distribution network operation optimization system with a distributed electricity-hydrogen coupling system, which uses the method as described above, and the system includes:

[0233] An operation model modeling unit is configured to construct operation models of electrolytic cells, hydrogen storage tanks, compressors and fuel cell devices in the distributed electricity-hydrogen coupling system DPHCS.

[0234] An optimization model modeling unit is configured to construct a power distribution network operation optimization model considering the DPHCS based on the operation model, and the operation optimization model includes an objective function, a constraint condition and a decision variable.

[0235] A conversion unit is configured to convert the optimization task of the operation optimization model into a Markov decision MDP model, and the MDP model includes a state space, an action space, a reward function, a state transition probability distribution and a power distribution network environment.

[0236] a solving unit configured to solve the MDP model using a DFPPO algorithm based on a DistFlow linearization equation and a proximal policy optimization PPO algorithm, the DFPPO algorithm being configured to design a deep reinforcement learning DRL agent based on the PPO algorithm and to construct an expert safety layer using the DistFlow linearization equation;

[0237] an online decision unit configured to design a running optimization model calculation optimization process based on the DFPPO algorithm, to train parameters of the DRL agent, and to save the trained parameters for online execution.

[0238] Embodiment 2:

[0239] With reference to Figure 3 , the embodiment provides a power distribution network running optimization method containing a distributed power-hydrogen coupling system, comprising:

[0240] Step 1: constructing a running model of an electrolyzer, a hydrogen storage tank, a compressor and a fuel cell device in the DPHCS;

[0241] As shown in a DPHCS schematic diagram in a power distribution network, for a node Figure 4 , , representing a node set containing the DPHCS in the power distribution network, running of the DPHCS needs to satisfy a node power balance constraint and device constraints:

[0242] 1) node power balance constraint: active power of the DPHCS at the node is determined by active power output of a fuel cell (FC) and active power consumption , of an electrolyzer (EL) and a compressor (CO).

[0243] (1)

[0244] 2) EL running model: in the embodiment, the EL obtains active power from a power grid or a photovoltaic system inside a DEHIS to convert into hydrogen, and the running constraint is as follows:

[0245] (2)

[0246] (3)

[0247] (4)

[0248] (5)

[0249] In the above formula: the active power consumption of the electrolytic cell Sources include power generated by self-built photovoltaic systems and grid power ;in As expressed by equation (3), For the actual power output of photovoltaic, For abandoned light power, The power output of the photovoltaic grid is given by equation (4); Equation (4) represents the active power consumption of the electrolyzer. With hydrogen production The relationship between them The higher heating value of hydrogen. For nodes The efficiency of the electrolytic cell; Equation (5) ensures Time Node The active power consumption of the electrolytic cell is within the allowable range. It is the minimum (maximum) active power consumption allowed by the electrolytic cell.

[0250] 3) FC Operation Model: In this invention, the FC modeling method is similar to that of an electrolytic cell, but the operation method is reversed. The FC operation constraints are as follows:

[0251] (6)

[0252] (7)

[0253] In the above formula: Equation (6) characterizes the active power output of the fuel cell. Hydrogen consumption The relationship between them This is the lower heating value of hydrogen. For nodes The efficiency of the fuel cell; Equation (7) ensures Time Node The active power output and hydrogen consumption of the fuel cell are within the specified limits. It is the minimum (maximum) active power allowed for a fuel cell.

[0254] 4) CO Operation Model: In the DPHCS, CO pumps the hydrogen produced by the electrolyzer into the hydrogen storage tank to improve its storage efficiency and meet the pressure requirements of the storage tank, facilitating transportation and use. The CO operation constraints are as follows:

[0255] (8)

[0256] Equation (8) describes Time Node The active power consumed by the compressor and the hydrogen production of EL ,temperature CO efficiency Universal gas constant Initial pressure Target pressure and isentropic index related.

[0257] 5) HSS Operating Model: The HSS is used to store hydrogen gas after CO compression. Its key state quantity is the hydrogen storage state (SOH). The constraints on the HSS are as follows:

[0258] (9)

[0259] (10)

[0260] Equation (9) defines Time Node Storage status of HSS It is from the previous one Storage state at any moment Hydrogen production of the electrolyzer Hydrogen consumption of fuel cells Hydrogen storage tank capacity and hydrogen demand The decision (10) ensures that the hydrogen storage state is within the specified limits. and These are the minimum and maximum limits for SOH, respectively;

[0261] 6) Hydrogen load model in DPHCS: The hydrogen load demand in DPHCS comes from the consumption of hydrogen fuel cell vehicles. This invention generates a hydrogen load characteristic curve by referring to the actual hydrogen refueling time and hydrogen charge. The average hydrogen charge of hydrogen fuel cell vehicles is 1.5 kg, and the daily refueling time is mainly distributed between 6 am and 6 pm.

[0262] Step 2: Construct a distribution network operation optimization model that takes into account DPHCS, including the objective function, constraints, and decision variables;

[0263] This invention considers a distribution network that includes distributed PV power plants, DPHCS (Diverterless Power Distribution System), and various types of electrical loads. It considers an active distribution network comprising a set of nodes. And a set of connecting nodes For nodes and lines The operation optimization model of the distribution network within one optimization cycle can be expressed as the following SOCP equation:

[0264] (9)

[0265] (10)

[0266] (11)

[0267] (12)

[0268] (13)

[0269] (14)

[0270] (15)

[0271] (16)

[0272] (17)

[0273] (18)

[0274] (19)

[0275] In an optimization cycle , the objective function is given by equation (9), which is composed of active power loss , operation cost and total light rejection cost as shown in equations (10)-(12). Equations (13)-(17) represent the power flow constraints of the distribution network based on SOCP, where equations (13) and (14) represent the active and reactive power balance constraints at node . Equation (18) limits the reactive power output of the PV , where represents the maximum capacity of the PV inverter at node . Equation (19) represents the power balance constraint of the DPHCS, i.e. the DPHCS power output at node is determined by the FC power generation and the power consumption of EL, CO , .

[0276] Step 3: Convert the distribution network operation optimization task with DPHCS in step 2 into an MDP model, including state space, action space, reward function, state transition probability distribution and distribution network environment;

[0277] 1) State space : The state reflects the perception of the DRL intelligent agent agent to the operation state of the distribution network at time t, , , and are the active / reactive power vector and the node voltage magnitude vector of all nodes; is the actual PV power generation vector at time t, is the hydrogen storage state vector of HSS; is the total active power loss value of the distribution network at the current time; is the hydrogen load vector.

[0278] 2) Action space : Action contains the decision variables made by the DRL agent in the operation environment of the distribution network. , , is the reactive power vector and the light abandoned amount of PV output; , is the active power vector output by DPHCS.

[0279] 3) Reward function : The reward function of the present application can be expressed as:

[0280] (20)

[0281] In the formula, is the immediate reward at time t, , is the voltage out-of-limit penalty, is the hydrogen load supply shortage and hydrogen production excess penalty, and are penalty weights, and the present application defines the voltage penalty of node at time t as:

[0282] (21)

[0283] wherein, is the rated voltage, and are the upper and lower limits of the voltage magnitude, and the formula represents that when the voltage of node at time t deviates from the rated voltage by more than a range, a positive penalty term is accepted, otherwise it is 0.

[0284] 4) State transition probability distribution P : In the present application, represents the state transition from s t to s t after the agent takes action a t+1 ​​​The probability of this. The core of the DRL method is to learn the dynamic characteristics of the distribution network and capture source-load uncertainties through data in this process.

[0285] 5) Distribution network environment: This invention constructs an environment for DRL agents to interact based on real node and line parameters, and at each decision step... Variables will be injected into the distribution network. , , (These electrical quantities can all be obtained through SCADA.) The injection of these random quantities leads to changes in the active power loss and voltage distribution of the distribution network. The DRL agent determines the changes based on the state at time t. Make the corresponding actions Input the data into the power distribution network simulation environment, and then perform power flow calculations to obtain the results. And the new trend distribution results.

[0286] Step 4: Propose a DFPPO algorithm to solve the MDP model in Step 3, design a DRL agent based on the PPO algorithm, and use DistFlow linearization equations to construct an expert safety layer to ensure the safety of the policy during the agent's exploration process.

[0287] In the PPO algorithm, the agent typically uses policy gradient theory to optimize the objective. The policy gradient is often estimated in the following form to determine its direction:

[0288] (twenty two)

[0289] In the formula, To calculate the gradient of target J, the strategy is... Indicates the state Take action below The probability, with parameter . This strategy is given through the Actor network. Indicates taking an action The expected cumulative return in the future is directly used. Using a target as a training objective often introduces significant variance during actual training. To reduce variance, a state-value function is typically introduced. And define an advantage function to reflect the action. How good or bad is relative to the average level under this condition:

[0290] (twenty three)

[0291] In the formula, Represents the dominant function;

[0292] The core of PPO algorithm is to limit the policy update range by introducing probability ratio and clipping mechanism, and guide the policy update by combining advantage function, so as to ensure the efficiency and stability of the training process. The objective function of PPO algorithm is:

[0293] (24)

[0294] where, is the probability ratio of the new and old policy, is the parameter of the Actor network, and the clipping function limits the policy ratio to a range not exceeding , is the clipping parameter to avoid excessive update each time. PPO algorithm estimates the value function

[0295] with the help of Critic network, and calculates the time difference error according to the following formula:

[0296] (25) The objective of Critic network is to make

[0297] more accurately approximate the actual value function by minimizing the loss function, and the loss function is defined as:

[0298] (26) The Actor and Critic parameters

[0299] and of PPO algorithm can be updated by the following formula:

[0300] (27)

[0301] (28)

[0302] The PPO policy security layer based on the DistFlow equation of distribution network is constructed according to the real-time state of distribution network, and the topology structure of distribution network can be defined by matrix:

[0303] (29)

[0304] (30)

[0305] (31)

[0306] ​In the above equation, the connection matrix F and T represent the start node and end node of the line, M0 is the incidence matrix of the power distribution network, and m0 is the column corresponding to the balanced node. According to the above representation, the power flow equation of the power distribution network can be represented as:

[0307] (32)

[0308] (33)

[0309] (34)

[0310] In the above equation, and are diagonal matrices, which are the resistance and reactance of the line, respectively, represents the vector of the square of the voltage of all nodes except the balanced node, represents the square of the voltage of the balanced node. represents the square of the line current, which is used to calculate the power loss, and in the power distribution network, can be ignored so that the power flow equation becomes linear. Therefore, after obtaining the active power and reactive power, the equation can be rewritten as:

[0311] (35)

[0312] where, , is the identity matrix, and the linear expression in the above equation shows the relationship between the power vector and the voltage vector v 2 . The action space a corresponding to the PPO strategy contains the active power instruction and the reactive power instruction . Therefore, the relationship between the action space a and the voltage v 2 can be further derived as:

[0313] (36)

[0314] Further, according to the above equation, an expert safety layer based on the DistFlow equation can be constructed, and the purpose of this safety layer is to minimize the Euclidean distance between the original unsafe action and the target safe action, thereby projecting the unsafe action of the PPO agent into the safe region. The specific projection process is realized by solving an optimization problem, and the optimization objective is to find the safe action closest to the original action a, which is mathematically expressed as:

[0315] (37)

[0316] (38)

[0317] (39)

[0318] In the above formula, , the relaxation parameter is used to manage the limit condition of the voltage amplitude. By adding the relaxation parameter, a certain buffer space can be provided for the operation constraint to adapt to the possible deviation between the prediction and the actual voltage amplitude, so as to ensure that the generated action is within the safe operation range, while minimizing the deviation caused by the strategy adjustment as much as possible.

[0319] Step 5: design a DPHCS collaborative power distribution network operation optimization process based on the DFPPO algorithm, train the parameters of the DFPPO agent, and save the trained parameters for online execution;

[0320] The process of the DFPPO algorithm of the application for performing the power distribution network optimization operation task can be divided into three stages: round experience exploration, offline parameter training and online decision execution.

[0321] In the exploration stage, the agent interacts with the power distribution network environment according to the initialized strategy, aiming to collect the experience trajectory of the current round. It is worth noting that in order to reflect the long-term flexible regulation and control capability of DPHCS, 4 months of continuous data are selected for interaction every round, including 3000 decision steps. The original action made by the agent at each decision step will be optimized by the safety layer, which will project the action that does not meet the voltage constraint to the safety domain as much as possible according to the real-time state of the power distribution network, to ensure the safety of the strategy.

[0322] In the training stage, after the round exploration, the agent trains the parameters of the Actor and Critic networks according to the collected historical experience trajectory. First, the advantage function is calculated according to formula (23), then the target function of PPO is calculated according to formula (24) and the advantage function, and finally the gradient update formula (27)-(28) is used to update the parameters and .

[0323] In the execution stage, the agent outputs the scheduling strategy of the power distribution network and the DPHCS through the trained Actor network, and feeds back the regulation and control effect through power flow calculation. The application trains the agent for 1000 rounds according to the three stages, saves the network parameters, and deploys them to the dispatching center of the power distribution network for online decision-making. Figure 5 The overall optimization framework and the exploration and training details of DFPPO are shown.

[0324] Please refer to Figure 6The structural schematic diagram of the computer device provided by the embodiment of the application is shown. The computer device 400 provided by the embodiment of the application comprises a processor 410 and a memory 420, the memory 420 stores a computer program executable by the processor 410, and the computer program is executed by the processor 410 to perform the method as above.

[0325] The embodiment of the application further provides a storage medium 430, the storage medium 430 stores a computer program, and the computer program is executed by the processor 410 to perform the method as above.

[0326] The storage medium 430 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk or an optical disk.

[0327] In the description of the application, the terms "first", "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" can explicitly or implicitly include one or more of the features. The meaning of "a plurality of" is two or more, unless otherwise specifically limited.

[0328] In the application, unless otherwise specifically defined and limited, the terms "mounting", "connection", "connection", "fixing" and the like should be understood in a broad sense, for example, it can be fixed connection, or detachable connection, or integrated; it can be mechanical connection, or electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, it can be the communication or interaction relationship between two elements. For those skilled in the art, the specific meaning of the above terms in the application can be understood according to the specific circumstances.

[0329] In the description of the specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific feature, structure, material or characteristic being described is included in at least one embodiment or example of the present application. The illustrative descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples. Furthermore, the different embodiments or examples described in the specification can be combined and combined with each other in any suitable manner without mutual contradiction.

[0330] Any process or method descriptions or descriptions of the flow diagrams in the specification can be understood as representing code modules, segments, or portions of code which include one or more executable instructions for implementing specific logic functions (or steps) in the process, and the various preferred embodiments of the application include additional or different code modules, segments, or portions of code for implementing the application. The various processes and methods described in the specification can be understood as representing a process or method, including the functions specified in the flow diagrams, and the preferred embodiments of the application include additional or different processes or methods, which can be implemented in hardware, software, firmware, or any combination thereof, as desired.

[0331] The logic and / or steps represented in the flow diagrams or otherwise described in the specification, for example, can be considered as a list of executable instructions for implementing the logic function, and the various preferred embodiments of the application include additional or different code modules, segments, or portions of code for implementing the application. The various processes and methods described in the specification can be understood as representing a process or method, including the functions specified in the flow diagrams, and the preferred embodiments of the application include additional or different processes or methods, which can be implemented in hardware, software, firmware, or any combination thereof, as desired. For the purposes of this specification, a "computer readable medium" can be any device or apparatus that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer readable medium include the following: an electrical connection having one or more wires (electrical devices), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). In addition, the computer readable medium can even be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example, by optically scanning the paper or other suitable medium, then electronically converted into a form that is suitable for use by the instruction execution system, apparatus, or device. The computer readable medium can also be a medium that can be programmed by the user or it can be a medium that is programmed to store data for use by or in connection with the instruction execution system, apparatus, or device.

[0332] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0333] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0334] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for optimizing the operation of a power distribution network containing a distributed electro-hydrogen coupling system, characterized in that, Includes the following steps: Construct an operational model for the electrolyzer (EL), hydrogen storage tank (HSS), compressor (CO), and fuel cell (FC) in a distributed electrohydrogen coupling system (DPHCS); Based on the aforementioned operating model, a distribution network operation optimization model considering DPHCS is constructed. The operation optimization model includes an objective function, constraints, and decision variables. The optimization task of the aforementioned operational optimization model is transformed into a Markov Decision Program (MDP) model, which includes a state space, action space, reward function, state transition probability distribution, and distribution network environment. The MDP model is solved using the DFPPO algorithm, which is based on the DistFlow linearized equation and the PPO algorithm for near-end policy optimization. The DFPPO algorithm designs a deep reinforcement learning DRL agent based on the PPO algorithm and uses the DistFlow linearized equation to construct an expert safety layer. Design the computational optimization process of the running optimization model based on the DFPPO algorithm, train the parameters of the DRL agent, and save the trained parameters for online execution; The construction of the distribution network operation optimization model based on the operating model and considering DHCCS includes: Consider an active distribution network, consisting of a set of nodes. And a set of connecting nodes For nodes and lines The operation optimization model of the distribution network within one optimization cycle includes: Objective function: ; ; ; ; In the formula, For active power loss, For operating costs, Total cost of curtailment; , , , , and These represent the active power loss of the distribution network, HSS operation, EL, CO, FC, and curtailment cost coefficient, respectively. For FC power generation, , Power consumption of EL and CO; Distribution network power flow constraints: ; ; ; ; ; In the formula, Indicates the node index. Inject active power into node j. This represents the active power of the DPHCS at node j. The active power output at node j is generated by the grid-connected photovoltaic inverter. Indicates load power. and For the active power flow on branch jk or ij, Let be the resistance of branch ij. Let be the square of the current in branch ij. Inject reactive power into node j. The reactive power output at node j is generated by the grid-connected photovoltaic inverter. Let j be the reactive power of the load at node j. and For reactive power flow on branch jk or ij, For the reactance of branch ij, and Let be the square of the voltage magnitude at nodes i and j. and These are the minimum and maximum allowable voltage amplitudes for the node; Solar reactive power output constraints: ; In the formula, To provide reactive power output for photovoltaic systems. Represents a node The maximum capacity of the PV inverter; DPHCS power balance constraint: ; The above formula represents the nodes. The DPHCS power output at the location is generated by FC power generation. and the power consumption of EL and CO , Decide.

2. The method for optimizing the operation of a distribution network containing a distributed electro-hydrogen coupling system according to claim 1, characterized in that, The operational model for the electrolyzer, hydrogen storage tank, compressor, and fuel cell device in the Distributed Electrohydrogen Coupling System (DPHCS) includes: For nodes , This represents the set of nodes in the distribution network that contain DPHCS. The operation of DPHCS needs to satisfy node power balance constraints and constraints of each device. The constraints for each piece of equipment include: electrolyzer (EL) operation constraints, fuel cell (FC) operation constraints, compressor (CO) operation constraints, and hydrogen storage tank (HSS) operation constraints.

3. The method for optimizing the operation of a distribution network containing a distributed electro-hydrogen coupling system according to claim 2, characterized in that, The operating constraints of the electrolytic cell EL include: ; ; ; ; In the formula, the active power consumption of the electrolytic cell is... Sources include power generated by self-built photovoltaic systems and grid power ; This represents the actual power output of the photovoltaic system. For abandoned light power, This refers to the grid-connected power of photovoltaic power. For the active power consumption of the electrolytic cell, Hydrogen production; The higher heating value of hydrogen. For nodes The efficiency of the electrolytic cell; yes Time Node Minimum / maximum active power consumption allowed in the electrolytic cell; The operating constraints of the fuel cell (FC) include: ; ; In the formula, For the active power output of the fuel cell, For hydrogen consumption, This is the lower heating value of hydrogen. For nodes The efficiency of fuel cells; yes Time Node The minimum / maximum active power allowed for the fuel cell; The compressor CO operating constraints include: ; The above formula describes Time Node The active power consumed by the compressor and the hydrogen production of the electrolyzer EL ,temperature CO efficiency Universal gas constant Initial pressure Target pressure and isentropic index related; The operating constraints of the hydrogen storage tank HSS include: ; ; The above formula defines Time Node Storage status of HSS It is based on the stored state from the previous moment. Hydrogen production of the electrolyzer at time t Hydrogen consumption of fuel cells Hydrogen storage tank capacity and hydrogen demand Decide; and These are the minimum and maximum limits for the storage state SOH, respectively.

4. The method for optimizing the operation of a distribution network containing a distributed electro-hydrogen coupling system according to claim 1, characterized in that, The state space Includes: status This reflects the DRL agent's perception of the distribution network's operating status at time t. , , and These are the active / reactive power vectors and node voltage magnitude vectors for all node loads, respectively. for The vector of actual photovoltaic power generation at any given time. Let be the hydrogen storage state vector of the HSS; This represents the total active power loss of the distribution network at the current moment. This is the hydrogen loading vector; The action space Includes: actions It includes the decision variables made by the DRL agent in the power distribution network operating environment; , , The reactive power vector and the amount of light wasted from the PV output; , This is the active power vector output by DPHCS; The reward function : ; In the formula, The instantaneous reward at time t. , Penalty for exceeding node voltage limits. As a penalty for insufficient hydrogen load supply and excessive hydrogen production, and As a penalty weight, Time Node The voltage penalty is defined as: ; In the formula, Rated voltage, and This formula represents the upper and lower limits of the voltage amplitude. Time Node voltage A positive penalty is applied if the voltage deviates from the rated voltage by more than a certain range; otherwise, the penalty is 0. The state transition probability distribution P includes: The agent takes action a at time t. t After power flow calculation, the state changes from s t Transfer to s t+1 The probability of; The power distribution network environment includes: constructing an environment for DRL agents to interact with based on real node and line parameters, and performing interactions at each decision step. Variables will be injected into the distribution network. , , ,variable , , The injection caused changes in the active power loss and voltage distribution of the distribution network; the DRL agent, based on the state at time t... Make the corresponding actions Input the data into the power distribution network simulation environment, and then perform power flow calculations to obtain the results. And the new trend distribution results.

5. The method for optimizing the operation of a distribution network containing a distributed electro-hydrogen coupling system according to claim 4, characterized in that, The DFPPO algorithm is based on the PPO algorithm to design a deep reinforcement learning DRL agent, including the following steps: In the PPO algorithm, the DRL agent uses policy gradient theory to optimize the objective. The policy gradient is estimated by the following form: ; In the formula, To calculate the gradient of target J, the strategy is... Indicates the state Take action below The probability, with parameter . This strategy is given through the Actor network; Indicates taking an action The expected cumulative return after the current state is introduced into the state value function. And define an advantage function to reflect the action. How good or bad is relative to the average level under this condition: ; In the formula, Represents the dominant function; The objective function of the PPO algorithm is: ; In the formula, New and old strategies The probability ratio, For the parameters of the Actor network, the shearing function By limiting the strategy ratio The range does not exceed , To truncate parameters and avoid excessive updates each time; The PPO algorithm estimates the value function using a critic network. The timing difference error is calculated according to the following formula: ; In the above formula, Let be the time-series difference error at time t. For instant rewards, As a discount factor, Let be the value function at time t+1; the goal of the Critic network is to minimize the loss function so that... A more accurate approximation of the actual value function The loss function is defined as: ; In the above formula, For loss function, Here, N is the number of empirical trajectories, and N is the parameter of the critic function. Let t be the actual value function at time t. Represents expectations, Let be the time-series difference error at time t; PPO algorithm actor and critic parameters and Update using the following formula: ; ; In the formula, The learning rate of the Actor network. is the learning rate of the Critic network.

6. The method for optimizing the operation of a power distribution network containing a distributed electro-hydrogen coupling system according to claim 5, characterized in that, The construction of the expert safety layer using DistFlow linearized equations includes the following steps: The topology of a power distribution network is defined by a matrix: ; ; ; In the formula, connection matrices F and T represent the starting and ending nodes of the line, M0 is the correlation matrix of the distribution network when calculating the slack node, and m0 is the column corresponding to the slack node. This is the correlation matrix of the distribution network without calculating the slack node. and It is a conditional function, meaning that if l is either i or j, then the value is 1; otherwise, it is 0. The power flow equations of a distribution network are expressed as follows: ; ; ; In the formula, and These are diagonal matrices, representing the line's resistance and reactance, respectively. P and Q are the active and reactive power matrices. A vector representing the squared voltage values ​​of all nodes except the slack node. The squared value of the voltage at the slack node is represented by the following dimensions: , This indicates the number of nodes in the distribution network excluding the balancing node; It represents the square of the line current and is used to calculate power loss in the distribution network. Neglecting this transforms the power flow equation into a linear form; therefore, after obtaining the active and reactive power, the equation is rewritten as: ; In the formula, , As the identity matrix, the linear expression in the above equation shows the power vector and voltage vector v. 2 The relationship between the PPO strategy and the action space 'a' contains active power commands and reactive power commands. Action space a and voltage v 2 The relationship between them is: ; In the above formula, and The components related to active and reactive power in the action are represented. Based on the above formula, an expert safety layer based on the DistFlow equation is constructed. The purpose of the expert safety layer is to minimize the Euclidean distance between the original unsafe action and the target safe action, and to project the unsafe action of the PPO agent into the safe area.

7. The method for optimizing the operation of a distribution network containing a distributed electro-hydrogen coupling system according to claim 6, characterized in that, The projection of unsafe actions of the PPO agent onto the safe area includes: The projection process is achieved by solving an optimization problem, the optimization objective of which is to find the safest action closest to the original action 'a'. Its mathematical form is: ; ; ; In the formula, relaxation parameters It is used to manage the limitations on voltage amplitude.

8. The method for optimizing the operation of a distribution network containing a distributed electro-hydrogen coupling system according to claim 1, characterized in that, The optimization process of the running optimization model based on the DFPPO algorithm includes the following steps: During the exploration phase, the agent interacts with the distribution network environment according to the initial strategy, aiming to collect the experience trajectory of the current round; each round selects continuous data of a set value for interaction, including several decision steps; the original action made by the agent in each decision step will be optimized by the expert safety layer, and actions that do not meet the voltage constraints will be projected to the safety domain as much as possible according to the real-time state of the distribution network. During the training phase, after each round of exploration, the agent trains the parameters of the Actor and Critic networks based on the collected historical experience trajectories; and updates the parameters using gradients. During the execution phase, the agent outputs the scheduling strategy of the distribution network and DPHCS through the trained Actor network, and feeds back the control effect through power flow calculation.

9. A power distribution network operation optimization system containing a distributed electro-hydrogen coupling system, characterized in that, Using the method of any one of claims 1 to 8, the system comprises: The operational modeling unit is used to construct operational models of the electrolyzer, hydrogen storage tank, compressor, and fuel cell device in the distributed electro-hydrogen coupling system (DPHCS). An optimization modeling unit is used to construct a distribution network operation optimization model considering DPHCS based on the operation model. The operation optimization model includes an objective function, constraints, and decision variables. The transformation unit is used to transform the optimization task of the running optimization model into a Markov decision MDP model, which includes a state space, action space, reward function, state transition probability distribution, and distribution network environment. The solving unit is used to solve the MDP model using the DFPPO algorithm, which is based on the DistFlow linearized equation and the PPO algorithm for near-end policy optimization. The DFPPO algorithm designs a deep reinforcement learning DRL agent based on the PPO algorithm and uses the DistFlow linearized equation to construct an expert safety layer. The online decision-making unit is used to design the computational optimization process of the running optimization model based on the DFPPO algorithm, train the parameters of the DRL agent, and save the trained parameters for online execution.

10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1-8.

11. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Distributed electricity-hydrogen coupling system coordination control method and device based on edge calculation and upper computer

    CN119401507A

  • Electricity-hydrogen coupling system risk scheduling method based on robust safety deep reinforcement learning

    CN119863051A