A fuel cell stack reliability assessment method based on strategy gradient and game

Through the multi-agent deep deterministic policy gradient method and Stackelberg game model, an intelligent agent system is constructed to optimize the operating parameters of fuel cell stacks and electricity market decisions, which solves the problem of insufficient dynamic response of fuel cell stack reliability assessment in existing technologies and realizes real-time monitoring of the health status of the stack and stability optimization of the power system.

CN119650772BActive Publication Date: 2025-09-26NORTH CHINA ELECTRIC POWER UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411695852.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-25
Publication Date
2025-09-26
Estimated Expiration
2044-11-25

AI Technical Summary

Technical Problem

Existing fuel cell stack reliability assessment methods cannot provide sufficient dynamic evaluation and timely response, resulting in low efficiency of stack health status monitoring and power optimization, and failing to fully cover the impact of the entire energy system.

Method used

Using the multi-agent deep deterministic policy gradient method and Stackelberg game model, three types of agents are constructed to monitor the health status of the fuel cell stack, adjust the electricity price and electricity purchase amount respectively, and optimize the operating parameters of the fuel cell stack and electricity market decision-making through policy gradient training and game model.

Benefits of technology

It improves the operating stability of the fuel cell stack and the reliability of the electric-hydrogen integrated energy system, optimizes the allocation of power resources, and ensures the continuity and reliability of power supply.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119650772B_ABST
    Figure CN119650772B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of fuel cell stack monitoring, and discloses a fuel cell stack reliability assessment method based on policy gradient and game theory, comprising: adopting a multi-agent deep deterministic policy gradient method to train a first agent, a second agent, and a third agent, and adopting a Stackelberg game model to divide the first agent, the second agent, and the third agent in training into leaders and followers, using the optimal strategy of the leader to adjust the optimal strategy of the follower to obtain the optimal first agent, the second agent, and the third agent, so as to realize real-time monitoring of the health status of the fuel cell stack, dynamically adjust the electricity price according to market demand, and establish a reliability assessment index system for evaluating the operating stability and power supply reliability of the fuel cell stack; the method improves economic benefits while ensuring the stability and reliability of the power supply.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fuel cell stack monitoring, and in particular to a fuel cell stack reliability assessment method based on strategy gradient and game theory. Background Art

[0002] As a key component of power supply, the reliability of fuel cell stacks plays a vital role in the stability and security of power systems. Existing assessment methods are mostly based on static models, which are inadequate for adapting to changing market and operating conditions. In particular, these models fail to provide adequate dynamic assessment and timely response to emergencies and system failures. The reliability of fuel cell stacks not only directly impacts their performance and lifespan but is also crucial for the continuity and reliability of power supply. Current stack health monitoring and optimization technologies primarily focus on local parameter adjustments and fail to fully account for the impact on the entire energy system. Summary of the Invention

[0003] In response to the above-mentioned deficiencies in the prior art, the present invention provides a fuel cell stack reliability assessment method based on strategy gradient and game theory, which is used to solve the problem in the prior art that the stack health status monitoring and power optimization efficiency are low due to the inability to provide sufficient dynamic assessment and timely response.

[0004] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:

[0005] A fuel cell stack reliability assessment method based on strategy gradient and game theory includes the following steps:

[0006] S1. Construct a first agent to evaluate the health status of the fuel cell stack by monitoring key operating parameters of the fuel cell stack, and establish the reward function, state space, and action space of the first agent;

[0007] S2. Build a second agent to dynamically adjust the electricity price based on the health status of the fuel cell stack and the power demand on the user side to maximize the electricity price profit, and establish the reward function, state space, and action space of the second agent;

[0008] S3. Build a third agent to make purchase decisions based on electricity prices and user-side electricity demand, obtain electricity purchase costs, minimize electricity costs, and establish the reward function, state space, and action space of the third agent.

[0009] S4. Use the multi-agent deep deterministic policy gradient method to train the first, second, and third agents to maximize the cumulative reward of each agent. Use the Stackelberg game model to divide the first, second, and third agents in training into leaders and followers. Use the leader's optimal strategy to adjust the follower's optimal strategy to obtain the optimal first, second, and third agents.

[0010] S5. Using the optimal first intelligent agent, adjust key operating parameters of the fuel cell stack to obtain an optimal health state of the fuel cell stack;

[0011] S6. Using the optimal second agent, adjust the electricity price to obtain the optimal electricity price profit;

[0012] S7. Using the optimal third agent, the optimal electricity cost is obtained by adjusting the electricity purchase amount on the user side;

[0013] S8. Based on the adjusted parameter information of the optimal first agent, the second agent, and the third agent, a reliability evaluation index system is established to evaluate the operational stability and power supply reliability of the fuel cell stack and the power system.

[0014] Furthermore, step S1 specifically includes:

[0015] S11. Construct a first intelligent agent to evaluate the health status of the fuel cell stack by monitoring the key operating parameters of the fuel cell stack, namely:

[0016]

[0017] Among them, H represents the health status coefficient of the fuel cell stack, P t represents the actual output power of the fuel cell stack at time t, η t represents the rated electrical efficiency of the fuel cell stack at the tth moment, c represents the attenuation ratio limit that the fuel cell stack can withstand, η0 represents the rated electrical efficiency of the fuel cell stack at the initial moment, e represents the exponential function, N start Indicates the number of startups of the fuel cell stack, T′ run represents the operating time of the fuel cell stack, and α, β, γ, and λ represent weight coefficients of key operating parameters of the fuel cell stack;

[0018] S12. The health status of the fuel cell stack is used as the reward function of the first agent, that is:

[0019] R1=H

[0020] Where R1 represents the reward function of the first agent;

[0021] S13. Establish the state space of the first agent, namely:

[0022] S1={P,η,N start ,T′ run ,T avg ,P′ max}

[0023] Where S1 represents the state space of the first agent, P represents the actual output power of the fuel cell stack, η represents the electrical efficiency of the fuel cell stack, T avg Represents the average temperature of the fuel cell stack during operation, P′ max Indicates the maximum pressure during the operation of the fuel cell stack;

[0024] S14. Establish the action space of the first agent, namely:

[0025]

[0026] Where A1 represents the action space of the first agent, T represents the temperature of the fuel cell stack, represents the hydrogen partial pressure, represents the oxygen partial pressure, R′ hum Relative humidity, Indicates the excess oxygen ratio.

[0027] Furthermore, step S2 specifically includes:

[0028] S21. Construct a second intelligent agent to dynamically adjust the electricity price based on the health status of the fuel cell stack and the power demand on the user side to obtain the goal of maximizing the electricity price profit, namely:

[0029]

[0030] Among them, O represents the profit maximization target of electricity price, P″ represents the electricity price, Q represents the power demand on the user side, C represents the total cost of the power station, and P″ base represents the basic electricity price, r represents the price adjustment rate, H threshold represents the preset health indicator threshold, D represents the basic electricity demand, σ represents the price elasticity coefficient, and τ represents the influence coefficient of the health status of the fuel cell stack on the electricity demand at the user side;

[0031] S22. The profit maximization goal of the second agent is used as the reward function of the second agent, that is:

[0032] R2=O

[0033] Where R2 represents the reward function of the second agent;

[0034] S23. Establish the state space of the second agent, namely:

[0035] S2={P″ base ,r,H,H threshold ,D,σ,τ}

[0036] Where S2 represents the state space of the second agent;

[0037] S24. Establish the action space of the second agent, namely:

[0038] A2={P″}

[0039] Among them, A2 represents the action space of the second agent.

[0040] Furthermore, step S3 specifically includes:

[0041] S31. Construct a third intelligent agent to make a purchase decision based on the electricity price and the electricity demand on the user side, and obtain the electricity purchase cost, that is:

[0042] C electricity =P″×Q

[0043] Among them, C electricity represents the cost of electricity purchase;

[0044] S32. The negative value of the electricity purchase cost is used as the reward function of the third agent, that is:

[0045] R3=-C electricity

[0046] Where R3 represents the reward function of the third agent;

[0047] S33. Establish the state space of the third agent, namely:

[0048] S3={P″,D,σ}

[0049] Among them, S3 represents the state space of the third agent;

[0050] S34. Establish the action space of the third agent, namely:

[0051] A3={K}

[0052] Among them, A3 represents the action space of the third agent, and K represents the electricity purchase amount on the user side.

[0053] Furthermore, step S4 specifically includes:

[0054] S41, establishing a strategy network for the first agent, the second agent, and the third agent respectively; wherein the strategy network includes an action network and a value network;

[0055] S42, inputting the reward function, state space, and action space corresponding to the first agent, the second agent, and the third agent into their respective policy networks for training, and maximizing the cumulative reward of each agent by obtaining the policy gradient of each agent;

[0056] S43. Using the Stackelberg game model, the first, second, and third agents in training are divided into leaders and followers, and the policy gradient of each agent is updated.

[0057] S44. Repeat steps S42-S43 until the set maximum number of iterations is reached, and the optimal first agent, second agent, and third agent are obtained.

[0058] Furthermore, the policy gradient of each agent in step S42 is:

[0059]

[0060]

[0061] in, represents the gradient, θ i represents the policy network parameters of agent i, J represents the cumulative reward of agent i, θ1, θ2, θ3 represent the policy network parameters of the first agent, the second agent, and the third agent respectively, represents the mean calculation, Q′ represents the value evaluation value of the policy network, s represents the state space of agent i, a1, a2, a3 represent the action space of the first agent, the second agent, and the third agent respectively, and R represents the reward of each agent. Represents the weight coefficient of the policy network, Q(S - ,a - ) represents the state space S of each agent at the previous moment - and action space a - .

[0062] Furthermore, the cumulative reward of each agent maximized in step S42 is the total reward obtained by each agent from the multi-agent environment when performing strategy optimization to maximize the action decision of each agent.

[0063] Furthermore, in step S43, the specific process of using the Stackelberg game model to divide the first agent, the second agent, and the third agent in training into leaders and followers is as follows:

[0064] The first intelligent agent is regarded as the leader, and the second and third intelligent agents are regarded as the first and second followers respectively; first, the leader makes a decision, and then the first and second followers optimize their own strategies according to the decision made by the leader.

[0065] Furthermore, the update formula of the policy gradient of each agent in step S43 is:

[0066]

[0067] Among them, θ′ i represents the updated policy network parameters of agent i, and α represents the learning rate of the policy network.

[0068] Furthermore, step S8 specifically includes:

[0069] S51, obtaining the optimal key operating parameters of the fuel cell stack after adjustment by the optimal first intelligent agent, and establishing an expected energy shortage indicator, namely:

[0070]

[0071] Among them, EENS stands for expected energy non-supply index;

[0072] S52. Obtain the optimal key operating parameters of the fuel cell stack after adjustment by the optimal first intelligent agent, and establish a system average interruption duration indicator, namely:

[0073]

[0074] Among them, SAIDI represents the system average interruption duration index;

[0075] S53. Establish the system average interruption frequency index indicator, namely:

[0076]

[0077] Among them, SAIFI represents the system average interruption frequency index;

[0078] S54. Evaluate the operational stability and power supply reliability of the fuel cell stack and the power system based on the expected energy non-supply index, the system average interruption duration index, and the system average interruption frequency index;

[0079] S55. Based on the optimal health status of the fuel cell stack, evaluate the operational stability of the fuel cell stack;

[0080] S56. Based on the optimal second agent, the optimal electricity price after adjustment by the third agent, and the predicted electricity purchase amount, the effectiveness of user demand side management is evaluated by assessing the impact of electricity price adjustment on electricity purchase amount.

[0081] The present invention has the following beneficial effects:

[0082] 1. This paper proposes a fuel cell stack reliability assessment method based on policy gradient and game theory. This method uses a multi-agent deep deterministic policy gradient method combined with a Stackelberg game model to train a first agent, a second agent, and a third agent. By obtaining the optimal first, second, and third agents, the method optimizes the strategy and adjusts it to enhance the operational stability of high-power fuel cell stacks and the reliability of the electric-hydrogen integrated energy system.

[0083] 2. Through detailed modeling of the fuel cell stack and its interaction with the power market, in-depth analysis of the impact of possible failure modes on performance, and comprehensive consideration of the operation of the fuel cell stack under different states, key indicators such as expected energy shortage, average system interruption duration and average system interruption frequency index are calculated to provide the power market with an accurate measure of the fuel cell stack reliability, thereby optimizing the allocation of power resources and ensuring the continuity of supply. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] Figure 1 This is a flow chart of a fuel cell stack reliability assessment method based on strategy gradient and game theory proposed in the present invention;

[0085] Figure 2 Schematic diagram of the interaction principle of three intelligent agents. DETAILED DESCRIPTION

[0086] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.

[0087] like Figure 1 As shown, a fuel cell stack reliability assessment method based on strategy gradient and game theory includes the following steps S1-S8:

[0088] S1. Construct a first intelligent agent to evaluate the health status of the fuel cell stack by monitoring the key operating parameters of the fuel cell stack, and establish the reward function, state space, and action space of the first intelligent agent.

[0089] In this embodiment, the first intelligent agent is constructed to be mainly responsible for monitoring the health status of the fuel cell stack, that is, through subsequent training, it will learn how to adjust the operating parameters according to the status of the fuel cell stack to maximize the health status of the fuel cell stack.

[0090] Specifically, step S1 includes S11-S14:

[0091] S11. Construct a first intelligent agent to evaluate the health status of the fuel cell stack by monitoring the key operating parameters of the fuel cell stack, namely:

[0092]

[0093] Among them, H represents the health status coefficient of the fuel cell stack, P t represents the actual output power of the fuel cell stack at time t, η t represents the rated electrical efficiency of the fuel cell stack at the tth moment, c represents the attenuation ratio limit that the fuel cell stack can withstand, η0 represents the rated electrical efficiency of the fuel cell stack at the initial moment, e represents the exponential function, N start Indicates the number of startups of the fuel cell stack, T′ run represents the operating time of the fuel cell stack, and α, β, γ, and λ represent the weight coefficients of the key operating parameters of the fuel cell stack.

[0094] S12. The health status of the fuel cell stack is used as the reward function of the first agent, that is:

[0095] R1=H

[0096] Where R1 represents the reward function of the first agent.

[0097] In this embodiment, the first intelligent agent is mainly responsible for monitoring and optimizing the health status of the high-power fuel cell stack; its core task is to evaluate and improve the health status coefficient of the fuel cell stack through real-time monitoring of the key operating parameters of the fuel cell stack, and directly set the health coefficient as the reward function of the first intelligent agent; this design ensures that every decision made by the first intelligent agent is aimed at maximizing the health status of the fuel cell stack, thereby indirectly improving the operating efficiency and life of the fuel cell stack.

[0098] S13. Establish the state space of the first agent, namely:

[0099] S1={P,η,N start ,T′ run ,T avg ,P′ max}

[0100] Where S1 represents the state space of the first agent, P represents the actual output power of the fuel cell stack, η represents the electrical efficiency of the fuel cell stack, R avg Represents the average temperature of the fuel cell stack during operation, P′ max Indicates the maximum pressure during fuel cell stack operation.

[0101] S14. Establish the action space of the first agent, namely:

[0102]

[0103] Among them, A1 represents the action space of the first agent, R represents the temperature of the fuel cell stack, represents the hydrogen partial pressure, represents the oxygen partial pressure, R′ hum Relative humidity, Indicates the excess oxygen ratio.

[0104] In this embodiment, the action space includes all possible actions that the first agent can take to adjust key operating parameters of the fuel cell stack to optimize its health. For example, the first agent might adjust the hydrogen partial pressure, oxygen partial pressure, operating temperature, etc. to affect the performance and lifespan of the fuel cell stack. In this way, the first agent can not only improve the performance and reliability of the fuel cell stack, but also play a coordinating role in the entire system, working with other agents to achieve the overall optimization goal of the system.

[0105] S2. Construct a second intelligent agent to dynamically adjust the electricity price according to the health status of the fuel cell stack and the electricity demand on the user side to maximize the electricity price profit, and establish the reward function, state space and action space of the second intelligent agent.

[0106] In this embodiment, the second agent dynamically adjusts electricity prices based on the health of the fuel cell stack and market supply and demand to maximize profits. Its core task is to formulate and adjust electricity pricing strategies in response to market changes and the health of the fuel cell stack to maximize power plant profits. Its goal is to balance supply and demand by optimizing electricity prices, improving the economic benefits of the power plant, and ensuring price stability and user acceptance.

[0107] Specifically, step S2 includes S21-S24:

[0108] S21. Construct a second intelligent agent to dynamically adjust the electricity price based on the health status of the fuel cell stack and the power demand on the user side to obtain the goal of maximizing the electricity price profit, namely:

[0109]

[0110] Among them, O represents the profit maximization target of electricity price, P″ represents the electricity price, Q represents the power demand on the user side, C represents the total cost of the power station, and P″ base represents the basic electricity price, r represents the price adjustment rate, H thresholdrepresents the preset health indicator threshold, D represents the basic electricity demand, σ represents the price elasticity coefficient, and τ represents the influence coefficient of the health status of the fuel cell stack on the electricity demand on the user side.

[0111] S22. The profit maximization goal of the second agent is used as the reward function of the second agent, that is:

[0112] R2=O

[0113] Where R2 represents the reward function of the second agent.

[0114] S23. Establish the state space of the second agent, namely:

[0115] S2={P″ base ,r,H,H threshold ,D,σ,τ}

[0116] Where S2 represents the state space of the second agent.

[0117] S24. Establish the action space of the second agent, namely:

[0118] A2={P″}

[0119] Among them, A2 represents the action space of the second agent.

[0120] In this example, the action space primarily involves adjusting electricity prices in response to market fluctuations and the health of the fuel cell stack. This allows the second agent to dynamically adjust electricity prices to maximize the plant's profits, taking into account the health of the fuel cell stack and market demand. This not only improves the plant's economic efficiency but also enhances its adaptability to market fluctuations.

[0121] S3. Build a third agent to make purchasing decisions based on electricity prices and user-side electricity demand, obtain electricity purchase costs, minimize electricity costs, and establish the reward function, state space, and action space of the third agent.

[0122] In this embodiment, the third intelligent agent is mainly used to simulate the market response on the user side and make purchasing decisions based on electricity prices and electricity demand; its core task is to simulate the market response on the user side and make purchasing decisions based on the electricity price set by the second intelligent agent and its own electricity demand; its goal is to reflect the market sensitivity to changes in electricity prices and the electricity demand on the user side, thereby providing market feedback to the second intelligent agent and helping it adjust its electricity price strategy.

[0123] Specifically, step S3 includes S31-S34:

[0124] S31. Construct a third intelligent agent to make a purchase decision based on the electricity price and the electricity demand on the user side, and obtain the electricity purchase cost, that is:

[0125] C electricity =P″″×Q

[0126] Among them, C electricity Indicates the cost of purchasing electricity.

[0127] S32. The negative value of the electricity purchase cost is used as the reward function of the third agent, that is:

[0128] R3=-C electricity

[0129] Where R3 represents the reward function of the third agent.

[0130] S33. Establish the state space of the third agent, namely:

[0131] S3={P″,D,σ}

[0132] Among them, S3 represents the state space of the third agent.

[0133] In this embodiment, the action space is mainly to adjust the amount of electricity purchased in response to changes in electricity prices and its own needs; in this way, the third agent can quantify its decision-making process and minimize its electricity purchase costs by adjusting the amount of electricity purchased, while taking into account the market's sensitivity to changes in electricity prices.

[0134] S34. Establish the action space of the third agent, namely:

[0135] A3={K}

[0136] Among them, A3 represents the action space of the third agent, and K represents the electricity purchase amount on the user side.

[0137] S4. Use the multi-agent deep deterministic policy gradient method to train the first, second, and third agents to maximize the cumulative reward of each agent, and use the Stackelberg game model to divide the first, second, and third agents in training into leaders and followers. Use the leader's optimal strategy to adjust the follower's optimal strategy to obtain the optimal first, second, and third agents.

[0138] In this embodiment, a multi-agent deep deterministic policy gradient method is adopted, and the three agents can learn and adapt to environmental changes and optimize their own strategies; the Stackelberg game model is used to coordinate the decision-making between the agents and realize the strategic interaction between the leader (fuel cell operator) and the follower (electricity market user); this method can not only improve the operating efficiency and market competitiveness of the fuel cell, but also realize the real-time monitoring and optimization of the health status of the fuel cell in the complex background of multi-energy coupling, as well as the dynamic pricing strategy of the electricity market, and promote the energy system to develop in a more efficient and reliable direction.

[0139] Specifically, step S4 includes S41-S44:

[0140] S41. Establish a strategy network for the first agent, the second agent, and the third agent respectively; wherein the strategy network includes an action network and a value network.

[0141] The core of this implementation is that each agent has its own action network and value network. The action network is responsible for outputting actions, while the value network is responsible for evaluating the value of actions. During training, the value network has access to information from all agents. However, during execution, each agent's action network can only make decisions based on its own local observations. Therefore, the goal of the multi-agent deep deterministic policy gradient method is to maximize the cumulative reward of each agent.

[0142] In addition, the value network of the first agent can access the information of the second and third agents to evaluate the impact of their actions on the entire system; the value networks of the second and third agents can also access the information of other agents; in this way, each agent can learn and adopt the best strategy to achieve its own goals in a multi-agent environment while taking into account the behavior of other agents, such as Figure 2 As shown in the figure, the first agent optimizes the performance of the battery stack by adjusting operating parameters. Changes in these parameters directly affect the health status of the battery stack, which in turn affects the electricity price decision of the second agent; the second agent adjusts the electricity price according to the health status of the battery stack and market demand, which in turn affects the electricity purchasing decision of the third agent; the market response of the third agent is fed back to the second agent to help it further optimize the electricity price strategy; this interaction and feedback mechanism enables the entire system to continuously learn and adapt in a dynamic environment to achieve overall performance optimization; that is, through the multi-agent deep deterministic policy gradient method, the agents can handle complex multi-agent interaction problems, thereby achieving collaborative optimization.

[0143] S42. Input the reward functions, state spaces, and action spaces corresponding to the first agent, the second agent, and the third agent into their respective policy networks for training, and maximize the cumulative reward of each agent by obtaining the policy gradient of each agent.

[0144] Specifically, the policy gradient of each agent in step S42 is:

[0145]

[0146]

[0147] in, represents the gradient, θ i represents the policy network parameters of agent i, J represents the cumulative reward of agent i, θ1, θ2, θ3 represent the policy network parameters of the first agent, the second agent, and the third agent respectively, represents the mean calculation, Q′ represents the value evaluation value of the policy network, s represents the state space of agent i, a1, a2, a3 represent the action space of the first agent, the second agent, and the third agent respectively, and R represents the reward of each agent. Represents the weight coefficient of the policy network, Q(S - ,a - ) represents the state space S of each agent at the previous moment - and action space a - .

[0148] Specifically, the cumulative reward of each agent maximized in step S42 is the total reward obtained by each agent from the multi-agent environment when performing strategy optimization to maximize the action decision of each agent.

[0149] S43. Use the Stackelberg game model to divide the first, second, and third agents in training into leaders and followers, and update the policy gradient of each agent.

[0150] Specifically, in step S43, the specific process of using the Stackelberg game model to divide the first agent, the second agent, and the third agent in training into leaders and followers is as follows:

[0151] The first intelligent agent is regarded as the leader, and the second and third intelligent agents are regarded as the first and second followers respectively; first, the leader makes a decision, and then the first and second followers optimize their own strategies according to the decision made by the leader.

[0152] Specifically, the update formula of the policy gradient of each agent in step S43 is:

[0153]

[0154] Among them, θ′ i represents the updated policy network parameters of agent i, and α represents the learning rate of the policy network.

[0155] S44. Repeat steps S42-S43 until the set maximum number of iterations is reached, and the optimal first agent, second agent, and third agent are obtained.

[0156] In this embodiment, through iterative learning in the above steps, the first agent, the second agent, and the third agent gradually find the optimal strategy, thereby achieving Stackelberg equilibrium. The Stackelberg game is a hierarchical decision-making model involving a leader and one or more followers. The leader makes the decision first, and the followers then optimize their strategies based on the leader's decision. This game model is often used to describe strategic interactions between participants with different decision-making powers and information levels. The leader's goal is to maximize its own utility function, while the followers choose their own optimal strategies given the leader's decision. At the same time, the leader takes the followers' reactions into consideration when making decisions. The specific process is as follows: the objective function of the first agent is to maximize the health status coefficient of the fuel cell stack, and its action space includes adjusting the stack operating parameters. The objective function of the second agent is to maximize profit, and its action space includes setting the electricity price. The objective function of the third agent is to minimize electricity cost, and its action space includes adjusting the amount of electricity purchased. In the Stackelberg game established in this embodiment, the first agent makes the decision first, and then the second and third agents optimize their strategies based on the decision of the first agent. The second agent dynamically adjusts the electricity price based on the health status of the fuel cell stack and market demand, while the third agent determines the amount of electricity purchased based on the electricity price and its own needs. At the same time, the multi-agent deep deterministic policy gradient method plays a key role in this process. It allows each agent to learn how to optimize its own strategy in a multi-agent environment. Specifically, the first agent learns to predict the responses of the second and third agents and adjusts its operating parameters accordingly. The second and third agents also learn to optimize their own strategies to respond to the first agent's actions and the strategies of the other agents. Through this interactive and learning process, the first, second, and third agents gradually reach a Stackelberg equilibrium, where each agent's strategy is the optimal response to the strategies of the others. This equilibrium helps the entire system achieve optimized fuel cell stack performance and improved market efficiency.

[0157] In addition, during the entire training process described above, the update of the value network (leader and follower) is: Q new =r1+γ1V(s′,θ L ,θ F ), loss=(Q current -Q new )2 , where Q new represents the updated value network, r1 represents the reward, γ1 represents the discount factor, V represents the value function, s ′ represents the updated state, θ L represents the leader’s strategy parameter, θ F represents the strategy parameter of the follower, loss represents the loss function of the value network, Q current represents the policy network before the update; the update of the action network (leader and follower) is: Among them, Q″ represents the value assessment, a L 、a F Represent the actions of the leader and the follower respectively, and s represents the state; finally, the follower updates the strategy by maximizing the expected return when the leader's strategy is given.

[0158] S5. Using the optimal first intelligent agent, adjust the key operating parameters of the fuel cell stack to obtain the optimal health state of the fuel cell stack.

[0159] S6. Use the optimal second intelligent agent to adjust the electricity price to obtain the optimal electricity price profit.

[0160] S7. Utilize the optimal third agent to adjust the amount of electricity purchased on the user side to obtain the optimal electricity cost.

[0161] S8. Based on the adjusted parameter information of the optimal first agent, the second agent, and the third agent, a reliability evaluation index system is established to evaluate the operational stability and power supply reliability of the fuel cell stack and the power system.

[0162] Specifically, step S8 includes S81-S86:

[0163] S81. Obtain the optimal key operating parameters of the fuel cell stack after adjustment by the optimal first intelligent agent, and establish an expected energy shortage indicator, namely:

[0164]

[0165] Among them, EENS stands for expected energy non-supply indicator.

[0166] In this embodiment, based on the key operating parameters provided by the optimal first agent, the ratio of energy that cannot be supplied within a specific time period to the total energy demand is calculated, i.e., the expected energy non-supply index is established. The planned energy supply is the total energy that the fuel cell stack should provide to the load according to the scheduling plan within a specific time period; the actual energy supply is the energy actually provided to the load by the fuel cell stack within a specific time period. Therefore, by calculating the expected energy non-supply index, it is possible to determine whether the fuel cell stack is operating stably. Specifically, if the EENS value remains at a low and stable level, it indicates that the fuel cell stack is able to provide energy as planned and is operating stably. If the EENS value increases abnormally or fluctuates significantly, it indicates that the fuel cell stack has performance issues or is undermaintained, resulting in unstable energy supply.

[0167] S82. Obtain the optimal key operating parameters of the fuel cell stack after adjustment by the optimal first agent, and establish a system average interruption duration indicator, namely:

[0168]

[0169] Among them, SAIDI represents the system average interruption duration index.

[0170] In this embodiment, the health status assessment results of the optimal first intelligent agent are used to record the duration of each power outage and calculate the average value, which is the system average outage duration index. The power outage duration is the length of time from the start of each power outage to the restoration of power to the power stack; the number of affected users is the number of independent loads or users affected by each power outage. Therefore, by calculating the system average outage duration index, it is possible to determine whether the power supply of the power system is reliable. Specifically, if the SAIDI value is low, that is, the outage duration is short and infrequent, it means that the power outage has little impact on users and the power supply reliability is high; if the SAIDI value is high, that is, the outage duration is long or frequent, it means that the power outage has a large impact on users and the power supply reliability is low.

[0171] S83. Establish the system average interruption frequency index indicator, namely:

[0172]

[0173] Among them, SAIFI represents the system average interruption frequency index.

[0174] In this embodiment, based on the data of the optimal first agent, the number of power outages within a specific time period is counted, and the average number of outages experienced by each user is calculated, i.e., the system average outage frequency index is established. The total number of power outages is the total number of power outage events for all fuel cell stacks that occurred within a specific time period, and the total number of users is the total number of independent loads or users that may be affected by the outage within a specific time period. Therefore, by calculating the system average outage frequency index, it is possible to determine whether the power supply system is stable. Specifically, if the SAIFI value is low, i.e., the number of outages is low, it indicates that the fuel cell stack is operating smoothly with few outages. If the SAIFI value is high, i.e., the number of outages is high, it indicates that the fuel cell stack may frequently fail or experience performance degradation, resulting in unstable operation.

[0175] S84. Evaluate the operational stability and power supply reliability of the fuel cell stack and the power system based on the expected energy non-supply index, the system average interruption duration index, and the system average interruption frequency index.

[0176] To sum up, when EENS, SAIDI and SAIFI are low, it indicates that the fuel cell stack operates stably and has high power supply reliability, that is, it has high stability and high reliability; when EENS, SAIDI and SAIFI are high, it indicates that the fuel cell stack operates unstable and has low power supply reliability, that is, it has low stability and low reliability.

[0177] S85. Evaluate the operational stability of the fuel cell stack based on the optimal health status of the fuel cell stack.

[0178] S86. Based on the optimal second agent, the optimal electricity price after adjustment by the third agent, and the predicted electricity purchase amount, the effectiveness of user demand-side management is evaluated by assessing the impact of electricity price adjustment on electricity purchase amount.

[0179] In summary, the fuel cell stack reliability assessment method based on strategy gradient and game theory proposed in the present invention aims to enhance the operational stability of high-power fuel cell stacks and the reliability of the electric-hydrogen integrated energy system. The system conducts in-depth analysis of the impact of possible failure modes on performance through fine modeling of the stack and its interaction with the electricity market, and comprehensively considers the operation of the stack under different states. By calculating key indicators such as expected energy non-supplied (EENS), system average interruption duration (SAIDI) and system average interruption frequency index (SAIFI), the present invention can provide the electricity market with an accurate measure of the reliability of the stack, thereby optimizing the allocation of power resources and ensuring the continuity of supply.

[0180] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.

[0181] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.

Claims

1. A fuel cell stack reliability assessment method based on strategy gradient and game theory, characterized in that: The following steps are involved: S1. Construct a first agent to evaluate the health status of the fuel cell stack by monitoring key operating parameters of the fuel cell stack, and establish the reward function, state space, and action space of the first agent; S2. Build a second agent to dynamically adjust the electricity price based on the health status of the fuel cell stack and the power demand on the user side to maximize the electricity price profit, and establish the reward function, state space, and action space of the second agent; S3. Build a third agent to make purchase decisions based on electricity prices and user-side electricity demand, obtain electricity purchase costs, minimize electricity costs, and establish the reward function, state space, and action space of the third agent. S4. Use the multi-agent deep deterministic policy gradient method to train the first, second, and third agents to maximize the cumulative reward of each agent. Use the Stackelberg game model to divide the first, second, and third agents in training into leaders and followers. Use the leader's optimal strategy to adjust the follower's optimal strategy to obtain the optimal first, second, and third agents. S5. Using the optimal first intelligent agent, adjust key operating parameters of the fuel cell stack to obtain an optimal health state of the fuel cell stack; S6. Using the optimal second agent, adjust the electricity price to obtain the optimal electricity price profit; S7. Using the optimal third agent, the optimal electricity cost is obtained by adjusting the electricity purchase amount on the user side; S8. Based on the adjusted parameter information of the optimal first agent, the second agent, and the third agent, a reliability evaluation index system is established to evaluate the operational stability and power supply reliability of the fuel cell stack and the power system.

2. The fuel cell stack reliability assessment method based on strategy gradient and game theory according to claim 1, characterized in that: Step S1 specifically includes: S11. Construct a first intelligent agent to evaluate the health status of the fuel cell stack by monitoring the key operating parameters of the fuel cell stack, namely: Among them, H represents the health status coefficient of the fuel cell stack, P t represents the actual output power of the fuel cell stack at time t, η t represents the rated electrical efficiency of the fuel cell stack at the tth moment, c represents the attenuation ratio limit that the fuel cell stack can withstand, η0 represents the rated electrical efficiency of the fuel cell stack at the initial moment, e represents the exponential function, N start Indicates the number of times the fuel cell stack is started, T r ′ un represents the operating time of the fuel cell stack, and α, β, γ, and λ represent weight coefficients of key operating parameters of the fuel cell stack; S12. The health status of the fuel cell stack is used as the reward function of the first agent, that is: R1=H Where R1 represents the reward function of the first agent; S13. Establish the state space of the first agent, namely: S1={P,η,N start ,T r r un ,T avg ,P m ′ ax } Where s1 represents the state space of the first agent, p represents the actual output power of the fuel cell stack, η represents the electrical efficiency of the fuel cell stack, T avg Indicates the average temperature of the fuel cell stack during operation, P m ′ ax Indicates the maximum pressure during the operation of the fuel cell stack; S14. Establish the action space of the first agent, namely: Where A1 represents the action space of the first agent, T represents the temperature of the fuel cell stack, represents the hydrogen partial pressure, represents the oxygen partial pressure, R′ hum Relative humidity, Indicates the excess oxygen ratio.

3. The fuel cell stack reliability assessment method based on strategy gradient and game theory according to claim 2, characterized in that: Step S2 specifically includes: S21. Construct a second intelligent agent to dynamically adjust the electricity price based on the health status of the fuel cell stack and the power demand on the user side to obtain the goal of maximizing the electricity price profit, namely: Among them, O represents the profit maximization target of electricity price, P″ represents the electricity price, Q represents the power demand on the user side, C represents the total cost of the power station, and P″ base represents the basic electricity price, r represents the price adjustment rate, H threshold represents the preset health indicator threshold, D represents the basic electricity demand, σ represents the price elasticity coefficient, and τ represents the influence coefficient of the health status of the fuel cell stack on the electricity demand at the user side; S22. The profit maximization goal of the second agent is used as the reward function of the second agent, that is: R2=O Where R2 represents the reward function of the second agent; S23. Establish the state space of the second agent, namely: S2={P″ base ,r,H,H threshold ,D,σ,τ} Where S2 represents the state space of the second agent; S24. Establish the action space of the second agent, namely: A2={P″} Among them, A2 represents the action space of the second agent.

4. The fuel cell stack reliability assessment method based on strategy gradient and game theory according to claim 3 is characterized in that: Step S3 specifically includes: S31. Construct a third intelligent agent to make a purchase decision based on the electricity price and the electricity demand on the user side, and obtain the electricity purchase cost, that is: C electricity =P″×Q Among them, C electricity represents the cost of electricity purchase; S32. The negative value of the electricity purchase cost is used as the reward function of the third agent, that is: R3=-C electricity Where R3 represents the reward function of the third agent; S33. Establish the state space of the third agent, namely: S3={P″,D,σ} Among them, S3 represents the state space of the third agent; S34. Establish the action space of the third agent, namely: A3={K} Among them, A3 represents the action space of the third agent, and K represents the electricity purchase amount on the user side.

5. The fuel cell stack reliability assessment method based on strategy gradient and game theory according to claim 4, characterized in that: Step S4 specifically includes: S41, establishing a strategy network for the first agent, the second agent, and the third agent respectively; wherein the strategy network includes an action network and a value network; S42, inputting the reward function, state space, and action space corresponding to the first agent, the second agent, and the third agent into their respective policy networks for training, and maximizing the cumulative reward of each agent by obtaining the policy gradient of each agent; S43. Using the Stackelberg game model, the first, second, and third agents in training are divided into leaders and followers, and the policy gradient of each agent is updated. S44. Repeat steps S42-S43 until the set maximum number of iterations is reached, and the optimal first agent, second agent, and third agent are obtained.

6. The fuel cell stack reliability assessment method based on strategy gradient and game theory according to claim 5, characterized in that: The policy gradient of each agent in step S42 is: in, represents the gradient, θ i represents the policy network parameters of agent i, J represents the cumulative reward of agent i, θ1, θ2, θ3 represent the policy network parameters of the first agent, the second agent, and the third agent respectively, represents the mean calculation, Q ′ represents the value evaluation value of the policy network, s represents the state space of agent i, a1, a2, a3 represent the action space of the first agent, the second agent, and the third agent respectively, and R represents the reward of each agent. Represents the weight coefficient of the policy network, Q(S - ,a - ) represents the state space S of each agent at the previous moment - and action space a _ .

7. The fuel cell stack reliability assessment method based on strategy gradient and game theory according to claim 6, characterized in that: In step S42, the cumulative reward of each agent is maximized, which is the total reward obtained by each agent from the multi-agent environment when performing strategy optimization to maximize the action decision of each agent.

8. The fuel cell stack reliability assessment method based on strategy gradient and game theory according to claim 7, characterized in that: The specific process of using the Stackelberg game model in step S43 to divide the first agent, the second agent, and the third agent in training into leaders and followers is as follows: The first intelligent agent is regarded as the leader, and the second and third intelligent agents are regarded as the first and second followers respectively; first, the leader makes a decision, and then the first and second followers optimize their own strategies according to the decision made by the leader.

9. The fuel cell stack reliability assessment method based on strategy gradient and game theory according to claim 8, characterized in that: The update formula of the policy gradient of each agent in step S43 is: Among them, θ i ′ represents the updated policy network parameters of agent i, and α represents the learning rate of the policy network.

10. The fuel cell stack reliability assessment method based on strategy gradient and game theory according to claim 9, characterized in that: Step S8 specifically includes: S51, obtaining the optimal key operating parameters of the fuel cell stack after adjustment by the optimal first intelligent agent, and establishing an expected energy shortage indicator, namely: Among them, EENS stands for expected energy non-supply index; S52. Obtain the optimal key operating parameters of the fuel cell stack after adjustment by the optimal first intelligent agent, and establish a system average interruption duration indicator, namely: Among them, SAIDI represents the system average interruption duration index; S53. Establish the system average interruption frequency index indicator, namely: Among them, SAIFI represents the system average interruption frequency index; S54. Evaluate the operational stability and power supply reliability of the fuel cell stack and the power system based on the expected energy non-supply index, the system average interruption duration index, and the system average interruption frequency index; S55. Based on the optimal health status of the fuel cell stack, evaluate the operational stability of the fuel cell stack; S56. Based on the optimal second agent, the optimal electricity price after adjustment by the third agent, and the predicted electricity purchase amount, the effectiveness of user demand side management is evaluated by assessing the impact of electricity price adjustment on electricity purchase amount.

Citation Information

Patent Citations

  • Self-adaptive optimal control method and device for linear system

    CN112149361A

  • MADDPG-based selling double-side decision optimization and operation method and device

    CN117391241A