Layered energy management method for multi-stack fuel cell hybrid power system

Through the decoupled feature clustering neural network and nested deep reinforcement learning method, the energy management and coordinated control problems of multi-stack fuel cell systems are solved, efficient and stable energy management and system optimization are achieved, the robustness and adaptability of the system are improved, and the automatic identification and optimized scheduling of equipment health status are supported.

CN120674668APending Publication Date: 2025-09-19SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510773550.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Multi-stack fuel cell systems face difficulties in energy management and collaborative control strategy design. Traditional methods are difficult to optimize the operating status of each stack and the overall efficiency of the system, and deep reinforcement learning methods have problems such as slow convergence and poor strategy stability.

Method used

By adopting a decoupled feature clustering neural network and a nested deep reinforcement learning method, and by constructing a unified feature vector set and a nested deep reinforcement learning controller, the regional division and optimal energy management of fuel cells, lithium batteries and supercapacitors are achieved. System-level coordinated scheduling is carried out through a central power coordination module, and the robustness and generalization ability are improved by combining an online update mechanism.

Benefits of technology

It achieves efficient operation and life protection of multi-stack fuel cell systems in complex energy environments, improves the global performance and long-term stability of the system, has strong generalization capabilities and adaptability, and supports automatic identification and optimized scheduling of equipment health status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120674668A_ABST
    Figure CN120674668A_ABST
Patent Text Reader

Abstract

The invention discloses a hierarchical energy management method for a multi-stack fuel cell hybrid power system. The method comprises the following steps: acquiring the recession state and real-time maximum efficiency of each fuel cell, each lithium battery and each super capacitor of the system; dividing a fuel cell, a lithium battery and a super capacitor with similar decline degrees and maximum efficiency into a region by using a decoupling feature clustering neural network, so as to divide the whole system into a plurality of regions; a nested deep reinforcement learning controller is constructed according to the instantaneous efficiency of the fuel cell, the equivalent hydrogen consumption of the system, the instantaneous decline rate of each power unit, the SOC of the lithium battery and the super capacitor and the operating temperature of the fuel cell, and optimal energy management of each region is realized; and performing system-level coordination scheduling on energy management results of all regions, and realizing cross-region power balance and boundary condition constraint through a central power coordination module. According to the invention, efficient operation and service life protection of a multi-pile hybrid power system in high-power and high-reliability application scenes such as rail transit vehicles are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of energy management of a multi-stack fuel cell hybrid power system, and in particular relates to a hierarchical energy management method for a multi-stack fuel cell hybrid power system. Background Art

[0002] With the large-scale integration of renewable energy and the gradual development of fuel cell technology, fuel cell hybrid systems are increasingly being used in high-energy, high-reliability scenarios such as rail transit, becoming a key driver of green transportation. Compared to traditional single-stack fuel cell systems, multi-stack fuel cell hybrid systems offer significant advantages in power redundancy, system reliability, and power level regulation, making them more suitable for high-power applications such as rail locomotives. However, they also present challenges such as energy management and the difficulty in designing coordinated control strategies.

[0003] In a multi-stack fuel cell system, individual fuel cell stacks exhibit differences in manufacturing tolerances, operational aging, and cooling system configurations, leading to significant variations in their electrical-thermal coupling characteristics, performance degradation behavior, and efficiency curves. This makes it difficult for traditional centralized or average power allocation strategies to simultaneously optimize both the operating status of each stack and overall system efficiency. Furthermore, the output response and aging process of fuel cell stacks exhibit significant nonlinear characteristics, making energy management methods based on linear models or fixed rules lack robustness and generalizability in complex operating scenarios.

[0004] To address the above challenges, intelligent methods such as deep learning and reinforcement learning have gradually been introduced into the energy management of multi-stack fuel cell systems in recent years. However, traditional deep reinforcement learning methods face problems such as slow convergence, poor strategy stability, and easy falling into local optimality, making it difficult to ensure the long-term, efficient and stable operation of multi-stack systems. Summary of the Invention

[0005] In order to solve the above problems, the present invention proposes a hierarchical energy management method for a multi-stack fuel cell hybrid system, which is a new energy management method with fuel cell structure identification and optimal energy management. The method takes into account operating efficiency, equipment health status and system control robustness, and is suitable for distributed, high-volatility hybrid vehicle application scenarios, so as to achieve efficient operation and life protection of multi-stack hybrid system systems in high-power, high-reliability application scenarios such as rail transit vehicles.

[0006] To achieve the above-mentioned object, the technical solution adopted by the present invention is: a hierarchical energy management method for a multi-stack fuel cell hybrid system, characterized by comprising the steps of:

[0007] Step 1: From the multi-stack fuel cell hybrid system, obtain the degradation state and real-time maximum efficiency of each fuel cell, each lithium battery, and each supercapacitor in the system, perform normalization processing, and construct a unified feature vector set;

[0008] Step 2: Using the constructed decoupled feature clustering neural network, fuel cells, lithium batteries, and supercapacitors with similar degradation levels and maximum efficiency are grouped into one region, thereby dividing the entire system into several regions;

[0009] Step 3: Build a nested deep reinforcement learning controller based on the instantaneous efficiency of the fuel cell, the equivalent hydrogen consumption of the system, the instantaneous degradation rate of each power unit, the SOC of the lithium battery and supercapacitor, and the operating temperature of the fuel cell to achieve optimal energy management within each zone;

[0010] Step 4: Perform system-level coordination and scheduling of energy management results for all regions. A central power coordination module is used to achieve cross-regional power balance and boundary condition constraints, ensuring power consistency between energy storage support and fuel cell operation requirements.

[0011] Step 5: Use long-term operating data to perform online updates and adaptive corrections on the decoupled feature clustering neural network and nested deep reinforcement learning to improve the system's robustness and generalization capabilities to non-stationary working conditions.

[0012] Furthermore, the multi-stack fuel cell hybrid system includes: multiple fuel cell power generation units, multiple lithium battery energy storage units, multiple supercapacitor energy storage units, locomotive loads, Boost converter units, Buck / Boost converter units, Buck converter units and inverter units. The multiple fuel cell power generation units are connected to the DC bus through the Boost converter unit, and the multiple lithium battery energy storage units and the multiple supercapacitor energy storage units are connected to the DC bus through the Buck / Boost converter unit. The DC bus provides energy to the locomotive load through the Buck converter unit and the inverter unit; the DC bus adopts a ring bus architecture and has a bidirectional power supply path.

[0013] Furthermore, the decoupled feature clustering neural network in step 2 mines the steady-state properties and dynamic response characteristics of fuel cells, lithium batteries, and supercapacitors in their operating states based on the static-dynamic decoupling and parallel channel collaborative extraction mechanism. The input vector is divided into the following after the decoupling module:

[0014] X=[X s ,X d ];

[0015] Among them, X is the neural network input vector; X s is the static characteristic subspace, which represents the rated parameters and structural characteristics of each power unit; X dIt is the dynamic feature subspace, representing the variables that change with time during the operation process.

[0016] Furthermore, the decoupled feature clustering neural network is optimized by a joint loss function.

[0017] The joint loss function is:

[0018] L total =L cluster +λ1L feature ;

[0019] Among them, L total represents the total clustering loss; L cluster represents the structural clustering loss, which is mainly used to drive the neural network to classify fuel cells, lithium batteries and supercapacitors with similar operating characteristics into the same category; L feature Represents feature preservation loss, which is mainly used to ensure that the clustering process does not lose the key information in the original input data; λ1 is the loss weighting coefficient.

[0020] Furthermore, each area of ​​the system is equipped with the nested deep reinforcement learning in step 3, and the nested deep reinforcement learning adopts a triple nested structure, including strategy nesting, time scale nesting and action space nesting, to achieve multi-objective optimal control of each area.

[0021] Furthermore, the objective function of multi-objective optimization control is:

[0022]

[0023] Among them, J is the overall fitness value; R t is the expected reward of the region at time t, which represents the immediate benefit obtained by the system after executing the action in this state; Var(P t ) is the average power fluctuation rate of the fuel cells in the region at time t; is the absolute value of the instantaneous average temperature change rate of the fuel cell in this area; α1 is the power fluctuation penalty factor, and α2 is the temperature change penalty factor.

[0024] Furthermore, the R t is the expected reward of the area at time t, and the calculation formula is:

[0025] R t =ω1η(t)+ω2C(t)+ω3D ec (t)+ω4D bat (t)+ω5D sc (t)

[0026] +ω6ΔSOC bat (t)+ω7ΔSOCsc (t);

[0027] Where η(t) represents the average instantaneous efficiency of the fuel cell in the region at time t; C(t) represents the equivalent hydrogen consumption in the region at time t; D ec (t), D ec (t) and D sc (t) represents the average instantaneous decay rate of the fuel cell, lithium battery and supercapacitor in the region at time t; ΔSOC bat (t) and ΔSOC sc (t) represents the average change in SOC of lithium batteries and supercapacitors in the area at time t; ω1, ω2, ω3, ω4, ω5, ω6, and ω7 represent the penalty coefficients of each variable, respectively.

[0028] Furthermore, the central power coordination module of step 4 adopts a master-slave structure, wherein the main strategy network is obtained by the output of each sub-strategy through an attention weighted method.

[0029] Furthermore, the attention weighting method is:

[0030] π global =Attention(π1,π2,...,π n );

[0031] Among them, π global is the global output strategy, π1, π2, π n They represent the output strategies of the 1st, 2nd, and nth regions respectively; Attention is the attention-weighted fusion function, which is used to calculate the influence of each regional strategy on the global strategy;

[0032] To further ensure that the scheduling results meet the operation constraints, a penalty function is constructed:

[0033] L penalty =β1max(0,P total -P limit )+β2max(0,SOC bat -SOC bat,max )

[0034] +β3max(0,SOC bat,min -SOC bat )+β4max(0,SOC sc -SOC sc,max )

[0035] +β5max(0,SOC sc,min -SOC sc )

[0036] Among them, L penaltyis the total penalty loss; P total With P limit The current total operating power and maximum operating power limit of the system; SOC bat , SOC bat,max With SOC bat,min The average instantaneous SOC, maximum SOC limit and minimum SOC limit of the system lithium battery; SOC sc , SOC sc,max With SOC sc,min is the average instantaneous SOC, maximum SOC limit and minimum SOC limit of the system supercapacitor; β1, β2, β3, β4 and β5 are the power weight limit factor, lithium battery maximum SOC weight limit factor, lithium battery minimum SOC weight limit factor, supercapacitor maximum SOC weight limit factor and supercapacitor minimum SOC weight limit factor respectively.

[0037] Furthermore, the online update mechanism of step 5 needs to be combined with the model uncertainty evaluation index σ model Dynamically prioritize the experience samples and use incremental gradient descent optimization. The expression is:

[0038]

[0039] Where Δθ represents the change in the model parameter θ; γ is the learning rate; Denotes the loss function L t (X t ,Y t ) with respect to the gradient of the model parameters θ; X t With Y t are state variables and output variables.

[0040] The beneficial effects of adopting this technical solution are:

[0041] The upper layer of the present invention models the static and dynamic operating characteristics of fuel cells by introducing a characteristic subspace decoupling mechanism, and realizes structured clustering between devices based on joint loss optimization; on this basis, the lower layer constructs a nested multi-layer reinforcement learning structure to take into account optimization goals such as system efficiency, operational stability and degradation suppression, and realizes collaborative optimization scheduling at the subsystem level and system level, thereby improving the global performance and long-term stability of the system in a complex energy environment.

[0042] The present invention can automatically identify different operating characteristic modes according to the dynamic characteristics and degradation trends of each power unit in a multi-stack fuel cell hybrid system, and perform structural clustering division on each power unit through a decoupled feature clustering neural network, thereby effectively improving the accuracy and adaptability of the hierarchical control strategy.

[0043] The present invention introduces a nested deep reinforcement learning method, which combines the modeling ability of deep learning for high-dimensional nonlinear state space with the processing ability of hierarchical reinforcement learning for multi-stage decision-making processes, realizes efficient solution of complex power scheduling tasks, and has strong generalization ability.

[0044] The present invention adopts a global coordination mechanism to integrate and constrain local strategies, thereby improving the coordination and robustness of the overall system under the condition of multi-source fluctuating energy supply.

[0045] The present invention supports an online adaptive update mechanism, which can dynamically adjust model parameters according to the evolution trend of the performance indicators of each power unit during the long-term operation of the system, alleviate the problem of strategic performance degradation caused by equipment aging or changes in operating conditions, and achieve the coordinated guarantee of multiple goals of hybrid power system energy efficiency optimization, rapid load response and extended equipment life. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 This is a schematic flow chart of a hierarchical energy management method for a multi-stack fuel cell hybrid system according to the present invention;

[0047] Figure 2 This is a topological diagram of a multi-stack hybrid power system in an embodiment of the present invention. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described below with reference to the accompanying drawings.

[0049] In this embodiment, see Figure 1 As shown, the present invention proposes a hierarchical energy management method for a multi-stack fuel cell hybrid system. The method is based on a decoupled feature clustering neural network and nested deep reinforcement learning, and includes the following steps:

[0050] Step 1: From the multi-stack fuel cell hybrid system, obtain the degradation state and real-time maximum efficiency of each fuel cell, each lithium battery, and each supercapacitor in the system, perform normalization processing, and construct a unified feature vector set;

[0051] Step 2: Using the constructed decoupled feature clustering neural network, fuel cells, lithium batteries, and supercapacitors with similar degradation levels and maximum efficiency are grouped into one region, thereby dividing the entire system into several regions;

[0052] Step 3: Build a nested deep reinforcement learning controller based on the instantaneous efficiency of the fuel cell, the equivalent hydrogen consumption of the system, the instantaneous degradation rate of each power unit, the SOC of the lithium battery and supercapacitor, and the operating temperature of the fuel cell to achieve optimal energy management within each zone;

[0053] Step 4: Perform system-level coordination and scheduling of energy management results for all regions. A central power coordination module is used to achieve cross-regional power balance and boundary condition constraints, ensuring power consistency between energy storage support and fuel cell operation requirements.

[0054] Step 5: Use long-term operating data to perform online updates and adaptive corrections on the decoupled feature clustering neural network and nested deep reinforcement learning to improve the system's robustness and generalization capabilities to non-stationary working conditions.

[0055] like Figure 2 As shown, the multi-stack fuel cell hybrid system includes: a multi-fuel cell power generation unit 100, a multi-lithium battery energy storage unit 200, a multi-supercapacitor energy storage unit 300, a locomotive load 400, a Boost converter unit 500, a Buck / Boost converter unit 600, a Buck converter unit 700 and an inverter unit 800. The multi-fuel cell power generation unit 100 is connected to the DC bus through the Boost converter unit 500, and the multi-lithium battery energy storage unit 200 and the multi-supercapacitor energy storage unit 300 are connected to the DC bus through the Buck / Boost converter unit 600. The DC bus provides energy to the locomotive load 400 through the Buck converter unit 700 and the inverter unit 800; the DC bus adopts a ring bus architecture and has a bidirectional power supply path.

[0056] As an optimization solution for the above embodiment, the decoupled feature clustering neural network in step 2 is based on the static-dynamic decoupling and parallel channel collaborative extraction mechanism to mine the steady-state properties and dynamic response characteristics of the fuel cell, lithium battery and supercapacitor in the operating state. The input vector is divided into:

[0057] X=[X s ,X d ];

[0058] Among them, X is the neural network input vector; X s is the static characteristic subspace, which represents the rated parameters and structural characteristics of each power unit; X d It is the dynamic feature subspace, representing the variables that change with time during the operation process.

[0059] Preferably, the decoupled feature clustering neural network is optimized by a joint loss function.

[0060] The joint loss function is:

[0061] L total =L cluster +λ1L feature ;

[0062] Among them, L totalrepresents the total clustering loss; L cluster represents the structural clustering loss, which is mainly used to drive the neural network to classify fuel cells, lithium batteries and supercapacitors with similar operating characteristics into the same category; L feature Represents feature preservation loss, which is mainly used to ensure that the clustering process does not lose the key information in the original input data; λ1 is the loss weighting coefficient.

[0063] As an optimization solution for the above embodiment, each area of ​​the system is equipped with the nested deep reinforcement learning in step 3, and the nested deep reinforcement learning adopts a triple nested structure, including strategy nesting, time scale nesting and action space nesting, to achieve multi-objective optimization control of each area.

[0064] The objective function of multi-objective optimization control is:

[0065]

[0066] Among them, J is the overall fitness value; R t is the expected reward of the region at time t, which represents the immediate benefit obtained by the system after executing the action in this state; Var(P t ) is the average power fluctuation rate of the fuel cells in the region at time t; is the absolute value of the instantaneous average temperature change rate of the fuel cell in this area; α1 is the power fluctuation penalty factor, and α2 is the temperature change penalty factor.

[0067] Among them, the R t is the expected reward of the area at time t, and the calculation formula is:

[0068] R t =ω1η(t)+ω2C(t)+ω3D ec (t)+ω4D bat (t)+ω5D sc (t)

[0069] +ω6ΔSOC bat (t)+ω7ΔSOC sc (t);

[0070] Where η(t) represents the average instantaneous efficiency of the fuel cell in the region at time t; C(t) represents the equivalent hydrogen consumption in the region at time t; D ec (t), D ec (t) and D sc (t) represents the average instantaneous decay rate of the fuel cell, lithium battery and supercapacitor in the region at time t; ΔSOC bat (t) and ΔSOC sc(t) represents the average change in SOC of lithium batteries and supercapacitors in the area at time t; ω1, ω2, ω3, ω4, ω5, ω6, and ω7 represent the penalty coefficients of each variable, respectively.

[0071] As an optimization solution of the above embodiment, the central power coordination module of step 4 adopts a master-slave structure, wherein the main strategy network is obtained by the output of each sub-strategy through an attention weighted method.

[0072] The attention weighting method is:

[0073] π global =Attention(π1,π2,...,π n );

[0074] Among them, π global is the global output strategy, π1, π2, π n They represent the output strategies of the 1st, 2nd, and nth regions respectively; Attention is the attention-weighted fusion function, which is used to calculate the influence of each regional strategy on the global strategy;

[0075] To further ensure that the scheduling results meet the operation constraints, a penalty function is constructed:

[0076] L penalty =β1max(0,P total -P limit )+β2max(0,SOC bat -SOC bat,max )

[0077] +β3max(0,SOC bat,min -SOC bat )+β4max(0,SOC sc -SOC sc,max )

[0078] +β5max(0,SOC sc,min -SOC sc )

[0079] Among them, L penalty is the total penalty loss; P total With P limit The system's current total operating power and maximum operating power limit; SOC bat , SOC bat,max With SOC bat,min The average instantaneous SOC, maximum SOC limit and minimum SOC limit of the system lithium battery; SOC sc , SOC sc,max With SOC sc,minis the average instantaneous SOC, maximum SOC limit and minimum SOC limit of the system supercapacitor; β1, β2, β3, β4 and β5 are the power weight limit factor, lithium battery maximum SOC weight limit factor, lithium battery minimum SOC weight limit factor, supercapacitor maximum SOC weight limit factor and supercapacitor minimum SOC weight limit factor respectively.

[0080] As an optimization solution of the above embodiment, the online update mechanism of step 5 needs to be combined with the model uncertainty evaluation index σ model Dynamically prioritize the experience samples and use incremental gradient descent optimization. The expression is:

[0081]

[0082] Where Δθ represents the change in the model parameter θ; γ is the learning rate; Denotes the loss function L t (X t ,Y t ) with respect to the gradient of the model parameters θ; X t With Y t are state variables and output variables

[0083] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. A hierarchical energy management method for a multi-stack fuel cell hybrid system, characterized in that: Including steps: Step 1: From the multi-stack fuel cell hybrid system, obtain the degradation state and real-time maximum efficiency of each fuel cell, each lithium battery, and each supercapacitor in the system, perform normalization processing, and construct a unified feature vector set; Step 2: Using the constructed decoupled feature clustering neural network, fuel cells, lithium batteries, and supercapacitors with similar degradation levels and maximum efficiency are grouped into one region, thereby dividing the entire system into several regions; Step 3: Build a nested deep reinforcement learning controller based on the instantaneous efficiency of the fuel cell, the equivalent hydrogen consumption of the system, the instantaneous degradation rate of each power unit, the SOC of the lithium battery and supercapacitor, and the operating temperature of the fuel cell to achieve optimal energy management within each zone; Step 4: Perform system-level coordination and scheduling of energy management results for all regions. A central power coordination module is used to achieve cross-regional power balance and boundary condition constraints, ensuring power consistency between energy storage support and fuel cell operation requirements. Step 5: Use long-term running data to perform online updates and adaptive corrections on the decoupled feature clustering neural network and nested deep reinforcement learning.

2. The hierarchical energy management method for a multi-stack fuel cell hybrid system according to claim 1, characterized in that: The multi-stack fuel cell hybrid power system comprises: a multi-fuel cell power generation unit (100), a multi-lithium battery energy storage unit (200), a multi-supercapacitor energy storage unit (300), a locomotive load (400), a Boost converter unit (500), a Buck / Boost converter unit (600), a Buck converter unit (700) and an inverter unit (800); the multi-fuel cell power generation unit (100) is connected to a DC bus via the Boost converter unit (500); the multi-lithium battery energy storage unit (200) and the multi-supercapacitor energy storage unit (300) are connected to the DC bus via the Buck / Boost converter unit (600); the DC bus provides energy to the locomotive load (400) via the Buck converter unit (700) and the inverter unit (800); the DC bus adopts a ring bus architecture and has a bidirectional power supply path.

3. The hierarchical energy management method for a multi-stack fuel cell hybrid system according to claim 1, characterized in that: In step 2, the decoupled feature clustering neural network is based on the static-dynamic decoupling and parallel channel collaborative extraction mechanism to mine the steady-state properties and dynamic response characteristics of the fuel cell, lithium battery and supercapacitor in the operating state. The input vector is divided into: X=[X s ,X d ]; Among them, X is the neural network input vector; X s is the static characteristic subspace, which represents the rated parameters and structural characteristics of each power unit; X d It is the dynamic feature subspace, representing the variables that change with time during the operation process.

4. The hierarchical energy management method for a multi-stack fuel cell hybrid system according to claim 3, characterized in that: The decoupled feature clustering neural network is optimized by a joint loss function. The joint loss function is: L total =L cluster +λ1L feature ; Among them, L total represents the total clustering loss; L cluster represents the structural clustering loss; L feature represents the feature preservation loss; λ1 is the loss weighting coefficient.

5. The hierarchical energy management method for a multi-stack fuel cell hybrid system according to claim 1, characterized in that: Each area of ​​the system is equipped with the nested deep reinforcement learning in step 3, and the nested deep reinforcement learning adopts a triple nested structure, including policy nesting, time scale nesting and action space nesting, to achieve multi-objective optimal control of each area.

6. The hierarchical energy management method for a multi-stack fuel cell hybrid system according to claim 5, characterized in that: The objective function of multi-objective optimization control is: Among them, J is the overall fitness value; R t is the expected reward of the region at time t, which represents the immediate benefit obtained by the system after executing the action in this state; Var(P t ) is the average power fluctuation rate of the fuel cells in the region at time t; is the absolute value of the instantaneous average temperature change rate of the fuel cell in this area; α1 is the power fluctuation penalty factor, and α2 is the temperature change penalty factor.

7. The hierarchical energy management method for a multi-stack fuel cell hybrid system according to claim 6, characterized in that: The R t is the expected reward of the area at time t, and the calculation formula is: R t =ω1η(t)+ω2C(t)+ω3D ec (t)+ω4D bat (t)+ω5D sc (t)+ω6ΔSOC bat (t)+ω7ΔSOC sc (t); Where η(t) represents the average instantaneous efficiency of the fuel cell in the region at time t; C(t) represents the equivalent hydrogen consumption in the region at time t; D ec (t), D ec (t) and D sc (t) represents the average instantaneous decay rate of the fuel cell, lithium battery and supercapacitor in the region at time t; ΔSOC bat (t) and ΔSOC sc (t) represents the average change in SOC of lithium batteries and supercapacitors in the area at time t; ω1, ω2, ω3, ω4, ω5, ω6, and ω7 represent the penalty coefficients of each variable, respectively.

8. The hierarchical energy management method for a multi-stack fuel cell hybrid system according to claim 1, characterized in that: The central power coordination module of step 4 adopts a master-slave structure, in which the main strategy network is obtained by the output of each sub-strategy through an attention weighted method.

9. The hierarchical energy management method for a multi-stack fuel cell hybrid system according to claim 7, characterized in that: The attention weighting method is: p global =Attention(π1,π2,...,π n ); Among them, π global is the global output strategy, π1, π2, π n They represent the output strategies of the 1st, 2nd, and nth regions respectively; Attention is the attention-weighted fusion function, which is used to calculate the influence of each regional strategy on the global strategy; To further ensure that the scheduling results meet the operation constraints, a penalty function is constructed: L penalty =β1max(0,P total -P limit )+β2max(0,SOC bat -SOC bat,max )+β3max(0,SOC bat,min -SOC bat )+β4max(0,SOC sc -SOC sc,max )+β5max(0,SOC sc,min -SOC sc ) Among them, L penalty is the total penalty loss; P total With P limit The current total operating power and maximum operating power limit of the system; SOC bat , SOC bat,max With SOC bat,min The average instantaneous SOC, maximum SOC limit and minimum SOC limit of the system lithium battery; SOC sc , SOC sc,max With SOC sc,min is the average instantaneous SOC, maximum SOC limit and minimum SOC limit of the system supercapacitor; β1, β2, β3, β4 and β5 are the power weight limit factor, lithium battery maximum SOC weight limit factor, lithium battery minimum SOC weight limit factor, supercapacitor maximum SOC weight limit factor and supercapacitor minimum SOC weight limit factor respectively.

10. The hierarchical energy management method for a multi-stack fuel cell hybrid system according to claim 1, characterized in that: The online update mechanism of step 5 needs to be combined with the model uncertainty evaluation index σ model Dynamically prioritize the experience samples and use incremental gradient descent optimization. The expression is: Where Δθ represents the change in the model parameter θ; γ is the learning rate; Denotes the loss function L t (X t ,Y t ) with respect to the gradient of the model parameters θ; X t With Y t are state variables and output variables.

Citation Information

Cited By

  • Energy distribution and scheduling control method and system for liquid cooling energy storage system

    CN120934038A