A multi-stack hybrid system adaptive cooperative energy management method

CN120657180BActive Publication Date: 2026-09-15SOUTHWEST JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510798460.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2026-09-15
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

并非所有故障都会导致电堆无法使用,电堆可能经历一个性能逐渐下降的过程,其在不同方面的性能表现也会有不同程度的劣化,如果不能准确评估这种故障的程度,而是直接隔离电堆,就无法充分利用故障电堆剩余的工作能力,增加不必要的成本损失

Benefits of technology

[0057]This invention discloses an adaptive cooperative energy management method for multi-stack hybrid systems based on intelligent fault severity assessment. This method fully considers the impact of differences in stack fault severity in multi-stack fuel cell hybrid systems. First, it analyzes the changing trends of the fuel cell's output polarization curve and efficiency curve as fuel cell performance deteriorates, extracting ten features, including open-circuit voltage and the slope of the ohmic polarization curve, as feature vectors reflecting the degree of fuel cell fault and assigning them category labels. Second, using the extracted ten-dimensional feature vectors and category labels as input, a dynamic adaptive attention neural network is constructed and trained. The dynamic attention mechanism is used to classify the stack fault severity into no fault, minor fault, moderate fault, and severe fault. Finally, an adaptive proximal policy optimization reinforcement learning intelligent system based on KL divergence constraints is constructed. The system employs a self-calibrating feedback neural network to identify the polarization and efficiency curves of fuel cell stacks online in real time for real-time fault assessment. The assessment results are aggregated into a comprehensive health status index and dynamically input into the agent. The agent adaptively adjusts the fuel cell performance degradation cost term in the reward function and autonomously determines the optimal power allocation law for the multi-stack fuel cell system and lithium batteries. Finally, based on the fault severity of each stack, if a stack is severely faulty, its fault assessment decision module will quickly isolate the stack, putting the system into a degraded operation mode. The calculated power demand of the multi-stack fuel cell system will be rationally allocated according to the fault severity of the remaining stacks, achieving optimal dynamic reconfiguration of the entire system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120657180B_ABST
    Figure CN120657180B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-stack hybrid system adaptive collaborative energy management methods, obtain the polarization and efficiency curve of fuel cell stack of different performance degradation state, extract the feature vector reflecting fault severity and assign class label;Adopt dynamic adaptive attention neural network, evaluate the most important feature and adaptively adjust, so as to divide the fault degree of stack;Build adaptive proximal policy optimization reinforcement learning agent based on KL divergence constraint, train self-correcting feedback neural network to identify fuel cell polarization and efficiency curve online, extract feature vector for real-time fault severity evaluation, aggregate the results into comprehensive health status index input agent, adaptively adjust reward function;According to the power distribution result of multi-stack fuel cell power generation system, if the stack is serious fault, then isolate the stack into degraded operation mode, the rest of the stack dynamically assumes power according to fault severity, improve system robustness, realize system dynamic optimal reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of fuel cell technology, and in particular relates to an adaptive cooperative energy management method for multi-stack hybrid systems. Background Technology

[0002] With the global energy crisis and environmental pollution intensifying, developing clean and efficient energy conversion technologies has become a key direction for scientific and technological development. Proton exchange membrane fuel cells (PEMFCs), due to their high energy conversion efficiency, zero emissions, low operating noise, and diverse fuel sources, are widely considered one of the ideal power sources to replace traditional internal combustion engines, demonstrating enormous application potential in transportation, stationary power plants, and portable power supplies.

[0003] To meet the requirements of power output, dynamic response, and system reliability in practical applications, fuel cell hybrid power systems have become a mainstream technology. Multi-stack fuel cell hybrid power systems not only flexibly expand power output by increasing the number of stacks to meet the needs of different application scenarios, but also improve the overall system performance, reliability, and fault tolerance through the collaborative work of each unit. However, the complex structure of multi-stack fuel cell hybrid power systems makes their control process extremely complex, with the core lying in the design of energy management methods. Currently, rule-based and optimization-based energy management methods are quite mature, and with the development of artificial intelligence technology, reinforcement learning-based energy management methods are also showing great potential. More importantly, the above-mentioned energy management methods often assume that system components are in a healthy or ideal state, or only consider a simplified, averaged aging model. However, in actual operation, fuel cell stacks inevitably experience performance degradation and even various failures. These failures not only reduce system performance and efficiency but also accelerate stack aging and may even lead to system shutdown.

[0004] Existing research on fuel cell failures largely focuses on failure diagnosis itself, namely detecting failure occurrence, identifying failure types, and isolating faulty units. At the energy management level, relatively simple fault-tolerant controls such as degraded operation are typically employed, lacking effective means to quantify the severity of failures with fine precision, and mechanisms to adaptively integrate this quantitative information into the collaborative energy management optimization process of multi-stack fuel cell hybrid power systems. Not all failures render the stack unusable; the stack may undergo a gradual performance degradation process, with varying degrees of deterioration in different aspects of its performance. If the severity of such failures cannot be accurately assessed, and the stack is directly isolated, the remaining working capacity of the faulty stack cannot be fully utilized, resulting in unnecessary cost losses.

[0005] Therefore, in multi-stack fuel cell hybrid power systems, it is necessary to study a method for detecting stack faults while assessing the severity of the faults, and to adaptively adjust the system's collaborative energy management strategy based on the assessment results, so as to achieve a dynamic optimization balance of multiple objectives while meeting basic operating requirements. Summary of the Invention

[0006] To address the aforementioned issues, this invention proposes an adaptive cooperative energy management method for multi-stack hybrid systems. This method simultaneously detects stack faults and assesses their severity, then adaptively adjusts the system's cooperative energy management strategy based on the assessment results. This achieves a dynamic optimization balance across multiple objectives while meeting basic operational requirements. The invention extracts features reflecting performance degradation from the polarization and efficiency curves of the fuel cell stack and employs a dynamic adaptive attention neural network to assess the fault severity of the fuel cell stack in real time. Based on the fault severity determination results, an adaptive proximal policy optimization reinforcement learning agent based on KL divergence constraints is trained to autonomously determine the optimal output power of the fuel cell at different fault levels under multi-objective optimization conditions, thus achieving optimal dynamic reconfiguration of the hybrid power system.

[0007] To achieve the above objectives, the technical solution adopted by this invention is: an adaptive cooperative energy management method for multi-reactor hybrid systems, comprising the following steps:

[0008] S100 periodically acquires polarization curve data and efficiency curve data of fuel cell stacks in different performance degradation states, extracts a set of feature vectors that can effectively reflect the severity of the fault, and divides them into four categories: no fault, minor fault, moderate fault, and severe fault.

[0009] S200: The feature vectors reflecting the degree of failure extracted in step S100 are combined with their category labels as input to construct and train a dynamic adaptive attention neural network. The attention mechanism is used to dynamically focus on the features most important to the current assessment and adaptively adjust them to classify the degree of failure of the fuel cell stack and output the discrete degree of failure assessment results of each fuel cell stack in real time.

[0010] S300 constructs an adaptive proximal policy optimization reinforcement learning agent based on KL divergence constraints. It uses a self-correcting feedback neural network to identify the polarization curve and efficiency curve of the fuel cell online for real-time fault assessment of each stack. The assessment results are aggregated into a comprehensive health index as an input state of the agent, and the reward function is adaptively changed accordingly to achieve optimal power allocation of multi-stack fuel cell system and lithium battery.

[0011] S400: Based on the power allocation result of the multi-stack fuel cell power generation system in step S300, if the stack is severely faulty, the stack is isolated and enters a degraded operation mode, and the remaining stacks dynamically take on power according to their respective fault severity, so as to achieve dynamic optimal reconfiguration of the entire system.

[0012] Furthermore, in step S100, the output characteristics of the fuel cell stack under different faults and fault degrees are simulated, and its polarization curve data and efficiency curve data are collected. After preprocessing these data, six features are extracted from the polarization curve, including the open-circuit voltage V. oc Voltage value V corresponding to rated current nom The slope R of the Ohmic polarization region curve ohm Slope S of the concentration polarization region curve conc limiting current density I lim and the maximum power point P max As a quantitative characteristic that sensitively reflects changes in the internal health state of the fuel cell stack, four features are extracted from the efficiency curve, including the maximum efficiency η. max The current I corresponding to the maximum efficiency point ηmax and the lower bound I of the efficient operating range lower and the upper boundary I upper The feature vector, consisting of 10 features, effectively reflects the severity of the fault:

[0013] F = [V] oc V nom ,R ohm ,S conc ,I lim ,P max ,η max ,I ηmax ,I lower ,I upper ];

[0014] Based on the actual operating conditions of the fuel cell stack, the feature vector is divided into four fault severity categories: no fault, minor fault, moderate fault, and severe fault.

[0015] Furthermore, in step S200, a dynamic adaptive attention neural network is constructed based on the feature vectors and category labels that reflect the degree of fault in the fuel cell stack, and the feature vectors are used as the state input of the neural network to output the discrete fault degree evaluation results of each fuel cell stack in real time.

[0016] The dynamic adaptive attention neural network includes an input layer, a feature embedding / transformation layer, a dynamic adaptive attention layer, a post-processing layer, and an output layer. The training steps include:

[0017] (1) Feature processing and network initialization: After normalizing the ten-dimensional feature vector, the training dataset is formed by combining it with the classification labels, and then the parameters of the network are initialized.

[0018] (2) Model forward propagation and feature transformation: The input feature vector is first transformed in a hierarchical nonlinear manner through a deep residual network structure. This process aims to project the original features into a high-dimensional latent embedding space to obtain the embedded features H. In this space, the intrinsic structure and semantic relationship of different failure modes are amplified, thereby providing a more discriminative input for the subsequent attention mechanism.

[0019] (3) Dynamic attention weight generation and application: Based on the deep feature representation H, a multi-head self-attention mechanism is used to calculate the context-related attention weights and generate the context vector C. This mechanism calculates h independent attention heads in parallel and concatenates the results to capture information from different subspaces.

[0020] (4) Calculation of predicted probability distribution and construction of loss function: Input the context vector C generated by the attention mechanism into the subsequent classifier network, and output the predicted probability distribution vector P of the four fault levels. Based on this, construct a weighted cross-entropy loss function with L2 regularization term to evaluate the difference between the real label and the predicted label and prevent overfitting.

[0021] (5) Parameter optimization based on adaptive moment estimation: The AdamW stochastic gradient descent optimizer is used to iteratively update the model parameters. By maintaining the exponential moving average of the first and second moments of the gradient, an independent and adaptive learning rate is calculated for each parameter to achieve better generalization effect. This process continues until the model's performance on the validation set reaches the convergence criterion.

[0022] Furthermore, in step S300, a self-calibrating feedback neural network is used to identify the polarization curve and efficiency curve of the fuel cell stack online in real time. Its inputs are the fuel cell output current and voltage, fuel cell efficiency, polarization voltage prediction error at the previous moment, and efficiency prediction error at the previous moment. The output is the fitting coefficient of the fuel cell polarization curve and efficiency curve. Based on the two curves identified online, feature vectors reflecting the degree of fault are extracted, and a dynamic adaptive attention neural network is used to evaluate the degree of fault of the stack in real time.

[0023] Furthermore, an adaptive proximal strategy optimization reinforcement learning algorithm based on KL divergence constraints is employed to achieve optimal power allocation between the multi-stack fuel cell power generation system and the lithium batteries. Based on the parameter identification results of the self-correcting feedback neural network and the real-time output results of the dynamic adaptive attention neural network, the fault levels of each stack are aggregated into a comprehensive health status index HI and the load demand power P. loadLithium-ion battery SOC and multi-stack fuel cell system output power P FCS As an intelligent agent, it perceives the state space from the environment and records the change in fuel cell output power ΔP. FCS The action space is the output of the intelligent agent.

[0024] Furthermore, the comprehensive health status index is defined as follows:

[0025]

[0026] Among them, P nom,n Let γ be the rated power of the nth fuel cell stack, γ be the attenuation effect exponent, and δ be the rated power of the nth fuel cell stack. n is the performance degradation factor of the nth fuel cell stack, where N is the total number of fuel cell stacks;

[0027]

[0028] Where α and β are adjustment factors, α controls the growth rate in the initial stage, and β controls the overall nonlinearity. max The highest fault level; δ n A value of 1 indicates that the fuel cell stack has completely failed, and the closer HI is to 1, the better the system availability.

[0029] Furthermore, the state space is represented as:

[0030] State = {HI, P} load SOC, P FCS};

[0031] The action space is represented as follows:

[0032] Action={ΔP FCS |ΔP FCS ∈[ΔP FCSmin ,ΔP FCSmax ]};

[0033] Where, ΔP FCSmin ΔP represents the lower limit of the output power variation of a multi-stack fuel cell. FCSmax This represents the upper limit of the output power variation of multiple fuel cell stacks.

[0034] Furthermore, in step S300, the reward function of the intelligent agent includes fuel cell hydrogen consumption cost, lithium battery equivalent hydrogen consumption cost, and lithium battery SOC fluctuation penalty term;

[0035] The hydrogen consumption cost of a fuel cell includes: basic hydrogen consumption, and performance degradation cost determined by a comprehensive health status index, expressed as:

[0036]

[0037] Among them, C deg Where μ is the performance degradation cost, HI is the overall health status index, and P is the scaling factor. FCS For the power of a multi-stack fuel cell system, P FCSnom Let ξ be the total rated power of the multi-stack fuel cell system, and let ξ be the power influence index.

[0038] The hydrogen consumption cost of a multi-stack fuel cell system is:

[0039]

[0040] Among them, C FC,n This represents the base hydrogen consumption for the nth fuel cell.

[0041] The final reward function expression is:

[0042] R = -(c1C) sys +c2C bat +E SOC );

[0043] Among them, C bat E is the equivalent hydrogen consumption of a lithium battery. SOC c1 and c2 are normalization coefficients, representing the penalty term for SOC fluctuations in lithium batteries.

[0044] Furthermore, in step S300, when training the reinforcement learning agent based on KL divergence constraints, the agent's policy and value function are each calculated and represented by a separate neural network, and the optimal policy is continuously updated through continuous interaction with the environment.

[0045] The PPO-KL objective function is defined to maximize the policy return, and its expression is:

[0046]

[0047] Where, r t (θ) represents the policy ratio, which measures the probability ratio of the new and old policies on the sampled action. t For the dominant function, Here, χ is the divergence penalty term, and E is the adaptive penalty coefficient. t This is the expected value at time step t;

[0048] The core of this algorithm is to introduce KL divergence constraints to limit the range of change between the old and new policies, thereby improving training stability. At the same time, the penalty coefficient of the KL divergence penalty term is dynamically adjusted. When the actual average KL divergence exceeds the preset value, the penalty coefficient is increased to limit the step size of subsequent updates; when the actual average KL divergence is less than the preset value, the penalty coefficient is decreased, allowing the policy to explore and update more extensively. This process is repeated until the policy converges.

[0049] Furthermore, in step S400, a fault assessment and decision module is set up for each fuel cell stack. If the fault level of the stack is a severe fault, an isolation command is generated to disconnect the faulty stack. Based on the multi-stack fuel cell system reference power obtained by the intelligent agent optimization, this power is adaptively and reasonably allocated among the available stacks according to their respective fault levels, so as to realize the fault-tolerant control and dynamic reconfiguration of the system.

[0050] In the fault assessment and decision-making module of each fuel cell stack, health factors and power allocation weights are calculated using the following formulas:

[0051] H n =f(s) n )=exp(-K·s n );

[0052] Among them, H n Let be the health factor of the nth fuel cell stack, which is a function of the degree of failure; K is the adjustment coefficient; and s n Let represent the failure level of the nth fuel cell stack. As the failure level of the fuel cell stack increases, its health factor decreases exponentially.

[0053] The formula for calculating power allocation weights is:

[0054]

[0055] Where, ε n Assign weights to the power of the nth stack, P nom,n S is the rated power of the nth fuel cell stack. active Let H be the set of all available electric piles, where p is the available electric pile number and H is the number of available electric piles. p Let be the health factor of the p-th stack.

[0056] The beneficial effects of adopting this technical solution are:

[0057] This invention discloses an adaptive cooperative energy management method for multi-stack hybrid systems based on intelligent fault severity assessment. This method fully considers the impact of differences in stack fault severity in multi-stack fuel cell hybrid systems. First, it analyzes the changing trends of the fuel cell's output polarization curve and efficiency curve as fuel cell performance deteriorates, extracting ten features, including open-circuit voltage and the slope of the ohmic polarization curve, as feature vectors reflecting the degree of fuel cell fault and assigning them category labels. Second, using the extracted ten-dimensional feature vectors and category labels as input, a dynamic adaptive attention neural network is constructed and trained. The dynamic attention mechanism is used to classify the stack fault severity into no fault, minor fault, moderate fault, and severe fault. Finally, an adaptive proximal policy optimization reinforcement learning intelligent system based on KL divergence constraints is constructed. The system employs a self-calibrating feedback neural network to identify the polarization and efficiency curves of fuel cell stacks online in real time for real-time fault assessment. The assessment results are aggregated into a comprehensive health status index and dynamically input into the agent. The agent adaptively adjusts the fuel cell performance degradation cost term in the reward function and autonomously determines the optimal power allocation law for the multi-stack fuel cell system and lithium batteries. Finally, based on the fault severity of each stack, if a stack is severely faulty, its fault assessment decision module will quickly isolate the stack, putting the system into a degraded operation mode. The calculated power demand of the multi-stack fuel cell system will be rationally allocated according to the fault severity of the remaining stacks, achieving optimal dynamic reconfiguration of the entire system.

[0058] This invention employs a dynamic adaptive attention neural network, which can deeply learn the complex nonlinear relationships hidden in large amounts of data. With the help of its unique attention mechanism, it dynamically identifies the performance characteristics that contribute the most to the fault severity judgment, accurately captures the weak signs in the early stage of the fault, and significantly improves the early warning capability of the system.

[0059] This invention constructs an aggregated health status index, which can macroscopically characterize the health loss of the entire multi-stack fuel cell system. This provides a clear understanding of the system health for high-level decision-making by reinforcement learning agents, thereby adaptively adjusting the reward function so that the power provided by the multi-stack fuel cell system always matches its availability. This maximizes system lifespan while pursuing system economy and protecting degraded fuel cell stacks.

[0060] This invention constructs a fault isolation and dynamic reconfiguration mechanism. Once a fuel cell stack experiences a severe fault, the system does not fail globally. Instead, the faulty unit is precisely isolated physically and in control, enabling online dynamic reconfiguration of the system topology. When a fuel cell stack experiences other fault levels, the system does not immediately isolate the stack. Instead, it dynamically weights and allocates power based on the fault level, ensuring that the remaining system resources are utilized in the most optimized way. This avoids unnecessary over-maintenance, significantly reduces the system's total lifecycle maintenance costs, and greatly improves the stability and fault tolerance of the hybrid power system. Attached Figure Description

[0061] Figure 1 This is a schematic diagram of the adaptive cooperative energy management method for a multi-stack hybrid system according to the present invention;

[0062] Figure 2 This is a schematic diagram of a multi-stack fuel cell hybrid power system in an embodiment of the present invention. Detailed Implementation

[0063] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described below with reference to the accompanying drawings.

[0064] In this embodiment, see Figure 1 As shown, this invention proposes an adaptive cooperative energy management method for multi-stack hybrid systems, comprising the following steps:

[0065] S100 periodically acquires polarization curve data and efficiency curve data of fuel cell stacks in different performance degradation states, extracts a set of feature vectors that can effectively reflect the severity of the fault, and divides them into four categories: no fault, minor fault, moderate fault, and severe fault.

[0066] S200: The feature vectors reflecting the degree of failure extracted in step S100 are combined with their category labels as input to construct and train a dynamic adaptive attention neural network. The attention mechanism is used to dynamically focus on the features most important to the current assessment and adaptively adjust them to classify the degree of failure of the fuel cell stack and output the discrete degree of failure assessment results of each fuel cell stack in real time.

[0067] S300 constructs an adaptive proximal policy optimization reinforcement learning agent based on KL divergence constraints. It uses a self-correcting feedback neural network to identify the polarization curve and efficiency curve of the fuel cell online for real-time fault assessment of each stack. The assessment results are aggregated into a comprehensive health index as an input state of the agent, and the reward function is adaptively changed accordingly to achieve optimal power allocation of multi-stack fuel cell system and lithium battery.

[0068] S400: Based on the power allocation result of the multi-stack fuel cell power generation system in step S300, if the stack is severely faulty, the stack is isolated and enters a degraded operation mode. The remaining stacks dynamically take on power according to their respective fault severity, thereby improving the robustness of the system and achieving dynamic optimal reconfiguration of the entire system.

[0069] In this embodiment, as Figure 2 As shown, the multi-fuel cell hybrid power system includes a multi-fuel cell power generation system, a lithium battery system, a DC bus, a DC / DC converter device, and an energy management and control system. Each fuel cell stack is connected to a unidirectional boost DC / DC converter in parallel to the DC bus to control the fuel cell output power and match the bus voltage. The lithium battery is connected to a bidirectional DC / DC converter connected to the DC bus to maintain bus voltage stability. The energy management and control system collects system operating parameter information and generates corresponding control signals based on this information to control the hybrid power system to achieve optimal power distribution.

[0070] The first step involves extracting features based on the fuel cell polarization and efficiency curves, and training a dynamic adaptive attention neural network to finely classify the degree of fuel cell failure, as detailed below:

[0071] In step S100, the output characteristics of the fuel cell stack under different faults and fault degrees are simulated, and its polarization curve data and efficiency curve data are collected. After preprocessing these data, six features are extracted from the polarization curve, including the open-circuit voltage V. oc Voltage value V corresponding to rated current nom The slope R of the Ohmic polarization region curve ohm Slope S of the concentration polarization region curve conc limiting current density I lim and the maximum power point P max As a quantitative characteristic that sensitively reflects changes in the internal health state of the fuel cell stack, four features are extracted from the efficiency curve, including the maximum efficiency η. max The current I corresponding to the maximum efficiency point ηmax and the lower bound I of the efficient operating range lower and the upper boundary I upper The feature vector, consisting of 10 features, effectively reflects the severity of the fault:

[0072] F = [V] oc V nom ,R ohm ,S conc ,I lim ,P max ,η max ,I ηmax ,I lower ,Iupper ];

[0073] Based on the actual operating conditions of the fuel cell stack, the feature vector is divided into four fault severity categories: no fault, minor fault, moderate fault, and severe fault.

[0074] In step S200, a dynamic adaptive attention neural network is constructed based on the feature vectors and category labels that reflect the degree of fault in the fuel cell stack. The feature vectors are used as the state input of the neural network to output the discrete fault degree evaluation results of each fuel cell stack in real time.

[0075] The dynamic adaptive attention neural network includes an input layer, a feature embedding / transformation layer, a dynamic adaptive attention layer, a post-processing layer, and an output layer. The training steps include:

[0076] (1) Feature processing and network initialization: After normalizing the ten-dimensional feature vector, the training dataset is formed by combining it with the classification labels, and then the parameters of the network are initialized.

[0077] (2) Model forward propagation and feature transformation: The input feature vector is first transformed in a hierarchical nonlinear manner through a deep residual network structure. This process aims to project the original features into a high-dimensional latent embedding space to obtain the embedded features H. In this space, the intrinsic structure and semantic relationship of different failure modes are amplified, thereby providing a more discriminative input for the subsequent attention mechanism.

[0078] (3) Dynamic attention weight generation and application: Based on the deep feature representation H, a multi-head self-attention mechanism is used to calculate context-related attention weights and generate a context vector C. This mechanism computes h independent attention heads in parallel and concatenates the results to capture information from different subspaces.

[0079] C = Concat(head1,...,head) h )·W0;

[0080] Where Concat is a vector concatenation operation, head represents the attention head, and W0 is the output projection matrix, which can fuse information from the left and right attention heads;

[0081] The formula for calculating each attention point is:

[0082]

[0083] in, Let d be the projection matrix of the query, key, and value of the i-th attention head. k Let be the dimension of the key vector, softmax be the activation function, and Attention be the attention mechanism function.

[0084] (4) Calculation of predicted probability distribution and construction of loss function: The context vector C generated by the attention mechanism is input into the subsequent classifier network, and the predicted probability distribution vectors P of the four fault levels are output. Based on this, a weighted cross-entropy loss function with L2 regularization is constructed to evaluate the difference between the true label and the predicted label and to prevent overfitting. Its expression is:

[0085]

[0086] Where ∵ represents the set of all learnable parameters of the model, M is the number of samples, and w j Y represents the weight of the j-th type of fault severity. j (k) For the k-th sample, the label represents the fault level of the j-th class. Let λ be the predicted probability of the j-th type of fault in the k-th sample, λ be the regularization coefficient, and θ be the model parameter.

[0087] (5) Parameter optimization based on adaptive moment estimation: The AdamW stochastic gradient descent optimizer is used to iteratively update the model parameters. By maintaining the exponential moving average of the first and second moments of the gradient, an independent and adaptive learning rate is calculated for each parameter to achieve better generalization effect. This process continues until the model's performance on the validation set reaches the convergence criterion.

[0088] The second step involves constructing an adaptive proximal policy optimization reinforcement learning agent based on KL divergence constraints. This agent adaptively optimizes the power allocation between the multi-stack fuel cell system and the lithium battery according to the fault level of the fuel cell stack. The details are as follows:

[0089] In step S300, a self-correcting feedback neural network is used to identify the polarization curve and efficiency curve of the fuel cell stack online in real time. Its inputs are the fuel cell output current and voltage, fuel cell efficiency, polarization voltage prediction error at the previous moment, and efficiency prediction error at the previous moment. The output is the fitting coefficient of the fuel cell polarization curve and efficiency curve. Based on the two curves identified online, feature vectors reflecting the degree of fault are extracted, and a dynamic adaptive attention neural network is used to evaluate the degree of fault of the stack in real time.

[0090] An adaptive proximal strategy optimization reinforcement learning algorithm based on KL divergence constraints is employed to achieve optimal power allocation between multi-stack fuel cell power generation systems and lithium batteries. Based on parameter identification results from a self-correcting feedback neural network and the real-time output of a dynamic adaptive attention neural network, the fault levels of each stack are aggregated into a comprehensive health status index HI and load demand power P. load Lithium-ion battery SOC and multi-stack fuel cell system output power P FCSAs an intelligent agent, it perceives the state space from the environment and records the change in fuel cell output power ΔP. FCS The action space is the output of the intelligent agent.

[0091] The comprehensive health status index is defined as follows:

[0092]

[0093] Among them, P nom,n Let γ be the rated power of the nth fuel cell stack, γ be the attenuation effect exponent, and δ be the rated power of the nth fuel cell stack. n is the performance degradation factor of the nth fuel cell stack, where N is the total number of fuel cell stacks;

[0094]

[0095] Where α and β are adjustment factors, α controls the growth rate in the initial stage, and β controls the overall nonlinearity. max The highest fault level; δ n A value of 1 indicates that the fuel cell stack has completely failed, and the closer HI is to 1, the better the system availability.

[0096] The state space is represented as follows:

[0097] State = {HI, P} load SOC, P FCS};

[0098] The action space is represented as follows:

[0099] Action={ΔP FCS |ΔP FCS ∈[ΔP FCSmin ,ΔP FCSmax ]};

[0100] Where, ΔP FCSmin ΔP represents the lower limit of the output power variation of a multi-stack fuel cell. FCSmax This represents the upper limit of the output power variation of multiple fuel cell stacks.

[0101] In step S300, the reward function of the intelligent agent includes the hydrogen consumption cost of fuel cell, the equivalent hydrogen consumption cost of lithium battery, and the SOC fluctuation penalty term of lithium battery.

[0102] The hydrogen consumption cost of a fuel cell includes: basic hydrogen consumption, and performance degradation cost determined by a comprehensive health status index, expressed as:

[0103]

[0104] Among them, C deg Where μ is the performance degradation cost, HI is the overall health status index, and P is the scaling factor.FCS For the power of a multi-stack fuel cell system, P FCSnom Let ξ be the total rated power of the multi-stack fuel cell system, and let ξ be the power influence index.

[0105] The hydrogen consumption cost of a multi-stack fuel cell system is:

[0106]

[0107] Among them, C FC,n This represents the base hydrogen consumption for the nth fuel cell.

[0108] The final reward function expression is:

[0109] R = -(c1C) sys +c2C bat +E SOC );

[0110] Among them, C bat E is the equivalent hydrogen consumption of a lithium battery. SOC c1 and c2 are normalization coefficients, representing the penalty term for SOC fluctuations in lithium batteries.

[0111] Preferably, when training a reinforcement learning agent based on KL divergence constraints, the agent's policy and value function are each calculated and represented by a separate neural network, and the optimal policy is continuously updated through continuous interaction with the environment.

[0112] The PPO-KL objective function is defined to maximize the policy return, and its expression is:

[0113]

[0114] Where, r t (θ) represents the policy ratio, which measures the probability ratio of the new and old policies on the sampled action. t For the dominant function, Here, χ is the divergence penalty term, and E is the adaptive penalty coefficient. t This is the expected value at time step t;

[0115] The core of this algorithm is to introduce KL divergence constraints to limit the range of change between the old and new policies, thereby improving training stability. At the same time, the penalty coefficient of the KL divergence penalty term is dynamically adjusted. When the actual average KL divergence exceeds the preset value, the penalty coefficient is increased to limit the step size of subsequent updates; when the actual average KL divergence is less than the preset value, the penalty coefficient is decreased, allowing the policy to explore and update more extensively. This improves stability while ensuring training convergence. This process is repeated until the policy converges.

[0116] In step S400, a fault assessment and decision module is set up for each fuel cell stack. If the fault level of the stack is a severe fault, an isolation command is generated to disconnect the faulty stack. Based on the multi-stack fuel cell system reference power obtained by the intelligent agent optimization, the power is adaptively and reasonably allocated among the available stacks according to their respective fault levels to realize the fault-tolerant control and dynamic reconfiguration of the system.

[0117] In the fault assessment and decision-making module of each fuel cell stack, health factors and power allocation weights are calculated using the following formulas:

[0118] H n =f(s) n )=exp(-K·s n );

[0119] Among them, H n Let be the health factor of the nth fuel cell stack, which is a function of the degree of failure; K is the adjustment coefficient; and s n Let represent the failure level of the nth fuel cell stack. As the failure level of the fuel cell stack increases, its health factor decreases exponentially.

[0120] The formula for calculating power allocation weights is:

[0121]

[0122] Where, ε n Assign weights to the power of the nth stack, P nom,n S is the rated power of the nth fuel cell stack. active Let H be the set of all available electric piles, where p is the available electric pile number and H is the number of available electric piles. p The health factor of the p-th fuel cell stack;

[0123] According to the above principles, the power is allocated so that stacks with high failure rates receive less power and stacks with low failure rates receive more power. This allows for the smooth handling of internal faults in multi-stack fuel cell systems without interrupting the power supply to the load, thereby improving the robustness of the system.

[0124] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. An adaptive cooperative energy management method for a multi-reactor hybrid system, characterized in that, Including the following steps: S100 periodically acquires polarization curve data and efficiency curve data of fuel cell stacks in different performance degradation states, extracts a set of feature vectors that can effectively reflect the severity of the fault, and divides them into four categories: no fault, minor fault, moderate fault, and severe fault. S200: The feature vectors reflecting the degree of failure extracted in step S100 are combined with their category labels as input to construct and train a dynamic adaptive attention neural network. The attention mechanism is used to dynamically focus on the features most important to the current assessment and adaptively adjust them to classify the degree of failure of the fuel cell stack and output the discrete degree of failure assessment results of each fuel cell stack in real time. S300 constructs an adaptive proximal policy optimization reinforcement learning agent based on KL divergence constraints. It uses a self-correcting feedback neural network to identify the polarization curve and efficiency curve of the fuel cell online for real-time fault assessment of each stack. The assessment results are aggregated into a comprehensive health index as an input state of the agent, and the reward function is adaptively changed accordingly to achieve optimal power allocation of multi-stack fuel cell system and lithium battery. In step S300, a self-correcting feedback neural network is used to identify the polarization curve and efficiency curve of the fuel cell stack online in real time. Its inputs are the fuel cell output current and voltage, fuel cell efficiency, polarization voltage prediction error at the previous moment, and efficiency prediction error at the previous moment. The output is the fitting coefficient of the fuel cell polarization curve and efficiency curve. Based on the two curves identified online, feature vectors reflecting the degree of fault are extracted, and a dynamic adaptive attention neural network is used to evaluate the degree of fault of the stack in real time. An adaptive proximal strategy optimization reinforcement learning algorithm based on KL divergence constraints is employed to achieve optimal power allocation between multi-stack fuel cell power generation systems and lithium batteries. Based on parameter identification results from a self-correcting feedback neural network and the real-time output of a dynamic adaptive attention neural network, the fault levels of each stack are aggregated into a comprehensive health index HI and load demand power P. load Lithium-ion battery SOC and multi-stack fuel cell system output power P FCS As an intelligent agent, it perceives the state space from the environment and records the change in fuel cell output power ∆P. FCS The action space as the output of the intelligent agent; S400: Based on the power allocation result of the multi-stack fuel cell power generation system in step S300, if the stack is severely faulty, the stack is isolated and enters a degraded operation mode, and the remaining stacks dynamically take on power according to their respective fault severity, so as to achieve dynamic optimal reconfiguration of the entire system.

2. The adaptive cooperative energy management method for a multi-stall hybrid system according to claim 1, characterized in that, In step S100, the output characteristics of the fuel cell stack under different faults and fault degrees are simulated, and its polarization curve data and efficiency curve data are collected. After preprocessing these data, six features are extracted from the polarization curve, including the open-circuit voltage V. oc Voltage value V corresponding to rated current nom The slope R of the Ohmic polarization region curve ohm Slope S of the concentration polarization region curve conc limiting current density I lim and the maximum power point P max As a quantitative characteristic that sensitively reflects changes in the internal health state of the fuel cell stack, four features are extracted from the efficiency curve, including the maximum efficiency η. max The current I corresponding to the maximum efficiency point ηmax and the lower bound I of the efficient operating range lower and the upper boundary I upper The feature vector, consisting of 10 features, effectively reflects the severity of the fault: ; Based on the actual operating conditions of the fuel cell stack, the feature vector is divided into four fault severity categories: no fault, minor fault, moderate fault, and severe fault.

3. The adaptive cooperative energy management method for a multi-stall hybrid system according to claim 2, characterized in that, In step S200, a dynamic adaptive attention neural network is constructed based on the feature vectors and category labels that reflect the degree of fault in the fuel cell stack. The feature vectors are used as the state input of the neural network to output the discrete fault degree evaluation results of each fuel cell stack in real time. The dynamic adaptive attention neural network includes an input layer, a feature embedding / transformation layer, a dynamic adaptive attention layer, a post-processing layer, and an output layer. The training steps include: (1) Feature processing and network initialization: After normalizing the ten-dimensional feature vector, the training dataset is formed by combining it with the classification labels, and then the parameters of the network are initialized. (2) Model forward propagation and feature transformation: The input feature vector is first transformed in a layered nonlinear manner through a deep residual network structure. This process aims to project the original features into a high-dimensional latent embedding space to obtain the embedded features H. In this space, the internal structure and semantic relationship of different fault modes are amplified, thereby providing a more discriminative input for the subsequent attention mechanism. (3) Dynamic attention weight generation and application: Based on the deep feature representation H, a multi-head self-attention mechanism is used to calculate the context-related attention weights and generate the context vector C. This mechanism calculates h independent attention heads in parallel and concatenates the results to capture information from different subspaces. (4) Calculation of predicted probability distribution and construction of loss function: The context vector C generated by the attention mechanism is input into the subsequent classifier network, and the predicted probability distribution vectors P of the four fault levels are output. Based on this, a weighted cross-entropy loss function with L2 regularization term is constructed to evaluate the difference between the real label and the predicted label and to prevent overfitting. (5) Parameter optimization based on adaptive moment estimation: The AdamW stochastic gradient descent optimizer is used to iteratively update the model parameters. By maintaining the exponential moving average of the first and second moments of the gradient, an independent and adaptive learning rate is calculated for each parameter to achieve better generalization effect. This process continues until the model's performance on the validation set reaches the convergence criterion.

4. The adaptive cooperative energy management method for a multi-stall hybrid system according to claim 1, characterized in that, The comprehensive health index is defined as follows: ; Among them, P nom,n Let γ be the rated power of the nth fuel cell stack, γ be the degradation effect exponent, and δ be the rated power of the nth fuel cell stack. n is the performance degradation factor of the nth fuel cell stack, where N is the total number of fuel cell stacks; ; Where α and β are adjustment factors, α controls the growth rate in the initial stage, and β controls the overall nonlinearity. max The highest fault level; δ n A value of 1 indicates that the fuel cell stack has completely failed, and the closer HI is to 1, the better the system availability.

5. The adaptive cooperative energy management method for a multi-stall hybrid system according to claim 4, characterized in that, The state space is represented as follows: ; The action space is represented as follows: ; Where, ∆P FCSmin ∆P represents the lower limit of the output power variation of a multi-stack fuel cell. FCSmax This represents the upper limit of the output power variation of multiple fuel cell stacks.

6. The adaptive cooperative energy management method for a multi-stall hybrid system according to claim 1, 4, or 5, characterized in that, In step S300, the reward function of the intelligent agent is set to include fuel cell hydrogen consumption cost, lithium battery equivalent hydrogen consumption cost, and lithium battery SOC fluctuation penalty term. The hydrogen consumption cost of a fuel cell includes: basic hydrogen consumption, and the performance degradation cost determined by comprehensive health indicators, expressed as: ; Among them, C deg Where μ is the cost of performance degradation, HI is the overall health index, and P is the cost of performance degradation. FCS For the power of a multi-stack fuel cell system, P FCSnom Let ξ be the total rated power of the multi-stack fuel cell system, and let ξ be the power influence index. The hydrogen consumption cost of a multi-stack fuel cell system is: ; Among them, C FC,n This represents the base hydrogen consumption for the nth fuel cell. The final reward function expression is: ; Among them, C bat E is the equivalent hydrogen consumption of a lithium battery. SOC c1 and c2 are normalization coefficients, representing the penalty term for SOC fluctuations in lithium batteries.

7. The adaptive cooperative energy management method for a multi-stall hybrid system according to claim 6, characterized in that, In step S300, when training a reinforcement learning agent based on KL divergence constraints, the agent's policy and value function are each calculated and represented by a separate neural network, and the optimal policy is continuously updated through continuous interaction with the environment. The PPO-KL objective function is defined to maximize the policy return, and its expression is: ; Where, r t (θ) represents the policy ratio, which measures the probability ratio of the new and old policies on the sampled action. t For the dominant function, For divergence penalty term, E is the adaptive penalty coefficient. t This represents the expected value at time step t. The core of this algorithm is to introduce KL divergence constraints to limit the range of change between the old and new policies, thereby improving training stability. At the same time, the penalty coefficient of the KL divergence penalty term is dynamically adjusted. When the actual average KL divergence exceeds the preset value, the penalty coefficient is increased to limit the step size of subsequent updates; when the actual average KL divergence is less than the preset value, the penalty coefficient is decreased, allowing the policy to explore and update more extensively. This process is repeated until the policy converges.

8. The adaptive cooperative energy management method for a multi-stall hybrid system according to claim 1, characterized in that, In step S400, a fault assessment and decision module is set up for each fuel cell stack. If the fault level of the stack is a severe fault, an isolation command is generated to disconnect the faulty stack. Based on the multi-stack fuel cell system reference power obtained by the intelligent agent optimization, the power is adaptively and reasonably allocated among the available stacks according to their respective fault levels to realize the fault-tolerant control and dynamic reconfiguration of the system. In the fault assessment and decision-making module of each fuel cell stack, health factors and power allocation weights are calculated using the following formulas: ; Among them, H n Let be the health factor of the nth fuel cell stack, which is a function of the degree of failure. s is the adjustment coefficient. n Let represent the failure level of the nth fuel cell stack. As the failure level of the fuel cell stack increases, its health factor decreases exponentially. The formula for calculating power allocation weights is: ; Where, ε n Assign weights to the power of the nth stack, P nom,n S is the rated power of the nth fuel cell stack. active Let H be the set of all available electric piles, where p is the available electric pile number and H is the number of available electric piles. p Let be the health factor of the p-th stack.

Citation Information

Patent Citations

  • Method and device for predicting dynamic performance of fuel cell system

    CN113506901A

  • Underwater detector cluster adaptive detection method and system based on distributed reinforcement learning

    CN119204155A