A fuel cell hydrogen circulation cooling integrated system
By constructing an integrated hydrogen cycle cooling system for fuel cells and utilizing multi-source state perception and reinforcement learning decision-making modules, the coupling control problem of the hydrogen cycle and cooling system of fuel cells was solved, achieving precise temperature stability and improved energy efficiency, and extending the life of the fuel cell stack.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- YANGZHOU JIAHE NEW ENERGY TECH CO LTD
- Filing Date
- 2025-08-21
- Publication Date
- 2026-07-24
AI Technical Summary
In the existing technology, the coupling control of hydrogen circulation and cooling system of fuel cells lacks predictability, making it difficult to quickly coordinate the actions of multiple actuators when the load changes abruptly, resulting in temperature overshoot or oscillation, and making it difficult to adapt to the instantaneous heat surge caused by sudden load changes in fuel cells.
An integrated hydrogen cycle cooling system for fuel cells is constructed, comprising a multi-source state perception module, a Markov decision process modeling module, a reinforcement learning decision module, and an execution control module. By collecting multi-dimensional state data in real time, a Markov decision process model is constructed to generate a collaborative control strategy for the hydrogen cycle-cooling system. Temperature fluctuations are rapidly suppressed through the joint regulation of the hydrogen cycle and cooling system.
Significantly improves the accuracy and stability of temperature control, reduces temperature overshoot, ensures the fuel cell stack operates within the optimal temperature range, extends service life and improves energy conversion efficiency, while also optimizing the energy consumption of the hydrogen circulation pump and cooling system, thereby improving overall energy efficiency.
Smart Images

Figure CN121011684B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fuel cell technology, and more specifically to an integrated system for cooling hydrogen circulation in a fuel cell. Background Technology
[0002] With increasing global focus on environmental protection and sustainable development, traditional gasoline-powered vehicles are gradually being replaced by electric vehicles. Hydrogen fuel cell vehicles, as another clean energy-powered mode of transportation, are increasingly seen as an effective means of addressing climate change and reducing greenhouse gas emissions. Hydrogen fuel cells not only reduce exhaust emissions but also provide longer driving ranges and faster refueling times, making them a key technology for solving energy and environmental problems. The core working principle of a hydrogen fuel cell is the reaction of hydrogen and oxygen to produce electricity and water. However, the operation of a hydrogen fuel cell requires stable operation within a certain temperature range. In practical applications, temperature management of fuel cells is particularly important.
[0003] In existing technologies, traditional PID control relies on fixed parameters. Due to the lack of predictability in the coupled control of the hydrogen cycle and cooling system, it is not only difficult to quickly coordinate the actions of multiple actuators during load changes, but also difficult to adapt to the instantaneous heat surge caused by sudden load changes in fuel cells, resulting in temperature overshoot or oscillation. Therefore, how to construct a Markov decision process model for joint control of hydrogen cycle and cooling with the goal of maximizing temperature stability, and then realize dynamic cooperative control based on reinforcement learning, is the problem to be solved by this invention. To this end, a fuel cell hydrogen cycle cooling integrated system is proposed. Summary of the Invention
[0004] The purpose of this invention is to provide an integrated cooling system for hydrogen circulation in fuel cells to solve the problems mentioned in the background art.
[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:
[0006] A fuel cell hydrogen cycle cooling integrated system includes a hydrogen cycle cooling management center, which is communicatively connected to the following modules:
[0007] The multi-source state sensing module is used to collect multi-dimensional state data of the hydrogen cycle-cooling system in real time, extract state features, and form a state feature sequence table.
[0008] The Markov Decision Process Modeling Module is used to construct Markov decision process models, define the state space and action space, and construct a multi-objective reward function with temperature stability as the core.
[0009] The reinforcement learning decision module trains a reinforcement learning agent based on a Markov decision process model, generates a coordinated control strategy for the hydrogen cycle-cooling system, and decomposes the continuous actions output by reinforcement learning into independent control commands for the hydrogen cycle and cooling system.
[0010] The hydrogen cycle execution control module is used to drive the hydrogen circulation pump frequency converter, adjust the hydrogen recirculation rate, and control the opening of the hydrogen injection valve according to the hydrogen cycle independent control command output by the reinforcement learning decision module, so as to optimize the hydrogen distribution inside the stack and achieve precise matching of hydrogen supply.
[0011] The cooling system execution control module is used to adjust the speed of the electronic water pump and the opening of the electric three-way valve according to the independent control commands of the cooling system, control the flow rate and direction of the coolant, and drive the radiator fan to adjust the speed using PWM to enhance convective heat transfer. Combined with PID control and feedforward compensation, it eliminates the overshoot phenomenon caused by the inertia of the cooling system.
[0012] A further improvement of the technical solution of the present invention is that the multi-source state perception module includes a state monitoring unit and a feature extraction unit;
[0013] The status monitoring unit is used to collect status data of the hydrogen cycle system and the cooling system, and preprocess the data to form a comprehensive dataset.
[0014] The feature extraction unit is used to perform feature analysis on the collected multi-dimensional state data, extract state features related to hydrogen cycle cooling, and integrate them to form a state feature sequence list. The state features include hydrogen supply fluctuation rate, stack temperature gradient, hydrogen-cold coupling efficiency, and latent heat utilization rate of phase change material.
[0015] A further improvement to the technical solution of the present invention is that the process of forming a comprehensive dataset in the status monitoring unit is as follows:
[0016] By deploying a sensor network in the hydrogen circulation system and cooling system of the fuel cell, the status data of the hydrogen circulation system and cooling system are collected in real time. The status data of the hydrogen circulation system includes hydrogen circulation pump speed, hydrogen inlet pressure / flow rate, hydrogen partial pressure inside the stack, and water vapor concentration, etc., to capture hydrogen supply fluctuations caused by load changes. The status data of the cooling system includes coolant temperature / flow rate, radiator fan speed, stack surface temperature distribution, and phase change material (PCM) phase change state, etc.
[0017] The raw state data is preprocessed, including filtering and denoising, outlier removal and time series alignment, to eliminate sensor drift and interference, and short-term statistical features are extracted using the sliding window method to form a standardized time series dataset.
[0018] The preprocessed state data is categorized into hydrogen cycle system and cooling system to form a comprehensive dataset.
[0019] A further improvement to the technical solution of the present invention is that the process of forming the state feature sequence table in the feature extraction unit is as follows:
[0020] The collected status data of the hydrogen cycle system and cooling system are synchronized in time to eliminate communication delay and sampling period differences. The data stream is divided by a sliding window so that the status data of each subsystem are on a unified time scale.
[0021] Feature analysis was performed on the state data of the pre-processed hydrogen cycle system and cooling system to extract state features. Specifically, for the hydrogen cycle system, the hydrogen supply fluctuation rate was calculated, and for the cooling system, the stack temperature gradient and the latent heat utilization rate of the phase change material were quantified. At the same time, the hydrogen-cooling coupling efficiency feature was extracted through the synergistic analysis of hydrogen utilization rate and cooling power.
[0022] The extracted hydrogen supply fluctuation rate, stack temperature gradient, hydrogen-cold coupling efficiency, and latent heat utilization rate of phase change material were normalized to eliminate dimensional differences and integrated into a state feature sequence table according to the time dimension. Each record contains a timestamp and four types of feature values.
[0023] A further improvement of the technical solution of the present invention is that the Markov decision process modeling module includes a state space definition unit and a reward function design unit;
[0024] The state space definition unit is used to abstract the dynamic coupling relationship of the hydrogen cycle-cooling system into a Markov decision process model, and to map the state features into discretized state vectors, while defining the state space and action space.
[0025] The reward function design unit is used to construct a multi-objective reward function with temperature stability as the core. It includes a temperature fluctuation suppression term, a system energy consumption optimization term, and an actuator action smoothing term. The weighted summation method is used to balance the conflicting objectives.
[0026] A further improvement to the technical solution of the present invention is that the state space definition unit specifically includes:
[0027] Using a state feature sequence list covering the state characteristics of the hydrogen cycle system and the cooling system as input, the dynamic dependencies between variables are quantified through covariance analysis and mutual information calculation, a state transition probability matrix is constructed, and a coupled dynamic framework of Markov decision process (MDP) is formed to ensure that the state evolution conforms to physical constraints.
[0028] An adaptive binning algorithm is used for the state features. Based on the feature distribution density and key working condition thresholds, discrete intervals are divided, and continuous values are mapped to discrete state symbols. At the same time, the dynamic sensitivity of the original state features is preserved, and a discretized state vector is generated to reduce the model complexity.
[0029] A state space is constructed based on discrete state vectors, and the action space is defined as discrete control commands for hydrogen circulation pump speed regulation and coolant flow control. An effective action set is generated by combining system safety boundary constraints to ensure that state transitions and action execution meet the real-time and robust requirements of fuel cell operation.
[0030] A further improvement to the technical solution of the present invention is that the reward function design unit specifically includes:
[0031] Based on the control requirements of the fuel cell system, the multi-objective reward function is decomposed into three reward terms: temperature fluctuation suppression term (quantifying thermal stability through stack temperature variance), system energy consumption optimization term (normalizing the energy consumption of hydrogen circulation pump and coolant flow), and actuator action smoothing term (penalizing the abrupt change in action between adjacent time moments). Each reward term is monotonically correlated with the corresponding objective.
[0032] Each reward item is normalized to eliminate dimensional differences. The temperature fluctuation suppression item is normalized based on the maximum variance of historical data. The system energy consumption optimization item is calibrated according to the rated power. The actuator motion smoothing item is linearly scaled by the motion range and the weights are dynamically allocated using the analytic hierarchy process. Among them, the temperature fluctuation suppression item has the highest weight (≥50%). The weights of the system energy consumption optimization item and the actuator motion smoothing item are adaptively adjusted according to the operating conditions.
[0033] A hard constraint penalty term is introduced, which forces the reward to be negative when the temperature exceeds the limit or the safety boundary is violated. Then, the normalized reward terms are summed according to their weights to construct the total reward function, which is then output to the reinforcement learning decision module.
[0034] A further improvement to the technical solution of the present invention is that the reinforcement learning decision module specifically includes:
[0035] The reinforcement learning agent is trained based on the constructed Markov Decision Process (MDP) model. The policy network parameters are optimized through environmental interaction, enabling it to learn and generate cooperative control actions that satisfy multi-objective constraints, and to perform a mapping from state to continuous action.
[0036] The continuous cooperative actions output by the reinforcement learning agent are mapped into independent control commands for the hydrogen cycle and cooling system through an action decomposition mechanism. The decomposition rules are designed based on the system coupling relationship, and the physical feasibility of the actions is ensured through weighted allocation. At the same time, the loyalty of the decomposed commands to the original cooperative target is maintained, and performance degradation due to action segmentation is avoided.
[0037] Independent control commands are input into the fuel cell system for execution, real-time status feedback is collected and instant rewards are calculated, the experience replay buffer of the Markov decision process is updated, the weights of the reinforcement learning strategy are dynamically adjusted based on new data, and the action decomposition rules are optimized to adapt to changes in operating conditions, thereby continuously improving the robustness and adaptability of the coordinated control of the hydrogen cycle-cooling system.
[0038] A further improvement to the technical solution of the present invention is that the hydrogen cycle execution control module specifically includes:
[0039] The hydrogen cycle execution control module receives the independent control command for hydrogen cycle output by the reinforcement learning decision module, performs precise analysis of the control command, and clarifies the drive parameters of the hydrogen cycle pump frequency converter and the set value of the hydrogen injection valve opening.
[0040] Based on the analyzed control commands, the hydrogen circulation pump frequency converter is driven to adjust its operating frequency according to the set parameters, thereby changing the hydrogen recirculation rate. At the same time, the hydrogen injection valve is controlled to achieve the set opening value, thus optimizing the hydrogen distribution inside the fuel cell stack.
[0041] During the adjustment process, the relevant parameters of the hydrogen cycle are continuously monitored, the matching of hydrogen supply is evaluated, and the actual effect is fed back to the reinforcement learning decision-making module so that it can adjust its strategy to achieve continuous and accurate matching of hydrogen supply.
[0042] A further improvement to the technical solution of the present invention is that the cooling system execution control module specifically includes:
[0043] It receives control commands output by the reinforcement learning module, analyzes the target speed of the electronic water pump, the opening degree of the electric three-way valve and the PWM duty cycle of the radiator fan, and dynamically trims them according to the physical limits of the equipment to keep the control commands within the safe operating range and avoid hardware overload or mechanical damage.
[0044] The pump speed and valve opening are adjusted based on the PID algorithm. Combined with the feedforward compensation model to predict changes in cooling demand, the control quantity is adjusted in advance to counteract system inertia. The fan speed is driven by the PWM signal to enhance the radiator's convective heat transfer efficiency and simultaneously optimize the coolant flow distribution and temperature uniformity.
[0045] The system monitors key parameters such as coolant temperature, flow rate, and pressure in real time, compares them with target values to calculate deviations, and if overshoot or response lag is detected, it adopts PID parameter adaptive adjustment or feedforward-feedback composite control strategy to fine-tune the actuator output, ensuring that the cooling system quickly and stably converges to the target state, and generates execution effect data to feed back to the reinforcement learning decision module.
[0046] Due to the adoption of the above technical solution, the technical progress achieved by this invention compared to the prior art is as follows:
[0047] 1. This invention provides an integrated hydrogen cycle cooling system for fuel cells. By collecting multi-dimensional state data of the hydrogen cycle-cooling system and combining it with a dynamic collaborative control strategy of a reinforcement learning decision module, the accuracy and stability of temperature control are significantly improved. The Markov decision process model is used to predict instantaneous heat generation changes caused by load mutations, and temperature fluctuations are quickly suppressed through joint regulation of the hydrogen cycle and cooling system. Compared with traditional PID control, the temperature overshoot is reduced under dynamic operating conditions, ensuring that the fuel cell stack always operates within the optimal temperature range, extending its service life and improving energy conversion efficiency.
[0048] 2. This invention provides an integrated hydrogen circulation cooling system for fuel cells. By constructing a multi-objective reward function with temperature stability as the core, the system ensures thermal management performance while optimizing the energy consumption of the hydrogen circulation pump and cooling system. Furthermore, the reinforcement learning agent dynamically adjusts the hydrogen circulation rate, coolant flow rate, and fan speed to reduce ineffective power consumption and lower overall parasitic power. In addition, the quantitative evaluation of the latent heat utilization rate of phase change materials further improves the heat recovery efficiency and significantly enhances the overall energy efficiency. Attached Figure Description
[0049] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0050] Figure 1 This is a schematic diagram of the working process of a fuel cell hydrogen cycle cooling integrated system according to the present invention;
[0051] Figure 2 This is a data flow diagram of the system functional modules of a fuel cell hydrogen cycle cooling integrated system according to the present invention. Detailed Implementation
[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0053] Example 1, such as Figure 1 , Figure 2 As shown, the present invention provides an integrated hydrogen cycle cooling system for fuel cells, including a hydrogen cycle cooling management center, which is communicatively connected to the following modules:
[0054] The multi-source state sensing module is used to collect multi-dimensional state data of the hydrogen cycle-cooling system in real time, extract state features, and form a state feature sequence table. The multi-source state sensing module includes a state monitoring unit and a feature extraction unit.
[0055] The condition monitoring unit is used to collect condition data from the hydrogen circulation system and the cooling system. The hydrogen circulation system's condition data includes hydrogen circulation pump speed, hydrogen inlet pressure / flow rate, internal hydrogen partial pressure, and water vapor concentration, capturing fluctuations in hydrogen supply caused by sudden load changes. The cooling system's condition data includes coolant temperature / flow rate, radiator fan speed, stack surface temperature distribution, and phase change material (PCM) phase change state, etc. These data are pre-processed to form a comprehensive dataset. A sensor network deployed in the fuel cell's hydrogen circulation and cooling systems collects the condition data of these systems in real time. The system's state data includes hydrogen circulation pump speed, hydrogen inlet pressure / flow rate, internal hydrogen partial pressure, and water vapor concentration, capturing hydrogen supply fluctuations caused by load abrupt changes. The cooling system's state data includes coolant temperature / flow rate, radiator fan speed, stack surface temperature distribution, and phase change material (PCM) phase change state. The raw state data is preprocessed, including filtering and noise reduction, outlier removal, and time-series alignment, to eliminate sensor drift and interference. Short-term statistical features are extracted using the sliding window method to form a standardized time-series dataset. The preprocessed state data is then categorized into hydrogen circulation system and cooling system to form a comprehensive dataset.
[0056] The specific tasks of the condition monitoring unit are as follows: Through a distributed sensor network deployed in the fuel cell hydrogen circulation system and cooling system, it achieves real-time acquisition of comprehensive condition data. For the hydrogen circulation system, it monitors dynamic parameters such as hydrogen circulation pump speed, hydrogen inlet pressure, and flow rate to capture instantaneous fluctuations in hydrogen supply caused by load changes. Simultaneously, it quantifies hydrogen utilization efficiency and internal humidity changes through sensors for hydrogen partial pressure and water vapor concentration within the fuel cell stack. On the cooling system side, the sensor network covers key indicators such as coolant temperature, flow rate, and radiator fan speed, reflecting the thermodynamic state of the cooling medium in real time. The surface temperature distribution of the fuel cell stack is acquired through an array of temperature sensors. Combined with phase change material (PCM) phase change state monitoring, a dynamic response map of the cooling system to fuel cell stack thermal management is constructed. The unit processes the collected raw condition data... Multi-stage preprocessing is performed to eliminate noise and interference. An adaptive filtering algorithm is used to denoise the sensor signals, suppressing random fluctuations caused by electromagnetic interference and mechanical vibration. Outliers are removed based on statistical thresholds to avoid data distortion caused by sensor failures or transient impacts. A time-series alignment algorithm is used to synchronize the timestamps of multi-sensor data to eliminate time-series misalignments caused by communication delays or differences in sampling periods. After preprocessing, the data enters the short-time feature extraction stage: statistical features, including mean, variance, maximum / minimum difference, and slope, are calculated using the sliding window method to quantify the instantaneous change trend of state parameters. At the same time, frequency domain features are extracted by combining Fourier transform or wavelet analysis to capture periodic fluctuation patterns and form a standardized time series dataset. The preprocessed data needs to be classified by subsystem and fused into a comprehensive dataset by time dimension.
[0057] The feature extraction unit is used to perform feature analysis on the collected multi-dimensional state data, extract state features related to hydrogen cycle cooling, and integrate them to form a state feature sequence table. The state features include hydrogen supply fluctuation rate, stack temperature gradient, hydrogen-cold coupling efficiency, and latent heat utilization rate of phase change material. The unit performs time-series synchronization on the collected state data of the hydrogen cycle system and cooling system to eliminate communication delay and sampling period differences. The data stream is segmented by a sliding window to ensure that the state data of each subsystem are on a unified time scale. Feature analysis is performed on the preprocessed state data of the hydrogen cycle system and cooling system to extract state features. For the hydrogen cycle system, the hydrogen supply fluctuation rate is calculated. For the cooling system, the stack temperature gradient and latent heat utilization rate of phase change material are quantified. At the same time, through the synergistic analysis of hydrogen utilization rate and cooling power, the hydrogen-cold coupling efficiency feature is extracted. The extracted hydrogen supply fluctuation rate, stack temperature gradient, hydrogen-cold coupling efficiency, and latent heat utilization rate of phase change material are normalized to eliminate dimensional differences and integrated into a state feature sequence table according to the time dimension. Each record contains a timestamp and four types of feature values.
[0058] The specific work of the feature extraction unit is as follows: It performs time synchronization of the state data of the hydrogen cycle system and the cooling system to eliminate timing misalignments caused by communication delays and sampling period differences. Since the sensor sampling periods and communication protocols of the two subsystems differ, hardware clock synchronization (IEEE 1990-2000) is required. The 1588 protocol and software interpolation algorithms eliminate time deviations. Hardware synchronization ensures that the clock error of all sensors is less than 1μs. The software layer uses cubic spline interpolation for low-frequency signals, resampling them to a high-frequency time base to achieve timestamp alignment. A sliding window method is used to segment the data stream under a unified time scale. The window length is set according to the dynamic characteristics of the system, and the step size is 50% of the window length, ensuring that feature extraction covers local dynamic changes while preserving temporal continuity. Based on temporal synchronization, state characteristics are calculated for the hydrogen circulation system and the cooling system separately. For the hydrogen circulation system, the dynamic fluctuation of hydrogen supply is quantified by calculating the ratio of the standard deviation to the mean of the hydrogen inlet flow rate within the sliding window to obtain the hydrogen supply fluctuation rate, reflecting the degree of instantaneous supply instability caused by load changes. For the cooling system, the stack temperature gradient is quantified by the ratio of the maximum temperature difference to the average temperature of the surface temperature sensor array to characterize the uniformity of heat distribution. Phase change material (PCM) The latent heat utilization rate is calculated by combining the phase change temperature threshold and temperature-time integral to evaluate the actual contribution of PCM in thermal management. Through covariance analysis of hydrogen utilization rate (calculated based on hydrogen partial pressure and flow rate) and cooling power (calculated based on coolant temperature difference and flow rate), the hydrogen-cooling coupling efficiency characteristics are extracted to reveal the dynamic synergistic relationship between hydrogen supply and thermal management. The hydrogen supply fluctuation rate, stack temperature gradient, hydrogen-cooling coupling efficiency and latent heat utilization rate of phase change material are normalized. Among them, the state characteristics of the hydrogen cycle system are normalized by Min-Max and mapped to the [0,1] interval to retain its relative change range. The state characteristics of the cooling system are dynamically normalized in combination with the design boundary to highlight the sensitivity under key operating conditions. The binary characteristics (PCM phase change state) are converted into numerical form through unique thermal encoding. After normalization, the four types of feature values are integrated according to the time dimension to generate a state feature sequence table. Each record contains a unified timestamp and a corresponding feature value vector.
[0059] The Markov Decision Process Modeling Module is used to construct Markov decision process models, define state space and action space, and construct a multi-objective reward function with temperature stability as the core. The Markov Decision Process Modeling Module includes a state space definition unit and a reward function design unit.
[0060] The state space definition unit is used to abstract the dynamic coupling relationship of the hydrogen cycle-cooling system into a Markov decision process model, and to map the state features into discretized state vectors. It defines the state space and action space, takes the state feature sequence list covering the state features of the hydrogen cycle system and the cooling system as input, and quantifies the dynamic dependence between variables through covariance analysis and mutual information calculation. It constructs a state transition probability matrix to form a coupled dynamic framework of Markov decision process (MDP) to ensure that the state evolution conforms to physical constraints. An adaptive binning algorithm is used for the state features to divide discrete intervals based on feature distribution density and key operating condition thresholds, and maps continuous values into discrete state symbols. At the same time, it retains the dynamic sensitivity of the original state features and generates discretized state vectors to reduce model complexity. The state space is constructed based on the discrete state vectors, and the action space is defined as discrete control commands for hydrogen cycle pump speed adjustment and coolant flow control. Combined with system safety boundary constraints, an effective action set is generated to ensure that state transition and action execution meet the real-time and robustness requirements of fuel cell operation.
[0061] The specific tasks of the state space definition unit are as follows: Integrating the state characteristic sequence lists of the hydrogen cycle system and the cooling system, and quantifying the dynamic dependencies between variables through covariance analysis and mutual information calculation. Covariance analysis focuses on linear correlations, calculating the covariance matrix between features and extracting principal components to reveal the coordinated changing trends of hydrogen supply volatility and cooling power. Mutual information measures the strength of nonlinear correlations through KL divergence, capturing the implicit coupling relationship between the stack temperature gradient and the PCM latent heat utilization rate, and then constructing a state transition probability matrix to define the conditional probability distribution of the system's transition from the current state to the next state, forming a coupled dynamics framework of a Markov decision process (MDP). The state space definition unit uses an adaptive binning algorithm to discretize the state characteristics, reducing the complexity of the continuous state space. The adaptive binning algorithm dynamically divides intervals based on feature distribution density and key operating condition thresholds: for high-density regions, it uses... Equal-frequency partitioning preserves local dynamic details. For low-density areas, forced partitioning is performed based on the design boundary (maximum withstand temperature of the fuel cell stack) to ensure sensitivity to critical operating conditions. The partitioning results are verified for interval independence through chi-square test to avoid information loss due to excessive discretization. Finally, each continuous state feature is mapped to a discrete state symbol, generating a discretized state vector. A finite set of states is constructed based on the discretized state vector, and the action space is defined as a set of discrete commands for hydrogen circulation pump speed regulation and coolant flow control. The action design of the action space follows the fuel cell control logic: the hydrogen circulation pump speed is divided into three levels: low speed, medium speed, and high speed, corresponding to different hydrogen supply requirements; the coolant flow control adopts three levels of regulation: minimum flow, rated flow, and maximum flow, matching the changes in fuel cell stack heat load. At the same time, the system safety boundary constraints, namely the lower limit of hydrogen concentration and the upper limit of coolant pressure, are integrated to filter out invalid action combinations and generate an effective action set.
[0062] The reward function design unit is used to construct a multi-objective reward function with temperature stability as its core. This function includes a temperature fluctuation suppression term, a system energy consumption optimization term, and an actuator action smoothing term. A weighted summation method is used to balance the conflicting objectives. Based on the control requirements of the fuel cell system, the multi-objective reward function is decomposed into three reward terms: a temperature fluctuation suppression term (quantifying thermal stability through stack temperature variance), a system energy consumption optimization term (normalizing the energy consumption of the hydrogen circulation pump and coolant flow), and an actuator action smoothing term (penalizing the abrupt changes in action between adjacent time steps). Each reward term is monotonically correlated with its corresponding objective. Specifically, the temperature fluctuation suppression term uses a negative exponential function to enhance sensitivity, and the system energy consumption optimization term uses a reciprocal form to avoid numerical overflow. The actuator motion smoothing term is achieved by integrating the absolute value of the motion difference. Each reward term is normalized to eliminate dimensional differences. The temperature fluctuation suppression term is normalized based on the maximum variance of historical data. The system energy consumption optimization term is calibrated according to the rated power. The actuator motion smoothing term is linearly scaled by the motion range and the weights are dynamically allocated using the analytic hierarchy process. Among them, the temperature fluctuation suppression term has the highest weight (≥50%). The weights of the system energy consumption optimization term and the actuator motion smoothing term are adaptively adjusted according to the operating conditions. A hard constraint penalty term is introduced. When the temperature exceeds the limit or the safety boundary is violated, the reward is forced to be negative. Then, the normalized reward terms are weighted and summed to construct the total reward function, which is output to the reinforcement learning decision module.
[0063] The specific work of the reward function design unit is as follows: The control objective of the fuel cell system is decomposed into three independent reward items, corresponding to thermal management stability, energy efficiency optimization, and actuator life protection, namely, temperature fluctuation suppression, system energy consumption optimization, and actuator action smoothing. The temperature fluctuation suppression item calculates the variance of the stack surface temperature sequence and maps it using a negative exponential function, making the reward highly sensitive to temperature fluctuations; even small deviations can trigger significant penalties. The system energy consumption optimization item is based on the real-time power consumption of the hydrogen circulation pump and cooling system, and normalized inversely to the rated power to avoid numerical truncation under high energy consumption conditions. The actuator action smoothing item suppresses mechanical wear caused by frequent action switching by integrating the absolute value of the difference between adjacent control commands. Each reward item is monotonically related to the objective, ensuring consistency in the optimization direction. Each reward item is then normalized. To eliminate dimensional differences, the temperature fluctuation suppression term is linearly scaled to the [0,1] interval based on the historical maximum temperature variance. The system energy consumption optimization term is directly calibrated with the rated power as the upper limit. The actuator action smoothing term is linearly normalized according to the physical range of the actuator. The weight allocation adopts the analytic hierarchy process, prioritizing temperature stability (fixed weight ≥50%), and the remaining weights are dynamically adjusted according to the operating conditions: the weight of the system energy consumption optimization term is increased under high load, and the action smoothness is emphasized under steady-state conditions. A hard constraint penalty term is introduced to ensure system safety. When the stack temperature exceeds the design threshold or the hydrogen concentration / coolant pressure violates the safety boundary, a large negative reward is directly assigned. This penalty term has the highest priority and covers all other reward terms. Finally, the total reward function integrates the normalized terms through weighted summation and outputs the total reward to the reinforcement learning decision module as a direct feedback signal for policy optimization.
[0064] The reinforcement learning decision module trains a reinforcement learning agent based on a Markov decision process model, generates a coordinated control strategy for the hydrogen cycle-cooling system, and decomposes the continuous actions output by reinforcement learning into independent control commands for the hydrogen cycle and cooling system.
[0065] The hydrogen cycle execution control module is used to drive the hydrogen circulation pump frequency converter, adjust the hydrogen recirculation rate, and control the opening of the hydrogen injection valve according to the hydrogen cycle independent control command output by the reinforcement learning decision module, so as to optimize the hydrogen distribution inside the stack and achieve precise matching of hydrogen supply.
[0066] The cooling system execution control module is used to adjust the speed of the electronic water pump and the opening of the electric three-way valve according to the independent control commands of the cooling system, control the flow rate and direction of the coolant, and drive the radiator fan to adjust the speed using PWM to enhance convective heat transfer. Combined with PID control and feedforward compensation, it eliminates the overshoot phenomenon caused by the inertia of the cooling system.
[0067] Example 2, as Figure 1 , Figure 2As shown, based on Embodiment 1, the present invention provides a technical solution: Preferably, the reinforcement learning decision module specifically includes:
[0068] A reinforcement learning agent is trained based on a constructed Markov Decision Process (MDP) model. The policy network parameters are optimized through environmental interaction, enabling the agent to learn and generate cooperative control actions that satisfy multi-objective constraints. A mapping is performed from state to continuous actions. The continuous cooperative actions output by the reinforcement learning agent are mapped into independent control commands for the hydrogen cycle and cooling system through an action decomposition mechanism. The decomposition rules are designed based on the system coupling relationship. Weighted allocation ensures the physical feasibility of the actions while maintaining the loyalty of the decomposed commands to the original cooperative objectives, avoiding performance degradation due to action segmentation. The independent control commands are input into the fuel cell system for execution. The state feedback is collected in real time and the instant reward is calculated. The experience replay buffer of the Markov Decision Process is updated. The weights of the reinforcement learning policy are dynamically adjusted based on new data, and the action decomposition rules are optimized to adapt to changes in operating conditions, continuously improving the robustness and adaptability of the cooperative control of the hydrogen cycle-cooling system.
[0069] The specific tasks of the reinforcement learning decision-making module are as follows: Based on the constructed Markov Decision Process (MDP) model, the module optimizes the policy network parameters through environmental interaction, learning the mapping relationship from system states to continuous control actions. The states include key parameters such as stack temperature, hydrogen concentration, and coolant pressure, while the action space consists of continuous cooperative control commands. During training, the reinforcement learning agent interacts with the fuel cell system environment through an explore-exploit mechanism, collecting state transition samples and calculating immediate rewards. An experience replay mechanism breaks down data correlations, and the gradient descent algorithm is used to update the policy network weights. The policy network employs a deep neural network structure, fitting complex state-action relationships through nonlinear transformations to output continuous cooperative actions that satisfy multi-objective constraints. An action decomposition mechanism maps the continuous cooperative actions output by the policy network into independent control commands for the hydrogen circulation system and the cooling system. The decomposition rules are designed based on the system coupling relationship, considering the physical weights of the hydrogen circulation pump speed on stack temperature and the contribution ratio of coolant flow rate to heat dissipation efficiency. To address constraints, a weighted allocation method is used to ensure the physical feasibility of the decomposed instructions. Simultaneously, the decomposition process must maintain loyalty to the original collaborative goals. Loyalty constraints are introduced to prevent performance degradation due to action segmentation. Furthermore, the decomposition rules support dynamic adjustment, allowing for real-time optimization of weight allocation based on changing operating conditions, ensuring effective collaboration of independent instructions even in complex scenarios. After independent control instructions are input into the fuel cell system and executed, real-time state feedback is collected, and immediate rewards are calculated based on a preset reward function to assess the degree to which the current action satisfies multi-objective constraints. New data is stored in an experience replay buffer to update the state transition probability distribution of the Markov decision process model, simultaneously triggering dynamic adjustments to the policy network weights. During optimization, a meta-learning mechanism dynamically adjusts the priority of each objective in the reward function based on operating condition characteristics, ensuring policy robustness. The optimization of the action decomposition rules is achieved through reinforcement learning subtasks, using the collaborative effect of the decomposed instructions as feedback signals to continuously iterate the rule parameters, ultimately forming an efficient collaborative control strategy adaptable to all operating conditions.
[0070] The hydrogen cycle execution control module specifically includes:
[0071] The hydrogen cycle execution control module receives independent control commands for hydrogen cycle output from the reinforcement learning decision module. It performs precise analysis of the control commands to determine the drive parameters of the hydrogen cycle pump frequency converter and the set value of the hydrogen injection valve opening. Based on the analyzed control commands, it drives the hydrogen cycle pump frequency converter to adjust its operating frequency according to the set parameters, thereby changing the hydrogen recirculation rate. At the same time, it controls the hydrogen injection valve to achieve the set value, optimizing the hydrogen distribution inside the fuel cell stack. During the adjustment process, it continuously monitors the relevant parameters of hydrogen cycle, evaluates the matching of hydrogen supply, and feeds back the actual effect to the reinforcement learning decision module so that it can adjust its strategy to achieve continuous and precise matching of hydrogen supply.
[0072] The specific functions of the hydrogen circulation execution control module are as follows: The module receives independent hydrogen circulation control commands output by the reinforcement learning decision module and converts them into executable low-level equipment control parameters. These commands include the target speed of the hydrogen circulation pump and the target opening degree of the hydrogen injection valve. Dynamic calibration is performed based on the current system state to ensure the physical feasibility of the commands. The analysis process must consider the safe operating range of the equipment; for example, the speed of the hydrogen circulation pump must be limited to the rated range, while the opening degree of the injection valve needs to be constrained and optimized in conjunction with the hydrogen flow rate requirements. After analysis, specific inverter drive signals and valve position control commands are generated and sent to the actuator. After command analysis, the inverter drives the hydrogen circulation pump, adjusting its speed according to the set frequency to change the hydrogen recirculation rate, ensuring that the anode hydrogen supply matches the stack load requirements. Simultaneously, the opening degree of the hydrogen injection valve is adjusted in real time according to the analyzed values to optimize the hydrogen distribution within the stack flow channels and avoid localized starvation. During the adjustment process… Real-time monitoring of key parameters, such as the actual speed of the circulating pump, the feedback opening of the injection valve, hydrogen flow rate, and pressure fluctuations, ensures execution accuracy. If a deviation between the actual value and the set value is detected, a feedforward-feedback composite control strategy is used for dynamic compensation. The inverter output is fine-tuned through a PID algorithm or the valve position command is corrected in advance using predictive control to reduce the impact of response lag. After the adjustment is completed, the hydrogen supply effect is comprehensively evaluated. Data on hydrogen circulation efficiency, hydrogen concentration distribution inside the fuel cell stack, and anode pressure fluctuations are collected through a sensor network to quantify the actual effect of the current control command. Evaluation indicators include hydrogen utilization rate, supply stability, and synergistic efficiency with the cooling system. The monitored data is compared and analyzed with the target value to generate an execution effect report, which is fed back to the reinforcement learning decision module as a key basis for strategy optimization. If a continuous mismatch between hydrogen supply and demand is found, i.e., flow rate fluctuation exceeds the limit or concentration is uneven, a real-time alarm is triggered, and the decision module is advised to adjust the control strategy.
[0073] The cooling system control module specifically includes:
[0074] The system receives control commands from the reinforcement learning module, analyzes the target speed of the electronic water pump, the opening of the electric three-way valve, and the PWM duty cycle of the radiator fan. It dynamically trims the commands based on the physical limits of the equipment to ensure that the control commands are within the safe operating range, avoiding hardware overload or mechanical damage. Based on the PID algorithm, it adjusts the water pump speed and valve opening, and combines a feedforward compensation model to predict changes in cooling demand. It adjusts the control quantity in advance to offset system inertia, and drives the fan speed through the PWM signal to enhance the radiator's convective heat transfer efficiency. It simultaneously optimizes the coolant flow distribution and temperature uniformity, monitors key parameters such as coolant temperature, flow rate, and pressure in real time, and calculates deviations by comparing them with target values. If overshoot or response lag is detected, it adopts PID parameter adaptive adjustment or a feedforward-feedback composite control strategy to fine-tune the actuator output, ensuring that the cooling system quickly and stably converges to the target state, and generates execution effect data to feed back to the reinforcement learning decision module.
[0075] The specific work of the cooling system execution control module is as follows: The cooling system execution control module receives continuous control commands output by the reinforcement learning decision module, including the target speed of the electric water pump, the opening degree of the electric three-way valve, and the PWM duty cycle of the radiator fan. During the parsing process, the physical feasibility of the commands is verified, and dynamic trimming is performed based on the safe operating range of the equipment. Specifically, the electric water pump speed must be limited within the rated range to avoid mechanical wear caused by overspeed; the opening degree of the electric three-way valve is optimized according to the coolant flow demand to prevent pressure surges caused by extreme opening degrees; the fan PWM duty cycle is adjusted based on the radiator heat load curve to ensure maximum convective heat transfer efficiency. The trimmed control commands must meet the dynamic constraints of the cooling system, namely the flow-pressure balance relationship and temperature gradient limitations. Simultaneously, a slope limiting algorithm is used to smooth command surges and avoid frequent actuator actions. After parsing, the corresponding inverter drive signal, valve position control command, and PWM waveform are generated and sent to the underlying actuator. Based on the parsed control commands, the cooling system execution control module uses a PID algorithm to adjust the electric water pump speed and the three-way valve opening to achieve precise control of the coolant flow. A feedforward compensation model is introduced based on the fuel cell stack heat generation rate... The system predicts changes in cooling demand and adjusts control parameters in advance. The radiator fan is driven by a PWM signal, with a duty cycle nonlinearly mapped to the heat load to enhance convective heat transfer efficiency. Simultaneously, a three-way valve proportionally distributes coolant flow to optimize fuel cell temperature uniformity. During execution, the system monitors coolant inlet and outlet temperatures, flow rates, and pressure parameters in real time, and dynamically corrects the control output through a feedback closed loop. If overshoot or response lag is detected, it triggers adaptive adjustment of PID parameters or a feedforward-feedback composite control strategy to ensure the system quickly and stably converges to the target state. The cooling system execution control module collects key parameters after execution in real time, including coolant temperature drop gradient, flow stability, and pressure fluctuation rate, quantifying the actual effect of the current control command. By comparing the deviation with the target value, it calculates the deviation and evaluates cooling power, temperature uniformity, and energy consumption indicators. If the deviation continues to exceed the limit, an abnormal alarm is generated and fed back to the reinforcement learning decision module to trigger strategy optimization. Execution effect data is uploaded in a standardized format, including control commands, actual outputs, and environmental state variables, providing high-confidence samples for strategy iteration. At the same time, the module supports dynamic adjustment of the control mode, switching to a preset safety strategy in case of hardware failure or extreme conditions to ensure system robustness.
[0076] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A fuel cell hydrogen cycle cooling integrated system, comprising a hydrogen cycle cooling management center, characterized in that, The hydrogen cycle cooling management center has the following communication modules: The multi-source state sensing module is used to collect multi-dimensional state data of the hydrogen cycle-cooling system in real time, extract state features, and form a state feature sequence table. The Markov Decision Process Modeling Module is used to construct Markov decision process models, define the state space and action space, and construct a multi-objective reward function with temperature stability as the core. The reinforcement learning decision module trains a reinforcement learning agent based on a Markov decision process model, generates a coordinated control strategy for the hydrogen cycle-cooling system, and decomposes the continuous actions output by reinforcement learning into independent control commands for the hydrogen cycle and cooling system. The hydrogen cycle execution control module is used to drive the hydrogen circulation pump frequency converter and control the opening degree of the hydrogen injection valve according to the hydrogen cycle independent control command output by the reinforcement learning decision module. The cooling system execution control module is used to adjust the speed of the electric water pump and the opening of the electric three-way valve according to the independent control commands of the cooling system, and drive the radiator fan PWM speed control. The multi-source state perception module includes a state monitoring unit and a feature extraction unit; The status monitoring unit is used to collect status data of the hydrogen cycle system and the cooling system, and preprocess the data to form a comprehensive dataset. The feature extraction unit is used to perform feature analysis on the collected multi-dimensional state data, extract state features related to hydrogen cycle cooling, and integrate them to form a state feature sequence list. The state features include hydrogen supply fluctuation rate, stack temperature gradient, hydrogen-cold coupling efficiency, and latent heat utilization rate of phase change material. The Markov decision process modeling module includes a state space definition unit and a reward function design unit; The state space definition unit is used to abstract the dynamic coupling relationship of the hydrogen cycle-cooling system into a Markov decision process model, and to map the state features into discretized state vectors, while defining the state space and action space. The reward function design unit is used to construct a multi-objective reward function with temperature stability as the core, which includes a temperature fluctuation suppression term, a system energy consumption optimization term, and an actuator action smoothing term.
2. The fuel cell hydrogen cycle cooling integrated system according to claim 1, characterized in that: In the status monitoring unit, the process of forming a comprehensive dataset is as follows: By deploying a sensor network in the hydrogen circulation system and cooling system of the fuel cell, the status data of the hydrogen circulation system and cooling system are collected in real time. The status data of the hydrogen circulation system includes the hydrogen circulation pump speed, hydrogen inlet pressure / flow rate, hydrogen partial pressure inside the stack, and water vapor concentration, capturing the hydrogen supply fluctuations caused by load changes. The status data of the cooling system includes coolant temperature / flow rate, radiator fan speed, stack surface temperature distribution, and phase change state of the phase change material. The raw state data is preprocessed, including filtering and denoising, outlier removal and time series alignment, and short-term statistical features are extracted using the sliding window method to form a standardized time series dataset. The preprocessed state data is categorized into hydrogen cycle system and cooling system to form a comprehensive dataset.
3. The fuel cell hydrogen cycle cooling integrated system according to claim 2, characterized in that: In the feature extraction unit, the process of forming the state feature sequence list is as follows: The collected status data of the hydrogen cycle system and cooling system are synchronized in time, and the data stream is divided by a sliding window so that the status data of each subsystem are on a unified time scale. Feature analysis was performed on the state data of the pre-processed hydrogen cycle system and cooling system to extract state features. Specifically, for the hydrogen cycle system, the hydrogen supply fluctuation rate was calculated, and for the cooling system, the stack temperature gradient and the latent heat utilization rate of the phase change material were quantified. At the same time, the hydrogen-cooling coupling efficiency feature was extracted through the synergistic analysis of hydrogen utilization rate and cooling power. The extracted hydrogen supply volatility, stack temperature gradient, hydrogen-cold coupling efficiency, and latent heat utilization rate of phase change material were normalized and integrated into a state feature sequence table according to the time dimension. Each record contains a timestamp and four types of feature values.
4. The fuel cell hydrogen cycle cooling integrated system according to claim 1, characterized in that: The state space definition unit specifically includes: Using a state feature sequence list covering the state characteristics of the hydrogen cycle system and the cooling system as input, the dynamic dependencies between variables are quantified through covariance analysis and mutual information calculation, a state transition probability matrix is constructed, and a coupled dynamics framework of the Markov decision process is formed. An adaptive binning algorithm is used for the state features. Based on the feature distribution density and key working condition thresholds, discrete intervals are divided, and continuous values are mapped to discrete state symbols to generate discretized state vectors. A state space is constructed based on discrete state vectors, and the action space is defined as discrete control commands for hydrogen circulation pump speed regulation and coolant flow control. The effective action set is generated by combining system safety boundary constraints.
5. The fuel cell hydrogen cycle cooling integrated system according to claim 1, characterized in that: The reward function design unit specifically includes: Based on the control requirements of the fuel cell system, the multi-objective reward function is decomposed into three reward items: temperature fluctuation suppression, system energy consumption optimization, and actuator action smoothing. Each reward item is monotonically correlated with its corresponding objective. Each reward item is normalized, the temperature fluctuation suppression item is normalized based on the maximum variance of historical data, the system energy consumption optimization item is calibrated according to the rated power, and the actuator motion smoothing item is linearly scaled by the motion range and the weights are dynamically allocated using the analytic hierarchy process. A hard constraint penalty term is introduced, which forces the reward to be negative when the temperature exceeds the limit or the safety boundary is violated. Then, the normalized reward terms are summed according to their weights to construct the total reward function, which is then output to the reinforcement learning decision module.
6. The fuel cell hydrogen cycle cooling integrated system according to claim 1, characterized in that: The reinforcement learning decision module specifically includes: The reinforcement learning agent is trained based on the constructed Markov decision process model. The policy network parameters are optimized through environmental interaction, enabling it to learn and generate cooperative control actions that satisfy multi-objective constraints, and to map states to continuous actions. The continuous cooperative actions output by the reinforcement learning agent are mapped into independent control commands for the hydrogen cycle and cooling system through an action decomposition mechanism. The decomposition rules are designed based on the system coupling relationship, while maintaining the loyalty of the decomposed commands to the original cooperative target. Independent control commands are input into the fuel cell system for execution, real-time status feedback is collected and instant rewards are calculated, the experience replay buffer of the Markov decision process is updated, the weights of the reinforcement learning strategy are dynamically adjusted based on new data, and the action decomposition rules are optimized to adapt to changes in operating conditions.
7. The fuel cell hydrogen cycle cooling integrated system according to claim 6, characterized in that: The hydrogen cycle execution control module specifically includes: The hydrogen cycle execution control module receives the independent control command for hydrogen cycle output by the reinforcement learning decision module, performs precise analysis of the control command, and clarifies the drive parameters of the hydrogen cycle pump frequency converter and the set value of the hydrogen injection valve opening. Based on the analyzed control instructions, drive the frequency converter of the hydrogen circulation pump, adjust its operating frequency according to the set parameters, change the hydrogen recirculation rate, and control the hydrogen injection valve to make its opening reach the set value. During the adjustment process, hydrogen cycle-related parameters are continuously monitored, hydrogen supply matching is assessed, and the actual results are fed back to the reinforcement learning decision-making module.
8. The fuel cell hydrogen cycle cooling integrated system according to claim 6, characterized in that: The cooling system control module specifically includes: It receives control commands output by the reinforcement learning module, analyzes the target speed of the electronic water pump, the opening degree of the electric three-way valve and the PWM duty cycle of the radiator fan, and dynamically trims them according to the physical limits of the equipment to keep the control commands within the safe operating range. The pump speed and valve opening are adjusted based on the PID algorithm. Combined with the feedforward compensation model to predict changes in cooling demand, the control quantity is adjusted in advance to counteract system inertia. The fan speed is driven by the PWM signal to enhance the radiator's convective heat transfer efficiency and simultaneously optimize the coolant flow distribution and temperature uniformity. The system monitors key parameters such as coolant temperature, flow rate, and pressure in real time, calculates deviations by comparing them with target values, and if overshoot or response lag is detected, it adopts PID parameter adaptive adjustment or feedforward-feedback composite control strategy to fine-tune the actuator output and generate execution effect data to feed back to the reinforcement learning decision module.