A method and system for constructing a proton exchange membrane fuel cell multi-physical field virtual sensor based on a spatial attention mechanism U-Net

By constructing a spatial attention mechanism U-Net network and combining it with prior information about the serpentine flow channel structure, high-precision, real-time virtual sensing of multi-physics fields inside the fuel cell is achieved. This solves the problems of insufficient reconstruction accuracy and unstable training in existing technologies, and realizes efficient and stable fuel cell state monitoring.

CN122330731APending Publication Date: 2026-07-03UNIV OF SCI & TECH OF CHINA

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UNIV OF SCI & TECH OF CHINA
Filing Date
2026-04-14
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing fuel cell condition monitoring methods suffer from problems such as intrusive detection affecting performance, excessively long physical model simulation time, insufficient accuracy of conventional deep learning reconstruction, and instability of adversarial network training, making it difficult to achieve high-resolution, high-real-time, and non-intrusive multiphysics reconstruction.

Method used

A U-Net network, a spatial attention mechanism integrating the prior of bipolar plate serpentine flow channel structure, is constructed. A three-level physical coupling series reconstruction architecture of pressure, water content, and current density is adopted to realize high-precision and high-real-time virtual sensing of multi-physics fields inside the fuel cell through external signals.

Benefits of technology

It achieves high-precision, real-time, and non-invasive virtual sensing of multi-physics fields inside fuel cells, with a single reconstruction time of less than 0.1 seconds, improved reconstruction accuracy, stable training process, high stability in engineering deployment, low cost, and no interference with power generation performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122330731A_ABST
    Figure CN122330731A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for constructing a multiphysics virtual sensor for proton exchange membrane fuel cells based on the spatial attention mechanism U-Net, belonging to the interdisciplinary field of fuel cell state monitoring and deep learning. This invention constructs a SAM-U-Net network embedding a spatial attention module and integrating prior knowledge of a binary mask in a bipolar serpentine flow channel. It employs a three-level physical coupling series virtual sensor architecture of pressure, water, and current density, using three easily measurable external signals—fuel cell inlet and outlet pressure and average current—as input to achieve 128×128 high-resolution reconstruction of the internal multiphysics distribution. Batch normalization layers are used during network training to improve convergence stability. This invention requires no additional sensors, with a single reconstruction time of less than 0.1 seconds. Accuracy at flow channel edges and turning areas is significantly improved, with MAPE reduced by 1.28%-1.47% compared to traditional U-Net, and SSIM significantly improved. It enables real-time, non-intrusive, and high-precision virtual sensing of the internal multiphysics field of fuel cells, suitable for online monitoring and digital twin applications in fuel cells.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of proton exchange membrane fuel cell state monitoring technology, and in particular to a multiphysics virtual sensing method and system for fuel cells based on the spatial attention mechanism U-Net. Background Technology

[0002] In the field of proton exchange membrane fuel cell (PEMFC) operating condition monitoring, existing technologies are mainly divided into two categories: hard measurement methods and soft measurement methods. Hard measurement technologies include non-destructive testing methods such as ultrasonic Lamb wave detection, magnetic resonance imaging, X-ray imaging, and charge-coupled device (CCD) imaging, as well as invasive or contact measurement methods such as embedded temperature and humidity sensors, printed circuit board-type local current sensors, and electrochemical impedance spectroscopy. These methods require direct installation inside the fuel cell stack or the addition of dedicated testing equipment, which is not only structurally complex and costly, but also prone to interfering with the internal flow field and power generation performance of the fuel cell.

[0003] The physical model-driven approach in soft sensing technology simulates the internal operating state of a battery by constructing a multiphysics coupled numerical model, including electrochemical models, multiphase flow models, and transmembrane mass transfer models. Existing fuel cell digital twins mostly employ a three-dimensional coupled one-dimensional simulation architecture. While this offers high accuracy, it suffers from high computational load, long processing time, and poor real-time performance. The excessively long simulation time per run makes it difficult to meet online monitoring requirements. Conversely, excessively simplifying the model to improve real-time performance leads to a significant decrease in monitoring accuracy.

[0004] In recent years, soft measurement methods based on artificial intelligence have been widely applied, with artificial neural networks and support vector machines used for parameter prediction, lifetime estimation, and control optimization. Deep learning structures such as U-Net (a convolutional neural network with an encoder-decoder structure), generative adversarial networks, and residual networks have shown great potential in image segmentation and field distribution reconstruction. However, conventional deep learning models do not specifically integrate prior structural information about the serpentine flow channel of a fuel cell bipolar plate, resulting in low reconstruction accuracy at key locations such as flow channel turning areas and edge areas, making it difficult to achieve high-resolution multiphysics synchronous reconstruction. Furthermore, generative adversarial networks rely on adversarial training between the generator and discriminator, which is prone to oscillations and instability during training, and the model structure and parameter tuning complexity are high.

[0005] In summary, achieving high-resolution, high-real-time, and high-stability reconstruction of internal multi-physics fields using only a small number of easily measurable external signals without adding embedded sensors or changing the battery structure is a technical problem that urgently needs to be solved in current engineering applications. Existing methods are difficult to simultaneously meet the requirements of accuracy, speed, and practicality. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method and system for constructing a multi-physics virtual sensor for proton exchange membrane fuel cells based on the spatial attention mechanism U-Net. This solves the technical problems existing in the internal state monitoring of fuel cells, such as the impact of invasive detection on performance, excessively long simulation time of physical models, insufficient accuracy of conventional deep learning reconstruction, and instability of adversarial network training.

[0007] This invention constructs a U-Net network, a spatial attention mechanism that integrates the prior knowledge of a bipolar plate serpentine flow channel structure, and adopts a three-level physical coupling series reconstruction architecture of pressure, water content, and current density to achieve high-precision, high-real-time, and non-invasive virtual sensing of multiple physical fields inside a fuel cell.

[0008] To achieve the above-mentioned objectives, the present invention adopts the following technical solution:

[0009] This invention provides a method for constructing a multiphysics virtual sensor for proton exchange membrane fuel cells based on the spatial attention mechanism U-Net. The specific steps are as follows:

[0010] S1. Establish a three-dimensional physical model of a proton exchange membrane fuel cell, perform computational fluid dynamics simulations under multiple operating conditions, generate a two-dimensional dataset of multiphysics fields containing gas, water, and heat distributions, and divide the monitoring plane parallel to the membrane electrode into 128×128 grid nodes.

[0011] S2. Construct the SAM-U-Net network, which is a U-Net architecture with an encoder-decoder structure. The spatial attention module (SAM) is embedded in the skip connection path, and the bipolar plate serpentine flow channel structure is incorporated into the network as prior knowledge in the form of a binary mask to improve the reconstruction accuracy of the flow channel turning and edge regions.

[0012] S3. Using externally measurable inlet gas pressure, outlet gas pressure, and average current density of the fuel cell stack as model inputs, and high-resolution distribution of multiphysics fields inside the fuel cell as output, the SAM-U-Net network is trained end-to-end.

[0013] S4. A three-stage series reconstruction mechanism of pressure distribution, water distribution, and current density distribution is adopted. The output of the previous stage virtual sensor is used as the additional input of the next stage. Based on the physical coupling relationship between pressure and membrane water content and membrane water content and current density, the multi-physics field reconstruction is completed in sequence.

[0014] S5. The flow channel design information, external sensor signals and load current are fused into a multi-channel image and input into the trained SAM-U-Net network to achieve high-resolution, high-real-time virtual perception of multi-physics fields inside the fuel cell, with a single reconstruction time of less than 0.1 seconds.

[0015] The SAM-U-Net network encoder includes three levels of downsampling modules, each consisting of two 3×3 convolutional layers, a batch normalization layer, and a max pooling layer, with feature channels numbered 32, 64, 128, and 256 respectively. The decoder includes three levels of upsampling modules, each consisting of a deconvolutional layer, two 3×3 convolutional layers, and a batch normalization layer. The spatial attention module calculates the maximum and average values ​​of the feature maps output by the encoder in the channel dimension, and obtains the spatial attention weight matrix through stacking, 3×3 convolutional fusion, and sigmoid activation.

[0016] During model training, the data was normalized to the 0-1 range using Min-Max, the loss function was mean squared error with L2 regularization, the optimizer was Adam, and the initial learning rate was 1×10⁻⁶. -4 The network architecture consists of an input convolutional layer (In_conv), a three-level downsampling module, a three-level upsampling module, an output convolutional layer (Out_conv), and four spatial attention modules (SAM). Specific parameter configurations are shown in Table 1.

[0017] Table 1. Parameter configuration of each layer of SAM-U-Net

[0018] Layer name Input dimensions Output size Convolution kernel configuration In_conv 3×128×128 32×128×128 3×3, 32 channels Downsampling_1 32×128×128 64×64×64 3×3, 64 channels Downsampling_2 64×64×64 128×32×32 3×3, 128 channels Downsampling_3 128×32×32 256×16×16 3×3, 256 channels Upsampling_1 256×16×16 128×32×32 3×3, 128 channels Upsampling_2 128×32×32 64×64×64 3×3, 64 channels Upsampling_3 64×64×64 32×128×128 3×3, 32 channels Out_conv 32×128×128 1×128×128 3×3, 1 channel SAM_1 32×128×128 1×128×128 3×3, 1 channel SAM_2 64×64×64 1×64×64 3×3, 1 channel SAM_3 128×32×32 1×32×32 3×3, 1 channel SAM_4 256×16×16 1×16×16 3×3, 1 channel

[0019] At the system level, this invention also provides a multiphysics virtual sensor construction system for proton exchange membrane fuel cells based on the spatial attention mechanism U-Net, including a data generation module, a network training module, and a multiphysics monitoring module. The data generation module is used for 3D fuel cell modeling, CFD simulation, and multiphysics dataset generation. The network training module is used to construct and train the aforementioned SAM-U-Net network. The multiphysics monitoring module is used for external signal fusion, three-level series reconstruction, and state assessment output. During system operation, the single multiphysics reconstruction time is less than 0.1 seconds, the MAPE (mean absolute percentage error) for pressure distribution reconstruction is 1.17%, the MAPE for water distribution reconstruction is 12.24%, and the MAPE for current density distribution reconstruction is 2.31%.

[0020] The beneficial effects of this invention are as follows:

[0021] (1) Excellent real-time performance: The deep generative model enables fast reasoning of multiphysics, avoiding iterative solution of complex multiphysics coupling equations, and the reconstruction time is less than 0.1 seconds.

[0022] (2) High reconstruction accuracy: The spatial attention mechanism combined with the binary mask prior of the serpentine flow channel enables the model to adaptively strengthen the feature weights of the flow channel, turning and edge regions, which significantly improves the problem of low local accuracy caused by the flow channel structure, and greatly improves the overall reconstruction accuracy and structural consistency.

[0023] (3) Strong physical rationality: The three-stage series architecture strictly follows the real physical coupling law of internal pressure-water content-current density of fuel cell. The preceding output provides physical constraints for the following output. The reconstruction result is more in line with the actual operation law and has a more accurate response capability to typical faults such as flooding.

[0024] (4) Stable and reliable training: It adopts a single encoder-decoder structure and relies on skip connections to achieve multi-scale feature fusion. The training process is smooth and easy to converge, avoiding the training oscillation and pattern collapse problems inherent in generative adversarial networks, and the engineering deployment is more stable.

[0025] (5) Excellent engineering practicality: It can realize internal state perception by relying only on the original external sensor signals of the battery system. It does not require embedded sensors, modification of the stack structure, or interference with normal power generation. It has advantages such as low cost, easy portability, and strong compatibility. Attached Figure Description

[0026] Figure 1 This is an overall flowchart of the method of the present invention; it shows the multiphysics virtual sensing process from external sensor signal input, through SAM-U-Net model reconstruction, to finally output pressure distribution, water content distribution and current density distribution.

[0027] Figure 2 This is a schematic diagram of the three-dimensional geometry and flow channel structure of the fuel cell of the present invention; it shows the serpentine flow channel of the anode / cathode of the PEM fuel cell and the layered structure of the membrane electrode assembly, including a gas diffusion layer, a catalyst layer, and a proton exchange membrane.

[0028] Figure 3 This is a diagram of the SAM-U-Net network architecture of the present invention; it depicts a symmetrical encoder-decoder structure, in which a spatial attention module (SAM) is embedded in the downsampling path, and weighted features are passed to the corresponding upsampling layer through skip connections to achieve end-to-end physical field reconstruction.

[0029] Figure 4 This is a structural diagram of the spatial attention module of the present invention; after the input feature map is max pooled and average pooled, a spatial weight map is generated by 3×3 convolution and Sigmoid to achieve focusing of key spatial regions.

[0030] Figure 5This is a diagram of the three-level physical coupling series virtual sensor architecture of the present invention; the three virtual sensors of pressure, water content and current density are connected in series in sequence, and the subsequent sensor integrates the output of the previous stage and the common input (physical sensor, flow channel design, load current) to reflect the multi-physical field coupling relationship.

[0031] Figure 6 The bar chart shows a comparison of the present invention with different models for multiphysics field reconstruction (MAPE). When comparing the present invention with traditional U-Net, attention-free U-Net, conventional CNN, variational autoencoder, DCGAN and other models for multiphysics field reconstruction (MAPE), SAM-U-Net has the lowest error in each field.

[0032] Figure 7 Line graphs comparing PSNR and SSIM of the reconstruction of this invention with different models are shown; the peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) of each model in three physical fields are displayed. SAM-U-Net achieves the highest PSNR and SSIM in most fields.

[0033] Figure 8 The graph shows the distribution of multiphysics field reconstruction error in the actual platform test of this invention. Based on real fuel cell test data, the R², MAPE, PSNR, and SSIM of the four models in the current density field and temperature field are compared, and the average value of each model is given as a piecewise curve to verify the effectiveness of SAM-U-Net in real-world scenarios. Detailed Implementation

[0034] To make the above-mentioned objectives, features, and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to examples. The following content is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the described specific embodiments or use similar methods to replace them, as long as they do not deviate from the concept of the invention, they should all fall within the protection scope of the present invention.

[0035] The preparation method of the present invention will be described below through specific embodiments.

[0036] Example

[0037] A three-dimensional physical model of a proton exchange membrane fuel cell was established in commercial computational fluid dynamics software, employing a serpentine flow channel structure. The membrane electrode assembly was placed between the hydrogen and air reaction channels, with hydrogen and air entering from opposite sides of the cell in a counter-current gas supply mode. The catalyst layer was not explicitly geometrically modeled but was represented by equivalent gas diffusion electrode boundary conditions. The overall computational mesh contained 286,210 domain elements, 85,734 boundary elements, and 7,542 edge elements.

[0038] Parallel numerical calculations were performed on the workstation, with the solution sequence as follows: first, component transport and electrochemical response were solved; then, pressure and velocity fields were solved; and finally, multiphysics coupling was performed. The average current density of the fuel cell stack was gradually increased through variable operating condition scanning to cover typical operating ranges. A total of 1000 sets of samples were generated through multi-condition simulations, and randomly divided into 600 training samples and 400 test samples in a 6:4 ratio.

[0039] The monitoring area parallel to the membrane electrode was divided into a 128×128 pixel grid. The model used inlet gas pressure, outlet gas pressure, and average current density of the fuel cell stack as inputs; and cathode and anode flow field pressure, humidity distribution, transmembrane water flux, and membrane electrode conductivity distribution as outputs.

[0040] The SAM-U-Net spatial attention mechanism U-Net model is constructed with a base convolutional channel count of 32 and an initial learning rate of 1×10⁻⁶. -4 The L2 regularization parameter is 10, a cosine annealing learning rate scheduling strategy is adopted, and the Adam optimizer is used for training until the validation set loss converges and stabilizes.

[0041] In the simulation dataset test, the pressure distribution reconstruction MAPE was 1.17%, the PSNR (peak signal-to-noise ratio) was 42.44 dB, the SSIM (structural similarity index) was 0.9949, and the coefficient of determination R0 was [missing value]. 2 The coefficient of determination was 0.9997; the water distribution reconstruction MAPE was 12.24%, the PSNR was 36.34 dB, the SSIM was 0.9225, and the R... 2 The value is 0.9961; the reconstructed current density distribution MAPE is 2.31%.

[0042] This invention compares with traditional U-Net, attention-free U-Net, conventional CNN, variational autoencoder, DCGAN and other models in Multiphysics Reconstruction (MAPE). Figure 6 As shown:

[0043] Compared to traditional U-Nets that do not incorporate prior flow channel information, the MAPE of this invention is reduced by 1.28%;

[0044] Compared to U-Net without a spatial attention module, the MAPE of this invention is reduced by 1.47%;

[0045] Compared to conventional convolutional neural networks, MAPE is reduced by 17.57%;

[0046] Compared with variational autoencoders and variational autoencoder residual networks, the reconstruction error of this invention is reduced by more than 50%.

[0047] Compared to deep convolutional generative adversarial networks, the average error is reduced by 22.7%.

[0048] Comparison of PSNR and SSIM indices reconstructed by each model Figure 7 As shown:

[0049] Compared to traditional U-Net, this invention improves PSNR and SSIM by 5.51 and 6.49, respectively;

[0050] Compared to the unattended U-Net, PSNR and SSIM are improved by 5.51 and 5.23, respectively;

[0051] Compared to conventional CNNs, PSNR and SSIM are improved by 14.51 and 22.56 respectively;

[0052] Compared to the variational autoencoder series models, SSIM has improved from level 0.7 to over 0.92.

[0053] In actual fuel cell test platform verification, the current density field reconstruction MAPE was 5.91%, PSNR was 33.00 dB, SSIM was 0.8936, and R... 2 The value is 0.9926; the temperature field MAPE is 9.28%, PSNR is 25.62 dB, SSIM is 0.8307, and R... 2 The average MAPE was 0.9606; the average SSIM was 0.8621.

[0054] The statistical distribution of reconstruction errors for each physical field in actual tests is as follows: Figure 8 As shown, among various comparison methods, the present invention has the lowest reconstruction error in both the current density field and the temperature field, demonstrating good engineering effectiveness and reliability under real-world conditions.

Claims

1. A method for constructing a proton exchange membrane fuel cell multi-physical field virtual sensor based on a spatial attention mechanism U-Net, characterized in that, Includes the following steps: S1. Establish a three-dimensional physical model of a proton exchange membrane fuel cell, perform computational fluid dynamics simulations under multiple operating conditions, generate a two-dimensional dataset containing multi-physics field distributions including gas, water, and heat, and divide the monitoring plane parallel to the membrane electrode into 128×128 grid nodes. S2. Construct a SAM-U-Net network, which is a U-Net with an encoder-decoder structure. Embed a spatial attention module in the skip connection path and integrate the bipolar plate serpentine flow channel structure as prior knowledge into the network in the form of a binary mask. S3. Using externally measurable inlet gas pressure, outlet gas pressure, and average current density of the fuel cell stack as model inputs, and high-resolution distribution of multiphysics fields inside the fuel cell as output, the SAM-U-Net network is trained end-to-end. S4. The reconstruction is carried out using a three-stage series method of pressure distribution, water distribution, and current density distribution. The output of the previous stage virtual sensor is used as the additional input of the next stage. The reconstruction is carried out in sequence according to the physical coupling relationship between pressure and membrane water content and membrane water content and current density. S5. The flow channel design information, external sensor signals and load current are fused into a multi-channel image and input into the trained SAM-U-Net network.

2. The construction method of claim 1, wherein, The encoder of the SAM-U-Net network includes a three-level downsampling module. Each downsampling module consists of two 3×3 convolutional layers, a batch normalization layer, and a max pooling layer, with the number of feature channels being 32, 64, 128, and 256, respectively. The decoder of the SAM-U-Net network includes three upsampling modules, each consisting of a deconvolutional layer, two 3×3 convolutional layers, and a batch normalization layer.

3. The construction method of claim 1, wherein, The spatial attention module calculates the maximum value feature and the average value feature in the channel dimension of the feature map output by the encoder. The two types of features are stacked and then fused through a 3×3 convolutional layer. Finally, the spatial attention weight matrix is ​​obtained by passing the sigmoid activation function.

4. The construction method of claim 1, wherein, The three-stage serial connection method is specifically as follows: The first stage takes external sensor signals, flow channel design information, and load current as inputs and outputs the gas pressure distribution within the flow channel. The second stage takes the above-mentioned input and pressure distribution as input and outputs the membrane water content distribution. The third stage takes the above-mentioned inputs, pressure distribution, and water distribution as inputs, and outputs the membrane electrode current density / conductivity distribution.

5. The construction method of claim 1, wherein, During model training: The data was normalized to the 0-1 range using Min-Max. The loss function uses mean squared error with L2 regularization; The optimizer employs Adam with an initial learning rate of 1 x 10 -4 .

6. A proton exchange membrane fuel cell multi-physical field virtual sensor construction system based on a spatial attention mechanism U-Net, characterized in that, include: The data generation module is used to construct a three-dimensional proton exchange membrane fuel cell model, perform CFD simulations, and generate a two-dimensional multiphysics dataset. A network training module is used to construct and train the SAM-U-Net network as described in any one of claims 1-5; The multiphysics monitoring module is used to fuse external signals with flow channel design information and realize real-time reconstruction and status assessment of internal multiphysics through a three-level series structure.

7. The construction system according to claim 6, characterized in that, The SAM-U-Net includes an input convolutional layer, a three-level downsampling module, a three-level upsampling module, an output convolutional layer, and four spatial attention modules. The configuration of each layer is as follows: In_conv: Input 3×128×128, output 32×128×128, convolution kernel 3×3, 32 channels; Downsampling_1: Input 32×128×128, output 64×64×64, convolution kernel 3×3, 64 channels; Downsampling_2: Input 64×64×64, output 128×32×32, convolution kernel 3×3, 128 channels; Downsampling_3: Input 128×32×32, output 256×16×16, convolution kernel 3×3, 256 channels; Upsampling_1: Input 256×16×16, output 128×32×32, convolution kernel 3×3, 128 channels; Upsampling_2: Input 128×32×32, output 64×64×64, convolution kernel 3×3, 64 channels; Upsampling_3: Input 64×64×64, output 32×128×128, convolution kernel 3×3, 32 channels; Out_conv: Input 32×128×128, output 1×128×128, convolution kernel 3×3, 1 channel; SAM_1 to SAM_4 correspond to feature maps of different scales, and the output of each is a 1-channel spatial weight map. The convolution kernels are all 3×3.