Cross-condition prediction method for remaining life of IGBT module

By collecting sensing data in the IGBT module to calculate transient thermal resistance and using the probability sparse self-attention structure and the maximum mean difference of multi-core, a prediction model is constructed, which solves the problem of poor generalization performance of life prediction of IGBT modules across operating conditions, and achieves accurate cross-operating conditions prediction.

CN116266249BActive Publication Date: 2025-07-22SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111523088.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-14
Publication Date
2025-07-22
Estimated Expiration
2041-12-14

AI Technical Summary

Technical Problem

The model generalization performance of the IGBT module in the prior art is poor in the model of remaining life prediction across operating conditions and cannot effectively adapt to the life prediction under different operating conditions.

Method used

By collecting the sensing data of the IGBT module under different operating conditions, calculating transient thermal resistance data and extracting effective features, further extracting information using the probabilistic sparse self-attention structure, measuring the depth feature difference in combination with the multi-core maximum mean difference, and constructing a prediction model for cross-operating conditions prediction.

Benefits of technology

The accuracy and stability of the IGBT module remaining life prediction under different operating conditions is achieved, the generalization ability of the model is improved, and the degradation process of the IGBT module can be effectively predicted across operating conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116266249B_ABST
    Figure CN116266249B_ABST
Patent Text Reader

Abstract

A method for predicting the remaining life of an IGBT module, which collects sensing data when the IGBT module performs power cycling under different working conditions, calculates transient thermal resistance data, and extracts and screens effective features therefrom; then uses source working condition data with remaining life labels and target working condition data whose remaining life is to be predicted as a training set to train a prediction model based on a probabilistic sparse self-attention mechanism; finally, in the online stage, the trained model is used to accurately predict the remaining life of the working condition. The present invention extracts and screens effective features from the transient thermal resistance data obtained from the power cycling of the IGBT module, uses a probabilistic sparse self-attention structure to further extract effective information, and measures the difference in deep features under different working conditions through the maximum mean discrepancy, so as to achieve cross-working condition prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology for the reliability of power electronic devices, specifically an IGBT module remaining life cross-condition prediction method based on probabilistic sparse self-attention and domain adaptation. Background Art

[0002] In order to improve the reliability of IGBT modules, it is necessary to predict their remaining life. Currently, the mainstream methods are mainly divided into physical model-based and analytical model-based methods. The former requires obtaining various detailed electrical parameters of the IGBT module, which is difficult to operate in practice, and there are large differences among different models, resulting in poor universality. The latter requires a large amount of experimental data to fit the mathematical model, which is difficult to construct, and it is necessary to balance the prediction accuracy and universality of the model. Recently, data-driven methods have gradually emerged, but usually, the remaining life is predicted using historical data before failure on IGBT modules under a single condition. However, in actual different conditions, these data vary greatly, resulting in significant problems with the universality of the model. Summary of the Invention

[0003] Aiming at the problems of poor generalization performance of the model in cross-condition remaining life prediction and the inability to predict the life of IGBT modules under different experimental conditions when the number of experimental conditions of IGBT modules is small and the span is large in the prior art, and the prediction effect under different conditions with large data differences has not been verified, the present invention proposes an IGBT module remaining life prediction method. Effective features are extracted and screened from the transient thermal resistance data obtained from the power cycle of the IGBT module, a probabilistic sparse self-attention structure is used to further extract effective information, and the difference in deep features under different conditions is measured by the maximum mean discrepancy, so as to achieve cross-condition prediction.

[0004] The present invention is realized through the following technical solutions:

[0005] The present invention relates to an IGBT module remaining life prediction method based on probabilistic sparse self-attention and domain adaptation. By collecting sensing data during power cycling of the IGBT module under different conditions, transient thermal resistance data is calculated, and effective features are extracted and screened from it; then, the source condition data with remaining life labels and the target condition data for which the remaining life is to be predicted are used as the training set to train a prediction model based on the probabilistic sparse self-attention mechanism; finally, the trained model is used in the online stage to accurately predict the remaining life of the condition.

[0006] The different conditions mentioned above refer to: the fluctuation range ΔT of the junction temperature when the IGBT module is working j and the average value T jm are different, and the two are represented by the maximum junction temperature T j_max and the minimum junction temperature Tj_min Denoted as ΔT j = T j_max - T j_min and

[0007] The power cycle mentioned above refers to: intermittently applying a thermal excitation to the IGBT module to cause fluctuations in the IGBT temperature, where the thermal excitation is achieved by increasing the voltage and current to make the IGBT module heat itself to reach the set temperature.

[0008] The sensing data mentioned above includes: junction temperature sensor data T j and case temperature sensor data T p , IGBT module voltage sensor data V ce_on and IGBT module current sensor data I.

[0009] The sampling rate of the sensing data mentioned above is to collect once every 0.1 s.

[0010] The transient thermal resistance data mentioned above is calculated based on the collected sensor data, specifically where: T j (t), T p (t) are the junction temperature and case temperature of the IGBT module at time t respectively, and V ce_on (t), I(t) are the voltage and current of the IGBT module collector-emitter at time t.

[0011] The prediction model based on the probabilistic sparse self-attention mechanism mentioned above includes: an embedding layer, a probabilistic sparse self-attention module, a feed-forward module, and a fully connected module, where: the input of the embedding layer is the thermal resistance features collected and calculated, and after performing a high-dimensional space transformation process, an embedding vector is obtained; the input of the probabilistic sparse self-attention module is the embedding vector output by the embedding layer, and through the process of reassigning weights to different moments of the input sequence, a high-dimensional feature vector is obtained; the feed-forward module and the fully connected module sequentially perform non-linear transformation processes on the high-dimensional feature vector output by the probabilistic sparse self-attention module, and finally a life prediction result is obtained.

[0012] The extraction mentioned above refers to: extracting 22 features including minimum value, maximum value, range, mean, median, first quartile, third quartile, variance, standard deviation, skewness, standard error, absolute median, kurtosis, mean square value, root mean square, absolute mean, root mean amplitude, impulse factor, waveform factor, margin factor, peak factor, and double logarithmic ratio based on the transient thermal resistance data within 5 s of each power cycle.

[0013] The screening mentioned above refers to: screening the extracted features according to the feature selection criteria and retaining the features with a value exceeding 0.5, where: correlation Let \(n\) be the total number of cycles, \(X(t)\) and \(RUL(t)\) be the characteristic value and remaining useful life ratio at time \(t\) respectively. and are the mean and monotonicity of all cycles \(\Delta X\) is the change value of the characteristic at adjacent times, \(N(\Delta X\gt0)\) is the number of change values greater than 0, and \(N(\Delta X\lt0)\) is the number of change values less than 0.

[0014] The source operating condition data with remaining useful life labels refers to: selecting one operating condition as the source operating condition, and each power cycle under this operating condition has a corresponding remaining useful life label.

[0015] The target operating condition data of the life to be predicted refers to: selecting one operating condition as the target operating condition, that is, the operating condition for which the remaining useful life of the IGBT needs to be predicted. Under this operating condition, only the transient thermal resistance characteristic data selected for each cycle is known, and the corresponding remaining useful life is unknown and needs to be predicted.

[0016] The training refers to: in each training batch, the source operating condition data and the target operating condition data are input sequentially. The labeled source operating condition data is used for supervised training, and the unlabeled target operating condition data is used for unsupervised training. At the same time, the multi-kernel maximum mean discrepancy is used to measure the difference between the deep features of the two operating condition data, and the model parameters are optimized based on the principle of reducing the difference.

[0017] The multi-kernel maximum mean discrepancy refers to: where \(f(\cdot)\) is the mapping from the original space to the reproducing Hilbert space, \(x\) i and \(y\) i are the \(i\)-th and \(j\)-th samples in the two distributions to be evaluated respectively, and \(m\) and \(n\) are the total number of samples in the two distributions respectively. is the norm in the mapped space, and this mapping can be calculated through the multi-kernel function. The multi-kernel function where \(k\) u is a single kernel function, \(\beta\) u is the weight of each kernel function, then the multi-kernel maximum mean discrepancy can be expressed as \(k\) is the multi-kernel function defined above, \(n\) s and \(n\) t are the number of samples in the source domain and the target domain respectively. are the feature vectors in the source domain and the target domain respectively. The multi-kernel maximum mean discrepancy can measure the difference in distribution between the two domains.

[0018] The difference between the described deep features refers to: using multi-kernel maximum mean discrepancy in the fully connected layer of the model to calculate the difference between the deep feature distributions of the source condition and target condition data. The front part of the model extracts more general features, while the fully connected layer extracts more unique features for each condition. The difference is greater when each condition is trained separately, and using multi-kernel maximum mean discrepancy here can help the model learn deep features with better generalization ability.

[0019] The principle of reducing the difference means: reducing the multi-kernel maximum mean discrepancy of the deep features in the fully connected layer of the source condition and target condition, which is achieved by integrating the maximum mean discrepancy into the overall optimization objective of the model. The optimization objective is jointly composed of the loss function and the multi-kernel maximum mean discrepancy: where: Θ is the weight parameter to be optimized by training the model, J(·) is the loss function on the labeled data set. Since only labeled data is provided in the source domain, this term is only calculated on the source domain data set. The second term is the multi-kernel maximum mean discrepancy of multiple layers, and λ is the weight coefficient.

[0020] Technical effects

[0021] Compared with the prior art, the present invention uses the labeled IGBT source condition thermal resistance data, that is, supervised training is carried out with the known remaining life in each cycle, and the unlabeled IGBT target condition thermal resistance data, that is, the remaining life is unknown and to be predicted, for unsupervised training, so as to realize the domain adaptation of high-dimensional deep features and finally realize the cross-condition prediction of the remaining life. The specific technical details with significant improvement compared with the existing conventional technical means are: aligning the deep features, using the information of the target domain data to adjust the model, rather than directly using the model trained with the source domain data for predicting the target domain. Brief description of the drawings

[0022] Figure 1 It is the flow chart of the present invention;

[0023] Figure 2 It is the schematic diagram of the experimental device in the embodiment;

[0024] Figure 3 It is the physical diagram of the experimental device in the embodiment. Detailed implementation manners

[0025] As Figure 1 shown, the present embodiment relates to an IGBT module remaining life prediction method based on probabilistic sparse self-attention and domain adaptation, including the following steps:

[0026] Step 1, as Figure 2As shown in the figure, according to the schematic diagram, an IGBT module accelerated aging test bench is built. A temperature sensor is installed on the surface of the IGBT chip to measure the junction temperature, a temperature sensor is installed outside the packaging shell of the IGBT module to measure the packaging temperature, and voltage and current sensors are installed on the power cycle circuit to measure the conduction voltage and current of the IGBT module. The sensor acquisition frequency is 10Hz.

[0027] Step 2: Calculate the transient thermal resistance of the IGBT module power cycle according to the data obtained by each sensor: Where T j (t) is the data of the junction temperature sensor, T p (t) is the data of the housing temperature sensor, is the data of the IGBT module voltage sensor, and I(t) is the data of the IGBT module current sensor.

[0028] Step 3: Divide the transient thermal resistance data by power cycle and extract and screen features: At the beginning of each cycle, the IGBT module is turned on, and the junction temperature T j continuously rises until it reaches the set maximum value T j_max ; the IGBT module is turned off, and the cooling system is turned on until the junction temperature T j drops to the set minimum value T j_min , and so on. Therefore, within one cycle, the transient thermal resistance data is constantly changing. According to the transient thermal resistance data within 5s of each power cycle, 22 features including minimum value, maximum value, range, mean, median, first quartile, third quartile, variance, standard deviation, skewness, standard error, absolute median, kurtosis, mean square value, root mean square, absolute mean, root mean amplitude, pulse factor, waveform factor, margin factor, peak factor, and double logarithmic ratio are extracted, and then the feature selection criterion is used to screen these 22 features, and the features with a value exceeding 0.5 are retained, where the correlation n is the total number of cycles, X(t) and RUL(t) are the feature values and remaining life ratios at time t respectively, and are the mean and monotonicity of all cycles ΔX is the change value of the feature at adjacent times, N(ΔX>0) is the number of change values greater than 0, and N(ΔX<0) is the number of change values less than 0.

[0029] Step 4: Build a prediction model based on the probabilistic sparse self-attention mechanism. The model includes: an embedding layer, a probabilistic sparse self-attention module, a feed-forward module, and a fully connected module.

[0030] The embedding layer mentioned above refers to: For an input feature sequence x=(x1,...,xL ) Let \(f\) be the dimension of the extracted features, and embed it into a high-dimensional space to obtain \(VE=(v_1,\cdots,v_f)\). L ) Let \(d\) be the dimension of the embedding space. For the position vector \(p=(0,\cdots,i,\cdots,L)\) of the input sequence, where \(i\) is the position index of each sample in the sequence, use the sine-cosine position encoding method to embed it into a high-dimensional space of the same dimension to obtain \(PE=(p_1,\cdots,p_d)\). L ) The output of the final embedding layer is

[0031] The probability sparse self-attention module represents the attention weight of an input sample as where \(q\) i , \(k\) i , \(v\) i are the \(i\)-th rows of matrices \(Q\), \(K\), and \(V\), and \(p(k|q)\) is the self-attention distribution. j | \(q\) i ) is the uniform distribution. By calculating according to the KL divergence formula and removing the variable-independent terms, the approximate evaluation function can be obtained To reduce the computational complexity, randomly select \(M = L\ln L\) dot product pairs for calculation and finally select the \(m\) query values with the highest scores. The corresponding distribution is considered the truly effective distribution, and the attention weight distributions corresponding to other distributions will be directly set to the uniform distribution.

[0032] The feed-forward module includes: layer normalization connecting two one-dimensional convolutional layers, and then connecting layer normalization, one one-dimensional convolutional layer, and one global pooling layer.

[0033] The fully connected module includes 3 fully connected layers with dimensions of 256×64, 64×32, and 32×1 respectively.

[0034] Step 5: Divide the source condition data and the target condition data into samples of the same length. In this embodiment, each sample length is selected to be 50, that is, the input is the transient thermal resistance characteristics of 50 consecutive cycles. The label of the source condition data is the remaining life corresponding to the end of the 50th cycle, and the target condition data has no label.

[0035] Step 6: In each batch training, input the source condition data samples and the target condition data samples in sequence to train and optimize the model parameters. First, after inputting the source condition data samples, record the corresponding feature vectors of the fully connected layer and calculate according to the model output and the label Then input the target condition data samples and record the corresponding feature vectors of the fully connected layer Calculate the maximum mean difference of multiple cores Finally, the optimization objective MSE + MK - MMD is obtained, and the model parameters are optimized by backpropagation.

[0036] Experimental data is collected under three different working conditions shown in Table 1, and mutual transfer prediction experiments are carried out between pairwise working conditions. Some existing technologies are compared, and two evaluation indexes, mean square error MSE and mean absolute error MAE, are used to measure the advantages and disadvantages of each method. The experimental results are shown in Table 2. It can be seen from the results that the model proposed in this paper achieves the best prediction effect in all transfer tasks and can better predict the remaining life throughout the degradation process of the IGBT module.

[0037] Table 1 Aging test conditions and results

[0038]

[0039] Table 2 Remaining life prediction results of IGBT modules with different technologies

[0040]

[0041]

[0042] Compared with the existing technologies, this method greatly improves the prediction accuracy in the initial stage of degradation in the remaining life prediction task of IGBT modules under different working conditions, and the stability in the whole life cycle prediction is also improved. At the same time, this invention extracts transient thermal resistance characteristics based on the junction temperature, package temperature, and power of the IGBT module, which has better robustness and higher reliability than single-variable characteristics; the existing technologies usually only focus on the mapping relationship between the selected characteristic variables and the remaining life at a certain moment. On this basis, this invention can also pay attention to the characteristic change law in time series, with higher information utilization rate; the existing technologies usually use historical data before failure to predict the remaining life under one working condition, and due to the large differences in data under different working conditions, it is difficult to directly achieve accurate prediction under other working conditions. This invention can migrate from the source working condition to the target working condition to achieve accurate prediction.

[0043] The above specific implementation can be locally adjusted by those skilled in the art in different ways without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific implementation. All implementation solutions within its scope are subject to the constraints of the present invention.

Claims

1. A method for predicting the remaining life of an IGBT module based on probabilistic sparse self-attention and domain adaptation, characterized in that, By collecting the sensing data of the IGBT module during power cycling under different working conditions, calculating the transient thermal resistance data, and extracting and screening effective features therefrom; then using the source working condition data with remaining life labels and the target working condition data with remaining life to be predicted as the training set to train the prediction model based on the probabilistic sparse self-attention mechanism; finally, using the trained model in the online stage to accurately predict the remaining life of the working condition; The different operating conditions refer to the fluctuation range ΔT of the junction temperature when the IGBT module is working j and the average value T jm are different, and the two are represented by the maximum junction temperature T j_max and the minimum junction temperature T j_min as ΔT j = T j_max - T j_min and The said sensing data includes: node temperature sensor data T j , housing temperature sensor data T p , IGBT module voltage sensor data V ce_on and IGBT module current sensor data I; The source working condition data with remaining life labels refers to: selecting a working condition as the source working condition, and each power cycle under this working condition has a corresponding remaining life label; The target working condition data with remaining life to be predicted refers to: selecting a working condition as the target working condition, that is, the working condition for which the remaining life of the IGBT needs to be predicted. Only the transient thermal resistance feature data selected for each cycle under this working condition is known, and the corresponding remaining life is unknown and needs to be predicted; The prediction model based on the probabilistic sparse self-attention mechanism includes: an embedding layer, a probabilistic sparse self-attention module, a feed-forward module, and a fully connected module, where: the input of the embedding layer is the thermal resistance features collected and calculated, and after performing high-dimensional space transformation processing, an embedding vector is obtained. The input of the probabilistic sparse self-attention module is the embedding vector output by the embedding layer. By reassigning weights to different moments of the input sequence, a high-dimensional feature vector is obtained. The feed-forward module and the fully connected module sequentially perform non-linear transformation processing on the high-dimensional feature vector output by the probabilistic sparse self-attention module, and finally obtain the life prediction result; The screening mentioned above refers to: according to the feature selection criteria screen the extracted features, and retain the features with a value exceeding 0.5, where: correlation n is the total number of cycles, X(t) and RUL(t) are the feature value and the remaining useful life ratio at time t respectively, and are the mean values of the feature values and the remaining useful life ratios for all cycles respectively, monotonicity ΔX is the change value of the feature at adjacent times, N(ΔX>0) is the number of change values greater than 0, and N(ΔX<0) is the number of change values less than 0.

2. The method for predicting the remaining life of an IGBT module based on probabilistic sparse self-attention and domain adaptation according to claim 1, wherein The power cycle refers to: intermittently applying a thermal excitation to the IGBT module to cause fluctuations in the IGBT temperature, where the thermal excitation is achieved by increasing the voltage and current to make the IGBT module heat itself to reach the set temperature.

3. The method for predicting the remaining life of an IGBT module based on probabilistic sparse self-attention and domain adaptation according to claim 1, characterized in that, The transient thermal resistance data is calculated based on the collected sensor data, specifically as follows where: T j (t), T p (t) are the junction temperature and case temperature of the IGBT module at time t, respectively, and V ce_on (t), I(t) are the collector-emitter voltage and current of the IGBT module at time t.

4. The method for predicting the remaining life of an IGBT module based on probabilistic sparse self-attention and domain adaptation according to claim 1, characterized in that, The extraction refers to: extracting 22 features including minimum value, maximum value, range, mean, median, first quartile, third quartile, variance, standard deviation, skewness, standard error, absolute median, kurtosis, mean square value, root mean square, absolute mean, root mean amplitude, impulse factor, waveform factor, margin factor, peak factor, and double logarithmic ratio according to the transient thermal resistance data within 5 s of each power cycle.

5. The method for predicting the remaining life of an IGBT module based on probabilistic sparse self-attention and domain adaptation according to claim 1, characterized in that, The training refers to: for each training batch, the source working condition data and the target working condition data are sequentially input. The source working condition data with labels is subjected to supervised training, and the target working condition data without labels is subjected to unsupervised training. At the same time, the multi-kernel maximum mean discrepancy is used to measure the difference between the deep features of the two working condition data, and finally the model parameters are optimized based on the principle of reducing the difference.

6. The method for predicting the remaining life of an IGBT module based on probabilistic sparse self-attention and domain adaptation according to claim 5, characterized in that, The so-called multi-kernel maximum mean discrepancy refers to: where \(f(\cdot)\) is a mapping from the original space to the reproducing Hilbert space, \(x\) i and \(y\) i are respectively the \(i\)-th and \(j\)-th samples in two distributions to be evaluated, \(m\) and \(n\) are the total number of samples in the two distributions respectively, is the norm in the mapped space, and this mapping can be calculated through a multi-kernel function. The multi-kernel function where \(k\) u is a single kernel function, and \(\beta\) u is the weight of each kernel function. Then the multi-kernel maximum mean discrepancy can be expressed as where \(k\) is the multi-kernel function defined above, \(n\) s , \(n\) t are respectively the number of samples in the source domain and the target domain, are respectively the feature vectors of the source domain and the target domain. The multi-kernel maximum mean discrepancy can measure the distribution difference between the two domains.

7. The method for predicting the remaining life of an IGBT module based on probabilistic sparse self-attention and domain adaptation according to claim 5, characterized in that The difference between the deep features refers to: using the multi-kernel maximum mean discrepancy to calculate the difference between the deep feature distributions of the source working condition and the target working condition data in the fully connected layer of the model. The features extracted in the front part of the model are more general features, while the fully connected layer is the more unique features of each working condition. The difference is greater when each working condition is trained separately, and using the multi-kernel maximum mean discrepancy here can help the model learn better generalizable deep features.

8. The method for predicting the remaining life of an IGBT module based on probabilistic sparse self-attention and domain adaptation according to claim 6, wherein The principle of minimizing the difference means: minimizing the multi-kernel maximum mean difference of the deep features of the fully connected layer between the source condition and the target condition, which is achieved by incorporating the maximum mean difference into the overall optimization objective of the model. The optimization objective consists of a loss function and the multi-kernel maximum mean difference: where Θ is the weight parameter to be optimized by training the model, J(·) is the loss function on the labeled data set. Since only labeled data is provided in the source domain, this term is only calculated on the source domain data set. The second term is the multi-kernel maximum mean difference of multiple layers, and λ is the weight coefficient.

Citation Information

Patent Citations

  • Underwater acoustic target radiation noise identification method based on domain adaptation

    CN111709315A

  • Non-intrusive load decomposition method based on Informer model coding structure

    CN113393025A