IGBT fault prediction method and device
The features of IGBT transistors are extracted through the CEEMDAN and KAN-Transformer models, and the problem of inaccurate IGBT fault prediction in the prior art is solved, and fault prediction with higher accuracy and robustness is achieved.
Patent Information
- Application Number
- CN202510847243.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-06-24
AI Technical Summary
The existing IGBT fault prediction methods cannot accurately reflect the failure mechanism of power electronic devices, and parameters and variables are difficult to accurately obtain, resulting in inaccurate prediction.
By acquiring the operating data of the IGBT transistor, using CEEMDAN for modal decomposition, combining local long and global long attention mechanisms to extract features, using the KAN-Transformer model for fault prediction, optimize the noise amplitude and integration quantity to improve prediction accuracy.
It realizes fault prediction from the perspective of physical detection, improves the accuracy and robustness of IGBT fault prediction, and enhances the practicality of prediction.
Smart Images

Figure CN120352751B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of smart grid technology, and in particular to an IGBT fault prediction method and device. Background Art
[0002] Insulated gate bipolar transistors (IGBTs) are widely used in power electronics systems due to their high input impedance and low on-state voltage drop. However, the high voltage and high current characteristics of their operating environment make them key components most prone to failure. Accurate and timely fault prediction is crucial to system reliability.
[0003] Existing IGBT failure prediction methods primarily rely on lifecycle-based prediction, primarily analytical lifecycle models and physical failure models. However, analytical lifecycle models are based on purely statistical analysis of test data and do not reflect the failure mechanisms of power electronic devices, thus presenting certain limitations. Furthermore, physical failure-based methods rely on analyzing the system's historical loads, environmental stresses, and operating conditions, making them difficult to accurately establish. Furthermore, the numerous parameters and variables involved in design are often difficult to accurately obtain. Consequently, lifecycle-based failure prediction methods are unable to accurately predict IGBT failures.
[0004] Therefore, it is urgent to propose an IGBT fault prediction method and device to solve the technical problems in the existing technology that the life cycle-based fault prediction method does not reflect the failure mechanism of power electronic devices, and the parameters and variables in physical detection are difficult to obtain accurately, resulting in the inability of the life cycle-based fault prediction method to accurately predict IGBT faults. Summary of the Invention
[0005] In view of this, it is necessary to provide an IGBT fault prediction method and device to solve the technical problem in the prior art that the life cycle-based fault prediction method cannot accurately predict IGBT faults.
[0006] In order to solve the above problems, in a first aspect, the present invention provides an IGBT fault prediction method, comprising:
[0007] Acquire operating data of the IGBT transistor within a preset time, and determine a time series signal of the preset time; the operating data includes a pole voltage;
[0008] After performing modal decomposition on the pole voltage according to the time series signal, feature extraction and slicing are performed on the IMF component data of the pole voltage according to a sliding window method to obtain data within each window;
[0009] Extract local features from the data in each window according to the local multi-head attention mechanism to obtain local features of each window, and concatenate the local features of all windows to obtain a global feature tensor;
[0010] The global dependency relationship between the time series signal and the global feature tensor is extracted according to the global multi-head attention mechanism. After obtaining the global feature sequence, the global feature sequence is predicted to obtain a predicted value, and the fault prediction result of the IGBT transistor is obtained based on the predicted value.
[0011] In a possible implementation, after obtaining the operating data of the IGBT transistor within a preset time and determining the time series signal of the preset time, the method further includes:
[0012] The preset noise amplitude and the preset integration number are optimized based on the Bayesian optimization algorithm to obtain the optimal parameters;
[0013] Decomposing the pole voltage according to the optimal parameters to obtain multiple IMF components;
[0014] The multiple IMF components are converted into a two-dimensional table according to the time series signal to obtain IMF component data.
[0015] In a possible implementation, the feature extraction and slicing of the IMF component data of the pole voltage according to the sliding window method to obtain data within each window includes:
[0016] Performing preliminary feature extraction on the IMF component data according to a one-dimensional convolution layer to obtain preliminary convolution features;
[0017] The preliminary convolution features are sliced according to the sliding window method to obtain data in each window.
[0018] In a possible implementation, predicting the global feature sequence to obtain a predicted value includes:
[0019] Mapping the global feature sequence to a high-dimensional latent space according to a feature embedding layer to obtain an embedding layer output feature matrix;
[0020] Performing global modeling on the embedding layer output feature matrix according to the encoder to obtain encoder output features;
[0021] interacting the encoder output features according to the decoder to obtain decoded output features;
[0022] The decoded output features are feature mapped and predicted according to the KAN layer to obtain a predicted value.
[0023] In a possible implementation, performing global modeling on the embedding layer output feature matrix according to the encoder to obtain encoder output features includes:
[0024] Performing global modeling on the embedding layer output feature matrix input to the first layer of the encoder to obtain a first output feature;
[0025] Inputting the first output feature into the second layer of the encoder for global modeling to obtain a second output feature;
[0026] When the second layer is the last layer of the encoder, the second output feature is determined as the encoder output feature.
[0027] In a possible implementation, performing global modeling on the embedding layer output feature matrix input to the first layer of the encoder to obtain the first output feature includes:
[0028] Calculating the embedding layer output feature matrix input to the first layer of the encoder based on a multi-head attention mechanism to obtain an attention score;
[0029] Performing a nonlinear transformation on the attention score through a feedforward network to obtain a nonlinear feature;
[0030] Perform residual connection and normalization processing on the embedding layer output feature matrix and the nonlinear feature to obtain the first output feature.
[0031] In a possible implementation, performing feature mapping and prediction on the decoded output features according to the KAN layer to obtain a predicted value includes:
[0032] Performing a basic linear mapping on the decoded output features input to the KAN layer to obtain a basic output relationship;
[0033] Performing basis function mapping on the decoded output features to obtain basis functions;
[0034] Performing weight calculation on the basis function according to a preset weight matrix to obtain basis function weights;
[0035] A predicted value is obtained according to the basic output relationship and the basis function weights.
[0036] In a possible implementation, the operating data further includes a measured value; and obtaining a fault prediction result of the IGBT transistor according to the predicted value includes:
[0037] Comparing the measured value with the predicted value to obtain a difference;
[0038] When the difference is greater than a preset threshold, it is determined that the IGBT transistor is in a fault state.
[0039] In one possible implementation, the calculation formula of the attention score is:
[0040]
[0041] Where, is the query matrix, ; is the bond matrix, ; is the value matrix, ; is the scaling factor; where 、 、 is the weight matrix of linear transformation; Output feature matrix for the embedding layer.
[0042] In a second aspect, the present invention further provides an IGBT fault prediction device, comprising:
[0043] A data acquisition module is used to acquire operating data of the IGBT transistor within a preset time and determine a time series signal of the preset time; the operating data includes a pole voltage;
[0044] a feature extraction module, configured to perform modal decomposition on the pole voltage according to the time series signal, and then perform feature extraction and segmentation on the IMF component data of the pole voltage according to a sliding window method to obtain data within each window;
[0045] A feature splicing module is used to extract local features of the data in each window according to the local multi-head attention mechanism to obtain local features of each window, and to splice the local features of all windows to obtain a global feature tensor;
[0046] A fault prediction module is used to extract the global dependency between the time series signal and the global feature tensor according to the global multi-head attention mechanism, obtain the global feature sequence, predict the global feature sequence, obtain the predicted value, and obtain the fault prediction result of the IGBT transistor according to the predicted value.
[0047] The beneficial effects of the present invention are: obtaining the operating data of the IGBT transistor within a preset time, the operating data including the pole voltage, which can be the collector-emitter voltage, so that fault prediction can be performed using the collector-emitter voltage as the fault characteristic parameter, so that fault prediction can be performed from the perspective of physical detection; the IMF component data obtained after modal decomposition of the pole voltage can also be subjected to local feature extraction through a local multi-head attention mechanism to obtain the local features of each window, so that the accuracy of feature detection can be increased by analyzing the local features of each window; the local features of all windows can also be spliced together, and the global dependency between the time series signal and the global feature tensor can be extracted according to the global multi-head attention mechanism to obtain a global feature sequence, so that a complete feature extraction result can be obtained, and then the global feature sequence can be predicted to improve the robustness of the prediction, thereby improving the practicality of IGBT fault prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A schematic flow chart of an embodiment of the IGBT fault prediction method provided by the present invention;
[0049] Figure 2 For the present invention Figure 1 A schematic flow chart of an embodiment after step S101;
[0050] Figure 3 A coordinate system diagram of an embodiment of the six IMF component graphs after CEEMDAN decomposition provided by the present invention;
[0051] Figure 4 For the present invention Figure 1 A schematic flow chart of an embodiment of step S102;
[0052] Figure 5 A schematic diagram of the structure of an embodiment of a time series feature extraction module provided by the present invention;
[0053] Figure 6 For the present invention Figure 1 A schematic flow chart of an embodiment of step S104;
[0054] Figure 7 For the present invention Figure 6 A schematic flow chart of an embodiment of step S602;
[0055] Figure 8 For the present invention Figure 6 A schematic flow chart of an embodiment of step S604;
[0056] Figure 9 A schematic diagram of the structure of an embodiment of a preset fault prediction model provided by the present invention;
[0057] Figure 10 A structural coordinate system diagram of an embodiment of the loss value in the training process provided by the present invention;
[0058] Figure 11 The present invention provides R² An embodiment of a structural coordinate system diagram of a value;
[0059] Figure 12 A schematic diagram of an embodiment of the model prediction effect provided by the present invention;
[0060] Figure 13 This is a schematic structural diagram of an embodiment of the IGBT fault prediction device provided by the present invention. DETAILED DESCRIPTION
[0061] The preferred embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, and are not used to limit the scope of the present invention.
[0062] An insulated gate bipolar transistor (IGBT) is a power switching transistor that combines the advantages of MOSFET and BJT. IGBT is mainly used in power supply and motor control circuits, and has the characteristics of high input impedance, high switching speed and low saturation voltage. IGBT is a three-terminal device, including gate, collector and emitter. Its internal structure is similar to a Darlington configuration, combining the insulated gate structure of MOSFET and the output performance characteristics of BJT. IGBT controls the current between the collector and emitter by controlling the gate voltage to achieve the switching function.
[0063] CEEMDAN (Complete Ensemble Empirical Mode Decomposition with Adaptive Noise) is an improved signal processing technique for processing nonlinear and nonstationary signals. The CEEMDAN algorithm, proposed by Torres et al. in 2011, aims to address the common modal aliasing problem in traditional EMD (Empirical Mode Decomposition) methods. Modal aliasing occurs when signal components of different frequencies interfere with each other, leading to inaccurate decomposition results. By adding adaptive noise to each iteration and performing multiple iterations, CEEMDAN effectively mitigates the impact of white noise on the decomposition results, improving the accuracy and stability of the decomposition.
[0064] The KAN-Transformer prediction model is a time series prediction model that combines the KAN (Kolmogorov-Arnold Networks) and Transformer architectures. KAN is a new type of neural network, inspired by the Kolmogorov-Arnold mathematical theorem. Its characteristic is the application of learnable activation functions to weights, enabling the network to process and learn complex relationships in input data in a more flexible manner. The KAN-Transformer model combines the flexibility and interpretability of KAN with the strong representation and sequence processing capabilities of the Transformer. By using learnable activation functions on the edges of the network, KAN enhances the model's ability to express nonlinear features and interpretability when processing high-dimensional data. The Transformer, on the other hand, handles long-range dependencies through a self-attention mechanism, improving the model's accuracy and efficiency.
[0065] like Figure 1 As shown, a specific embodiment of the present invention discloses an IGBT fault prediction method, comprising:
[0066] S101 , obtaining operating data of an IGBT transistor within a preset time, and determining a time series signal of the preset time; the operating data includes a pole voltage.
[0067] IGBTs, with their excellent high-voltage blocking capability, low on-state voltage drop, and fast switching dynamics, are widely used in various fields, such as smart grids, new energy vehicles, rail transit, and industrial control. In smart grids, IGBTs, as core switching elements, dominate the energy regulation process in key devices such as inverter topologies, motor drive systems, and high-voltage direct current (HVDC). In the power architecture of new energy vehicles, IGBTs, integrated into motor control units (MCUs), form the physical foundation for power conversion in electric drive systems. In rail transit, IGBTs are used for traction control and power conversion, improving the efficiency and reliability of train operations. In industrial control systems, IGBTs are used in motor drives, inverters, and switching power supplies to achieve efficient energy conversion and control. The embodiments of the present invention can be applied to an IGBT fault prediction system. The IGBT fault prediction system can be a software system running on a terminal device. The terminal device can be a server, a tablet computer, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a mobile phone, and other terminal devices. The embodiments of the present application do not impose any restrictions on the specific type of the terminal device.
[0068] It should be understood that the method of obtaining the operating data of the IGBT transistor in step S101 can be a sensor data set obtained by a sensor device, or a sensor data set stored historically in a storage medium. The obtained operating data may include the electrode voltage, that is, the collector-emitter peak voltage. Among them, the IGBT accelerated aging test measurement data is shared on NASA's official website in the form of an open database. This data set records multiple IGBT characteristic parameter variables, such as collector-emitter voltage , gate-emitter voltage , collector current , package temperature, etc. From the perspective of online measurement, versatility, calibration, accuracy, linearity and sensitivity, The comprehensive performance is the best. Extract the collector-emitter peak voltage And it is processed by wavelet noise reduction. The processing results show that during the IGBT degradation process There is an obvious downward trend, which is suitable as a fault parameter for fault prediction. Therefore, in the subsequent process, the fault prediction of IGBT is based on the collector-emitter peak voltage. Processing for fault parameters.
[0069] In a specific embodiment of the present invention, the operating data of the IGBT transistor within a preset time can be obtained. The preset time can be a time determined according to actual conditions, such as 1 month, or 24 hours, etc. In order to be able to measure the electrode-emitter peak voltage To analyze and process according to different time periods, it is necessary to decompose the preset time period to obtain a time series signal. .in, for The time corresponding to the sequence.
[0070] S102 , after performing modal decomposition on the pole voltage according to the time series signal, feature extraction and segmentation are performed on the IMF component data of the pole voltage according to the sliding window method to obtain data within each window.
[0071] A complementary ensemble empirical mode decomposition with adaptive noise (CEEMDAN) module was established to perform modal decomposition on the pole voltage in the time series signal. The pole voltage can have a length of 418, resulting in IMF component data consisting of multiple IMF components and a residual trend term. Based on the acquired IMF component data, a feature extraction module based on a local-global attention mechanism was designed. This feature extraction module combines convolutional layers with local and global attention mechanisms. A sliding window approach was used to extract features and segment the pole voltage IMF component data, thus obtaining the data within each window.
[0072] S103. Perform local feature extraction on the data in each window according to the local multi-head attention mechanism to obtain the local features of each window, and concatenate the local features of all windows to obtain a global feature tensor.
[0073] Among them, local and global multi-head attention mechanisms can be used to extract local and global features respectively, significantly enhancing the expressiveness of time series features. Therefore, by designing local and global multi-head attention mechanisms in the feature extraction module, local features can be extracted and spliced on the IMF component data output by CEEMDAN to obtain a global feature sequence.
[0074] Specifically, the data in each window can be input into the local multi-head attention mechanism to extract the local features of each window. The local features of all windows can be spliced together to form a global feature tensor.
[0075] S104. Extract the global dependency between the time series signal and the global feature tensor according to the global multi-head attention mechanism, obtain the global feature sequence, predict the global feature sequence, obtain the predicted value, and obtain the fault prediction result of the IGBT transistor based on the predicted value.
[0076] Among them, the global dependencies of time series signals can be extracted through the global multi-head attention mechanism, and the extracted global feature sequence can be saved for subsequent fault prediction tasks. Specifically, a KAN-Transformer preset fault prediction model is designed, and the Kolmogorov-Arnold Network is introduced into the traditional Transformer architecture to enhance the model's ability to capture nonlinear relationships. The preset fault prediction model can include an input embedding layer, an encoder (Encoder), a decoder (Decoder) and a KAN layer. The global feature sequence can be input into the preset fault prediction model, and the global feature sequence can be processed starting from the input embedding layer until the KAN layer outputs the predicted value, completing the prediction of the operating data of the IGBT transistor. Then, the fault prediction result of the IGBT transistor can be obtained through the predicted value.
[0077] Compared with the prior art, this embodiment provides the acquisition of operating data of the IGBT transistor within a preset time, and the operating data includes the pole voltage, which can be the collector-emitter voltage, so that fault prediction can be performed using the collector-emitter voltage as the fault characteristic parameter, so that fault prediction can be performed from the perspective of physical detection; the IMF component data obtained after modal decomposition of the pole voltage can also be subjected to local feature extraction through a local multi-head attention mechanism to obtain the local features of each window, so that the accuracy of feature detection can be increased by analyzing the local features of each window; the local features of all windows can also be spliced together, and the global dependency between the time series signal and the global feature tensor can be extracted according to the global multi-head attention mechanism to obtain a global feature sequence, so that a complete feature extraction result can be obtained, and then the global feature sequence can be predicted to improve the robustness of the prediction, thereby improving the practicality of IGBT fault prediction.
[0078] In some embodiments of the present invention, Figure 2 As shown, after step S101, the following steps are further included:
[0079] S201. Optimize the preset noise amplitude and the preset integration quantity based on the Bayesian optimization algorithm to obtain optimal parameters.
[0080] The noise amplitude and the number of integrations can be set. The noise amplitude and the number of integrations can be set according to actual conditions and are not limited in this embodiment of the present invention. To make the noise amplitude and the number of integrations more accurate, a Bayesian optimization algorithm can be used to optimize the noise amplitude and the number of integrations, thereby obtaining optimal parameters for the noise amplitude and the number of integrations. The noise amplitude is used to represent the magnitude of the noise amplitude in the pole voltage, and the number of integrations can be used to represent the number of pole voltage decompositions.
[0081] S202 : Decompose the pole voltage according to the optimal parameters to obtain multiple IMF components.
[0082] Among them, the pole voltage can be decomposed using optimal parameters to obtain the intrinsic mode functions (IMFs) of the signal, that is, multiple IMF components.
[0083] S203 , converting the multiple IMF components into a two-dimensional table according to the time series signal to obtain IMF component data.
[0084] Among them, the multiple IMF components obtained by decomposition can be organized into array form and converted into two-dimensional table data to obtain IMF component data for subsequent analysis. Figure 3 As shown, Figure 3 Figure 6 shows the six IMF components after CEEMDAN decomposition. Each sub-graph corresponds to a single IMF component. The x-axis represents the sampling time step, and the y-axis represents the amplitude of the IMF component. IMF_1 and IMF_2 are the highest-frequency components, with the highest oscillation frequency and relatively large amplitude, which gradually decay over time. IMF_3 and IMF_4 have moderate frequencies, reduced oscillation amplitude, and a more stable oscillation pattern. IMF_5 and IMF_6 are low-frequency components. IMF_6 primarily exhibits long-term trends, with minimal amplitude variation, representing the trend term of the original data signal.
[0085] In some embodiments of the present invention, Figure 4 As shown, step S102 includes:
[0086] S401: Perform preliminary feature extraction on the IMF component data according to the one-dimensional convolution layer to obtain preliminary convolution features.
[0087] Among them, the IMF component data after CEEMDAN decomposition can be input into the one-dimensional convolution layer for preliminary feature extraction to obtain preliminary convolution features.
[0088] S402: Slice the preliminary convolution features according to the sliding window method to obtain data in each window.
[0089] Among them, the sliding window method (window size is 64, step size is 16) can be used to slice the preliminary convolution features to obtain the data in each window.
[0090] In this embodiment of the present invention, all steps of the feature extraction module based on the local-global attention mechanism are as follows: Figure 5 As shown, Figure 5 This is a schematic diagram of the feature extraction module for time series. The specific process is as follows:
[0091] (1) Data loading and preprocessing: The extracted IMF data are integrated into IMF component data and imported into the model, and a tensor containing multiple input sequences is constructed with a shape of [1, 418, 6].
[0092] (2) Convolutional feature extraction: Apply a one-dimensional convolution operation to the input data to extract local features. The shape of the convolution output is [1, 128, 418].
[0093] (3) Local attention extraction: The features of the convolution output are segmented into sliding windows. The size of each window is set to window_size (64), and local features are gradually extracted with a step size of 16. The data dimension in each window is [1, 128, 64]. For the data in each window, the local multi-head attention mechanism is applied to extract the feature distribution of the window. The input data dimension is transformed into [64, 1, 128] as query, key, and value. The attention score is calculated, the features in the window are extracted, and the mean of the features in each window is used as the local feature representation, with a shape of [1, 64, 128]. The features of the time dimension in the window are averaged and pooled and aggregated to obtain the combined local features with a shape of [1, 23, 128].
[0094] (4) Global attention extraction: After concatenating the local features of all windows, the global features are input into the global attention mechanism to further extract the global features of the entire time series. The input shape is [1, 23, 128], which needs to be converted to [23, 1, 128] before calculating the attention. The global dependency relationship between windows is learned through the multi-head attention mechanism. The output is the window feature processed with global dependency, with the shape of [1, 23, 128]. Global attention aggregates the features of each local window in a weighted manner, and the output global features are used as the final representation of the model.
[0095] (5) Feature output: The features extracted by local and global attention are output as two-dimensional feature representations for subsequent model training and performance evaluation. The final global features can be directly applied to subsequent tasks to improve the expressive power of time series features.
[0096] In some embodiments of the present invention, Figure 6As shown, step S104 includes:
[0097] S601. Map the global feature sequence to a high-dimensional latent space according to the feature embedding layer to obtain an embedding layer output feature matrix.
[0098] The model includes an input embedding layer, an encoder, a decoder, and a KAN layer. The processed global feature sequence is used as the input of the encoder, and the input global feature sequence is converted into (in B is the batch size, T is the length of the time series, F is the feature dimension) is mapped to the high-dimensional latent space to obtain the embedding layer output feature matrix, as shown in formula (1):
[0099] (1)
[0100] Where, is the embedding weight matrix, It is the bias vector that adapts the input requirements of the Transformer to obtain the output feature matrix of the embedding layer (in D is the dimension of the hidden layer).
[0101] S602: Globally model the embedding layer output feature matrix according to the encoder to obtain encoder output features.
[0102] Among them, the embedding layer outputs the feature matrix The input is sent to the Transformer encoder to perform global modeling on the embedded features, capture the dependencies between time steps, and obtain the encoder output features.
[0103] S603: Interact with the encoder output features according to the decoder to obtain decoded output features.
[0104] The decoder is similar to the encoder in structure, but with an additional "encoder-decoder attention" module for capturing target features. The decoder interacts its input features with the encoder output features, and learns the relationship between historical features and predicted targets through the attention mechanism. The input of each layer is the features of the previous time step, and after processing by the multi-head attention and feedforward network, the output is the updated feature representation, which is the decoded output feature. .
[0105] S604: Perform feature mapping and prediction on the decoded output features according to the KAN layer to obtain a predicted value.
[0106] The decoded output features may be input to the KAN layer, and the decoded output features may be feature mapped and predicted by the KAN layer, so that a predicted value may be output.
[0107] In some embodiments of the present invention, Figure 7 As shown, step S602 includes:
[0108] S701: Perform global modeling on the embedding layer output feature matrix input to the first layer of the encoder to obtain a first output feature.
[0109] The encoder can include multiple layers, each of which is implemented using a multi-head attention mechanism, a feed-forward network (FFN), residual connections, and layer normalization. The embedding layer output feature matrix input to the encoder can be used as the input feature of the first layer for global modeling, resulting in the first output feature.
[0110] S702: Input the first output feature into the second layer of the encoder for global modeling to obtain a second output feature.
[0111] Among them, the first output feature of the first layer can be used as the input feature of the second layer and then globally modeled again through the multi-head attention mechanism, feedforward network (FFN), residual connection and layer normalization to obtain the second output feature.
[0112] S703: When the second layer is the last layer of the encoder, determine the second output feature as the encoder output feature.
[0113] In a specific embodiment of the present invention, if the second layer is the last layer of the encoder, the second output feature can be determined as the encoder output feature; if the second layer is not the last layer of the encoder, the second output feature of the second layer can be used as the input feature of the third layer and globally modeled again through a multi-head attention mechanism, a feedforward network (FFN), a residual connection, and layer normalization, and so on, until the encoder output feature of the last layer of the encoder is obtained.
[0114] In some embodiments of the present invention, step S701 includes:
[0115] Based on the multi-head attention mechanism, the output feature matrix of the embedding layer input to the first layer of the encoder is calculated to obtain the attention score.
[0116] Among them, the multi-head attention mechanism first calculates the attention score based on the embedding layer output feature matrix input to the first layer of the encoder, as shown in formula (2):
[0117] (2)
[0118] Where, is the query matrix, ; is the bond matrix, ; is the value matrix, ;in, 、 、 is the weight matrix of the linear transformation, Output feature matrix for the embedding layer; is a scaling factor to prevent the softmax gradient from disappearing due to excessive dot product values. , is the number of heads of multi-head attention. In order to capture different types of dependencies, Transformer uses multiple independent attention heads (Multi-HeadAttention) for parallel calculation. O 、 K 、 V Split into h Each head performs the self-attention machine operation independently, and then concatenates the outputs of each head and projects them back to the original dimension through the matrix.
[0119] The attention scores are transformed nonlinearly through the feedforward network to obtain nonlinear features.
[0120] Among them, the attention score Input to the feedforward network for nonlinear transformation, and the nonlinear characteristics can be obtained, as shown in formula (3):
[0121] (3)
[0122] Where, 、 is the linear layer weight; 、 is the bias vector; It is a nonlinear feature.
[0123] The output feature matrix of the embedding layer is residually connected and normalized with the nonlinear features to obtain the first output feature.
[0124] Among them, the embedding layer output feature matrix and nonlinear characteristics Perform residual connection and normalization to obtain the first output feature, as shown in formulas (4) and (5):
[0125] (4)
[0126] (5)
[0127] Where, is the characteristic mean; is the characteristic variance; is a vector; A very small number used to avoid division by zero; is the first output feature.
[0128] Furthermore, after obtaining the first output feature, the first output feature can be used as the input feature of the second layer, and then the input feature is processed in the second layer through all the processes in step S701 to obtain the second output feature of the second layer. The processing process of subsequent encoding layers is similar until all layers of the encoder are processed and the encoder output feature is obtained.
[0129] In some embodiments of the present invention, Figure 8 As shown, step S604 includes:
[0130] S801: Perform basic linear mapping on the decoded output features input to the KAN layer to obtain a basic output relationship.
[0131] Among them, the decoding output features of the input to the KAN layer are obtained through the basic linear mapping Capturing the linear relationship, as shown in formula (6):
[0132] (6)
[0133] Where, is the basic output relationship; is the weight matrix, is the offset of the linear mapping, C is the feature dimension of the output.
[0134] S802: Perform basis function mapping on the decoded output features to obtain basis functions.
[0135] S803: Calculate the weights of the basis functions according to a preset weight matrix to obtain basis function weights.
[0136] Among them, the B-spline basis function mapping can be performed on the decoded output features based on the KAN network model to obtain the basis function and weight, as shown in formulas (7) and (8):
[0137] (7)
[0138] (8)
[0139] Where, is the weight to be learned; is the decoded output feature Basis functions of is the weight matrix; is the basis function weight; is the number of basis functions, which is [1, M ], ,in, is the number of spline networks, is the order of the basis function.
[0140] S804: Obtain a predicted value according to the basic output relationship and basis function weights.
[0141] Among them, the basic output relationship and basis function weight can be calculated to obtain the predicted value, as shown in formula (9):
[0142] (9)
[0143] Where, is the predicted value. The KAN layer can capture the complex dynamic characteristics of time series through adaptive parameter learning, and the generated predicted value represents the future state of the IGBT fault parameters.
[0144] The present invention is implemented as follows Figure 9 As shown, Figure 9 This is the structure diagram of the preset fault prediction model, and the specific structure is as follows:
[0145] (1) Input mapping layer, which uses linear transformation to map input features to a fixed latent space dimension , which can encode the original features into the input format [N, T, F] acceptable to Transformer, where N represents the number of samples, T represents the length of the time series, F Represents the feature dimension, and the shape of the embedded data is [23, 128, 32].
[0146] (2) Transformer encoder, which is composed of multiple stacked Transformer layers. Each layer is first implemented by a multi-head attention mechanism to learn the similarity between time steps. The Query, Key, and Value matrices of the input data are generated through linear transformation, and the attention score is calculated by formula (2). h Parallel attention calculations enhance feature extraction capabilities, and the results are then concatenated and linearly transformed into output features. The features at each time step are then transformed nonlinearly using Formula (3) through two fully connected layers in a feedforward neural network (FFN), yielding nonlinear features. The nonlinear features output by each module are then normalized after adding residual connections to ensure gradient stability and accelerate training. The encoder output shape remains [23, 128, 32].
[0147] (3) Transformer decoder. The decoder has a similar structure to the encoder, but with an additional encoder-decoder attention module to capture target features. The decoder interacts its input features with the encoder output features, learning the relationship between historical features and predicted targets through the attention mechanism. The input of each layer is the features of the previous time step. After multi-head attention and feedforward network processing, the output is an updated feature representation with an output shape of [23, 128, 32].
[0148] (4) KAN layer replaces the traditional fully connected layer and inputs the output features of the last layer of the decoder into KAN. KAN maps the input features through a nonlinear combination model based on the cubic B-spline function and obtains the final prediction value through formula (9). The KAN layer can capture the complex dynamic characteristics of time series through adaptive parameter learning, and the generated prediction value represents the future state of the IGBT fault parameters.
[0149] In some embodiments of the present invention, the operating data also includes measured values; step S104 includes:
[0150] Compare the measured value with the predicted value to obtain the difference;
[0151] When the difference is greater than a preset threshold, it is determined that the IGBT transistor is in a fault state.
[0152] In a specific embodiment of the present invention, the actual measured value of the IGBT transistor at the current moment can be detected, and the predicted value is the value predicted by the IGBT transistor at the next moment. The actual measured value and the predicted value can be compared. If the difference between the actual measured value and the predicted value is greater than a preset threshold, it means that the collector-emitter peak voltage will suddenly drop or rise at the next moment, which is an abnormal situation, so the IGBT transistor is determined to be in a fault state.
[0153] Furthermore, the CEKAT model was established. Its implementation and execution were based on the PyTorch framework, using an Intel(R) Core(TM) i7-1165G7 CPU, 16 GB of RAM, and an NVIDIA GeForceMX450Ti (11 GB) GPU. Code was written to load and preprocess the dataset. The model structure consisted of a feature extraction module and a KAN-Transformer prediction module, using classes and functions to implement the model architecture. After loading the data, the model was assembled and deployed to the GPU device. The Adam optimizer was set up, using the mean squared error (MSE) as the loss function, in preparation for feature extraction and training. The CEKAT model utilizes a pipelined architecture with tightly interconnected modules, seamlessly interoperating through standardized data interfaces. Data is directly passed to the CEEMDAN decomposition module as input for subsequent signal decomposition, ensuring a consistent numerical range for the input signal and minimizing the impact of data fluctuations on decomposition accuracy. The IMF components output by the CEEMDAN decomposition module serve as input to the feature extraction module and are stored in a two-dimensional table, with each column corresponding to an IMF component and each row representing a sampling point in the time series. This unified format facilitates subsequent processing. The local and global features generated by the feature extraction module are output as tensors and directly fed into the KAN-Transformer prediction module. The high-dimensional representation of the features fully captures key patterns in the time series. This modular combination allows each component to independently perform its own functions while enabling efficient data transfer and processing through standard interfaces, ensuring the robustness and efficiency of the overall process.
[0154] The CEKAT model is built on the PyTorch framework, utilizing a modular design for efficient integration. First, a feature extraction module and a KAN-Transformer module are created. The construction process is implemented using classes and functions to ensure modularity and reusability. Each module of the model is connected via standardized data interfaces, and input tensors are uniformly mapped to latent space dimensions to facilitate subsequent modeling. Using PyTorch's module encapsulation, all submodules are integrated into the complete CEKAT framework. After model initialization, the model is deployed to a GPU device. The resulting model is highly scalable and efficient, laying the foundation for subsequent training and evaluation.
[0155] The model is trained using the IGBT accelerated aging dataset, and the features and target values are normalized. The data is divided into a training set and a test set using the sliding window method, with the training set accounting for 80% and the test set accounting for 20%. The prediction length is 80. The model supports multiple architecture options and parameter configurations (such as hidden space dimensions, number of attention heads, number of encoder layers, etc.). The mean square error is used as the loss function, and the model parameters are optimized as shown in formula (10):
[0156] (10)
[0157] Where, is the actual value; is the predicted value of the model.
[0158] The Adam optimizer updates the parameters and records the loss value in each iteration loss and R² Coefficient of determination. R² The calculation is shown in formula (11):
[0159] (11)
[0160] Finally, the visualization module is used to draw a comparison chart of training and test results, a training curve, and save the model performance indicators to a file for subsequent analysis, as shown in the attached file. Figure 10 For the training process loss Loss value, attached Figure 11 for R² value, Figure 10 The x-axis is the number of training rounds, and the y-axis is the loss value. Figure 11 The x-axis is the number of training rounds, and the y-axis is R² value.
[0161] Load the best model weights saved in the training phase into the model framework to ensure that the optimized parameters are used during the test. Input the test set into the model for prediction and generate the predicted value. , draw the time series curves of the predicted value and the true value in the same figure, and intuitively show the prediction effect of the model on the test set as shown in the attached figure. Figure 12 , the solid line in the figure is the true value, and the dotted line is the predicted value. Figure 12 It can be seen that the overall changes in the predicted data are roughly similar to the true values, and the errors between the numerical values and data values of each specific point are also small, indicating that this method can effectively predict the next changes in the fault parameters, thereby giving an early warning when the IGBT is about to fail, reminding management personnel to repair or replace it in time, thereby ensuring the safe operation of the overall system.
[0162] Once the model's training and testing results meet the expected performance indicators, save the model's parameters and structure to a file. When predicting IGBT faults later, use the same model structure code used for training and restore the model by loading the saved weight file. The loaded model does not require retraining and can be used directly for prediction.
[0163] In order to better implement the IGBT fault prediction method in the embodiment of the present invention, based on the IGBT fault prediction method, the embodiment of the present invention also provides an IGBT fault prediction device, such as Figure 13As shown, the IGBT fault prediction device 1300 includes:
[0164] The data acquisition module 1301 is used to acquire the operating data of the IGBT transistor within a preset time and determine the time series signal of the preset time; the operating data includes the electrode voltage;
[0165] A feature extraction module 1302 is configured to perform modal decomposition on the pole voltage according to the time series signal, and then perform feature extraction and segmentation on the IMF component data of the pole voltage according to a sliding window method to obtain data within each window;
[0166] A feature splicing module 1303 is used to extract local features from the data in each window according to the local multi-head attention mechanism to obtain local features of each window, and to splice the local features of all windows to obtain a global feature tensor;
[0167] The fault prediction module 1304 is used to extract the global dependency between the time series signal and the global feature tensor according to the global multi-head attention mechanism, obtain the global feature sequence, predict the global feature sequence, obtain the predicted value, and obtain the fault prediction result of the IGBT transistor according to the predicted value.
[0168] The IGBT fault prediction device 1300 provided in the above embodiment can implement the technical solution described in the above IGBT fault prediction method embodiment. The specific implementation principles of the above modules or units can be found in the corresponding contents of the above IGBT fault prediction method embodiment, which will not be repeated here.
[0169] The IGBT fault prediction method and device provided by the present invention are introduced in detail above. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.
Claims
1. A method for predicting IGBT failure, characterized in that: include: Acquire operating data of the IGBT transistor within a preset time, and determine a time series signal of the preset time; The operating data includes pole voltage; After performing modal decomposition on the pole voltage according to the time series signal, feature extraction and slicing are performed on the IMF component data of the pole voltage according to a sliding window method to obtain data within each window; Extract local features from the data in each window according to the local multi-head attention mechanism to obtain local features of each window, and concatenate the local features of all windows to obtain a global feature tensor; The global dependency relationship between the time series signal and the global feature tensor is extracted according to the global multi-head attention mechanism. After obtaining the global feature sequence, the global feature sequence is predicted to obtain a predicted value, and the fault prediction result of the IGBT transistor is obtained based on the predicted value.
2. The IGBT fault prediction method according to claim 1, characterized in that: After obtaining the operating data of the IGBT transistor within a preset time and determining the time series signal of the preset time, the method further includes: The preset noise amplitude and the preset integration number are optimized based on the Bayesian optimization algorithm to obtain the optimal parameters; Decomposing the pole voltage according to the optimal parameters to obtain multiple IMF components; The multiple IMF components are converted into a two-dimensional table according to the time series signal to obtain IMF component data.
3. The IGBT fault prediction method according to claim 1, wherein: The feature extraction and slicing of the IMF component data of the pole voltage according to the sliding window method to obtain data within each window includes: Performing preliminary feature extraction on the IMF component data according to a one-dimensional convolution layer to obtain preliminary convolution features; The preliminary convolution features are sliced according to the sliding window method to obtain data in each window.
4. The IGBT fault prediction method according to claim 1, wherein: The predicting of the global feature sequence to obtain a predicted value includes: Mapping the global feature sequence to a high-dimensional latent space according to a feature embedding layer to obtain an embedding layer output feature matrix; Performing global modeling on the embedding layer output feature matrix according to the encoder to obtain encoder output features; interacting the encoder output features according to the decoder to obtain decoded output features; The decoded output features are feature mapped and predicted according to the KAN layer to obtain a predicted value.
5. The IGBT fault prediction method according to claim 4, characterized in that: The globally modeling the embedding layer output feature matrix according to the encoder to obtain the encoder output feature includes: Performing global modeling on the embedding layer output feature matrix input to the first layer of the encoder to obtain a first output feature; Inputting the first output feature into the second layer of the encoder for global modeling to obtain a second output feature; When the second layer is the last layer of the encoder, the second output feature is determined as the encoder output feature.
6. The IGBT fault prediction method according to claim 5, characterized in that: The globally modeling the embedding layer output feature matrix input to the first layer of the encoder to obtain a first output feature includes: Calculating the embedding layer output feature matrix input to the first layer of the encoder based on a multi-head attention mechanism to obtain an attention score; Performing a nonlinear transformation on the attention score through a feedforward network to obtain a nonlinear feature; Performing residual connection and normalization processing on the embedding layer output feature matrix and the nonlinear feature to obtain the first output feature.
7. The IGBT fault prediction method according to claim 4, characterized in that: The performing feature mapping and prediction on the decoded output features according to the KAN layer to obtain a predicted value includes: Performing a basic linear mapping on the decoded output features input to the KAN layer to obtain a basic output relationship; Performing basis function mapping on the decoded output features to obtain basis functions; Performing weight calculation on the basis function according to a preset weight matrix to obtain basis function weights; A predicted value is obtained according to the basic output relationship and the basis function weights.
8. The IGBT fault prediction method according to claim 1, characterized in that: The operating data also includes measured values; and the fault prediction result of the IGBT transistor obtained according to the predicted value includes: Comparing the measured value with the predicted value to obtain a difference; When the difference is greater than a preset threshold, it is determined that the IGBT transistor is in a fault state.
9. The IGBT fault prediction method according to claim 6, characterized in that: The calculation formula of the attention score is: Where, is the query matrix, ; is the bond matrix, ; is the value matrix, ; is the scaling factor; where 、 、 is the weight matrix of linear transformation; Output feature matrix for the embedding layer.
10. An IGBT fault prediction device, characterized in that: include: A data acquisition module is used to acquire the operating data of the IGBT transistor within a preset time and determine a time series signal of the preset time; The operating data includes pole voltage; a feature extraction module, configured to perform modal decomposition on the pole voltage according to the time series signal, and then perform feature extraction and segmentation on the IMF component data of the pole voltage according to a sliding window method to obtain data within each window; A feature splicing module is used to extract local features of the data in each window according to the local multi-head attention mechanism to obtain local features of each window, and to splice the local features of all windows to obtain a global feature tensor; A fault prediction module is used to extract the global dependency between the time series signal and the global feature tensor according to the global multi-head attention mechanism, obtain the global feature sequence, predict the global feature sequence, obtain the predicted value, and obtain the fault prediction result of the IGBT transistor according to the predicted value.
Citation Information
Patent Citations
IGBT fault prediction method and system based on inverted Transform network
CN118761444A
Complex device fault diagnosis method and system based on multi-dimensional features
US12314149B1