Reverse conducting IGBT intelligent power module fault automatic diagnosis method and system
By using multi-source data analysis and adaptive clustering algorithms to identify faults in inverse-conducting IGBT intelligent power modules, the problem of insufficient real-time performance and accuracy in existing technologies is solved, enabling efficient fault identification and adaptive diagnosis, and improving the reliability of power electronic systems.
Patent Information
- Application Number
- CN202511483224.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-10-17
AI Technical Summary
Existing fault diagnosis methods for reverse-conducting IGBT smart power modules rely on traditional offline detection or single-parameter monitoring, which are difficult to meet the requirements of real-time performance, accuracy, and adaptability. They cannot effectively identify fault modes under complex operating conditions and lack a systematic fault feature library.
By acquiring multi-source monitoring data, a dynamic feature extraction model is established to perform joint analysis in the time and frequency domains, a fault feature space is constructed, and an adaptive clustering algorithm is used to identify abnormal feature clusters. Combined with a pattern matching engine, fault type identification is performed, and the fault feature library is updated in real time to adapt to new fault modes.
It enables efficient and accurate identification of faults in reverse-conducting IGBT intelligent power modules, improves the real-time performance and adaptability of fault diagnosis, reduces misjudgments and omissions, and ensures the stable operation of power electronic systems.
Smart Images

Figure CN120948950A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power electronic device diagnostic technology, specifically to an automatic fault diagnosis method and system for reverse-conducting IGBT intelligent power modules. Background Technology
[0002] In power electronic systems, reverse-conducting IGBT intelligent power modules serve as core power conversion units, widely used in new energy power generation, industrial frequency conversion, rail transportation, and smart grids. Their operational stability directly determines the reliability of the entire system. With the increasing demands for power density and switching frequency in application scenarios, reverse-conducting IGBT intelligent power modules operate under high voltage, high current, and complex thermal cycling conditions for extended periods. This makes them prone to faults such as gate oxide layer damage, collector-emitter leakage, and increased package thermal resistance. If these faults are not diagnosed and addressed promptly, they can lead to module performance degradation, or even system shutdown, equipment burnout, or safety accidents.
[0003] Current fault diagnosis methods for reverse-conducting IGBT smart power modules mostly rely on traditional offline testing or single-parameter monitoring. Offline testing requires disassembling the module from the system, which is not only time-consuming and labor-intensive, but also fails to capture dynamic fault characteristics during module operation, making it difficult to meet real-time diagnostic needs. Single-parameter monitoring often focuses on a single indicator such as case temperature or collector current, ignoring the correlation characteristics between multiple physical quantities under different fault modes. For example, gate failure may simultaneously cause gate voltage waveform distortion and collector current overshoot; monitoring only a single parameter is prone to misjudgment or missed detection. In addition, in existing diagnostic methods, feature extraction is often limited to a single dimension in time-domain or frequency-domain analysis, making it difficult to comprehensively characterize fault information in multi-source monitoring data, resulting in insufficient targeting of the generated feature parameters. The fault identification process often uses fixed thresholds or traditional clustering algorithms, which cannot adapt to feature drift caused by fluctuations in operating conditions during module operation, resulting in low accuracy in identifying abnormal feature clusters under complex operating conditions.
[0004] Traditional diagnostic methods rely heavily on manually preset, simple feature matching rules for fault comparison, lacking a systematic fault feature database. This results in poor adaptability to novel fault modes and an inability to accurately identify fault types. These issues lead to significant shortcomings in the real-time performance, accuracy, and adaptability of existing fault diagnosis methods for reverse-conducting IGBT intelligent power modules. These methods fail to meet the high demands of modern power electronic systems for fault early warning and health management. Therefore, there is an urgent need for an automated diagnostic method that can integrate multi-source monitoring data, achieve dynamic feature extraction, and adaptive fault identification to improve the efficiency and reliability of module fault diagnosis and ensure the stable operation of power electronic systems. Summary of the Invention
[0005] The purpose of this invention is to provide an automatic fault diagnosis method and system for reverse-conducting IGBT intelligent power modules to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides an automatic fault diagnosis method for reverse-conducting IGBT intelligent power modules, the method comprising: Acquire multi-source monitoring data during the operation of the power module, including gate voltage waveform, collector current waveform, and case temperature change curve; A dynamic feature extraction model is established to perform joint time-domain and frequency-domain analysis on the multi-source monitoring data, generating a set of feature parameters. Construct a fault feature space and map the set of feature parameters to a high-dimensional space to form a feature vector distribution map; An adaptive clustering algorithm is used to divide the feature vector distribution map into regions and identify areas of anomalous feature clusters. The abnormal feature clusters are compared with a preset fault feature database using a pattern matching engine, and the fault type identification result is output.
[0007] Preferably, the establishment of the dynamic feature extraction model includes: The rising and falling edge features of the gate voltage waveform are extracted, and the voltage change rate during the switching process is calculated. Analyze the conduction and turn-off characteristics of the collector current waveform to extract the current overshoot amplitude and oscillation frequency; Monitor the slope change of the shell temperature change curve and calculate the dynamic response parameters of thermal resistance; The voltage change rate, current overshoot amplitude, oscillation frequency, and thermal resistance dynamic response parameters are combined to form a set of characteristic parameters.
[0008] Preferably, the construction of the fault feature space includes: The set of feature parameters is standardized to eliminate dimensional differences; A nonlinear dimensionality reduction method is used to project the standardized feature parameters onto a three-dimensional space. Based on the distribution patterns of historical fault data, the boundaries of the normal operating condition area are marked in three-dimensional space; Calculate the distance between the current feature vector and the boundary of the normal operating condition area, and generate a feature vector distribution map.
[0009] Preferably, the adaptive clustering algorithm includes: The number of cluster centers is automatically determined based on the feature vector density distribution. Set the dynamic neighborhood radius parameter to adjust the clustering region boundaries; Calculate the feature vector dispersion of each cluster region and filter out abnormal feature clusters with dispersion exceeding the threshold.
[0010] Preferably, the comparison using a pattern matching engine includes: A multi-level fault feature tree is established, which includes three types of fault modes: device aging, gate failure, and thermal runaway. Extract the geometric feature parameters of the abnormal feature cluster region, including region area, shape factor, and center position; The similarity between the geometric feature parameters and the standard patterns in the fault feature tree is calculated. The fault type corresponding to the standard pattern with the highest similarity was selected as the identification result.
[0011] Preferably, establishing a multi-level fault feature tree includes: Collect typical fault sample data and extract common features to form basic fault patterns; Based on the failure development process, severity levels are classified, and failure mode evolution paths are established; Set feature weight coefficients to distinguish the influence of primary and secondary features.
[0012] Preferably, the method further includes: The fault feature database is updated in real time, and when a new cluster of abnormal features is detected: Record the characteristic parameters and operating environment data of new anomaly clusters; The hazard level of newly emerging anomalous feature clusters is assessed using an expert system. The newly identified fault modes are added to the fault feature tree.
[0013] Preferably, the real-time updating of the fault feature database includes: Set up a feature change monitoring window to track the long-term evolution trend of feature parameters; Establish a feature drift detection mechanism to identify gradual changes in parameter distribution patterns; When a significant feature drift is detected, the fault feature library reconstruction process is triggered.
[0014] Preferably, the method further includes: constructing a fault prediction model based on the evolution trend of feature parameters: Calculate the degradation rate of key characteristic parameters; predict remaining service life; generate preventative maintenance recommendations.
[0015] Preferably, the method further includes an automatic fault diagnosis system for a reverse-conducting IGBT intelligent power module, comprising a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of the above-described automatic fault diagnosis method for a reverse-conducting IGBT intelligent power module.
[0016] Compared with the prior art, the beneficial effects of the present invention are: The automatic fault diagnosis method for reverse-conducting IGBT smart power modules provided by this invention effectively solves the shortcomings of traditional diagnostic methods in terms of real-time performance, accuracy, and adaptability by integrating multi-source monitoring data and constructing a systematic diagnostic process, providing a better solution for fault diagnosis of reverse-conducting IGBT smart power modules. In the data acquisition stage, this method selectively acquires multi-source monitoring data, including gate voltage waveforms, collector current waveforms, and case temperature change curves, rather than relying on a single parameter. This comprehensively captures the electrical and thermal characteristics of the module during operation, avoiding the omission of fault features due to limitations in monitoring dimensions. It provides a more comprehensive foundation for subsequent diagnosis from the data source, enabling the diagnostic process to cover the multi-physical quantity correlation characteristics under different fault modes of the module, reducing misjudgments or omissions caused by incomplete information.
[0017] In the feature extraction stage, this method establishes a dynamic feature extraction model to perform joint time-domain and frequency-domain analysis on multi-source monitoring data, rather than being limited to feature mining in a single dimension. This allows for a more thorough extraction of fault information contained in the monitoring data. Time-domain analysis can capture dynamic transient features such as voltage change rate and current overshoot amplitude, while frequency-domain analysis can identify implicit features such as oscillation frequency. The feature parameter set generated by the combination of the two is more targeted and representative, and can accurately reflect subtle changes in the module's operating status. This provides a high-quality feature foundation for subsequent fault feature space construction and anomaly identification, avoiding a decrease in diagnostic accuracy due to insufficient feature representation.
[0018] In the fault feature space construction and anomaly identification stages, this method maps the feature parameter set to a high-dimensional space to form a feature vector distribution map, and uses an adaptive clustering algorithm for region division, rather than relying on a fixed threshold or traditional clustering methods. The adaptive clustering algorithm can automatically determine the number of cluster centers and adjust the neighborhood radius parameter based on the feature vector density distribution, effectively adapting to feature drift caused by fluctuations in operating conditions during module operation, accurately dividing normal and abnormal feature regions, precisely identifying abnormal feature clusters, avoiding misjudgments of abnormal regions caused by changes in operating conditions, and improving adaptability to complex operating conditions.
[0019] In the fault type identification stage, this method compares the abnormal feature clusters with the preset fault feature library through a pattern matching engine, rather than relying on simple matching rules preset by humans. The systematic fault feature library can provide standardized comparison basis for different fault modes, ensuring the standardization and accuracy of fault type identification. At the same time, the architecture of the preset fault feature library also provides a foundation for the subsequent expansion of fault modes, making the method scalable when facing new faults. Attached Figure Description
[0020] Figure 1This is a schematic diagram illustrating the working principle of the automatic fault diagnosis method for reverse-conducting IGBT intelligent power modules described in this invention. Figure 2 A schematic diagram illustrating the working principle of the dynamic feature extraction model; Figure 3 A schematic diagram of the working principle for constructing the fault feature space; Figure 4 This is an adaptive clustering selection graph. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Please see Figure 1 This invention provides an automatic fault diagnosis method and system for reverse-conducting IGBT intelligent power modules. The method includes: during power module operation, real-time acquisition of multi-source monitoring data via a sensor array, including key parameters such as gate voltage waveform, collector current waveform, and case temperature change curve; establishing a dynamic feature extraction model to perform joint time-domain and frequency-domain analysis on the multi-source monitoring data. The time-domain analysis focuses on the time-series characteristics of the waveforms, while the frequency-domain analysis uses Fast Fourier Transform to extract harmonic components, thereby generating a comprehensive set of feature parameters; constructing a fault feature space, mapping the feature parameter set to a high-dimensional space through standardization and nonlinear dimensionality reduction techniques to form a visualized feature vector distribution map, which clearly reflects the correlation between parameters; employing an adaptive clustering algorithm to intelligently divide the feature vector distribution map into regions, with the algorithm automatically adjusting the cluster boundaries based on vector density to identify abnormal feature clusters corresponding to potential fault modes; comparing the abnormal feature clusters with a preset fault feature library using a pattern matching engine, which contains data on historical fault modes, and calculating the similarity and outputting the fault type identification result, such as device aging or gate failure, thus achieving automated diagnosis.
[0023] Example 1: See Figure 2The implementation of the dynamic feature extraction model begins with the fine processing of the gate voltage waveform. Its core lies in capturing subtle changes during the switching transient process. The system captures the voltage signal at an extremely high sampling rate through a high-voltage differential probe deployed on the gate drive circuit. This sampling rate is usually set to be accurate enough to distinguish the leading edge of a nanosecond-level pulse, thereby ensuring the authenticity of the original data. The acquired original voltage waveform is first digitally filtered to suppress high-frequency noise interference. Then, a sliding window algorithm is used to identify the key turning points of the waveform, namely the starting point of the rising edge and the ending point of the falling edge. After accurately defining these time boundaries, the voltage change rate during the switching process is calculated. This calculation is not a simple endpoint difference, but a linear fitting of the waveform edge based on the least squares method to obtain its slope as a quantitative indicator of the change rate. This method can effectively reduce the error caused by random fluctuations and reflect the average rate characteristics of voltage change. The analysis of collector current waveform focuses on its dynamic response characteristics. Rogowski coils or Hall effect current sensors are coupled to the main power circuit to acquire current signals in a non-invasive manner. The conduction characteristic analysis of the current waveform focuses on the range from zero to steady-state value, emphasizing the extraction of the absolute value of its rise time and the smoothness of the rise edge. The analysis of turn-off characteristics is more complex, requiring attention to current tailing phenomena and current spikes that may occur at the moment of turn-off. The extraction of current overshoot amplitude is achieved by comparing the theoretical current value at the moment of turn-off with the actual peak current acquired. Its calculation involves the detection of the zero-crossing point of the first derivative of the current waveform. The extraction of oscillation frequency requires local spectral analysis of the current waveform after turn-off. Short-time Fourier transform or wavelet transform is usually used to identify the dominant oscillation frequency components and their decay characteristics. These parameters together reveal the stress state of the power switching device under hard-switching or soft-switching conditions.
[0024] Monitoring the shell temperature change curve relies on high-precision thermocouples or thermistors mounted on the module substrate or shell surface. Although the temperature signal acquisition frequency is much lower than that of electrical signals, its long-term trend contains crucial information. Calculating the thermal resistance dynamic response parameter requires correlating the shell temperature change with simultaneous power loss data. Power loss can be calculated from the acquired voltage and current waveforms. By analyzing the shell temperature response curve under a specific power step change, a first-order thermal network model is used to fit the heating and cooling process, thereby extracting the thermal resistance and thermal capacity parameters characterizing heat dissipation capability. The time constant exponent of the thermal resistance dynamic response parameter is an indicator of the module's temperature response speed; it is obtained from fitting the shell temperature change curve to the first-order thermal network model. Specifically, this exponent is the product of the thermal resistance and thermal capacity values obtained from the model fitting. During the model fitting process, the thermal resistance and thermal capacity are adjusted through optimization algorithms to ensure the theoretical heating curve output by the model best matches the measured curve. The product of the thermal resistance and thermal capacity determined at this point is the time constant. This index quantifies the module's thermal inertia; a larger index value indicates a slower temperature response, which helps assess the module's thermal behavior under different power change frequencies. This dynamic parameter reflects the module's thermal behavior under actual operating conditions better than the static thermal resistance value. Effectively combining the characteristic parameters extracted from different physical quantities is a key step in forming the characteristic parameter set. This combination process is not a simple splicing but requires consideration of the dimensional differences and physical meanings of each parameter. First, the voltage change rate, current overshoot amplitude, oscillation frequency, and thermal resistance dynamic response parameters are pre-normalized to convert them into dimensionless relative values. The reference values for normalization are derived from the rated values specified in the device datasheet or statistical averages under a large number of historical normal operating conditions. Then, appropriate weighting coefficients are assigned to each parameter based on its sensitivity to faults. The weighting allocation for characteristic parameter fusion refers to assigning a coefficient value reflecting the relative importance of different characteristic parameters such as the voltage change rate and current overshoot amplitude. The weighting allocation is based on the sensitivity of each parameter to fault mode discrimination. Sensitivity is determined by analyzing historical fault datasets, calculating the degree of deviation of each parameter from the normal state before and after a fault occurs. Parameters with large and stable deviations receive higher weights. In the technical solution, the normalized value of each feature parameter is multiplied by its corresponding weighting coefficient. Then, all weighted parameter values are summed or combined into a comprehensive feature index. This process highlights the contribution of key parameters to the overall feature. For example, an abnormal voltage change rate may directly indicate a gate drive fault and is therefore given a higher weight, while a relatively slow change in thermal resistance may have a slightly lower weight. The final set of feature parameters is a structured data vector, which serves as a digital fingerprint of the power module's health status.
[0025] At the hardware implementation level, the aforementioned feature extraction functions are typically integrated into a dedicated signal conditioning and processing unit. This unit includes a multi-channel synchronous acquisition circuit, a high-performance floating-point digital signal processor, and a static memory for temporarily storing waveform data. The signal conditioning circuit is responsible for amplifying, level-shifting, and anti-aliasing filtering the raw signal output from the sensor to ensure the signal quality fed into the ADC. The embedded software running on the digital signal processor implements the aforementioned feature extraction algorithms. The software adopts a modular design, encapsulating voltage analysis, current analysis, and temperature analysis into independent task modules, which are scheduled by a real-time operating system to ensure timely processing. The extracted feature parameter set is uploaded to the main control unit via a communication interface (such as CAN bus or Ethernet) for subsequent analysis. For example, under different conditions such as module startup, light load operation, or overload, the normal range of certain feature parameters (such as current overshoot amplitude) will differ. Therefore, a working condition identification module is preset in the system. This module classifies the current working condition based on operating parameters such as average current and switching frequency. The discrimination threshold parameter of the working condition identification module is a boundary value used to classify the current operating state, such as the current or power threshold for distinguishing between light load, rated load, and overload. These parameters are obtained by analyzing the equipment's historical operating data and combining it with its rated specifications. Specifically, the typical distribution range of average current and switching frequency under different load conditions is statistically analyzed, and boundary values that clearly distinguish different operating conditions are selected. In the technical solution, these parameters are used for real-time comparison: the currently measured average current and switching frequency are compared with preset threshold parameters to determine the current operating condition category and select the corresponding normal reference range for the evaluation of characteristic parameters. This dynamic adjustment mechanism enhances the robustness and accuracy of the feature extraction model under different application scenarios, enabling subsequent fault diagnosis to be based on contextual information that is closer to the actual situation.
[0026] Example 2: See Figure 3The process of constructing the fault feature space begins with the standardization preprocessing of the feature parameter set. This step aims to eliminate the numerical differences caused by different physical dimensions, so that subsequent analysis is not affected by the absolute value of the parameters. The Z-score standardization method is used to perform a linear transformation on each feature parameter, calculate its deviation from the mean of its sample set, and measure it in terms of standard deviation. After processing, each feature parameter will follow a standard normal distribution with a mean of zero and a standard deviation of one. This processing enables parameters with different dimensions, such as voltage change rate and current overshoot amplitude, to be compared and calculated on the same scale. After standardization, the nonlinear dimensionality reduction stage begins. In this stage, the t-distributed random neighborhood embedding algorithm is chosen as the main tool. This algorithm is particularly suitable for processing the local structural relationships of high-dimensional data. During implementation, an appropriate perplexity parameter needs to be set. This parameter is essentially a smooth measure of the number of neighbors considered by the algorithm for each point. Through iterative optimization, the probability distribution of similarity between data points in the high-dimensional and low-dimensional spaces is continuously adjusted to minimize the KL divergence between them. Finally, the set of feature parameters with dozens or even more dimensions is projected onto the visualized three-dimensional space. Each point in this three-dimensional space uniquely corresponds to an original feature vector, and the relative distance between points reflects the degree of similarity between the original features. Based on the three-dimensional projection space, it is necessary to mark the boundary of the normal operating condition area. This work relies on the accumulation of a large amount of historical normal operation data. First, the feature vectors generated by the system's long-term operation under known normal conditions are collected. These normal data are also standardized and dimensionality reduced before being projected onto the same three-dimensional space. Then, a density-based spatial clustering method is used to analyze the distribution of these normal data points. The algorithm will automatically identify the region with the densest normal data points and generate a minimum convex hull surface that can enclose the vast majority of normal points. This surface is the boundary of the normal operating condition area. Any new data point falling outside the boundary may indicate an abnormal state.
[0027] The process of calculating the distance between the current feature vector and the boundary of the normal operating condition region involves spatial geometric operations. For each newly generated feature vector point in three-dimensional space, the system calculates its shortest Euclidean distance to the boundary surface of the normal region. The sign of this distance value is used to distinguish the positional relationship of the point: a positive sign indicates that the point is outside the normal region, and a negative sign indicates that the point is inside. The absolute value of the distance reflects the degree of abnormality or normality. All the spatial positions and boundary distance information of the current feature vectors together constitute a dynamic feature vector distribution map, which is updated in real time as new data continues to flow in. Parameter tuning of the nonlinear dimensionality reduction algorithm is an important step. The learning rate parameter in the t-SNE algorithm needs to be appropriately adjusted according to the data scale to prevent the optimization process from getting trapped in local optima. The number of iterations needs to be balanced between computational accuracy and efficiency. Usually, hundreds to thousands of iterations are required to obtain stable projection results. For large-scale datasets, stratified sampling methods can be used to preprocess the data to improve the efficiency of dimensionality reduction calculations. The accuracy of the normal area boundary marking directly affects the sensitivity and specificity of fault detection. If the boundary is too loose, it will lead to an increase in missed detections, while if the boundary is too strict, it may produce false alarms. Therefore, the tightness of the boundary needs to be adjusted according to the actual application scenario. In the early stage of system operation, the normal area boundary can be set relatively loosely. As normal operation data is accumulated, the boundary range can be gradually tightened. This gradual adjustment strategy can improve the adaptability of the system in the early stage of deployment.
[0028] The generation of feature vector distribution maps is not a one-time process but requires a continuously updating mechanism. The system maintains a sliding time window, retaining only the feature vector data from the most recent period for map generation. This ensures the distribution map reflects the latest system state while controlling memory usage. The visualization of the distribution map helps maintenance personnel intuitively understand the system's health status and provides a direct reference for fault diagnosis. The entire fault feature space construction process is designed as a closed-loop system. As new normal operation data accumulates, the normal region boundaries and feature vector distribution maps are updated accordingly. This adaptive mechanism allows the system to automatically adjust the normal state baseline as the equipment ages, avoiding false alarms caused by natural equipment aging while maintaining fault detection sensitivity. In engineering implementation, the construction of the fault feature space is typically completed by a separate computation module. This module receives the feature parameter set from the feature extraction module, performs a series of processes such as standardization, dimensionality reduction, boundary calculation, and distance measurement, and finally outputs a feature vector distribution map with distance annotations. The distribution map data is then fed into a subsequent clustering analysis module for further processing.
[0029] Example 3: The implementation of the adaptive clustering algorithm automatically determines the number of cluster centers based on the density distribution of feature vectors in space. This method identifies potential core regions by analyzing the local concentration of points. The density distribution is calculated using kernel density estimation technology. A Gaussian kernel function is used to perform weighted smoothing on the neighborhood of each feature vector point, thereby obtaining a continuous density surface. The bandwidth parameter of the kernel function is adaptively adjusted according to the overall size of the dataset to avoid over-smoothing or under-smoothing problems. The process of automatically determining the number of cluster centers depends on the detection of density peaks. The algorithm scans the entire density surface and identifies those points whose local density is significantly higher than the surrounding points and have sufficient distance from higher density points as candidate centers. The distance threshold is dynamically set by multiplying the average distance between data points by a scaling factor to ensure that the selection of center points is neither too dense nor too sparse. This automatic mechanism eliminates the subjectivity of manually setting the number of clusters and adapts to the changes in feature vector distribution under different operating conditions.
[0030] Setting the dynamic neighborhood radius parameter is a crucial step in the clustering process. The neighborhood radius directly affects the boundary delineation of clustered regions and the sensitivity of outlier identification. The initial radius value is calculated based on the global distribution characteristics of feature vectors, such as using the median distance between all point pairs as a reference. The dynamic adjustment mechanism is implemented by monitoring real-time changes in the data flow. For example, when new feature vectors continuously flow in, the system periodically recalculates the average density of the vectors. If the density increases, the radius is appropriately reduced to capture more refined local structures; conversely, the radius is expanded to avoid misclassifying normal points as outliers. The dynamic update of the radius adopts a sliding window method, and the window size is synchronized with the data sampling frequency to ensure timely adjustment. In addition, the radius parameter is also related to the stability of the cluster centers. If the position of the center point changes too much in continuous iterations, the system will trigger a recalibration of the radius to maintain the consistency of clustering.
[0031] When calculating the feature vector dispersion of each cluster region, a quantitative index is introduced to measure the degree of dispersion of the vector within the cluster. The dispersion formula is defined as:
[0032] in: Represents clustering regions The greater the value of the dispersion, the more dispersed the distribution of feature vectors within that region. It is clustering The total number of feature vectors contained in it; Representative clustering The first in 1 eigenvector; It is clustering The center vector is obtained by calculating the arithmetic mean of all eigenvectors within the region; The Euclidean norm of a vector is used to measure the geometric distance between vectors. The calculation of the dispersion is performed immediately after the initial clustering, with each cluster calculating its own dispersion independently. The vector data required for the calculation is retrieved from a real-time cache to ensure the timeliness of the results.
[0033] The filtering of anomalous feature clusters with dispersion exceeding a threshold relies on a preset threshold mechanism. The threshold value is determined based on the statistical distribution of cluster dispersion in historical normal operation data. For example, the 95th percentile of the dispersion value under normal conditions is used as a baseline. The system periodically updates the threshold to reflect the impact of equipment aging or environmental changes. The filtering process traverses all identified cluster regions and compares each region. The system compares the current value with the current threshold. If the dispersion of a certain region is significantly higher than the threshold, it is marked as an abnormal cluster area. The geometric features of the abnormal region, such as its area and shape, are extracted for subsequent analysis. In addition, the trend of dispersion is also monitored. If the dispersion of a region continues to rise within a continuous time window, the system will increase its abnormal priority to ensure that potential faults can be detected as early as possible.
[0034] The clustering algorithm is embedded in a dedicated data processing unit equipped with a multi-core processor to process the computation of high-dimensional feature vectors in parallel. During algorithm iteration, an incremental update strategy is employed; when a new feature vector arrives, the system does not recalculate the entire cluster but only adjusts the boundaries of the affected regions. This significantly improves computational efficiency to meet real-time requirements. Density distribution calculation utilizes spatial index structures such as KD-trees to accelerate neighborhood queries, enabling the processing of large-scale vector data to be completed in milliseconds. Dynamic neighborhood radius parameter adjustment is achieved through closed-loop control, with feedback signals from clustering quality evaluation metrics such as the silhouette coefficient, ensuring the radius value remains within the optimal range. The identification results of anomalous feature clusters are combined with the system's learning mechanism. When a new anomalous pattern appears, the algorithm records its discrete characteristics and occurrence context. This information is used to optimize threshold settings and radius parameters, forming a self-improving diagnostic capability. The entire adaptive clustering process is designed with robustness in mind; for example, a noise tolerance mechanism is introduced to avoid misjudgments caused by temporary interference. A visualization tool for cluster boundaries allows maintenance personnel to visually inspect the vector distribution, providing a reference for algorithm optimization. Finally, a list of anomalous clusters and their attribute data are output, seamlessly connecting to the next stage's pattern matching engine.
[0035] See Figure 4In the automatic fault diagnosis process of the reverse-guided IGBT intelligent power module, adaptive clustering is a key step in identifying anomalous feature clusters. This figure illustrates the distribution of cluster dispersion and the anomaly screening results: the horizontal axis represents cluster dispersion, and the vertical axis represents the number of clusters. The algorithm uses a dispersion threshold of 2.5 as a boundary. The light gray area on the left represents the "normal clustering region (dispersion < threshold)," reflecting the clustering state of feature vectors during normal module operation; the dark gray area on the right represents the "abnormal feature clustering region (dispersion > threshold)," representing clusters deviating from the normal distribution. Through this density-based dispersion analysis and dynamic threshold screening, the adaptive clustering algorithm can automatically distinguish between normal and anomalous feature regions, overcoming the shortcomings of traditional methods in adapting to fluctuations in operating conditions. This provides reliable anomalous region input for the subsequent pattern matching engine to accurately identify fault types, ensuring high accuracy in identifying abnormal features of the module under complex operating conditions.
[0036] Example 4: The comparison operation of the pattern matching engine is based on a multi-level fault feature tree. This tree structure categorizes known faults into three main branches: device aging, gate failure, and thermal runaway. Each branch is further subdivided into more specific fault modes. For example, the device aging branch may include subcategories such as gate oxide degradation and bond wire fatigue, while the gate failure branch covers specific issues such as abnormal drive resistance and Miller capacitance changes. The construction of the fault feature tree is not a static process; it originates from in-depth analysis of a large number of historical fault samples. Standardized reference templates are formed by mining the common patterns of feature parameters when different faults occur. Extracting the geometric feature parameters of abnormal feature clusters is a prerequisite for accurate comparison. The system first obtains the spatial contour information of the abnormal regions output by the adaptive clustering algorithm. Taking an abnormal cluster region marked as "Region Alpha" as an example, its geometric features are quantized to form a set of computable data. These parameters include the projected area of the region in the three-dimensional feature space, the shape factor describing its shape complexity (such as calculated by the ratio of the contour perimeter to the area), and the precise coordinates of the centroid of the region in space, as shown in Table 1.
[0037] Table 1: Geometric characteristic parameters of Alpha region in the abnormal clustering area
[0038] After acquiring the quantified geometric features of the abnormal region, the pattern matching engine initiates the similarity calculation process. This process compares the geometric feature parameters of the abnormal region with the standard fault patterns pre-stored in the fault feature tree. The similarity calculation is not a simple Euclidean distance metric, but rather uses a weighted cosine similarity algorithm. This algorithm assigns different weight coefficients to different geometric feature parameters. For example, the center position of the region may have a higher weight than the area of the region because it better reflects the essential characteristics of the fault. The calculation process first vectorizes the feature vector of the abnormal region and the feature vector of the standard fault pattern, then calculates the cosine value of the angle between the two vectors in the direction, and multiplies it by the corresponding weight coefficient matrix, finally obtaining a comprehensive similarity score between 0 and 1. The closer the score is to 1, the higher the matching degree.
[0039] The process of building the fault feature tree reflects a profound understanding of the fault evolution law. When collecting typical fault sample data, the system not only records a snapshot of the data at the moment the fault occurs, but also traces the data sequence during the fault development process. By analyzing the changing trend of parameters over time, common evolutionary features are extracted and solidified into basic fault modes. For example, for progressive faults such as gate oxide degradation, the features include not only the final parameter anomaly, but also the rate and trajectory of parameter drift. According to the severity of the fault development process, a clear hierarchical division is established within the feature tree, classifying the same type of fault into different stages such as initial warning, intermediate development, and late severity, and defining the evolution path between each stage. When constructing the feature tree, the system also sets a dynamic weight coefficient for each feature parameter. The allocation of the weight coefficient is based on the discriminative ability of the parameter in distinguishing different fault modes. The discriminative ability is evaluated through statistical methods (such as information gain or chi-square test) to ensure that the main features play a decisive role in the matching process.
[0040] The execution of the pattern matching engine is accompanied by a decision-making logic. When the similarity scores of an abnormal region and multiple standard patterns in the fault feature tree are high, the engine does not simply select the highest score. Instead, it introduces a confidence evaluation mechanism. This mechanism checks the difference between the second-highest score and the highest score. If the difference is less than a preset threshold, it is judged as a fuzzy matching result and triggers a more in-depth data analysis request. For example, it may require the system to trace the evolution history of the abnormal region over a period of time, or to perform cross-validation by combining other sensor data (such as the module's real-time operating current and ambient temperature). This design effectively reduces the risk of misjudgment. The entire pattern matching process is designed as an interactive loop. After the matching result is output, the system allows experienced operation and maintenance personnel to review and confirm the result. If the operation and maintenance personnel believe that there is a deviation in the matching result based on experience, they can feed the data of the abnormal region and the final fault type judgment as new samples back to the system. The system will use these high-quality samples that have been manually confirmed to incrementally learn the fault feature tree, optimize existing standard patterns or add new fault patterns. This mechanism enables the fault feature library to continuously evolve and become more and more accurate with the accumulation of actual operating experience.
[0041] Example 5: The mechanism of real-time updating of the fault feature library is the core of the system's long-term diagnostic accuracy. This mechanism tracks the long-term evolution trend of key parameters by setting a feature change monitoring window. The monitoring window usually adopts a sliding time window mechanism, and the window size is dynamically adjusted according to the data update frequency and the real-time requirements of the application scenario. For example, in a wind power converter application, the system may maintain a data window with a length of three months. The feature parameter data within the window is organized into a time series and trend fitting is performed. The key parameters tracked include the baseline value of the gate voltage change rate, the average value of the collector current overshoot amplitude, and the slow drift of the case temperature thermal resistance. By performing linear regression or more complex seasonal decomposition on these series, the system can capture the gradual changes of parameters with operating time or number of switching operations. This trend tracking provides a direct basis for judging whether the device has entered the aging stage. The feature drift detection mechanism is established to identify gradual changes in the parameter distribution. The core of this mechanism is to continuously compare the difference between the statistical distribution of the current feature parameters and the baseline distribution established during the initial learning phase of the system. It uses concepts from statistical process control, such as calculating the Wasserstein distance or the difference between the maximum mean and the historical baseline set, to quantify the degree of distribution drift. The system will periodically (e.g., weekly) calculate these statistics and compare them with their control limits. When the system detects that the statistics continuously exceed the control limits, it indicates that the basic distribution of the parameters has changed significantly. This change may be due to the performance degradation of the power module itself or to long-term changes in the operating environment (e.g., ambient temperature, load conditions). After triggering drift detection, the system will start the root cause analysis process to try to correlate the drift with specific work log events (e.g., continuous overload operation, cooling system efficiency decline records).
[0042] When a significant feature drift is confirmed, the system triggers a reconstruction process for the fault feature library. Reconstruction is not simply about replacing the old data with new data, but rather a cautious incremental learning process. First, the system creates a backup of the current fault feature tree. Then, it fine-tunes the existing classification model using recent feature data (such as the period before and after the drift is detected). During the fine-tuning process, new data is given higher weights, while valuable generalization knowledge from the original model is preserved through regularization techniques. For the fault feature tree itself, reconstruction may involve adjusting the boundary thresholds of the normal operating condition region, updating the feature weight coefficients used in fault mode matching, or adding new branches to the corresponding nodes of the feature tree after confirming a new fault evolution path. The entire reconstruction process is carried out in an isolated sandbox environment. Only after the performance of the new model is verified to be superior to that of the old model will an online switch be performed.
[0043] The fault prediction model is a natural extension of the evolution trend of characteristic parameters. This model focuses on calculating the degradation rate of key characteristic parameters. For example, by analyzing the slow rising trend of the gate threshold voltage, fitting its change curve and calculating the voltage increment per thousand hours of operation, it serves as the degradation rate of the parameter. When predicting the remaining service life, the model combines the current parameter value, its degradation rate, and the preset failure threshold, and uses a proportional risk model or a physics-based degradation model to estimate the expected time from the current state to functional failure. When generating preventive maintenance recommendations, the system not only considers the predicted remaining service life, but also comprehensively evaluates the module's criticality in the system, historical maintenance costs, and current load conditions, and outputs actionable recommendations, such as "It is recommended to check the gate drive circuit at the next planned shutdown" or "Preventive replacement of the module should be arranged within the next two months."
[0044] When the system detects a new cluster of abnormal features, the update process enters an interactive mode requiring manual intervention. The system will fully record all raw data of characteristic parameters and corresponding operating environment data for a period of time before and after the emergence of the new cluster, such as DC bus voltage, output current frequency, and radiator inlet temperature. This data is packaged into a case to be analyzed and submitted to the domain engineer through the expert system interface. The evaluation module built into the expert system will provide preliminary analysis, such as by comparing the early signs of known faults, giving an assessment of the possible hazard level of the new anomaly (such as "potential risk", "needs attention", "high risk"). The engineer reviews and confirms the system evaluation based on their own experience and feeds back the final conclusion to the system. The confirmed new fault mode and its characteristic parameters will be assigned a unique mode code and then formally added to the corresponding level of the fault feature tree. At the same time, the complete context information of the first discovery and confirmation of the mode will be recorded. The update of the feature library is not a passive accumulation of data, but an active, evidence-based knowledge accumulation process. The introduction of the fault prediction model elevates diagnosis from "post-event judgment" to "pre-event warning". The discovery and confirmation mechanism of the new pattern ensures that the system can adapt to unknown fault types. This design makes the intelligent diagnostic system no longer a static tool, but a learning system that can evolve with the equipment and continuously accumulate experience.
[0045] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0046] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An automatic fault diagnosis method for reverse-conducting IGBT intelligent power modules, characterized in that, include: Acquire multi-source monitoring data during the operation of the power module, including gate voltage waveform, collector current waveform, and case temperature change curve; A dynamic feature extraction model is established to perform joint time-domain and frequency-domain analysis on the multi-source monitoring data, generating a set of feature parameters. Construct a fault feature space and map the set of feature parameters to a high-dimensional space to form a feature vector distribution map; An adaptive clustering algorithm is used to divide the feature vector distribution map into regions and identify areas of anomalous feature clusters. The abnormal feature clusters are compared with a preset fault feature database using a pattern matching engine, and the fault type identification result is output.
2. The automatic fault diagnosis method for a reverse-conducting IGBT intelligent power module according to claim 1, characterized in that, The establishment of the dynamic feature extraction model includes: The rising and falling edge features of the gate voltage waveform are extracted, and the voltage change rate during the switching process is calculated. Analyze the conduction and turn-off characteristics of the collector current waveform to extract the current overshoot amplitude and oscillation frequency; Monitor the slope change of the shell temperature change curve and calculate the dynamic response parameters of thermal resistance; The voltage change rate, current overshoot amplitude, oscillation frequency, and thermal resistance dynamic response parameters are combined to form a set of characteristic parameters.
3. The automatic fault diagnosis method for a reverse-conducting IGBT intelligent power module according to claim 2, characterized in that, The construction of the fault feature space includes: The set of feature parameters is standardized to eliminate dimensional differences; A nonlinear dimensionality reduction method is used to project the standardized feature parameters onto a three-dimensional space. Based on the distribution patterns of historical fault data, the boundaries of the normal operating condition area are marked in three-dimensional space; Calculate the distance between the current feature vector and the boundary of the normal operating condition area, and generate a feature vector distribution map.
4. The automatic fault diagnosis method for a reverse-conducting IGBT intelligent power module according to claim 3, characterized in that, The adaptive clustering algorithm employed includes: The number of cluster centers is automatically determined based on the feature vector density distribution. Set the dynamic neighborhood radius parameter to adjust the clustering region boundaries; Calculate the feature vector dispersion of each cluster region and filter out abnormal feature clusters with dispersion exceeding the threshold.
5. The automatic fault diagnosis method for a reverse-conducting IGBT intelligent power module according to claim 4, characterized in that, The comparison performed using a pattern matching engine includes: A multi-level fault feature tree is established, which includes three types of fault modes: device aging, gate failure, and thermal runaway. Extract the geometric feature parameters of the abnormal feature cluster region, including region area, shape factor, and center position; The similarity between the geometric feature parameters and the standard patterns in the fault feature tree is calculated. The fault type corresponding to the standard pattern with the highest similarity was selected as the identification result.
6. The automatic fault diagnosis method for a reverse-conducting IGBT intelligent power module according to claim 5, characterized in that, The establishment of the multi-level fault feature tree includes: Collect typical fault sample data and extract common features to form basic fault patterns; Based on the failure development process, severity levels are classified, and failure mode evolution paths are established; Set feature weight coefficients to distinguish the influence of primary and secondary features.
7. The automatic fault diagnosis method for a reverse-conducting IGBT intelligent power module according to claim 6, characterized in that, The method further includes: The fault feature database is updated in real time, and when a new cluster of abnormal features is detected: Record the characteristic parameters and operating environment data of new anomaly clusters; The hazard level of newly emerging anomalous feature clusters is assessed using an expert system. The newly identified fault modes are added to the fault feature tree.
8. The automatic fault diagnosis method for a reverse-conducting IGBT intelligent power module according to claim 7, characterized in that, The real-time updating of the fault feature database includes: Set up a feature change monitoring window to track the long-term evolution trend of feature parameters; Establish a feature drift detection mechanism to identify gradual changes in parameter distribution patterns; When a significant feature drift is detected, the fault feature library reconstruction process is triggered.
9. The automatic fault diagnosis method for a reverse-conducting IGBT intelligent power module according to claim 8, characterized in that, The method further includes: constructing a fault prediction model based on the evolution trend of feature parameters. Calculate the degradation rate of key characteristic parameters; predict remaining service life; generate preventative maintenance recommendations.
10. An automatic fault diagnosis system for a reverse-conducting IGBT intelligent power module, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the automatic fault diagnosis method for a reverse-conducting IGBT intelligent power module as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Solder layer degradation identification method of high-voltage multi-chip power module
CN118191539A
Distribution box fault detection method
CN118818202A
Power cable fault detection method
CN119535106A
Power supply fault rapid detection system and method based on intelligent diagnosis engine
CN120085217A
IGBT driver fault prediction method and safety system
CN120492856A
Cited By
Safety monitoring intelligent management system based on big data Internet of Things
CN120856463A
Power equipment operation and maintenance method and system fusing inspection data
CN121301894A
Abnormity monitoring method and device for IGBT power module
CN121633761A