Fan fault diagnosis method based on multi-working condition data synthesis and enhanced bayesian anomaly probability

By synthesizing multi-operating condition data and enhancing the Bayesian anomaly probability algorithm, combined with a regularized Gaussian mixture model and a random forest classifier based on expert rules, the diagnosis problem of wind turbines under multi-fault coupling and operating condition fluctuations is solved, achieving efficient and accurate fault detection and location, and improving the safety and reliability of the equipment.

CN120030454BActive Publication Date: 2025-10-10LIAONING DONGKE ELECTRIC POWER
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510475036.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-10-10
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

Existing wind turbine fault diagnosis methods are difficult to maintain efficient, accurate and flexible diagnostic performance when faced with multiple fault coupling, covariance matrix singularity and operating condition fluctuations, resulting in insufficient equipment safety and reliability.

Method used

A three-level diagnostic system is formed by synthesizing multi-operating condition data and enhancing the Bayesian anomaly probability (BAP) algorithm, combining regularized Gaussian mixture model (RGMM) and sliding window filtering technology with a random forest classifier based on expert rules to achieve fault detection, classification, and location.

Benefits of technology

It significantly improves the accuracy and stability of fault detection, enhances the safety and reliability of wind turbine units, reduces the false alarm rate, and improves the adaptability of equipment and fault repair efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030454B_ABST
    Figure CN120030454B_ABST
Patent Text Reader

Abstract

A fan fault diagnosis method based on multi-working condition data synthesis and enhanced Bayesian anomaly probability solves the detection problem of the fan unit under multiple fault conditions.The simulation model containing multiple fault working conditions such as sensor zero drift,process parameter linear drift,asymmetric noise interference and actuator jamming is constructed,high-accuracy fault data is generated through dynamic feature selection and noise asymmetric injection strategy,the Bayesian anomaly probability BAP algorithm is improved,regularization of covariance matrix and sliding window filtering technology are introduced,thus the singularity problem of covariance matrix is solved and the stability and robustness of fault detection are improved.Through the three-level diagnosis system of "detection-classification-location",combined with isolated forest data cleaning,dynamic Gaussian mixture model GMM and random forest classifier,the end-to-end fault analysis under different working conditions is realized and the fault detection ability and diagnosis precision of the fan unit of the coal-fired unit under the condition of multiple fault coupling are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to an intelligent fault diagnosis technology for a fan unit of a coal-fired unit, which combines multi-working condition data synthesis, a Bayesian anomaly probability (BAP) algorithm and a hierarchical diagnosis method to improve the accuracy and adaptability of fault detection. BACKGROUND

[0002] Coal-fired units play a core role in the power supply system in China, and their safe and stable operation is crucial to the reliability of power supply and the stability of the power grid. As a key auxiliary equipment of coal-fired units, fan units are responsible for providing air required for the combustion process and discharging exhaust gas after combustion, and their operating conditions directly affect the overall performance and safety of the unit.

[0003] In the actual operation of fan units, due to long-term operation in complex and variable working conditions, they face harsh conditions such as high temperature, high pressure and strong vibration, and the probability of fan unit failure is high. Once a fan unit fails, it may lead to a decrease in unit efficiency, an increase in energy consumption, and cause huge economic losses to power companies, and seriously affect the stable operation of the power grid.

[0004] With the continuous improvement of the intelligent level of the power industry in China, traditional fan unit fault diagnosis methods have been unable to effectively cope with the complex working conditions of multiple fault coupling in modern industry, resulting in low diagnosis accuracy and being easily disturbed by frequent fluctuations in working conditions, which cannot timely discover and warn potential faults. Therefore, there is an urgent need for an efficient, accurate and highly adaptable fan unit fault diagnosis method to cope with multiple fault coupling and complex working conditions to ensure the safe and stable operation of the fan unit of the coal-fired unit.

[0005] In existing technologies, wind turbine fault diagnosis typically relies on simulation models based on a single fault mode to construct diagnostic systems. These traditional methods are designed to simulate a specific fault type and construct a corresponding diagnostic model. However, in real-world industrial applications, wind turbines often face a variety of complex operating conditions and the interplay and interaction of multiple faults. In such a volatile environment, simulations of a single fault mode cannot fully and accurately reflect the complexity and diversity of wind turbine faults. Therefore, diagnostic models relying on a single fault mode exhibit poor adaptability and diagnostic capabilities in practical applications. In particular, they often fail to effectively detect and diagnose the interactions of multiple faults when faced with coupled multiple faults. Furthermore, existing fault diagnosis methods are mostly based on data-driven models, with Bayesian posterior probability inference being widely used. While this method can infer system states based on observed data to a certain extent, it has certain limitations when dealing with variable real-world operating conditions. In practical applications, Bayesian posterior probability inference often relies on the covariance matrix for calculation. However, when the covariance matrix becomes singular, this often leads to instability in the inference results, which in turn affects the accuracy and reliability of fault diagnosis. The existence of the covariance matrix singularity problem makes traditional Bayesian inference unable to provide accurate and reliable assessment of the actual operating status of the equipment, thereby affecting the early warning and fault location of wind turbine unit failures, and further reducing the efficiency and accuracy of fault diagnosis.

[0006] Traditional methods typically use fixed-structure models, which lack sufficient flexibility to adapt to varying wind turbine operating conditions. During actual operation, wind turbine operating conditions frequently fluctuate with factors such as turbine load, ambient temperature, and humidity. Fixed-structure models are unable to adjust their internal parameters and model structure in a timely manner to these changes. This makes it difficult for the model to maintain good adaptability in the face of complex and changing operating conditions, resulting in a decline in diagnostic performance and an inability to meet the increasingly stringent requirements for equipment reliability and safety in modern power production.

[0007] Existing fault diagnosis methods for wind turbines have significant shortcomings and limitations when dealing with complex scenarios such as multiple fault coupling and changing operating conditions. Existing diagnostic methods struggle to effectively address issues such as multiple fault coupling, singular covariance matrices, and poor adaptability. Therefore, a more efficient, accurate, and adaptable fault diagnosis method is urgently needed that can effectively address the multiple fault coupling characteristics of wind turbines under complex operating conditions, improve the accuracy, stability, and adaptability of fault detection, and thus ensure the safe and stable operation of wind turbines in coal-fired units. Summary of the Invention

[0008] The purpose of this invention is to provide an intelligent fault detection and diagnosis method that addresses the challenge of detecting multiple coupled faults in coal-fired turbines. By integrating multi-condition data synthesis, enhancing the Bayesian anomaly probability (BAP) algorithm, and layered diagnosis technology, this method aims to overcome the shortcomings of traditional fault diagnosis methods and improve the accuracy, stability, and adaptability of fault diagnosis. Through the above-mentioned technical solution, the present invention can solve the following technical problems existing in the prior art:

[0009] 1) Multiple-fault coupling detection: Multiple-fault coupling is common in wind turbine operation, and traditional methods struggle to cope with its complexity. This invention combines multi-condition data synthesis with an enhanced Bayesian Anomaly Probability (BAP) algorithm to provide more accurate fault detection and diagnosis.

[0010] 2) Singular covariance matrices affect the stability and accuracy of fault diagnosis. This invention uses regularization and sliding window filtering techniques to improve the stability of diagnostic results, detect anomalies in a timely manner, and prevent unplanned downtime.

[0011] 3) Adaptability to fluctuating operating conditions: Traditional methods are unable to flexibly cope with fluctuations in wind turbine operating conditions, resulting in reduced diagnostic performance. This invention combines a multi-condition fault simulation model with a dynamic feature selection strategy to effectively improve the adaptability of the diagnostic model.

[0012] 4) Inadequate fault detection efficiency and accuracy: Traditional methods often fail to detect potential faults in a timely manner, impacting unit operating efficiency. This invention significantly improves fault detection efficiency and accuracy through a three-level diagnostic system ("detection-classification-location") and advanced algorithms.

[0013] 5) Equipment safety and reliability: Traditional methods are prone to misjudgment under changing operating conditions and fail to promptly identify potential risks. This invention uses multi-dimensional data fusion and dynamic analysis to accurately identify potential faults and improve the safety and reliability of wind turbines.

[0014] In order to achieve the above objectives, the technical solution adopted by the present invention is: a fan fault diagnosis method based on multi-condition data synthesis and enhanced Bayesian abnormality probability, including two stages: "offline training" and "online monitoring".

[0015] A offline training phase:

[0016] A1) Generate synthetic fault data by using historical data of normal operating conditions and simulating four typical faults, namely offset, drift, noise, and stuck, to expand the training set.

[0017] The method for synthesizing multi-condition fault data is:

[0018] A1.1) Zero point offset

[0019] Sensor zero offset is a common fault that manifests as a fixed offset of the measured value near the baseline. Specifically, 1 to 3 features are randomly selected and a fixed offset is added. Its mathematical expression is:

[0020] ;in, is the measured value after offset, is the original measurement value, △ represents the offset, which is usually a random uniformly distributed value in the range of [2, 5], simulating the zero drift of the sensor.

[0021] A1.2) Linear Drift

[0022] One common fault is when a process parameter changes linearly over time, such as a temperature that continuously rises and falls. In this model, a single feature increases linearly with a randomly generated slope. The mathematical expression is:

[0023] ; is the measured value after the offset at time t, is the original measurement value at time t, k is the drift slope, a uniformly distributed random value in the range of [0.1, 0.5], and t is the time variable, which is simplified to 1 in single sample generation. This linear drift model simulates faults that change over time.

[0024] A1.3) Asymmetric noise

[0025] It is a common phenomenon in actual operation that a signal is interfered with by asymmetric noise, such as sensor vibration or electromagnetic interference. In the present invention, a single feature increases linearly under noise interference, and the slope is randomly generated. Its mathematical expression is:

[0026] ∈ is the noise intensity, typically a uniformly distributed random value in the range [1, 3]. The probability of the noise direction is 70% positive and 30% negative. The proposed model injects noise of varying intensities through probabilistic injection, realistically simulating the noise interference experienced in actual operations and enhancing data diversity and complexity.

[0027] A1.4) Actuator stuck

[0028] Actuator stuck failure means that the actuator loses its response and stays at a fixed value, such as the valve opening cannot change. In this model, a single feature is fixed to an extreme value, and the mathematical expression is:

[0029] ; and are the minimum and maximum values ​​of this feature in normal data, respectively. This setting can accurately simulate the fault condition where the actuator is stuck in the limit position.

[0030] This paper designs a feature selection mechanism to evenly distribute the sensitive feature ratios. By dynamically adjusting the ratios, the impact of faults on features is controlled, data diversity and accuracy are enhanced, and feature differences under different fault scenarios are simulated.

[0031] A2) Use normal data to train the regularized Gaussian mixture model (RGMM) to establish the probability distribution of normal operating conditions, and combine the expert knowledge base with synthetic data to train a random forest classifier.

[0032] The regularized Gaussian mixture model RGMM is established as follows:

[0033] Normal operating condition data set , where n is the number of samples and d is the feature dimension. When processing small samples or noisy data, the covariance matrix of the traditional Gaussian mixture model may become ill-conditioned, reducing robustness. To solve this problem, this paper adopts the regularized Gaussian mixture model (RGMM), whose probability density function is:

[0034] ; where K is the number of Gaussian components, π k is the mixing coefficient of the kth Gaussian component, which needs to satisfy π k ≥0, is the probability density function of the kth Gaussian distribution, whose mean is μ k , the covariance matrix is ​​(∑ k +λΙ), where λ is the regularization parameter used to regularize the covariance matrix, and Ι is the d-order identity matrix. By adding the regularization term λΙ to the covariance matrix, we can avoid the situation where the covariance matrix may become singular or nearly singular when the number of samples is small. After adding the regularization term, the covariance matrix can maintain good properties, thereby ensuring the stability and accuracy of the model. B. Online monitoring stage:

[0035] B1) After preprocessing, the real-time collected wind turbine data is used to calculate the Bayesian anomaly probability through the Regularized Gaussian Mixture Model (RGMM), and compared with the set threshold to trigger an alarm.

[0036] The enhanced Bayesian anomaly probability calculation method is:

[0037] Traditional Bayesian posterior probability calculation is susceptible to singular covariance matrices, leading to instability and affecting diagnostic accuracy. To enhance stability, this paper performs a double regularization on the covariance matrix.

[0038] First, the covariance matrix is ​​avoided from being singularized by a dynamic threshold: min =max(10 -4 λ max ,10 -5 )

[0039] Among them, λ min and λ max are the minimum and maximum eigenvalues ​​of the covariance matrix, respectively, through the dynamic threshold λ min =max(10 -4 λ max ,10 -5 ) Avoid singularization of the covariance matrix by replacing eigenvalues ​​smaller than the threshold with λ min , get the regularized eigenvalue matrix to enhance numerical stability. Then, reconstruct the covariance matrix:

[0040] ; V is an orthogonal matrix consisting of the eigenvectors of the covariance matrix, is a diagonal matrix composed of regularized eigenvalues, c is a regularization constant, and I is the identity matrix. After reconstructing the covariance matrix, the identity matrix regularization term cI is superimposed. This dual regularization strategy can alleviate overfitting in high-dimensional data and improve model robustness. Unlike traditional single regularization methods such as L2 regularization, this method uses eigenvalue decomposition and truncation techniques to preserve the principal component directional information while suppressing noise.

[0041] Based on the covariance matrix after double regularization, the Bayesian anomaly probability BAP under the double enhanced Bayesian posterior probability is:

[0042] ;in, Represents x t to the mean μ i The Mahalanobis distance, g i and h i The scaling factor and degree of freedom correction terms are as follows:

[0043] ; Dynamically calculate the degree of freedom parameter g through trace operation i and h i , solve the traditional chi-square distribution Χ 2 (·) Model bias caused by assuming fixed degrees of freedom. In addition, introducing the interaction between the covariance matrix and the inverse covariance matrix into the degree of freedom calculation also expands the application boundaries of the chi-square distribution in anomaly detection.

[0044] Based on the RGMM posterior probability π i The abnormal probabilities of different Gaussian components are weighted to achieve probabilistic fusion of multi-operating condition data, which is superior to the traditional single-model detection method. At the same time, the contribution of each component to the final score is clarified through probabilistic decomposition, which facilitates subsequent model diagnosis and tuning.

[0045] In addition, in order to further smooth the fluctuation of BAP and reduce the occurrence of false positives, this paper adopts the sliding window filtering technology. The sliding window size ω=8, and the specific filtering formula is as follows:

[0046] ;in, is the BAP index after sliding window filtering, BAP t The original BAP index is filtered through a sliding window, combining the current BAP index with the BAP index at the past ω moments, and taking the maximum value as the filtered result. By combining this with time series analysis, the stability and accuracy of fault detection can be improved, helping to suppress the impact of transient abnormal fluctuations on the score, making it suitable for online detection of non-steady-state data streams.

[0047] B2) If an anomaly is detected, the random forest classifier further identifies the fault type and locates the key fault parameters through contribution variable analysis, forming a closed-loop diagnosis.

[0048] B2.1) The random forest classification method based on expert rules is:

[0049] In the fault classification stage, in order to fully utilize the knowledge and experience of domain experts and improve the accuracy and robustness of classification, this paper integrates 137 empirical rules from 7 domain experts through the Delphi method to form four types of rule templates.

[0050] The amplitude constraint rule is mainly used to detect sudden over-limit situations of pressure sensors. Its mathematical expression is:

[0051] ;x i is the current measured value, μ i and σ i are the mean and standard deviation of normal data, respectively, and k is the threshold coefficient. When the measured value exceeds a certain range, it is judged as a sudden over-limit failure of the pressure sensor.

[0052] B2.1.2) Association constraints

[0053] The association constraint rule is used for the pressure association of the inlet and outlet, and the mathematical expression is:

[0054] ;in, and are the measured values ​​of inlet and outlet pressures, △ th The inlet and outlet pressure difference threshold under normal circumstances. When the inlet and outlet pressure difference exceeds the threshold, it is judged as an inlet and outlet pressure difference abnormal fault.

[0055] B2.1.3) Temporal continuity

[0056] The temporal persistence rule is used to detect when vibration exceeds the limit continuously. The mathematical expression is:

[0057] ; is the vibration measurement value at time i, is the vibration threshold, T is the time window size, n is the threshold for the number of vibrations exceeding the limit within the time window, and I(·) is the indicator function. When the number of vibrations exceeding the limit reaches a certain threshold within a certain time window, it is determined to be a continuous vibration exceeding limit fault.

[0058] B2.1.4) Combination constraints

[0059] Combined constraint rules are used to jointly judge surge and flow. The mathematical expression is relatively complex and requires comprehensive consideration of multiple parameters. For example:

[0060] ; are multiple measurement parameters related to surge and flow, and f(·) is a comprehensive judgment function determined based on expert experience. When the function value is greater than 0, it is judged to be a combined surge and flow fault.

[0061] According to the above expert rules, a random forest classifier based on expert rules is designed. For the random forest classifier, there are classification decisions:

[0062] ; represents the final predicted category of the input sample x, c represents all possible categories, Tk(x) represents the predicted category of the kth decision tree for sample x, I(·) is the indicator function, which is 1 when Tk(x)=c and 0 otherwise; It is used to count the number of times that the predicted category is c in all trees. The prediction results of multiple decision trees are integrated through the "majority voting" mechanism to improve the accuracy and robustness of classification. The final random forest classification result based on expert rules is:

[0063] ; and are the prediction results of expert rules and random forest classifier respectively. It represents a hyperparameter ranging from 0 to 1. Through rule-driven feature space segmentation, it relies on random forest output under steady-state conditions and combines expert rules under transient conditions to reduce the false alarm rate under variable load conditions.

[0064] B2.2) The fault variable location method is:

[0065] Accurately locating the fault variable is crucial for fault repair and equipment maintenance. The formula for calculating the contribution of the jth variable to the fault is as follows:

[0066] ; wherein, CI j represents the contribution index of the jth feature, π k is the weight of the kth Gaussian distribution, V k is the feature vector of the kth Gaussian distribution, x is the feature vector at the current moment, μ k is the mean of the kth Gaussian distribution, k is the number of Gaussian distributions; by calculating the feature contribution index, the influence of each feature on the fault is evaluated, and then the fault location is located.

[0067] The beneficial effects of the present application are:

[0068] The present application designs a three-level diagnosis method of "detection-classification-location", which starts from fault detection, judges the occurrence of fault by enhancing the Bayesian abnormal probability; then uses a random forest classifier combined with expert rules for fault type classification; finally determines the fault location through variable positioning, forming a complete fault diagnosis closed loop. BRIEF DESCRIPTION OF DRAWINGS

[0069] Figure 1 It is an online automatic checking technical roadmap based on Ethernet remote transmission data;

[0070] Figure 2 It is a confusion matrix result graph of different methods for different data type classification;

[0071] Figure 3 It is a fault location result (top 7 variables) graph of the RGMM-BAP method;

[0072] Figure 4 It is the k value selection result in the RGMM model. DETAILED DESCRIPTION

[0073] The specific embodiments of the present application will be further described in detail below in combination with the drawings and examples. The following examples are used to illustrate the present application, but are not used to limit the scope of the present application.

[0074] As Figure 1The technical roadmap for wind turbine fault diagnosis based on multi-condition data synthesis and enhanced Bayesian anomaly probability is shown. This technology can be divided into two phases: offline training and online monitoring. During the offline training phase, synthetic fault data is generated by simulating historical normal operating data and four typical faults (offset, drift, noise, and stuck) to expand the training set. Next, a regularized Gaussian mixture model (RGMM) is trained using the normal data to establish a probability distribution for normal operating conditions. A random forest classifier is trained using the expert knowledge base and the synthetic data. During the online monitoring phase, real-time wind turbine data is preprocessed and then calculated using the regularized Gaussian mixture model (RGMM) to calculate the Bayesian anomaly probability. This probability is then compared with a set threshold to trigger an alarm. If an anomaly is detected, the random forest classifier further identifies the fault type and locates the key fault parameters through contribution variable analysis, forming a closed-loop diagnosis.

[0075] Functions of each part of the device of the present invention:

[0076] 1) SCADA system: used to collect the operating data of the wind turbine in real time and transmit the data.

[0077] 2) Data storage system: used to store real-time data and historical data collected from the SCADA system.

[0078] 3) High-performance computing platform: used to run fault diagnosis algorithms, such as the Bayesian Anomaly Probability (BAP) algorithm.

[0079] 4) Fault diagnosis system: Based on the data processing platform, it can detect, classify and locate the faults of the fan group.

[0080] 5) Remote monitoring equipment: allows operators to monitor and diagnose the status of the fan unit through remote equipment.

[0081] The method of the present invention comprises the following steps: (1) synthesis of multi-operating condition fault data; (2) establishment of a regularized Gaussian mixture model (RGMM); (3) enhanced Bayesian anomaly probability calculation; (4) random forest classification based on expert rules; and (5) fault variable location.

[0082] Step 1: Synthesis of multi-condition fault data

[0083] 1) Zero point offset

[0084] Sensor zero offset is a common fault that manifests as a fixed offset of the measured value near the baseline. Specifically, 1 to 3 features are randomly selected and a fixed offset is added. Its mathematical expression is:

[0085] ; is the measured value after offset, is the original measurement value, Δ represents the offset, usually a random uniformly distributed value in the range of [2, 5], simulating the sensor zero drift.

[0086] 2) Linear drift

[0087] Linear change of process parameters over time is one of the common faults, such as continuous temperature rise and fall. In this model, a single feature is linearly increasing, and the slope is randomly generated. The mathematical expression is:

[0088] ; is the measurement value after offset at time t, is the original measurement value at time t, k is the drift slope, which is a uniformly distributed random value in the range of [0.1, 0.5], t is the time variable, which is simplified to 1 in single sample generation. This linear drift model simulates the fault that changes over time.

[0089] 3) Asymmetric noise

[0090] Signal interference by asymmetric noise is a common phenomenon in actual operation, such as sensor vibration or electromagnetic interference. In this invention, a single feature is linearly increasing under noise interference, and the slope is randomly generated. The mathematical expression is:

[0091] ; ∈ is the noise intensity, usually a uniformly distributed random value, ranging between [1, 3]. The noise direction probability is 70% positive and 30% negative. The model of this invention injects different intensity noise by probability, simulates the noise interference in actual operation, and enhances the diversity and complexity of data.

[0092] 4) Actuator stuck

[0093] Actuator stuck fault means that the actuator loses response and stays at a fixed value, such as valve opening cannot change. In this model, a single feature is fixed as an extreme value, and the mathematical expression is:

[0094] ; and are the minimum and maximum values of the feature in normal data, respectively. This setting can accurately simulate the fault condition of actuator stuck at the limit position.

[0095] This invention designs a feature selection mechanism, with a uniform distribution of sensitive feature proportion interval. By dynamically adjusting the proportion, the influence of faults on features is controlled, the data diversity and accuracy are enhanced, and the feature differences under different fault scenarios are simulated.

[0096] Step 2: Regularized Gaussian Mixture Model (RGMM) establishment

[0097] Normal operating condition data set , where n is the number of samples and d is the feature dimension. When processing small samples or noisy data, the covariance matrix of the traditional Gaussian mixture model may become ill-conditioned, reducing robustness. To address this problem, this paper adopts the regularized Gaussian mixture model (RGMM), whose probability density function is:

[0098] ; where K is the number of Gaussian components, π k is the mixing coefficient of the kth Gaussian component, which needs to satisfy π k ≥0, is the probability density function of the kth Gaussian distribution, whose mean is μ k , the covariance matrix is ​​(∑ k +λΙ), where λ is the regularization parameter used to regularize the covariance matrix, and Ι is the d-order identity matrix. By adding the regularization term λΙ to the covariance matrix, we can avoid the situation where the covariance matrix may become singular or nearly singular when the number of samples is small. After adding the regularization term, the covariance matrix can maintain good properties, thereby ensuring the stability and accuracy of the model. B. Online monitoring stage:

[0099] Step 3: Enhanced Bayesian anomaly probability calculation

[0100] Traditional Bayesian posterior probability calculation is susceptible to singular covariance matrices, leading to instability and affecting diagnostic accuracy. To enhance stability, this paper performs a double regularization on the covariance matrix.

[0101] First, the covariance matrix is ​​avoided from being singularized by a dynamic threshold: min =max(10 -4 λ max ,10 -5 )

[0102] Among them, λ min and λ max are the minimum and maximum eigenvalues ​​of the covariance matrix, respectively, through the dynamic threshold λ min =max(10 -4 λ max ,10 -5 ) Avoid singularization of the covariance matrix by replacing eigenvalues ​​smaller than the threshold with λ min , get the regularized eigenvalue matrix to enhance numerical stability. Then, reconstruct the covariance matrix:

[0103] ; V is an orthogonal matrix consisting of the eigenvectors of the covariance matrix, is a diagonal matrix consisting of regularized eigenvalues, c is a regularization constant, and I is the identity matrix. After reconstructing the covariance matrix, the identity matrix regularization term cI is superimposed. This dual regularization strategy can alleviate overfitting in high-dimensional data and improve model robustness. Unlike traditional single regularization methods (such as L2 regularization), this method uses eigenvalue decomposition and truncation techniques to preserve the principal component directional information while suppressing noise.

[0104] Based on the covariance matrix after double regularization, the Bayesian anomaly probability (BAP) under the double enhanced Bayesian posterior probability is:

[0105] ;in, Represents x t to the mean μ i The Mahalanobis distance, g i and h i The scaling factor and degree of freedom correction terms are as follows:

[0106] ; Dynamically calculate the degree of freedom parameter g through trace operation i and h i , solve the traditional chi-square distribution Χ 2 (·) Model bias caused by assuming fixed degrees of freedom. In addition, introducing the interaction between the covariance matrix and the inverse covariance matrix into the degree of freedom calculation also expands the application boundaries of the chi-square distribution in anomaly detection.

[0107] Based on the RGMM posterior probability π i The abnormal probabilities of different Gaussian components are weighted to achieve probabilistic fusion of multi-operating condition data, which is superior to the traditional single-model detection method. At the same time, the contribution of each component to the final score is clarified through probabilistic decomposition, which facilitates subsequent model diagnosis and tuning.

[0108] In addition, in order to further smooth the fluctuation of BAP and reduce the occurrence of false positives, this paper adopts the sliding window filtering technology. The sliding window size ω=8, and the specific filtering formula is as follows:

[0109] ;in, is the BAP index after sliding window filtering, BAP t The original BAP index is filtered through a sliding window, combining the current BAP index with the BAP index at the past ω moments, and taking the maximum value as the filtered result. By combining this with time series analysis, the stability and accuracy of fault detection can be improved, helping to suppress the impact of transient abnormal fluctuations on the score, making it suitable for online detection of non-steady-state data streams.

[0110] Step 4: Random Forest Classification Based on Expert Rules

[0111] In the fault classification stage, in order to make full use of the knowledge and experience of domain experts and improve the accuracy and robustness of classification, this paper integrates 137 experience rules of 7 domain experts through the Delphi method to form four types of rule templates.

[0112] 1) Amplitude constraint

[0113] The amplitude constraint rule is mainly used to detect sudden over-limit situations of pressure sensors. Its mathematical expression is:

[0114] ;in, is the current measured value, μ i and σ i are the mean and standard deviation of normal data respectively, and k is the threshold coefficient. When the measured value exceeds a certain range, it is judged as a sudden over-limit failure of the pressure sensor.

[0115] 2) Association constraints

[0116] The association constraint rule is used for the pressure association of the inlet and outlet, and the mathematical expression is:

[0117] ;in, and are the measured values ​​of inlet and outlet pressures, △ th The inlet and outlet pressure difference threshold under normal circumstances. When the inlet and outlet pressure difference exceeds the threshold, it is judged as an inlet and outlet pressure difference abnormal fault.

[0118] 3) Temporal persistence

[0119] The temporal persistence rule is used to detect when vibration exceeds the limit continuously. The mathematical expression is:

[0120] ; is the vibration measurement value at time i, is the vibration threshold, T is the time window size, n is the threshold for the number of vibrations exceeding the limit within the time window, and I(·) is the indicator function. When the number of vibrations exceeding the limit reaches a certain threshold within a certain time window, it is determined to be a continuous vibration exceeding limit fault.

[0121] 4) Combination constraints

[0122] Combined constraint rules are used to jointly judge surge and flow. The mathematical expression is relatively complex and requires comprehensive consideration of multiple parameters. For example:

[0123] ; are multiple measurement parameters related to surge and flow, and f(·) is a comprehensive judgment function determined based on expert experience. When the function value is greater than 0, it is judged to be a combined surge and flow fault.

[0124] According to the above expert rules, a random forest classifier based on expert rules is designed. For the random forest classifier, there are classification decisions:

[0125] ; represents the final predicted category of the input sample x, c represents all possible categories, Tk(x) represents the predicted category of the kth decision tree for sample x, I(·) is the indicator function, which is 1 when Tk(x)=c and 0 otherwise; It is used to count the number of times that the predicted category is c in all trees. The prediction results of multiple decision trees are integrated through the "majority voting" mechanism to improve the accuracy and robustness of classification. The final random forest classification result based on expert rules is:

[0126] ; and are the prediction results of expert rules and random forest classifier respectively. It represents a hyperparameter ranging from 0 to 1. Through rule-driven feature space segmentation, it relies on random forest output under steady-state conditions and combines expert rules under transient conditions to reduce the false alarm rate under variable load conditions.

[0127] Step 5: Fault variable location

[0128] Accurately locating the fault variable is crucial for fault repair and equipment maintenance. The formula for calculating the contribution of each variable to the fault is as follows:

[0129] Among them, CI j Represents the contribution index of the jth feature, π k is the weight of the kth Gaussian distribution, V k is the eigenvector of the kth Gaussian distribution, x is the eigenvector at the current moment, μ k is the mean of the kth Gaussian distribution, and k is the number of Gaussian distributions. By calculating the feature contribution index, the impact of each feature on the fault is evaluated, and the fault location is then located.

[0130] Based on this, this paper designed a three-stage diagnostic system: "Detection-Classification-Location." The overall process is shown in Figure 1. This system begins with fault detection, using enhanced Bayesian anomaly probability to determine fault occurrence. Next, a random forest classifier combined with expert rules is used to classify the fault type. Finally, variable localization is used to determine the fault location, forming a complete fault diagnosis closed loop. The present invention is further described in detail below using specific examples.

[0131] 1) Experimental scenario and data configuration

[0132] The data used in this experiment comes from SCADA data collected from a 300MW wind turbine unit at a power plant during the 2020-2021 period. The data, sampled at a 1Hz frequency, covers key operating parameters (such as temperature, pressure, and speed) under both normal and various fault conditions, providing an important basis for fault diagnosis. To ensure the reliability of the experiment, the data was meticulously cleaned and preprocessed to remove outliers and fill in missing values ​​to ensure data integrity.

[0133] In the experiment, PCA, SVM, and CNN-LSTM were used as comparison methods. PCA is primarily used for data dimensionality reduction and feature extraction. SVM is a classic classification algorithm with good generalization capabilities. CNN-LSTM combines convolutional neural networks and long-short-term memory networks, making it suitable for processing wind turbine data with spatiotemporal characteristics.

[0134] 2) Comparative test and analysis of experimental results

[0135] The specific experimental data statistics are shown in Table 1, which shows the F1 scores of different methods for fault detection. The table can intuitively compare the performance differences of different methods in fault detection.

[0136]

[0137] Table 2 shows the comprehensive classification performance of different methods for different data types, expressed as precision (%). The data with the highest performance in each category is in bold:

[0138] Figure 2 Figure 2 shows the confusion matrices of different methods for classifying different data types. (a) to (d) are PCA, SVM, CNN-LSTM, and the proposed method RGMM-BAP, respectively. Analysis of these confusion matrices clearly demonstrates that the proposed method has significant advantages in fault classification and can more accurately classify different types of faults.

[0139] Figure 3The horizontal axis represents the model's fault location results. The random forest classifier's contribution to the fault classification model is shown on the horizontal axis. Fault location helps technicians quickly identify the key factors that cause faults in actual model applications, improving fault repair efficiency.

[0140] Figure 4 shows the k-value selection results for the RGMM model in this paper. Based on the data scatter plot, the optimal k-value is 5, circled in red. Choosing a reasonable k-value is crucial to the performance of the RGMM model. It ensures that the model maintains good performance under different operating conditions and improves the accuracy and reliability of fault diagnosis.

[0141] Based on the experimental results (Table 1-2, Figures 2-4 The RGMM-BAP method proposed in this paper demonstrates significant advantages in fault detection, classification accuracy, and localization effectiveness. Compared to traditional methods, its F1 score reaches 90.4%, significantly outperforming PCA (50.6%), SVM (76.3%), and CNN-LSTM (74.9%). Through a double-regularized covariance matrix and sliding window filtering, detection robustness is improved, reducing the false alarm rate to 0.3%. In a multi-fault classification task, RGMM-BAP achieves 99.7% accuracy in identifying normal data and effectively improves the classification accuracy of linear drift (53.8%) and actuator stuck (60.3%) faults, outperforming traditional methods (both 0%). However, drift and stuck faults still suffer from a 21% cross-classification error, which can be further optimized through temporal correlation constraints to improve feature separability. Furthermore, the contribution backpropagation mechanism accurately locates the fault source. The HFE30CT201 pressure sensor contributes over 85% of 60% of stuck faults, consistent with expert annotations. Overall, RGMM-BAP, leveraging its regularized covariance matrix, bimodal classification architecture, and contribution backpropagation mechanism, establishes a closed-loop theoretical and practical approach to multi-condition wind turbine fault diagnosis, providing reliable support for intelligent operation and maintenance systems. Future research will focus on feature decoupling under multi-fault coupling to improve diagnostic accuracy for mixed fault modes.

[0142] The above description is merely an illustration of the preferred embodiments of the present disclosure and the technical principles employed. Those skilled in the art should understand that the scope of the invention encompassed by the embodiments of the present disclosure is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the aforementioned inventive concept. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A wind turbine fault diagnosis method based on multi-condition data synthesis and enhanced Bayesian abnormal probability is characterized by: It includes two stages: offline training and online monitoring. A offline training phase: A1) Generate synthetic fault data by using historical data of normal operating conditions and simulating four typical faults, namely offset, drift, noise, and stuck, to expand the training set; A1.1) Zero point offset Sensor zero offset is a common fault that manifests as a fixed offset of the measured value near the baseline. Randomly select 1 to 3 features and add a fixed offset. The mathematical expression is: ;in, is the measured value after offset, is the original measurement value, △ represents the offset, a random uniformly distributed value in the range of [2, 5], simulating the zero drift of the sensor; A1.2) Linear Drift One of the common faults is that the process parameters change linearly over time. In this model, a single feature increases linearly with a randomly generated slope. The mathematical expression is: ;in, is the measured value after the offset at time t, is the original measurement value at time t, k is the drift slope, a uniformly distributed random value in the range [0.1, 0.5], t is the time variable, which is simplified to 1 in single sample generation. This linear drift model simulates faults that change over time; A1.3) Asymmetric noise Asymmetric noise interference is a common phenomenon in actual operation. A single feature increases linearly under noise interference, with a randomly generated slope. Its mathematical expression is: Where ∈ is the noise intensity, which is usually a uniformly distributed random value in the range [1,3]. The probability of the noise direction is 70% positive and 30% negative. This model injects noise of different intensities by probability to simulate the noise interference in actual operation. A1.4) Actuator stuck Actuator stuck failure means that the actuator loses response and stays at a fixed value. In this model, the single feature is fixed to an extreme value, and the mathematical expression is: ; and are the minimum and maximum values ​​of the feature in normal data, respectively. This setting can simulate the fault situation where the actuator is stuck in the limit position; A2) Use normal data to train a regularized Gaussian mixture model (RGMM) to establish the probability distribution of normal operating conditions, and combine the expert knowledge base with the synthetic data to train a random forest classifier. The regularized Gaussian mixture model RGMM is established as follows: Normal operating condition data set , where n is the number of samples and d is the feature dimension; using the regularized Gaussian mixture model RGMM, the probability density function is: ; where K is the number of Gaussian components, π k is the mixing coefficient of the kth Gaussian component, which needs to satisfy π k ≥0, is the probability density function of the kth Gaussian distribution, whose mean is μ k , the covariance matrix is ​​(∑ k +λΙ), λ is the regularization parameter used to regularize the covariance matrix, and Ι is the d-order unit matrix; B Online monitoring stage: B1) After preprocessing, the real-time collected wind turbine data is used to calculate the Bayesian anomaly probability using the Regularized Gaussian Mixture Model (RGMM) and compared with the set threshold to trigger an alarm; In B1), the enhanced Bayesian anomaly probability calculation method is: The covariance matrix is ​​double-regularized. First, a dynamic threshold is used to avoid the covariance matrix from becoming singular: ; Among them, λ min and λ max are the minimum and maximum eigenvalues ​​of the covariance matrix, respectively, through the dynamic threshold λ min =max(10 -4 λ max ,10 -5 ) Avoid singularization of the covariance matrix and replace eigenvalues ​​smaller than the threshold with λ min , get the regularized eigenvalue matrix; then, reconstruct the covariance matrix: ; V is an orthogonal matrix consisting of the eigenvectors of the covariance matrix, It is a diagonal matrix composed of regularized eigenvalues, c is a regularization constant, I is the identity matrix, and the identity matrix regularization term cI is superimposed after reconstructing the covariance matrix; Based on the covariance matrix after double regularization, the Bayesian anomaly probability BAP under the double enhanced Bayesian posterior probability is: ;in, Represents x t to the mean μ i The Mahalanobis distance, g i and h i The scaling factor and degree of freedom correction terms are as follows: ; Based on the RGMM posterior probability π i Weight the abnormal probabilities of different Gaussian components to achieve probabilistic fusion of multi-operating condition data, and clarify the contribution of each component to the final score through probabilistic decomposition; To further smooth the fluctuation of BAP, a sliding window filtering technique is used with a sliding window size of ω=8. The specific filtering formula is as follows: ;in, is the BAP index after sliding window filtering, BAP t is the original BAP index. Through sliding window filtering, the BAP index at the current moment and the BAP index at the past ω moments are comprehensively considered, and the maximum value is taken as the filtered result. B2) If an anomaly is detected, the random forest classifier further identifies the fault type and locates the key fault parameters through contribution variable analysis, forming a closed-loop diagnosis.

2. The wind turbine fault diagnosis method based on multi-operating condition data synthesis and enhanced Bayesian abnormal probability according to claim 1 is characterized in that: The specific method of B2) is as follows: B2.1) The random forest classification method based on expert rules is: In the fault classification stage, the knowledge and experience of domain experts are utilized and the Delphi method is used to integrate the empirical rules of domain experts to form four types of rule templates; B2.1.1) Amplitude Constraints The amplitude constraint rule is used to detect sudden over-limit situations of the pressure sensor. The mathematical expression is: ;x i is the current measured value, μ i and σ i are the mean and standard deviation of normal data respectively, and k is the threshold coefficient. When the measured value exceeds a certain range, it is judged as a sudden over-limit failure of the pressure sensor; B2.1.2) Association constraints The association constraint rule is used for the pressure association of the inlet and outlet, and the mathematical expression is: ;in, and are the measured values ​​of inlet and outlet pressures, △ th The inlet and outlet pressure difference threshold value under normal circumstances. When the inlet and outlet pressure difference exceeds the threshold value, it is judged as an inlet and outlet pressure difference abnormal fault; B2.1.3) Temporal continuity The temporal persistence rule is used to detect when vibration exceeds the limit continuously. The mathematical expression is: ; is the vibration measurement value at time i, is the vibration threshold, T is the time window size, n is the threshold of the number of times the vibration exceeds the standard within the time window, and I(·) is the indicator function. When the number of times the vibration exceeds the standard within a certain time window reaches the threshold, it is judged as a continuous vibration exceeding the standard fault. B2.1.4) Combination constraints Combined constraint rules are used to jointly judge surge and flow, taking multiple parameters into consideration. ; are multiple measurement parameters related to surge and flow, f(·) is a comprehensive judgment function determined based on expert experience. When the function value is greater than 0, it is judged as a combined surge and flow fault; According to the expert rules, a random forest classifier based on expert rules is designed. For the random forest classifier, there are classification decisions: ; represents the final predicted category of the input sample x, c represents all possible categories, Tk(x) represents the predicted category of the kth decision tree for sample x, I(·) is the indicator function, which is 1 when Tk(x)=c and 0 otherwise; It is used to count the number of times that the predicted category is c in all trees, and integrate the prediction results of multiple decision trees through the "majority voting" mechanism. The final random forest classification result based on expert rules is: ; and are the prediction results of expert rules and random forest classifier respectively. represents a hyperparameter, ranging from 0 to 1, which is driven by rule-driven feature space segmentation, relying on random forest output under steady-state conditions and combining expert rules under transient conditions; B2.2 Fault variable location The calculation formula for the contribution of the Kth variable to the fault is as follows: Among them, CI j Represents the contribution index of the jth feature, π k is the weight of the kth Gaussian distribution, V k is the eigenvector of the kth Gaussian distribution, x is the eigenvector at the current moment, μ k is the mean of the kth Gaussian distribution, where k is the number of Gaussian distributions. By calculating the feature contribution index, the impact of each feature on the fault is evaluated and the fault location is located.

Citation Information

Patent Citations

  • Key information infrastructure asset identification method combined with mixed random forest

    CN110245693A

  • Lightweight satellite fault diagnosis method based on self-supervised learning

    CN118861710A

  • Multi-working-condition process industrial fault detection and diagnosis method based on deep transfer learning

    WO2023071217A1