Fan fault diagnosis method based on multi-working-condition data synthesis and Bayesian anomaly probability enhancement

Through multi-condition data synthesis and enhancement Bayesian anomaly probability algorithm, combined with regularized Gaussian hybrid model and sliding window filtering technology, the problems of multi-failure coupling, singular covariance matrix and poor adaptability in the fan unit fault diagnosis are solved, and efficient and accurate fault detection and diagnosis are achieved.

CN120030454AActive Publication Date: 2025-05-23LIAONING DONGKE ELECTRIC POWER

Patent Information

Application Number
CN202510475036.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-23
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The existing blower unit fault diagnosis methods are difficult to effectively deal with the problems of multi-fault coupling, singular covariance matrix and poor adaptability, resulting in insufficient diagnostic accuracy and reliability, and the inability to timely detect and warning of potential faults.

Method used

Multi-case data synthesis, enhanced Bayesian anomaly probability (BAP) algorithm and hierarchical diagnostic method are used to improve the accuracy and stability of fault detection through regularized Gaussian hybrid model (RGMM) and sliding window filtering technology, and fault type classification and location are combined with the random forest classifier of expert rules.

Benefits of technology

It significantly improves the accuracy, stability and adaptability of fault detection, can promptly detect and warn of potential faults, reduce false alarm rates, and improve the safety and reliability of the fan unit.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030454A_ABST
    Figure CN120030454A_ABST
Patent Text Reader

Abstract

The invention discloses a fan fault diagnosis method based on multi-working-condition data synthesis and Bayesian anomaly probability enhancement. The problem of detection of a fan set under the multi-fault condition is solved. According to the method, a simulation model including various fault working conditions such as sensor zero drift, process parameter linear drift, asymmetric noise interference and actuator jamming is constructed, and high-authenticity fault data is generated through dynamic feature selection and a noise asymmetric injection strategy. The Bayesian anomaly probability BAP algorithm is improved, the covariance matrix regularization and sliding window filtering technology is introduced, the problem of covariance matrix singularity is solved, and the stability and robustness of fault detection are improved. According to the method, end-to-end fault analysis under different working conditions is realized through a three-stage diagnosis system of'detection-classification-positioning 'in combination with isolated forest data cleaning, the dynamic Gaussian mixture model GMM and the random forest classifier, and the fault detection capability and diagnosis precision of a coal-fired unit fan unit under the multi-fault coupling condition are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an intelligent fault diagnosis technology for a coal-fired unit fan group, which combines multi-operating condition data synthesis, a Bayesian abnormal probability BAP algorithm and a hierarchical diagnosis method to improve the accuracy and adaptability of fault detection. Background Art

[0002] Coal-fired units occupy a core position in my country's power supply system. Their safe and stable operation is crucial to ensuring the reliability of power supply and the stability of the power grid. As a key auxiliary equipment of coal-fired units, wind turbines are responsible for providing the air required for the combustion process and discharging the exhaust gas after combustion. Their operating conditions directly affect the overall performance and safety of the units.

[0003] In the actual operation of wind turbines, the probability of wind turbine failure is high due to the long-term complex and changeable working environment, high temperature, high pressure, strong vibration and other harsh conditions. Once a wind turbine fails, it may lead to a decrease in unit efficiency and an increase in energy consumption, causing huge economic losses to power companies and seriously affecting the stable operation of the power grid.

[0004] With the continuous improvement of the intelligence level of my country's power industry, the traditional wind turbine fault diagnosis method can no longer effectively cope with the complex working conditions of multiple fault coupling in modern industry, resulting in low diagnostic accuracy and easy interference from frequently fluctuating working conditions, and unable to timely detect and warn potential faults. Therefore, there is an urgent need for an efficient, accurate and highly adaptable wind turbine fault diagnosis method to cope with multiple fault coupling and complex working conditions, and to ensure the safe and stable operation of coal-fired unit wind turbines.

[0005] In the prior art, wind turbine fault diagnosis usually relies on a simulation model based on a single fault mode to build a diagnostic system. The design idea of ​​these traditional methods is to simulate a certain type of fault and build a corresponding diagnostic model. However, in actual industrial applications, wind turbines are usually faced with changes in multiple complex working conditions and the mutual coupling and interaction of multiple faults. In this changing environment, the simulation of a single fault mode cannot fully and accurately reflect the complexity and diversity of wind turbine faults. Therefore, the diagnostic model that relies on a single fault mode shows poor adaptability and diagnostic ability in practical applications, especially in the case of multiple fault coupling, it is often unable to effectively detect and diagnose the interaction of multiple faults. Further, most of the existing fault diagnosis methods are based on data-driven models, among which the Bayesian posterior probability inference method is widely used. This method can infer the state of the system based on the observed data to a certain extent, but it has certain limitations when dealing with variable actual working conditions. In practical applications, Bayesian posterior probability inference often relies on the covariance matrix for calculation, but when the covariance matrix is ​​singular, it often leads to instability of the inference result, which in turn affects the accuracy and reliability of fault diagnosis. The existence of the covariance matrix singularity problem makes it impossible for traditional Bayesian inference to provide accurate and reliable evaluation of the actual operating status of the equipment, thereby affecting the early warning and fault location of wind turbine failures, and further reducing the efficiency and accuracy of fault diagnosis.

[0006] Traditional methods usually use fixed-structure models, which lack sufficient flexibility when facing changes in wind turbines under different operating conditions. Since the operating conditions of wind turbines change frequently during actual operation due to fluctuations in factors such as unit load, ambient temperature, and humidity, fixed-structure models cannot adjust their internal parameters and model structures in a timely manner according to these operating condition changes. This makes it difficult for the model to maintain good adaptability when facing complex and changing operating conditions, resulting in a decrease in diagnostic performance and failing to meet the increasingly stringent requirements for equipment reliability and safety in modern power production.

[0007] Existing fault diagnosis methods for wind turbines have significant deficiencies and limitations when dealing with complex scenarios such as multi-fault coupling and operating condition changes of wind turbines. Existing diagnostic methods are difficult to effectively solve problems such as multi-fault coupling, singular covariance matrix, and poor adaptability. Therefore, there is an urgent need for a more efficient, accurate and highly adaptable fault diagnosis method that can effectively deal with the multi-fault coupling characteristics of wind turbines under complex working conditions, improve the accuracy, stability and adaptability of fault detection, and thus ensure the safe and stable operation of wind turbines in coal-fired units. Summary of the invention

[0008] The purpose of the present invention is to provide an intelligent fault detection and diagnosis method for the problem of multi-fault coupling detection of coal-fired unit fan unit. By integrating multi-condition data synthesis, enhancing the Bayesian abnormal probability BAP algorithm and hierarchical diagnosis technology, it aims to overcome the shortcomings of traditional fault diagnosis methods and improve the accuracy, stability and adaptability of fault diagnosis. The present invention can solve the following technical problems existing in the prior art through the above technical solution: 1) Multiple fault coupling detection problem: Multiple fault coupling is common in wind turbine operation, and traditional methods are difficult to cope with its complexity. The present invention provides more accurate fault detection and diagnosis by integrating multi-condition data synthesis and enhancing the Bayesian anomaly probability (BAP) algorithm.

[0009] 2) The singularity problem of the covariance matrix affects the stability and accuracy of fault diagnosis. The present invention improves the stability of the diagnosis results through regularization and sliding window filtering technology, detects anomalies in time, and prevents unplanned downtime.

[0010] 3) Adaptability to operating condition fluctuations: Traditional methods cannot flexibly cope with wind turbine operating condition fluctuations, resulting in reduced diagnostic performance. The present invention combines a multi-operating condition fault simulation model with a dynamic feature selection strategy to effectively improve the adaptability of the diagnostic model.

[0011] 4) Insufficient fault detection efficiency and accuracy: Traditional methods often fail to detect potential faults in a timely manner, affecting the operating efficiency of the unit. The present invention significantly improves the efficiency and accuracy of fault detection through the three-level diagnosis system of "detection-classification-location" and advanced algorithms.

[0012] 5) Equipment safety and reliability issues: Traditional methods are prone to misjudgment under variable working conditions and fail to detect potential risks in a timely manner. This invention accurately identifies potential fault hazards and improves the safety and reliability of wind turbines through multi-dimensional data fusion and dynamic analysis.

[0013] In order to achieve the above objectives, the technical solution adopted by the present invention is: a fan fault diagnosis method based on multi-condition data synthesis and enhanced Bayesian abnormal probability, including two stages: "offline training" and "online monitoring".

[0014] A offline training phase: A1) Generate synthetic fault data through historical data of normal operating conditions and simulate four typical faults, namely offset, drift, noise, and jamming, to expand the training set.

[0015] The method for synthesizing multi-condition fault data is: A1.1) Zero point offset Sensor zero offset is a common fault that manifests as a fixed offset of the measured value near the baseline. Specifically, 1 to 3 features are randomly selected and a fixed offset is added. The mathematical expression is: ;in, is the measured value after offset, is the original measurement value, △ represents the offset, which is usually a random uniformly distributed value in the range of [2, 5], simulating the zero drift of the sensor.

[0016] A1.2) Linear drift One of the common faults is that the process parameters change linearly over time, such as the temperature continues to rise and fall. In this model, a single feature increases linearly, and the slope is randomly generated. Its mathematical expression is: ; is the measured value after the offset at time t, is the original measurement value at time t, k is the drift slope, a uniformly distributed random value in the range of [0.1, 0.5], and t is the time variable, which is simplified to 1 in single sample generation. This linear drift model simulates faults that change over time.

[0017] A1.3) Asymmetric noise It is a common phenomenon in actual operation that the signal is interfered by asymmetric noise, such as sensor vibration or electromagnetic interference. In the present invention, a single feature increases linearly under noise interference, and the slope is randomly generated. Its mathematical expression is: ∈ is the noise intensity, which is usually a uniformly distributed random value in the range of [1,3]. The probability of the noise direction is 70% positive and 30% negative. The model of the present invention injects noise of different intensities through probability, truly simulates the noise interference in actual operation, and enhances data diversity and complexity.

[0018] A1.4) Actuator stuck Actuator stuck fault means that the actuator loses response and stays at a fixed value, such as the valve opening cannot change. In this model, the single feature is fixed to the extreme value, and the mathematical expression is: ; and are the minimum and maximum values ​​of the feature in normal data, respectively. This setting can accurately simulate the fault condition where the actuator is stuck in the limit position.

[0019] The present invention designs a feature selection mechanism to evenly distribute the sensitive feature ratio interval. By dynamically adjusting the ratio, the impact of faults on features is controlled, data diversity and accuracy are enhanced, and feature differences under different fault scenarios are simulated.

[0020] A2) Use normal data to train the regularized Gaussian mixture model RGMM to establish the probability distribution of normal operating status, and combine the expert knowledge base with synthetic data to train the random forest classifier.

[0021] The regularized Gaussian mixture model RGMM is established as follows: Normal operating data set , where n is the number of samples and d is the feature dimension. When processing small samples or noisy data, the covariance matrix of the traditional Gaussian mixture model may become ill-conditioned, reducing robustness. To solve this problem, this paper adopts the regularized Gaussian mixture model RGMM, whose probability density function is: ; where K is the number of Gaussian components, π k is the mixing coefficient of the kth Gaussian component, which needs to satisfy π k ≥0, is the probability density function of the kth Gaussian distribution, whose mean is μ k , the covariance matrix is ​​(∑ k +λΙ), λ is the regularization parameter used to regularize the covariance matrix, and Ι is the d-order unit matrix. By adding the regularization term λΙ to the covariance matrix, it is possible to avoid the situation where the covariance matrix may be singular or nearly singular when the number of samples is small. After adding the regularization term, the covariance matrix can maintain good properties, thereby ensuring the stability and accuracy of the model. B Online monitoring stage: B1) After preprocessing, the fan data collected in real time is used to calculate the Bayesian anomaly probability through the regularized Gaussian mixture model (RGMM), and compared with the set threshold to trigger an alarm.

[0022] The enhanced Bayesian anomaly probability calculation method is: Traditional Bayesian posterior probability calculation is easily affected by the singularity of the covariance matrix, which leads to instability and affects the diagnostic accuracy. To enhance stability, this paper performs double regularization on the covariance matrix.

[0023] First, we avoid singularization of the covariance matrix by using a dynamic threshold: min =max(10 -4 λ max ,10 -5 ) Among them, λ min and λ max are the minimum and maximum eigenvalues ​​of the covariance matrix, respectively, through the dynamic threshold λ min =max(10 -4 λ max ,10 -5 ) avoids the covariance matrix from becoming singular by replacing eigenvalues ​​smaller than the threshold with λ min, and obtain the regularized eigenvalue matrix to enhance numerical stability. Next, reconstruct the covariance matrix: ; V is an orthogonal matrix consisting of the eigenvectors of the covariance matrix, It is a diagonal matrix composed of regularized eigenvalues, c is a regularization constant, I is the unit matrix, and the unit matrix regularization term cI is superimposed after reconstructing the covariance matrix. The double regularization strategy can alleviate the overfitting problem under high-dimensional data and improve the robustness of the model. Different from traditional single regularization methods such as L 2 Regularization,This method uses eigenvalue decomposition and truncation techniques to retain the direction information of the principal component while suppressing noise.

[0024] Based on the covariance matrix after double regularization, the Bayesian anomaly probability BAP under the double enhanced Bayesian posterior probability is: ;in, Represents x t to the mean μ i The Mahalanobis distance, g i and h i The scaling factor and degree of freedom correction terms are as follows: ; Dynamically calculate the degree of freedom parameter g through trace operation i and h i , solving the traditional chi-square distribution Χ 2 (·) Model bias caused by assuming fixed degrees of freedom. In addition, introducing the interaction between the covariance matrix and the inverse covariance matrix into the calculation of degrees of freedom also expands the application boundary of the chi-square distribution in anomaly detection.

[0025] Based on RGMM posterior probability π i The abnormal probabilities of different Gaussian components are weighted to achieve probabilistic fusion of multi-condition data, which is better than the traditional single model detection method. At the same time, the contribution of each component to the final score is clarified through probabilistic decomposition, which is convenient for subsequent model diagnosis and tuning.

[0026] In addition, in order to further smooth the fluctuation of BAP and reduce the occurrence of false alarms, this paper adopts sliding window filtering technology. The sliding window size ω=8, and the specific filtering formula is as follows: ;in, is the BAP index after sliding window filtering, BAP tis the original BAP index. Through sliding window filtering, the BAP index at the current moment and the BAP index at the past ω moments are comprehensively considered, and the maximum value is taken as the filtered result. By combining time series analysis, the stability and accuracy of fault detection can be improved, which helps to suppress the impact of instantaneous abnormal fluctuations on the score, and is suitable for online detection of non-steady-state data streams.

[0027] B2) If an anomaly is detected, the random forest classifier further identifies the fault type and locates the key fault parameters through contribution variable analysis, forming a closed-loop diagnosis.

[0028] B2.1) The random forest classification method based on expert rules is: In the fault classification stage, in order to make full use of the knowledge and experience of domain experts and improve the accuracy and robustness of classification, this paper integrates 137 empirical rules of 7 domain experts through the Delphi method to form four types of rule templates. B2.1.1) Amplitude constraint The amplitude constraint rule is mainly used to detect the sudden over-limit situation of the pressure sensor, and its mathematical expression is: ;x i is the current measured value, μ i and σ i are the mean and standard deviation of normal data respectively, and k is the threshold coefficient. When the measured value exceeds a certain range, it is judged as a sudden over-limit fault of the pressure sensor.

[0029] B2.1.2) Association constraints The association constraint rule is used for the pressure association of the inlet and outlet, and the mathematical expression is: ;in, and are the measured values ​​of inlet and outlet pressures, △ th It is the inlet and outlet pressure difference threshold under normal circumstances. When the inlet and outlet pressure difference exceeds the threshold, it is judged as an inlet and outlet pressure difference abnormal fault.

[0030] B2.1.3) Temporal continuity The timing continuity rule is used to detect the situation where the vibration exceeds the limit continuously. The mathematical expression is: ; is the vibration measurement value at time i, is the vibration threshold, T is the time window size, n is the threshold of the number of vibration exceeding the standard within the time window, and I(·) is the indicator function. When the number of vibration exceeding the standard within a certain time window reaches a certain threshold, it is judged as a vibration continuous exceeding standard fault.

[0031] B2.1.4) Combination constraints The combined constraint rule is used for joint judgment of surge and flow. The mathematical expression is relatively complex and requires comprehensive consideration of multiple parameters. For example: ; are multiple measurement parameters related to surge and flow, and f(·) is a comprehensive judgment function determined based on expert experience. When the function value is greater than 0, it is judged as a combined surge and flow fault.

[0032] According to the above expert rules, a random forest classifier based on expert rules is designed. For the random forest classifier, there are classification decisions: ; represents the final predicted category of the input sample x, c represents all possible categories, Tk(x) represents the predicted category of the sample x by the kth decision tree, I(·) is the indicator function, which is 1 when Tk(x)=c, otherwise it is 0; It is used to count the number of times the predicted category c is in all trees, and integrate the prediction results of multiple decision trees through the "majority voting" mechanism to improve the accuracy and robustness of classification. Finally, the random forest classification result based on expert rules is obtained as follows: ; and are the prediction results of expert rules and random forest classifier, respectively. It represents a hyperparameter, ranging from 0 to 1. It uses rule-driven feature space segmentation, relies on random forest output under steady-state conditions, and combines expert rules under transient conditions to reduce the false alarm rate under variable load conditions.

[0033] B2.2) The fault variable location method is: Accurately locating the fault variable is crucial for fault repair and equipment maintenance. The formula for calculating the contribution of the jth variable to the fault is as follows: Among them, CI j Represents the contribution index of the jth feature, π k is the weight of the kth Gaussian distribution, V k is the eigenvector of the kth Gaussian distribution, x is the eigenvector at the current moment, μ k is the mean of the kth Gaussian distribution, and k is the number of Gaussian distributions. By calculating the feature contribution index, the impact of each feature on the fault is evaluated, and then the fault location is located.

[0034] The beneficial effects of the present invention are: The present invention designs a three-level diagnosis method of "detection-classification-location", which starts with fault detection and determines the occurrence of faults by enhancing the Bayesian abnormal probability; then uses a random forest classifier combined with expert rules to classify the fault type; finally, determines the fault location by variable location, forming a complete fault diagnosis closed loop. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 Provide a technical roadmap for online automatic verification of telecontrol transmission data based on Ethernet; Figure 2 The confusion matrix results for different methods for different data types are shown below; Figure 3 This is the fault location result (first 7 variables) of the RGMM-BAP method; Figure 4 Results for selecting the value of k in the RGMM model. DETAILED DESCRIPTION

[0036] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0037] like Figure 1 The figure shows the technical roadmap for wind turbine fault diagnosis based on multi-condition data synthesis and enhanced Bayesian anomaly probability, which can be divided into two stages: "offline training" and "online monitoring". In the offline training stage, synthetic fault data is generated through normal operating historical data and simulation of four typical faults (offset, drift, noise, and jamming) to expand the training set. Then, the regularized Gaussian mixture model (RGMM) is trained with normal data to establish the probability distribution of normal operating status, and the random forest classifier is trained in combination with the expert knowledge base and synthetic data. In the online monitoring stage, the real-time collected wind turbine data is preprocessed, and the Bayesian anomaly probability is calculated by the regularized Gaussian mixture model (RGMM), and the alarm is triggered by comparing it with the set threshold. If an anomaly is detected, the random forest classifier further identifies the fault type and locates the key fault parameters through contribution variable analysis to form a closed-loop diagnosis.

[0038] Functions of each part of the device of the present invention: 1) SCADA system: used to collect the operating data of the fan unit in real time and transmit the data.

[0039] 2) Data storage system: used to store real-time data and historical data collected from the SCADA system.

[0040] 3) High-performance computing platform: used to run fault diagnosis algorithms, such as the Bayesian Anomaly Probability (BAP) algorithm.

[0041] 4) Fault diagnosis system: Based on the data processing platform, fault detection, classification and location of the fan unit are carried out.

[0042] 5) Remote monitoring equipment: allows operators to monitor and diagnose the status of the fan unit through remote equipment.

[0043] The method of the present invention comprises: (1) synthesis of multi-operating condition fault data; (2) establishment of a regularized Gaussian mixture model (RGMM); (3) enhanced Bayesian abnormal probability calculation; (4) random forest classification based on expert rules; and (5) fault variable location.

[0044] Step 1: Synthesis of multi-condition fault data 1) Zero offset Sensor zero offset is a common fault that manifests as a fixed offset of the measured value near the baseline. Specifically, 1 to 3 features are randomly selected and a fixed offset is added. The mathematical expression is: ; is the measured value after offset, is the original measurement value, △ represents the offset, which is usually a random uniformly distributed value in the range of [2, 5], simulating the zero drift of the sensor.

[0045] 2) Linear drift One of the common faults is that the process parameters change linearly over time, such as the temperature continues to rise and fall. In this model, a single feature increases linearly, and the slope is randomly generated. Its mathematical expression is: ; is the measured value after the offset at time t, is the original measurement value at time t, k is the drift slope, a uniformly distributed random value in the range of [0.1, 0.5], and t is the time variable, which is simplified to 1 in single sample generation. This linear drift model simulates faults that change over time.

[0046] 3) Asymmetric noise It is a common phenomenon in actual operation that the signal is interfered by asymmetric noise, such as sensor vibration or electromagnetic interference. In the present invention, a single feature increases linearly under noise interference, and the slope is randomly generated. Its mathematical expression is: ∈ is the noise intensity, which is usually a uniformly distributed random value in the range of [1,3]. The probability of the noise direction is 70% positive and 30% negative. The model of the present invention injects noise of different intensities through probability, truly simulates the noise interference in actual operation, and enhances data diversity and complexity.

[0047] 4) The actuator is stuck Actuator stuck fault means that the actuator loses response and stays at a fixed value, such as the valve opening cannot change. In this model, the single feature is fixed to the extreme value, and the mathematical expression is: ; and are the minimum and maximum values ​​of the feature in normal data, respectively. This setting can accurately simulate the fault condition where the actuator is stuck in the limit position.

[0048] The present invention designs a feature selection mechanism to evenly distribute the sensitive feature ratio interval. By dynamically adjusting the ratio, the impact of faults on features is controlled, data diversity and accuracy are enhanced, and feature differences under different fault scenarios are simulated.

[0049] Step 2: Regularized Gaussian mixture model (RGMM) establishment Normal operating data set , where n is the number of samples and d is the feature dimension. When processing small samples or noisy data, the covariance matrix of the traditional Gaussian mixture model may become ill-conditioned, reducing robustness. To solve this problem, this paper adopts the regularized Gaussian mixture model (RGMM), whose probability density function is: ; where K is the number of Gaussian components, π k is the mixing coefficient of the kth Gaussian component, which needs to satisfy π k ≥0, is the probability density function of the kth Gaussian distribution, whose mean is μ k , the covariance matrix is ​​(∑ k +λΙ), λ is the regularization parameter used to regularize the covariance matrix, and Ι is the d-order unit matrix. By adding the regularization term λΙ to the covariance matrix, it is possible to avoid the situation where the covariance matrix may be singular or nearly singular when the number of samples is small. After adding the regularization term, the covariance matrix can maintain good properties, thereby ensuring the stability and accuracy of the model. B Online monitoring stage: Step 3: Enhanced Bayesian anomaly probability calculation Traditional Bayesian posterior probability calculation is easily affected by the singularity of the covariance matrix, which leads to instability and affects the diagnostic accuracy. To enhance stability, this paper performs double regularization on the covariance matrix.

[0050] First, we avoid singularization of the covariance matrix by using a dynamic threshold: min =max(10 -4 λ max ,10 -5 ) Among them, λ min and λ maxare the minimum and maximum eigenvalues ​​of the covariance matrix, respectively, through the dynamic threshold λ min =max(10 -4 λ max ,10 -5 ) avoids the covariance matrix from becoming singular by replacing eigenvalues ​​smaller than the threshold with λ min , and obtain the regularized eigenvalue matrix to enhance numerical stability. Next, reconstruct the covariance matrix: ; V is an orthogonal matrix consisting of the eigenvectors of the covariance matrix, is a diagonal matrix composed of regularized eigenvalues, c is a regularization constant, I is the identity matrix, and the identity matrix regularization term cI is superimposed after reconstructing the covariance matrix. The double regularization strategy can alleviate the overfitting problem under high-dimensional data and improve the robustness of the model. Different from the traditional single regularization method (such as L 2 Regularization) This method uses eigenvalue decomposition and truncation techniques to retain the direction information of the principal component while suppressing noise.

[0051] Based on the covariance matrix after double regularization, the Bayesian anomaly probability (BAP) under the double enhanced Bayesian posterior probability is: ;in, Represents x t to the mean μ i The Mahalanobis distance, g i and h i The scaling factor and degree of freedom correction terms are as follows: ; Dynamically calculate the degree of freedom parameter g through trace operation i and h i , solving the traditional chi-square distribution Χ 2 (·) Model bias caused by assuming fixed degrees of freedom. In addition, introducing the interaction between the covariance matrix and the inverse covariance matrix into the calculation of degrees of freedom also expands the application boundary of the chi-square distribution in anomaly detection.

[0052] Based on RGMM posterior probability π i The abnormal probabilities of different Gaussian components are weighted to achieve probabilistic fusion of multi-condition data, which is better than the traditional single model detection method. At the same time, the contribution of each component to the final score is clarified through probabilistic decomposition, which is convenient for subsequent model diagnosis and tuning.

[0053] In addition, in order to further smooth the fluctuation of BAP and reduce the occurrence of false alarms, this paper adopts sliding window filtering technology. The sliding window size ω=8, and the specific filtering formula is as follows: ;in, is the BAP index after sliding window filtering, BAP t is the original BAP index. Through sliding window filtering, the BAP index at the current moment and the BAP index at the past ω moments are comprehensively considered, and the maximum value is taken as the filtered result. By combining time series analysis, the stability and accuracy of fault detection can be improved, which helps to suppress the impact of instantaneous abnormal fluctuations on the score, and is suitable for online detection of non-steady-state data streams.

[0054] Step 4: Expert rule-based random forest classification In the fault classification stage, in order to make full use of the knowledge and experience of domain experts and improve the accuracy and robustness of classification, this paper integrates 137 empirical rules of 7 domain experts through the Delphi method to form four types of rule templates.

[0055] 1) Amplitude constraint The amplitude constraint rule is mainly used to detect the sudden over-limit situation of the pressure sensor, and its mathematical expression is: ;in, is the current measured value, μ i and σ i are the mean and standard deviation of normal data respectively, k is the threshold coefficient. When the measured value exceeds a certain range, it is judged as a sudden over-limit failure of the pressure sensor.

[0056] 2) Association constraints The association constraint rule is used for the pressure association of the inlet and outlet, and the mathematical expression is: ;in, and are the measured values ​​of inlet and outlet pressures, △ th It is the inlet and outlet pressure difference threshold under normal circumstances. When the inlet and outlet pressure difference exceeds the threshold, it is judged as an inlet and outlet pressure difference abnormal fault.

[0057] 3) Temporal continuity The timing continuity rule is used to detect the situation where the vibration exceeds the limit continuously. The mathematical expression is: ; is the vibration measurement value at time i, is the vibration threshold, T is the time window size, n is the threshold of the number of vibration exceeding the standard within the time window, and I(·) is the indicator function. When the number of vibration exceeding the standard within a certain time window reaches a certain threshold, it is judged as a vibration continuous exceeding standard fault.

[0058] 4) Combination constraints The combined constraint rule is used for joint judgment of surge and flow. The mathematical expression is relatively complex and requires comprehensive consideration of multiple parameters. For example: ; are multiple measurement parameters related to surge and flow, and f(·) is a comprehensive judgment function determined based on expert experience. When the function value is greater than 0, it is judged as a combined surge and flow fault.

[0059] According to the above expert rules, a random forest classifier based on expert rules is designed. For the random forest classifier, there are classification decisions: ; represents the final predicted category of the input sample x, c represents all possible categories, Tk(x) represents the predicted category of the sample x by the kth decision tree, I(·) is the indicator function, which is 1 when Tk(x)=c, otherwise it is 0; It is used to count the number of times the predicted category is c in all trees, and integrate the prediction results of multiple decision trees through the "majority voting" mechanism to improve the accuracy and robustness of classification. The final random forest classification result based on expert rules is: ; and are the prediction results of expert rules and random forest classifier, respectively. It represents a hyperparameter, ranging from 0 to 1. It uses rule-driven feature space segmentation, relies on random forest output under steady-state conditions, and combines expert rules under transient conditions to reduce the false alarm rate under variable load conditions.

[0060] Step 5: Fault variable location Accurately locating the fault variable is crucial for fault repair and equipment maintenance. The calculation formula for the contribution of the first variable to the fault is as follows: Among them, CI j Represents the contribution index of the jth feature, π k is the weight of the kth Gaussian distribution, V k is the eigenvector of the kth Gaussian distribution, x is the eigenvector at the current moment, μ k is the mean of the kth Gaussian distribution, and k is the number of Gaussian distributions. By calculating the feature contribution index, the impact of each feature on the fault is evaluated, and then the fault position is located.

[0061] Based on this, this paper designs a three-level diagnosis system of "detection-classification-location", and the overall process is shown in Figure 1. The system starts with fault detection, and judges the occurrence of faults by enhancing the Bayesian abnormal probability; then uses the random forest classifier combined with expert rules to classify the fault type; finally, the fault location is determined by variable location, forming a complete fault diagnosis closed loop. The present invention is further described in detail below through specific embodiments.

[0062] 1) Experimental scenario and data configuration The data used in this experiment comes from the SCADA data of the wind turbine of a 300MW unit in a power plant from 2020 to 2021. The sampling frequency is 1Hz, covering the key operating parameters of the wind turbine under normal and various fault conditions (such as temperature, pressure, speed, etc.), providing an important basis for fault diagnosis. In order to ensure the reliability of the experiment, the data has been carefully cleaned and preprocessed to remove outliers and fill in missing values ​​to ensure the integrity of the data.

[0063] In the experiment, PCA, SVM, and CNN-LSTM were selected as comparison methods. PCA is mainly used for data dimension reduction and feature extraction, SVM is a classic classification algorithm with good generalization ability, and CNN-LSTM combines convolutional neural network and long short-term memory network, which is suitable for processing wind turbine data with spatiotemporal characteristics.

[0064] 2) Comparative test and analysis of experimental results The specific experimental data statistics are shown in Table 1, which shows the F1 scores of different methods for fault detection. From the table, we can intuitively compare the performance differences of different methods in fault detection.

[0065]

[0066] Table 2 shows the comprehensive classification performance of different methods for different data types, expressed as precision (%). The data with the highest performance in each category is in bold:

[0067] Figure 2 The confusion matrices of different methods for classifying different data types are shown in Figure 1. (a) to (d) are PCA, SVM, CNN-LSTM, and the method proposed in this paper, RGMM-BAP. Through the analysis of these confusion matrices, it can be clearly seen that the method proposed in this paper has significant advantages in fault classification and can more accurately classify different types of faults.

[0068] Figure 3 The horizontal axis represents the positioning result of the model for fault detection, and the horizontal axis reflects the contribution of the random forest classifier to the fault classification model. Fault location helps technicians quickly find the key factors of fault occurrence in actual model applications and improve the efficiency of fault repair.

[0069] Figure 4 shows the k value selection results in the RGMM model of this paper. According to the distribution of the data scatter plot, it can be seen that the optimal k value is 5, which is circled by a red ellipse in the figure. Reasonable selection of k value is crucial to the performance of the RGMM model. It can ensure that the model can maintain good performance under different working conditions and improve the accuracy and reliability of fault diagnosis.

[0070] Based on the experimental results (Table 1-2, Figure 2-4 ), the RGMM-BAP method proposed in this paper shows significant advantages in fault detection, classification accuracy and positioning effectiveness. Compared with traditional methods, its F1 score reaches 90.4%, which is significantly better than PCA (50.6%), SVM (76.3%) and CNN-LSTM (74.9%). Through the double regularization covariance matrix and sliding window filtering, the detection robustness is improved and the false alarm rate is reduced to 0.3%. In the multi-fault classification task, RGMM-BAP reaches 99.7% in normal data recognition, and effectively improves the classification accuracy of linear drift (53.8%) and actuator stuck (60.3%) faults, which is better than the traditional method (both 0%). However, there is still 21% cross misclassification of drift and stuck faults, and the feature separability can be optimized through temporal correlation constraints in the future. In addition, the contribution back-propagation mechanism accurately locates the fault source, and the contribution of pressure sensor HFE30CT201 exceeds 85% in 60% stuck faults, which is consistent with the expert annotation. In general, RGMM-BAP builds a closed loop of theory and practice for fault diagnosis of wind turbines under multiple working conditions by means of regularized covariance matrix, bimodal classification architecture and contribution back propagation mechanism, providing reliable support for intelligent operation and maintenance system. In the future, the focus will be on feature decoupling under multiple fault coupling to improve the diagnostic accuracy of mixed fault modes.

[0071] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the above features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure to form a technical solution.

Claims

1. A fan fault diagnosis method based on multi-condition data synthesis and enhanced Bayesian abnormal probability, characterized in that: It includes two stages: offline training and online monitoring. A offline training phase: A1) Generate synthetic fault data and expand the training set by using historical data of normal working conditions and simulating four typical faults, namely offset, drift, noise, and stuck; A2) Use normal data to train the regularized Gaussian mixture model RGMM to establish the probability distribution of normal operating status, and combine the expert knowledge base with synthetic data to train the random forest classifier; The regularized Gaussian mixture model RGMM is established as follows: Normal operating data set , where n is the number of samples and d is the feature dimension; the regularized Gaussian mixture model RGMM is used, and the probability density function is: ; where K is the number of Gaussian components, π k is the mixing coefficient of the kth Gaussian component, which needs to satisfy π k ≥0, is the probability density function of the kth Gaussian distribution, whose mean is μ k , the covariance matrix is ​​(∑ k +λΙ), λ is the regularization parameter used to regularize the covariance matrix, and Ι is the d-order unit matrix; By adding the regularization term λΙ to the covariance matrix, the covariance matrix can be prevented from being singular or nearly singular when the number of samples is small. The covariance matrix can maintain its properties after adding the regularization term. B Online monitoring stage: B1) After preprocessing, the real-time collected wind turbine data is used to calculate the Bayesian abnormal probability through the regularized Gaussian mixture model (RGMM), and then compared with the set threshold to trigger an alarm; B2) If an anomaly is detected, the random forest classifier further identifies the fault type and locates the key fault parameters through contribution variable analysis, forming a closed-loop diagnosis.

2. The wind turbine fault diagnosis method based on multi-condition data synthesis and enhanced Bayesian abnormal probability according to claim 1 is characterized in that: In A1), the method for synthesizing multi-condition fault data is: A1.1) Zero point offset Sensor zero offset is a common fault that manifests as a fixed offset of the measured value near the baseline; randomly select 1 to 3 features and add a fixed offset, the mathematical expression is: ;in, is the measured value after offset, is the original measurement value, △ represents the offset, a random uniformly distributed value in the range of [2, 5], simulating the zero drift of the sensor; A1.2) Linear drift One of the common faults is that the process parameters change linearly with time. In this model, a single feature increases linearly, and the slope is randomly generated. The mathematical expression is: ;in, is the measured value after the offset at time t, is the original measurement value at time t, k is the drift slope, a uniformly distributed random value in the range of [0.1, 0.5], t is the time variable, which is simplified to 1 in single sample generation. This linear drift model simulates faults that change over time; A1.3) Asymmetric noise It is a common phenomenon in actual operation that a signal is interfered by asymmetric noise. In the present invention, a single feature increases linearly under noise interference, and the slope is randomly generated. Its mathematical expression is: ; Wherein, ∈ is the noise intensity, which is usually a uniformly distributed random value in the range of [1,3], and the probability of the noise direction is 70% positive and 30% negative; the model of the present invention injects noise of different intensities through probability to truly simulate the noise interference in actual operation; A1.4) Actuator stuck The actuator stuck fault means that the actuator loses response and stays at a fixed value. In this model, the single feature is fixed to the extreme value, and the mathematical expression is: ; and are the minimum and maximum values ​​of the feature in normal data respectively. This setting can simulate the fault situation where the actuator is stuck in the limit position.

3. The wind turbine fault diagnosis method based on multi-condition data synthesis and enhanced Bayesian abnormal probability according to claim 1 is characterized in that: In B1), the enhanced Bayesian anomaly probability calculation method is: The covariance matrix is ​​double regularized. First, a dynamic threshold is used to avoid singularization of the covariance matrix: ; Among them, λ min and λ max are the minimum and maximum eigenvalues ​​of the covariance matrix, respectively, through the dynamic threshold λ min =max(10 -4 λ max ,10 -5 ) To avoid the covariance matrix from becoming singular, replace the eigenvalues ​​smaller than the threshold with λ min , get the regularized eigenvalue matrix to enhance numerical stability; then, reconstruct the covariance matrix: ; V is an orthogonal matrix consisting of the eigenvectors of the covariance matrix, It is a diagonal matrix composed of regularized eigenvalues, c is a regularization constant, I is the identity matrix, and the identity matrix regularization term cI is superimposed after reconstructing the covariance matrix. The double regularization strategy can alleviate the overfitting problem under high-dimensional data. Through eigenvalue decomposition and truncation technology, the principal component direction information is retained while suppressing noise; Based on the covariance matrix after double regularization, the Bayesian anomaly probability BAP under the double enhanced Bayesian posterior probability is: ;in, Represents x t to the mean μ i The Mahalanobis distance, g i and h i The scaling factor and degree of freedom correction terms are as follows: ; Dynamically calculate the degree of freedom parameter g through trace operation i and h i , solving the traditional chi-square distribution Χ 2 (·) Model bias caused by assuming fixed degrees of freedom; introduce the interaction between the covariance matrix and the inverse covariance matrix into the degree of freedom calculation to expand the application boundary of the chi-square distribution in anomaly detection; Based on RGMM posterior probability π i Weight the abnormal probabilities of different Gaussian components to achieve the probability fusion of multi-condition data, and clarify the contribution of each component to the final score through probability decomposition; To further smooth the fluctuation of BAP and reduce the occurrence of false alarms, a sliding window filtering technique is used with a sliding window size of ω=8. The specific filtering formula is as follows: ;in, is the BAP index after sliding window filtering, BAP t The original BAP index is filtered by sliding window, and the BAP index at the current moment is comprehensively considered with the BAP index at the past ω moments. The maximum value is taken as the filtered result. By combining time series analysis, the stability and accuracy of fault detection are improved, and the influence of instantaneous abnormal fluctuations on the score is suppressed. It is suitable for online detection of non-steady-state data streams.

4. The wind turbine fault diagnosis method based on multi-operating condition data synthesis and enhanced Bayesian abnormal probability according to claim 1 is characterized in that: The specific method of B2) is as follows: B2.1) The random forest classification method based on expert rules is: In the fault classification stage, the knowledge and experience of domain experts are used to integrate the empirical rules of domain experts through the Delphi method to form four types of rule templates; B2.1.1) Amplitude constraints The amplitude constraint rule is used to detect the sudden over-limit situation of the pressure sensor, and the mathematical expression is: ;x i is the current measured value, μ i and σ i are the mean and standard deviation of normal data respectively, k is the threshold coefficient, when the measured value exceeds a certain range, it is judged as a sudden over-limit fault of the pressure sensor; B2.1.2) Association constraints The association constraint rule is used for the pressure association of the inlet and outlet, and the mathematical expression is: ;in, and are the measured values ​​of inlet and outlet pressures, △ th It is the inlet and outlet pressure difference threshold under normal circumstances. When the inlet and outlet pressure difference exceeds the threshold, it is judged as an inlet and outlet pressure difference abnormal fault; B2.1.3) Temporal continuity The timing continuity rule is used to detect the situation where the vibration exceeds the limit continuously. The mathematical expression is: ; is the vibration measurement value at time i, is the vibration threshold, T is the time window size, n is the threshold of the number of vibration exceeding the standard within the time window, and I(·) is the indicator function; when the number of vibration exceeding the standard within a certain time window reaches the threshold, it is judged as a vibration continuous exceeding standard fault; B2.1.4) Combination constraints The combined constraint rule is used for joint judgment of surge and flow. The mathematical expression is relatively complex and multiple parameters are considered comprehensively. ; are multiple measurement parameters related to surge and flow, f(·) is a comprehensive judgment function determined according to expert experience, when the function value is greater than 0, it is judged as a combined surge and flow fault; According to the expert rules, a random forest classifier based on expert rules is designed. For the random forest classifier, there are classification decisions: ; represents the final predicted category of the input sample x, c represents all possible categories, Tk(x) represents the predicted category of the sample x by the kth decision tree, I(·) is the indicator function, which is 1 when Tk(x)=c, otherwise it is 0; It is used to count the number of times the predicted category is c in all trees, and integrate the prediction results of multiple decision trees through the "majority voting" mechanism. The final random forest classification result based on expert rules is: ; and are the prediction results of expert rules and random forest classifier, respectively. It represents a hyperparameter, ranging from 0 to 1. It uses rule-driven feature space segmentation, relies on random forest output under steady-state conditions, and combines expert rules under transient conditions to reduce the false alarm rate under variable load conditions. B2.2 Fault variable location Locating fault variables is crucial for fault repair and equipment maintenance. The calculation formula for the contribution of the first variable to the fault is as follows: Among them, CI j Represents the contribution index of the jth feature, π k is the weight of the kth Gaussian distribution, V k is the eigenvector of the kth Gaussian distribution, x is the eigenvector at the current moment, μ k is the mean of the kth Gaussian distribution, and k is the number of Gaussian distributions. By calculating the feature contribution index, the impact of each feature on the fault is evaluated and the fault location is located.

Citation Information

Patent Citations

  • Key information infrastructure asset identification method combined with mixed random forest

    CN110245693A

  • Fault online diagnosis and fault-tolerant control method for underwater robot propeller

    CN115576184A

  • Synchronous positioning and mapping system and method for dynamic environment

    CN118816844A

  • Lightweight satellite fault diagnosis method based on self-supervised learning

    CN118861710A

  • Multi-working-condition process industrial fault detection and diagnosis method based on deep transfer learning

    WO2023071217A1

Cited By

  • Small fault detection method based on dynamic covariance Markov kernel OCSVM

    CN120892901A

  • Industrial fault detection method and system based on dynamic drift perception and diffusion enhancement

    CN120995184A

  • Communication machine room environment adaptive adjustment management system based on machine learning

    CN121115961A

  • Equipment operation data anomaly detection method and system oriented to extremely cold environment

    CN122065020A

  • A method and system for detecting abnormal equipment operation data in extremely cold environments

    CN122065020B