Wind driven generator fault diagnosis method based on missing SCADA (supervisory control and data acquisition) data and adaptive time-sensitive multi-head attention mechanism

The method addresses data loss and computational challenges in wind turbine fault diagnosis by using an adaptive time-sensitive multi-head attention mechanism to model data loss patterns, enhancing fault prediction accuracy and real-time capability.

CN120316632APending Publication Date: 2025-07-15HARBIN UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510314480.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

When the existing wind turbine fault diagnosis methods face the lack of SCADA data, the error accumulation and timing dependence are insufficient, resulting in a decrease in diagnostic accuracy and high computational complexity, making it difficult to meet the real-time requirements.

Method used

The fault diagnosis method based on missing SCADA data and adaptive time-sensitive multi-head attention mechanism is adopted. By generating missing value mapping matrix and dynamic timing dependency mapping matrix, combining GRU and improved Transformer structure, the data missing mode is directly used to predict faults, reduce interpolation errors, and improve the accuracy and real-timeness of the model in the missing data environment.

Benefits of technology

It effectively reduces the impact of data missing on predicted results, improves the accuracy and real-time nature of fault diagnosis, especially in the case of high missing rates, can still maintain high fault diagnosis accuracy, reduces false alarm rates, and is suitable for large-scale deployment of wind farms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316632A_ABST
    Figure CN120316632A_ABST
Patent Text Reader

Abstract

The invention discloses a wind driven generator fault diagnosis method based on missing SCADA data and a self-adaptive time-sensitive multi-head attention mechanism, and relates to the technical field of wind driven generator fault diagnosis and prediction. In order to solve the technical defects in the prior art that errors and error accumulation are caused by data missing, only data interpolation is concerned, and time sequence dependence modeling is insufficient in the existing fan fault diagnosis work, the technical scheme provided by the invention comprises the following steps: collecting monitoring data in a wind turbine generator SCADA system to form an SCADA data matrix; carrying out numerical value standardization processing on the SCADA data; performing segmentation extraction on the time sequence data, constructing a missing indication matrix, and generating a missing value mapping matrix; extracting time sequence dependence characteristics of the SCADA data, and generating a dynamic time sequence dependence mapping matrix; and generating a fan fault prediction result according to the dynamic time sequence dependence mapping matrix. The method is suitable for intelligent operation and maintenance of a wind power plant, SCADA data analysis, equipment health management, predictive maintenance and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fault diagnosis and prediction of wind turbines. Background Art

[0002] As an important part of renewable energy, wind power plays a key role in the global energy transition. With the expansion of wind power generation scale, the stable operation and efficient maintenance of wind turbines have become the focus of the industry. However, wind turbines are usually deployed in remote geographical locations with complex environmental conditions, making on-line monitoring and fault diagnosis face many challenges. To improve the operation efficiency of wind turbines, wind farms generally adopt a Supervisory Control and Data Acquisition (SCADA) system for real-time status monitoring, and predict equipment faults through data analysis methods to reduce maintenance costs. However, the integrity and reliability of SCADA data have a decisive impact on the accuracy of fault diagnosis.

[0003] 1. Research Status of Existing Technologies

[0004] Currently, the fault diagnosis methods for wind turbines are mainly divided into the following categories:

[0005] Methods Based on Physical Models

[0006] Traditional wind turbine fault diagnosis methods rely on physical models, such as finite element analysis (FEA), vibration signal analysis, etc. These methods require accurate equipment parameters and usually rely on high-quality sensor data. For example, gearbox fault diagnosis uses a feature extraction method based on vibration signals, and analyzes vibration data through Fourier transform or wavelet transform to judge the degree of gear damage. However, the applicability of such methods is limited by data quality and working condition changes. Once the sensor data is missing or noisy, the reliability of the diagnosis results will drop significantly.

[0007] Methods Based on Machine Learning

[0008] In recent years, machine learning (ML) methods have been widely used in wind turbine fault diagnosis. For example, supervised learning models such as support vector machine (SVM), random forest (RF), and extreme gradient boosting (XGBoost) are used to construct fault classifiers. These methods require a large amount of historical fault data for training and have high requirements for data integrity. When there are missing values in the SCADA data, the classification accuracy of these models will be affected, and even misdiagnosis may occur.

[0009] Methods Based on Deep Learning

[0010] Deep learning methods, especially recurrent neural networks (RNNs), long short-term memory networks (LSTMs), gated recurrent units (GRUs), and Transformers, have shown superiority in the time series data analysis of wind turbines. For example:

[0011] There is a prior art that proposes an LSTM-based wind turbine fault prediction method, which uses wind power SCADA data to establish a time series prediction model, improving the recognition ability of complex faults. However, this method requires complete time series data, and the handling of missing data depends on interpolation strategies, which may lead to the propagation of data noise and affect the final diagnosis result.

[0012] There is also a prior art that studies a wind turbine fault detection method based on the Transformer model, which introduces an attention mechanism to improve the fault classification accuracy. However, this method is not optimized for the problem of missing SCADA data, and the prediction accuracy drops significantly when there is a large amount of missing data.

[0013] Data interpolation techniques

[0014] Since data missing may occur during the SCADA data acquisition process, researchers have proposed a series of data interpolation methods, such as:

[0015] K-nearest neighbor (KNN) interpolation: Filling missing data by finding similar sample points. It is suitable for small-scale data sets, but has a high computational complexity and cannot guarantee the rationality of time series data.

[0016] Mean interpolation, linear interpolation, and K-means interpolation: Using statistical methods to fill missing values, but it is difficult to capture the dynamic characteristics of wind turbine time series data, affecting the prediction accuracy.

[0017] Interpolation methods based on GAN (Generative Adversarial Network): Using adversarial training to learn the data distribution and enhancing the authenticity of the interpolated data. However, it has a high computational cost, a complex training process, and great difficulty in practical applications.

[0018] 2. Limitations of the prior art

[0019] Although the above methods have improved the accuracy of wind turbine fault diagnosis to a certain extent, there are still the following technical problems:

[0020] Error accumulation caused by data missing

[0021] Existing methods usually adopt the strategy of "interpolate first, then predict", but the interpolation method itself may introduce errors, which will accumulate continuously in the subsequent prediction process and affect the final diagnosis accuracy.

[0022] The data missing pattern is not fully utilized

[0023] Traditional methods only focus on data imputation while ignoring the information contained in data missing itself. In fact, the missing of certain specific patterns (such as the long-term missing of certain sensors) may be an early sign of equipment abnormality, but the existing methods fail to effectively utilize this information for fault identification.

[0024] Insufficient modeling of temporal dependence

[0025] Existing methods (such as SVM, RF) often ignore the temporal characteristics of SCADA data. Although some deep learning methods (such as LSTM) consider temporal dependence, they fail to make adaptive adjustments for different time scales, resulting in a decline in model performance during long-term prediction.

[0026] Computational complexity and real-time issues

[0027] Traditional neural network methods often require large-scale computing resources and do not fully consider real-time requirements, making it difficult to meet the needs of online fault warning for wind turbines. Summary of the invention

[0028] To solve the technical defects existing in the prior art, namely, in the existing wind turbine fault diagnosis work, the errors and error propagation caused by data missing, only focusing on data imputation and insufficient modeling of temporal dependence, the technical solution provided by the present invention is as follows:

[0029] A wind turbine fault diagnosis method based on missing SCADA data and an adaptive time-sensitive multi-head attention mechanism, comprising:

[0030] The step of collecting monitoring data in the SCADA system of the wind turbine to form a SCADA data matrix;

[0031] The step of performing numerical standardization processing on the SCADA data;

[0032] The step of segmenting and extracting the time series data, constructing a missing indication matrix, and generating a missing value mapping matrix;

[0033] The step of extracting the temporal dependence features of the SCADA data and generating a dynamic temporal dependence mapping matrix;

[0034] The step of generating a wind turbine fault prediction result according to the dynamic temporal dependence mapping matrix.

[0035] Furthermore, a preferred implementation manner is provided. The measurement parameters of the sensors of each important component of the unit are recorded every 10 minutes, the key operating variables are screened out, and a SCADA data matrix is formed.

[0036] Further, a preferred embodiment is provided, which uses an outlier detection algorithm to identify and remove abnormal data, and uses a normalization method to perform numerical standardization processing on SCADA data.

[0037] Further, a preferred embodiment is provided, which uses a sliding time window method to segment and extract time series data, constructs a missing indication matrix, and generates a missing value mapping matrix in combination with missing time step information.

[0038] Further, a preferred embodiment is provided, which uses a gated recurrent unit to extract the temporal dependence features of SCADA data, calculates an adaptive adjustment factor, and combines the adaptive adjustment matrix with the missing value mapping matrix to generate a dynamic temporal dependence mapping matrix.

[0039] Further, a preferred embodiment is provided, which calculates attention weights through an improved Transformer structure, and performs fault diagnosis in combination with the dynamic temporal dependence mapping matrix to generate a wind turbine fault prediction result.

[0040] A wind turbine fault diagnosis device based on missing SCADA data and an adaptive time-sensitive multi-head attention mechanism is also provided, including:

[0041] A module for collecting monitoring data in the SCADA system of a wind turbine to form a SCADA data matrix;

[0042] A module for performing numerical standardization processing on SCADA data;

[0043] A module for segmenting and extracting time series data, constructing a missing indication matrix, and generating a missing value mapping matrix;

[0044] A module for extracting the temporal dependence features of SCADA data and generating a dynamic temporal dependence mapping matrix;

[0045] A module for generating a wind turbine fault prediction result according to the dynamic temporal dependence mapping matrix.

[0046] A computer storage medium is also provided for storing a computing program, and when the computer program is read by a computer, the computer executes the method.

[0047] A computer is also provided, including a processor and a storage medium, and when the processor reads the computer program stored in the storage medium, the computer executes the method.

[0048] A computer program product is also provided, which, as a computer program, implements the method when the computer program is executed.

[0049] Compared with the prior art, the beneficial effects of the technical solution provided by the present invention are as follows:

[0050] In this solution, by introducing a missing value mapping matrix to directly encode the data missing pattern, the model can distinguish real observed values from missing values, rather than simply relying on data imputation. Compared with traditional methods such as KNN imputation and mean imputation, this method can reduce the accumulation of imputation errors, improve the prediction accuracy, and still maintain a high fault diagnosis accuracy in the case of a high missing rate.

[0051] The sliding time window method is used to segment and extract the time-step data, and combined with the GRU module to learn the temporal features, enabling the model to capture short-term and long-term fault patterns. Compared with traditional LSTM or RNN methods, this method can improve the performance of the model in complex temporal data by parallel computing the dynamic dependencies at different time scales, and reduce the vanishing gradient phenomenon caused by long-term dependence problems.

[0052] Using the adaptive time-sensitive multi-head attention mechanism, combined with the dynamic temporal dependence mapping matrix, can adjust the attention weights according to the distance between different time steps, enabling the model to accurately extract key information even in the case of data missing or uneven distribution of observation points. Compared with the traditional Transformer's way of directly modeling time series, this method still has a high fault diagnosis accuracy in an environment with severe data missing, and effectively reduces the false alarm rate.

[0053] The isolation forest algorithm is used for outlier detection, so that the data is denoised before entering the model, avoiding the interference of abnormal data on model learning. Compared with the MAD method or IQR method that only removes outliers based on statistical distribution, the isolation forest can more flexibly adapt to the complex distribution of SCADA data, improve the effectiveness of data cleaning, and ensure the quality of input data.

[0054] This solution uses GRU to generate an adaptive adjustment factor to dynamically adjust the attention distribution, enabling the model to optimize the weight calculation of key variables according to the feature distribution of different fault types. Compared with the LSTM with fixed weight allocation or the global attention mechanism, this method can improve the sensitivity to faults of different components of the fan, making the model more accurate in dealing with different types of faults, especially in early fault warning.

[0055] Combined with the improved Transformer structure, the model can maintain a high computational efficiency when processing large-scale SCADA data. Compared with the combined method of GAN data generation and completion + LSTM prediction, this solution reduces the overhead of data preprocessing and secondary calculation, improves the online real-time prediction ability, and is more suitable for large-scale deployment environments in wind farms.

[0056] It is applicable to work such as intelligent operation and maintenance of wind farms, SCADA data analysis, equipment health management, and predictive maintenance. Brief Description of the Drawings

[0057] Figure 1 is a flowchart of the method;

[0058] Figure 2 are 59 key operating variables in the SCADA dataset;

[0059] Figure 3 is a schematic diagram of missing information extraction and indication matrix generation;

[0060] Figure 4 is a schematic diagram of learning temporal features based on GRU and generating a dynamic temporal dependence mapping matrix;

[0061] Figure 5 is a schematic diagram of the overall model structure of the proposed adaptive time-sensitive multi-head attention mechanism. Detailed Embodiments

[0062] To make the advantages and beneficial effects of the technical solution provided by the present invention more clearly manifested, the technical solution provided by the present invention will now be further described in detail with reference to the accompanying drawings. Specifically:

[0063] Embodiment 1. This embodiment provides a wind turbine fault diagnosis method based on missing SCADA data and an adaptive time-sensitive multi-head attention mechanism, including:

[0064] The step of collecting monitoring data in the SCADA system of the wind turbine to form a SCADA data matrix;

[0065] The step of performing numerical standardization processing on the SCADA data;

[0066] The step of segmenting and extracting the time series data, constructing a missing indication matrix, and generating a missing value mapping matrix;

[0067] The step of extracting the temporal dependence features of the SCADA data and generating a dynamic temporal dependence mapping matrix;

[0068] The step of generating a wind turbine fault prediction result according to the dynamic temporal dependence mapping matrix.

[0069] Record the measurement parameters of the sensors of each important component of the unit every 10 minutes, screen out the key operating variables, and form a SCADA data matrix.

[0070] Use an outlier detection algorithm to identify and eliminate abnormal data, and use a normalization method to perform numerical standardization processing on the SCADA data.

[0071] The sliding time window method is used to segment and extract time series data, construct a missing indication matrix, and combine the missing time step information to generate a missing value mapping matrix.

[0072] The gated recurrent unit is used to extract the temporal dependence features of SCADA data, and an adaptive adjustment factor is calculated. The adaptive adjustment matrix is combined with the missing value mapping matrix to generate a dynamic temporal dependence mapping matrix.

[0073] The attention weights are calculated through an improved Transformer structure, and fault diagnosis is performed in combination with the dynamic temporal dependence mapping matrix to generate the wind turbine fault prediction result.

[0074] Embodiment 2: This embodiment is a further explanatory description of the technical solution provided in Embodiment 1. Specifically: A wind turbine fault diagnosis method based on missing SCADA data and an adaptive time-sensitive multi-head attention mechanism. Its core technical solution includes five steps: data preprocessing, missing value modeling, temporal feature extraction, adaptive attention calculation, and fault diagnosis. The input and output logic of each step is clear, which can ensure a complete fault diagnosis process and improve the health status prediction ability of the wind turbine. The following will describe the specific implementation steps in detail.

[0075] Step 1: SCADA data collection and key feature extraction

[0076] The goal of this step is to obtain the wind turbine operation data and screen out the key operation variables to provide a data basis for subsequent fault diagnosis.

[0077] Obtain the monitoring data from the wind turbine SCADA system, record the measurement values of the sensors of important components every 10 minutes, and screen out the key variables that can comprehensively reflect the real-time health status of the wind turbine.

[0078] Detailed description:

[0079] Collect the operation data of each key component of the wind turbine, including operation variables such as wind wheel speed, wind direction angle, temperature, cable torsion position, and power.

[0080] After data screening, 59 important measurement parameters are selected, including but not limited to:

[0081] Wind wheel speed, blade drive current, temperature of the IGBT power module of the converter

[0082] Cabinet temperature, ambient temperature, generator A / B / C phase current, gearbox oil temperature

[0083] Pitch system data, gearbox cooling system data, vibration sensor data, etc.

[0084] These variables can truly reflect the operating status of the fan and provide basic data for subsequent fault diagnosis.

[0085] Output:

[0086] Generate a structured SCADA dataset containing 59 key operating variables to form a complete data input matrix.

[0087] Step 2: Data cleaning and outlier handling

[0088] The goal of this step is to remove abnormal data and normalize the SCADA data to ensure the quality and consistency of data input.

[0089] Perform outlier detection and denoising on the SCADA dataset. Use the Isolation Forest algorithm to detect abnormal data and the Z-score normalization method to normalize the data.

[0090] Detailed description:

[0091] Outlier detection:

[0092] Use the Isolation Forest algorithm to identify and remove abnormal data. This method can automatically identify isolated data points based on the distribution characteristics of the data, avoiding the influence of outliers on the training effect of the model.

[0093] Data normalization:

[0094] Since the units and magnitudes of SCADA data variables are different, the Z-score normalization method is used for normalization to make the data distribution tend to zero mean and unit variance, ensuring that all variables are within a similar numerical range.

[0095] Output:

[0096] The SCADA data matrix after cleaning and normalization has removed outliers and ensured the numerical consistency of variables.

[0097] Step 3: Missing value modeling and missing pattern learning

[0098] The goal of this step is to model the missing situations in the SCADA data and construct a missing value mapping matrix so that the data missing characteristics can be directly utilized in subsequent analysis instead of simple imputation.

[0099] Extract time series data in segments through the sliding time window method and construct a missing indicator matrix. Combine the missing time step information to generate a missing value mapping matrix for the model to distinguish between real observed values and missing values.

[0100] Detailed description:

[0101] Using the Sliding Time Window method, the SCADA time series data is segmented into multiple subsequences, each subsequence containing multiple time steps.

[0102] Calculate the missing indication matrix:

[0103] Record the number of time steps elapsed since the last observation for each variable and mark the data missing situation.

[0104] Generate a missing pattern matrix to reflect the missing degree and distribution of each variable.

[0105] Construct a missing value mapping matrix:

[0106] By combining the missing time information and the indication matrix, a missing value mapping matrix is generated, enabling the model to adaptively adjust the attention calculation and reduce the impact of missing data on the prediction result.

[0107] Output:

[0108] Generate a missing value mapping matrix D for dynamically adjusting the attention calculation weights during subsequent time series feature learning.

[0109] Step 4: Time Series Feature Extraction and Dynamic Weight Adjustment

[0110] The goal of this step is to utilize GRU (Gated Recurrent Unit) to extract the time series dependence features of SCADA data and calculate the dynamic time series dependence mapping values in combination with the missing value mapping matrix.

[0111] Adopt a GRU model to extract time series features, calculate an adaptive adjustment factor, and generate a dynamic time series dependence mapping matrix in combination with the missing value mapping matrix for adjusting the attention calculation weights.

[0112] Detailed description:

[0113] Extract the time series dependence features of SCADA data through the GRU module, calculate the hidden state at each time step, enabling the model to understand long-term and short-term information.

[0114] Calculate the adaptive adjustment factor:

[0115] Combining the hidden state information generated by GRU, calculate the dynamic adjustment factor, enabling the model to adaptively adjust the attention distribution between different time steps.

[0116] Calculate the dynamic time series dependence mapping values:

[0117] By combining the missing value mapping matrix and the adjustment factor, the model can dynamically adjust the weights between time steps according to the data missing situation, thereby reducing the impact of missing data and improving the time series feature extraction ability.

[0118] Output:

[0119] Generate a dynamic time - series dependency mapping matrix P for subsequent attention calculation.

[0120] Step 5: Fault diagnosis based on the adaptive time - sensitive multi - head attention mechanism

[0121] The goal of this step is to use the improved Transformer structure for fault diagnosis and improve the model's prediction ability in the missing data environment through adaptive attention weight calculation.

[0122] Calculate the attention weights based on the dynamic time - series dependency mapping matrix and combine the multi - head attention mechanism for wind turbine fault diagnosis.

[0123] Detailed description:

[0124] Calculate the attention weights:

[0125] Combine the dynamic time - series dependency mapping matrix and the multi - head attention mechanism to calculate the attention weights of the wind turbine operating state and assign adaptive weighting coefficients to different time steps.

[0126] Normalization processing:

[0127] Normalize the calculated attention weights to ensure a reasonable distribution of the weights and optimize the stability of attention calculation.

[0128] Calculate the final output:

[0129] Combine the normalized attention weights to calculate the final wind turbine fault prediction result and classify the possible fault types.

[0130] Final output:

[0131] Predict the fault categories of the wind turbine and improve the stability and accuracy of the model in the missing data environment.

[0132] Embodiment 3. Combination Figures 1-5 Describe this embodiment. This embodiment further describes the above - provided technical solution in detail through specific embodiments. Specifically:

[0133] This embodiment provides a wind turbine fault diagnosis method based on missing SCADA data and the adaptive time - sensitive multi - head attention mechanism, including the following steps:

[0134] Step 1: Extract key operating variables from the wind farm SCADA system that can comprehensively reflect the real - time health status of the wind turbine, including wind wheel speed, wind direction angle, temperature, cable - twisting position, power, etc. These operating data truly reflect the actual operating state of the wind turbine.

[0135] Step 2: Determine five representative types of faults, covering typical and diverse operating states.

[0136] Step 3: Perform outlier and normalization processing on the data in the SCADA system.

[0137] Step 4: Use the method of a sliding time window to segment and extract the time-step data. After extracting the subsequences, calculate the time steps for each subsequence to generate a temporary matrix. Missing indicator matrix Δ ij By combining the number of time steps in the temporary matrix and the missing indicator information, it comprehensively reflects the missing patterns and degrees of each feature in the dataset. Based on the indicator matrix, a missing value mapping matrix D is constructed, which can effectively distinguish real observed values and missing values.

[0138] Step 5: Use multiple parallel gated recurrent unit (GRU) modules to learn the temporal features of the input data and generate an adaptive adjustment factor λ ij , and combine the adaptive adjustment matrix with the missing value mapping matrix D to generate a dynamic temporal dependence mapping matrix.

[0139] The key feature variables in the said Step 1 are 59, including: time, wind turbine speed, generator F-side temperature, generator R-side temperature, generator phase A current, generator phase B current, generator phase C current, gearbox cooling water temperature, gearbox oil temperature, gearbox inlet oil temperature, gearbox high-speed shaft intermediate temperature, blade 1 motor encoder angle, blade 2 motor encoder angle, blade 3 motor encoder angle, cable twisting position, nacelle position, active power, reactive power, nacelle cabinet temperature, outdoor temperature, instantaneous wind speed, wind direction angle, 30-second average wind speed, 10-minute sliding average wind speed, 60-second average wind direction angle, grid phase A voltage, grid phase B voltage, grid phase C voltage, grid phase A current, grid phase B current, grid phase C current, generator cooler internal air path inlet temperature, generator cooler external air path outlet temperature, main bearing front bearing temperature, converter 1 grid-side IGBT power module maximum temperature, gearbox high-speed shaft generator-side temperature, V1-phase generator temperature, converter 1 machine-side IGBT power module maximum temperature, real-time values of built-in left and right vibration sensors, ambient temperature, gearbox high-speed shaft wind turbine-side temperature, blade 2 driver current, U1-phase generator temperature, generator cooler non-drive-end internal air path inlet temperature, gearbox inlet oil pressure, real-time values of external front and rear vibration sensors, nacelle temperature, generator speed, blade 3 driver current, hydraulic station system pressure, W1-phase generator temperature, generator cooler drive-end internal air path outlet temperature, gearbox oil pump outlet oil pressure, real-time values of built-in front and rear vibration sensors, 30s average outdoor temperature, gearbox intermediate shaft wind turbine-side temperature, blade 1 driver current, main bearing rear bearing temperature, generator cooler drive-end internal air path inlet temperature.

[0140] The five failure modes in step 2 include pitch failure, generator failure, converter failure, gearbox failure, and hydraulic system failure. Among them, the occurrence frequency of pitch failure is 40.1% (41.9% for pitch safety chain disconnection, 23.7% for pitch axis 1 driver failure, 17.7% for pitch axis 2 driver failure, and pitch axis 3 driver failure), the occurrence frequency of generator failure is 27.7% (93.5% for carbon brush failure at the generator shaft extension end, 6.5% for generator cooling fan protection), the occurrence frequency of gearbox failure is 14.3% (34.9% for abnormal gearbox cooling water pressure, 25.2% for gearbox lubricating oil pump motor protection, 21.3% for gearbox oil heater protection, 18.6% for high gearbox lubricating oil outlet pressure), the occurrence frequency of converter failure is 11.6% (79.3% for converter ISU smoke alarm, 20.7% for converter machine side phase A overcurrent), and the occurrence frequency of hydraulic system failure is 6.7% (60.9% for too low main pressure of the hydraulic system, 39.1% for hydraulic oil pump motor protection).

[0141] Step 3 is specifically as follows:

[0142] First, the Isolation Forest algorithm is used to detect and remove outliers in the data. However, due to the differences in units and magnitudes of different sensor channels, SCADA variables with smaller magnitudes often fail to play their due roles. Therefore, in this embodiment, the Z-score normalization method is adopted to adjust the data distribution of each variable to zero mean and unit variance, so that all variables are within a similar numerical range. The formula is as follows.

[0143]

[0144] Among them, is the original SCADA data variable, N is the number of sampling moments, I is the number of variables, and x and σ are its mean and standard deviation respectively.

[0145] Step 4 is specifically as follows:

[0146] The temporary matrix is an m×j matrix, and each element in it records the number of time steps elapsed since the corresponding feature was last observed with a valid value. The variable e is used to indicate the missing situation of data points, where e = 0 indicates that the data point is not missing, and e = 1 indicates that the data point is missing. There is a column vector a of m×1 on the right side of the temporary matrix, and each element a f corresponds to the f-th row of the temporary matrix. The assignment is carried out sequentially from bottom to top, starting from 0. The calculation method is to select the rows in the temporary matrix where at least one e fj = 0, and define the set:

[0147]

[0148] For each i = 1, 2, …, |S - 1| and each j, the indicator matrix is defined as:

[0149]

[0150] where Δ is a matrix with the same dimension as the input matrix, v is the original data, i ≤ m, and f = S(i).

[0151] Based on the indicator matrix Δ, a missing value mapping matrix D is constructed, and the calculation process is as follows:

[0152] d ij = exp(-ReLU(Δ ij ))

[0153] where ReLU(x) = max(0, x), and d ij as the missing value mapping value enables the model to dynamically adjust the attention weights according to the distance between time steps, effectively distinguishing real observations from missing values.

[0154] The specific content of step 5 is as follows:

[0155] First, the input sequence passes through the GRU module to generate a hidden state:

[0156] h m = GRU(x m , h m-1 )

[0157] where x m is the input at the m-th time step, and h m is the corresponding hidden state.

[0158] Then, the adaptive adjustment factor is calculated using the hidden state, and the formula is as follows:

[0159] λ ij = σ(W λ h m + b λ )

[0160] where W λ and b λ are learnable parameters, σ(·) is the activation function of Sigmoid, ensuring that λ i is within the range of (0, 1).

[0161] The dynamic temporal dependence mapping value is defined as:

[0162] p ij = exp(-λ ij ·d ij )

[0163] After adjustment based on the dynamic time-series dependence mapping matrix, the calculation formula for the attention weight is as follows:

[0164]

[0165] After obtaining the weights, perform normalization processing:

[0166]

[0167] The calculation formula for the final output is as follows, and the fan fault diagnosis result is output:

[0168]

[0169] The missing value mapping matrix in step 4 and the dynamic time-series dependence mapping in step 5 act together on the Transformer structure. The subsequence processed by the sliding window is input into the multi-head attention mechanism of the improved Transformer, enabling the model to exhibit better fault prediction performance and lower false alarm rate in scenarios of dealing with missing data and complex operating environments, being able to capture key fault features more effectively and improve the diagnostic accuracy.

[0170] Advantages of this embodiment:

[0171] (1) It can directly utilize the missing features and model the missing patterns, avoiding relying on additional filling steps, thus more accurately capturing the key trends and operating states and generating more practical prediction results.

[0172] (2) It can flexibly capture the time-series dependence at different time scales, thus effectively improving the accuracy and reliability of the fan fault early warning.

[0173] (3) Based on the real SCADA dataset, a comprehensive verification and comparative analysis of the model of this embodiment was carried out, and the results show that this model performs excellently in the comprehensive performance of various fault types.

[0174] The above further describes the technical solutions provided by the present invention through several specific embodiments to highlight the advantages and beneficial effects of the technical solutions provided by the present invention. However, the above-mentioned several specific embodiments are not used as limitations on the present invention. Any reasonable modifications and improvements, combinations of embodiments, and equivalent replacements based on the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A fault diagnosis method for wind turbines based on missing SCADA data and an adaptive time-sensitive multi-head attention mechanism, characterized in that, Including: The step of collecting monitoring data in the SCADA system of a wind turbine to form a SCADA data matrix; The step of performing numerical standardization processing on the SCADA data; The step of segmenting and extracting time series data, constructing a missing indication matrix, and generating a missing value mapping matrix; The step of extracting the time series dependence features of the SCADA data and generating a dynamic time series dependence mapping matrix; The step of generating a wind turbine fault prediction result according to the dynamic time series dependence mapping matrix.

2. The fault diagnosis method of a wind turbine based on missing SCADA data and an adaptive time-sensitive multi-head attention mechanism according to claim 1, wherein, Record the measurement parameters of the sensors of each important component of the unit every 10 minutes, screen out the key operating variables, and form a SCADA data matrix.

3. A fault diagnosis method for a wind turbine based on missing SCADA data and an adaptive time-sensitive multi-head attention mechanism according to claim 1, characterized in that, Use an outlier detection algorithm to identify and remove abnormal data, and use a normalization method to perform numerical standardization processing on the SCADA data.

4. A fault diagnosis method for a wind turbine based on missing SCADA data and an adaptive time-sensitive multi-head attention mechanism according to claim 1, characterized in that, Adopt a sliding time window method to segment and extract time series data, construct a missing indication matrix, and combine the missing time step information to generate a missing value mapping matrix.

5. A fault diagnosis method for a wind turbine based on missing SCADA data and an adaptive time-sensitive multi-head attention mechanism according to claim 1, characterized in that, Use a gated recurrent unit to extract the time series dependence features of the SCADA data, calculate an adaptive adjustment factor, combine the adaptive adjustment matrix with the missing value mapping matrix, and generate a dynamic time series dependence mapping matrix.

6. A fault diagnosis method for a wind turbine based on missing SCADA data and an adaptive time-sensitive multi-head attention mechanism according to claim 1, characterized in that, Calculate the attention weights through an improved Transformer structure, and perform fault diagnosis in combination with the dynamic time series dependence mapping matrix to generate a wind turbine fault prediction result.

7. A wind turbine fault diagnosis device based on missing SCADA data and an adaptive time-sensitive multi-head attention mechanism, characterized in that, Including: A module for collecting monitoring data in the SCADA system of a wind turbine to form a SCADA data matrix; A module for performing numerical standardization processing on the SCADA data; A module for segmenting and extracting time series data, constructing a missing indication matrix, and generating a missing value mapping matrix; A module for extracting the time series dependence features of the SCADA data and generating a dynamic time series dependence mapping matrix; A module for generating a wind turbine fault prediction result according to the dynamic time series dependence mapping matrix.

8. A computer storage medium for storing a computing program, characterized in that, When the computer program is read by a computer, the computer executes the method described in claim 1.

9. A computer, comprising a processor and a storage medium, characterized in that, When the processor reads the computer program stored in the storage medium, the computer executes the method described in claim 1.

10. A computer program product, as a computer program, characterized in that, When the computer program is executed, the method described in claim 1 is implemented.

Citation Information

Cited By

  • Wind turbine fault diagnosis method and system based on mechanism data fusion

    CN120974282A