Early fault detection method for doubly-fed induction generator of wind power system based on multi-scale latent variable regression
By applying a multi-scale latent variable regression method in DFIG fault detection, combining mutual information technology and multivariate variational modal decomposition, the problem of early DFIG fault diagnosis in the existing technology is solved, and efficient fault detection is achieved in multiple fault types and complex coupling scenarios.
Patent Information
- Application Number
- CN202510294347.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to accurately diagnose early failures of double-feed induction generators (DFIG) in multiple fault types and complex coupling effect scenarios, and the fault detection method based on SCADA data has problems such as overfitting and sensitivity to extreme data.
The fault detection method based on multi-scale latent variable regression (MSLVR) is adopted to analyze the coupling between DFIG and the system through mutual information technology, and combine multivariate modal decomposition (MVMD) and serial structure latent variable regression model to gradually eliminate interference from different time scales, and build Hotelling's T2 monitoring statistics to achieve early fault detection.
This method can keenly and robustly capture early abnormalities caused by DFIG in the context of multi-time scale data, improves the robustness and accuracy of fault detection, and is suitable for comprehensive detection of multiple faults.
Smart Images

Figure CN120216973A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a method for diagnosing early faults of a doubly-fed induction generator (DFIG) in scenarios of multiple fault types and complex coupling effects, belonging to the technical field of fault detection. Background Art
[0002] As a clean and mature power generation method, wind power generation technology has been widely applied and promoted in the energy field. According to statistics, the installed capacity of wind turbines (WT) globally in 2023 increased by 50% compared to the previous year. With a large number of WTs put into operation, the failure rate of the equipment has also increased accordingly. Among them, doubly-fed wind turbines, as the mainstream model, are widely deployed in onshore and offshore wind farms and account for a large proportion. As the core of energy conversion in the entire wind power system, the operating state of the doubly-fed induction machine (DFIG) has an important impact on the performance of the unit. The relatively high proportion of fault downtime and the high cost of equipment make it particularly important to detect DFIG faults promptly and accurately.
[0003] In the current field of DFIG fault detection, many experts and scholars have carried out relevant technical research based on current and vibration signals and achieved certain results. Such signals have high sampling accuracy, enabling fault characteristics to be detected promptly and accurately, but they require additional deployment of expensive sensor systems, which is not conducive to the economic efficiency of wind turbine operation and maintenance. Since temperature signals are included in the data collected by the basic monitoring, control, and data acquisition system (SCADA) of the unit, and a large amount of historical data has been accumulated over the years, DFIG fault detection based on SCADA data temperature signals and combined with data-driven technology has gradually become the research focus in the field.
[0004] In the research on DFIG fault detection based on SCADA data, constructing a normal behavior model (NBM) using artificial intelligence has become the mainstream FD technology. Common methods include using improved machine learning algorithms and neural networks, etc. Such models often require a large amount of high-quality data for training and are mostly "black box models" with poor interpretability and complex structures. The prediction target often corresponds to the component temperature characteristics of a single DFIG fault, which is not conducive to the comprehensive detection of multiple faults. Although multi-target monitoring schemes for WTs have been developed, they are essentially a combination of single targets and lack comprehensive capabilities in the face of equipment-level FD tasks.
[0005] In the comprehensive monitoring research of WT key equipment, common methods include using unsupervised networks or prediction models such as non - linear state estimation and autoencoders (AE). These methods conduct comprehensive fault detection of equipment by predicting multiple relevant features of the equipment and combining the prediction error square (SPE) statistic. However, such methods, like supervised normal behavior models, have the risk of overfitting, and the SPE statistic is constructed based on Euclidean distance, which is easily affected by a certain extreme data during the calculation process. Therefore, some studies propose a fault detection method based on Mahalanobis distance (MD). Although the MD method overcomes the deficiencies of traditional normal behavior models to a certain extent, it depends on the covariance matrix, resulting in the phenomenon of matrix non - invertibility easily occurring in high - dimensional data or in the presence of outliers.
[0006] To address the above problems, multivariate statistical process control (MSPC) technology provides a good solution. MSPC technology features stable and efficient calculation, clear models, and the ability to uniformly monitor multi - dimensional data. Among them, principal component analysis (PCA) and its variants have been successfully applied to WT fault detection. However, the PCA model is in an unsupervised form, unable to effectively consider the influence of the fan operation mechanism and working conditions on the temperature signal, and not applicable to collinear data. To solve these problems, some studies use the generalized additive model (GAM) to eliminate the influence of covariates (such as wind speed, rotational speed, power, ambient temperature) on the fan temperature signal and combine the statistics in MSPC technology to monitor the fan state. However, this method mainly focuses on the overall fan, with less discussion on external factors affecting DFIG, and GAM is limited by the additive structure, which is not conducive to the expansion of complex mechanisms.
[0007] As a complex system with highly coupled sub - components, the interior of the nacelle of a wind turbine is a closed space, and the generator temperature is inevitably affected by linked sub - components, making it difficult to directly characterize faults. This influence exists in a form that cannot be directly observed, namely the form of latent variables. Therefore, eliminating such influence helps to capture early anomalies caused by faults from the generator temperature signal. And the latent variable regression model (LVR) in supervised MSPC technology provides good inspiration. This method extracts the potential influence between two sets of multi - variables by minimizing the projection error and has strong scalability. Compared with PCA, LVR can establish the connection between the equipment and the system, making the monitoring more targeted. However, there has been no study applying LVR to WT state monitoring based on SCADA data.
[0008] In addition, due to the constraints of the harsh operating conditions of the wind turbine and data storage conditions, SCADA data usually exhibits the characteristics of being dynamic, non-stationary, and having a low sampling rate (such as a 10-minute sampling interval), which poses higher requirements for the robustness of detection. By extending the recorded time-series data to multiple time scales, hidden dynamic patterns can be revealed and non-stationarity can be improved. However, current extended research on LVR has focused more on dynamic and non-linear extensions, and there has been no development suitable for multi-scale data. If latent variables at different time scales can be fully extracted, it will be more helpful to eliminate the impact of system operation on DFIG. Summary of the Invention
[0009] Based on the above deficiencies and inspirations, the present application relates to a fault detection method based on multi-scale latent variable regression (MSLVR). This method can fully eliminate the interference of system latent variables on the DFIG temperature signal in the context of multi-time scale data, and develops monitoring indicators for comprehensive early detection of various faults of DFIG.
[0010] The technical solution adopted in the present application is as follows:
[0011] An early fault detection method for a doubly-fed induction generator in a wind power system based on multi-scale latent variable regression, which includes the following steps:
[0012] a. Extract the operation data with a 10-minute sampling interval within five years from the wind turbine SCADA system. For the temperature characteristics of the generator, use the mutual information technology to calculate the coupled variable matrix related to it, and preferably select the coupled variables as the input;
[0013] b. Process the coupled variable matrix using multivariate variational mode decomposition (MVMD), decompose the variables into low-frequency, medium-frequency, and high-frequency modal components, and form a multi-scale data set with physical meanings. Among them, the low-frequency component reflects the long-term trend, the medium-frequency component represents the periodic change, and the high-frequency component reflects the short-term fluctuation;
[0014] c. Use a serial structure latent variable regression (LVR) model to sequentially extract the potential influence of each modal component on the temperature signal, and gradually eliminate the interference of this modal until a residual subspace that eliminates all time-scale interferences of the coupled variables is obtained;
[0015] d. Orthogonally decompose the residual subspace in combination with principal component analysis (PCA) to remove noise and correlation, and construct Hotelling's T 2 monitoring statistic;
[0016] e. Use the kernel density estimation method to determine the control limit of the monitoring statistic, with a 0.01 confidence level as the standard;
[0017] f. In actual operation, when the real-time temperature signal of the generator exceeds the control limit, a fault alarm is output.
[0018] The most significant advantage of the present invention is:
[0019] (1) Considering the influence of the WT system, the information measurement technology is adopted to analyze the coupling situation between the DFIG and the system, and the supervised MSPC technology is introduced to obtain the potential relationship, aiming to sensitively and robustly capture the early anomalies caused by faults in the DFIG.
[0020] (2) Develop the MSLVR model, which combines the multivariate variational mode decomposition (MVMD) to expand the time scale and enhance the data quality. And it has a serial structure, which can gradually extract and eliminate the interference of different modal components to improve the detection robustness.
[0021] (3) To comprehensively detect different types of faults of the DFIG, for the DFIG temperature residual subspace after eliminating the influence of latent variables, the monitoring statistic is developed by combining PCA. Description of the Drawings
[0022] Figure 1 It is the result graph of the mutual information analysis between the temperature signal and the system variables;
[0023] Figure 2 It is the principle block diagram of the multi-scale latent variable regression model (MSLVR);
[0024] Figure 3 It is the fault detection flow chart based on MSLVR;
[0025] Figure 4 The detection result graph of the bearing fault in Case 1;
[0026] Figure 5 It is the detection result graph of the stator winding turn-to-turn short circuit fault in Case 2;
[0027] Figure 6 It is the detection result graph of the rotor slip ring wear fault in Case 3;
[0028] Figure 7 It is the confusion matrix graph for comprehensively analyzing the calculation results of the 3 cases; Detailed Implementation Manner
[0029] The specific implementation steps of the solution of the present invention include:
[0030] 1. Problem Description and Variable Analysis
[0031] This application analyzes the SCADA data recorded by the 20th wind turbine in a wind farm in northern China in 2022. The generator type in this wind farm is a 1.5MW doubly-fed wind turbine. The data sampling interval is 10 minutes, and the wind speed, power, rotation speed of 50 wind turbines, and the temperature information of internal equipment such as the main shaft, gearbox, and generator are recorded for a total of 36 features in the five years from 2018 to 2022. In order to comprehensively detect generator faults as much as possible, the temperature signals related to the generator (drive-end, non-drive-end bearings, three-phase windings) are used to form the monitoring matrix Y = {y1~y5}. For the remaining variables X * ={x1~x 31}, the mutual information (MI) technique is used to calculate the comprehensive MI coefficients with Y respectively. The MI technique is a method in information theory for measuring the degree of information dependence between two variables, defined as
[0032]
[0033] where p(x,y) is the joint probability distribution of X * and Y, and p(x) and p(y) are their respective marginal probability distributions.
[0034] The coefficient is between {0,1}. When the coefficient is 0, it means that the variables are independent of each other. When it is not 0, it means that there is information interaction. The closer it is to 1, the greater the amount of interactive information. Then the comprehensive MI coefficient for Y can be defined as the average of the mutual information of the features for all target variables Y j , expressed as
[0035]
[0036] where m = 5 is the dimension of Y.
[0037] The calculation results are as Figure 1 shown. It is not difficult to find that the variables with high information interaction with the generator temperature are all temperature variables of the coupling components, indicating that the generator temperature is mainly affected by the information of these variables. Although from the perspective of the wind turbine operation mechanism, signals such as wind speed and power will affect the generator temperature, according to the information amount results, the generator temperature is significantly more affected by the coupling components (the variable coefficients of wind direction, oil pressure, etc. are all less than 0.2, so they are not shown). Therefore, the temperature signals of the coupling components are used to form the observed variable matrix X = {x1~x7}, and further, the method proposed in this invention is used to eliminate the influence of the latent variables generated by the system coupling mechanism.
[0038] 2. Proposed detection method
[0039] (1) Fault detection principle based on the latent variable regression model
[0040] LVR is a newly proposed supervised MSPC technique, similar to the classical latent variable models PLS and CCA. Suppose there is input data and target data where n is the number of samples, b and m are the number of features of the corresponding data respectively. LVR also maps high-dimensional data to a low-dimensional space through linear projection to construct latent variables in the data. The mathematical formula can be expressed as t i = Xw i , u i = Yq i , i = 1, 2,..., l. Where t i and u i are latent variables, w i and q i are the corresponding weight vectors respectively. The weight matrices formed are W and Q, and l is the number of latent variables extracted. However, different from the former two, the LVR model selects latent variables and weight vectors by minimizing the error between latent variables. The objective functions of its internal and external models are expressed as:
[0041]
[0042] where b is the regression coefficient.
[0043] It is not difficult to find that LVR has the characteristic of consistent internal and external models, which enables the algorithm to avoid the influence of noise in the process of extracting weight vectors. The analysis of geometric properties and experimental comparison also prove that LVR maximizes the projection of u in the latent space, considering both the correlation between X and Y and the variance structure of Y. This enables LVR to extract latent variables between X and Y better than traditional methods. To avoid the ill-conditioning of the model caused by collinearity problems, a regularization term is introduced in the external model. Then the external model can be equivalently expanded as
[0044]
[0045] where γ is the regularization parameter.
[0046] After solving by the eigenvalue decomposition method and the nonlinear iterative algorithm, X and Y can be decomposed into
[0047]
[0048] where, T = [t1, t2,..., t l is the latent variable matrix composed of the extracted latent variables, extracted by the external model; P = [p1, p2,..., p l is the loading matrix, and Q = [q1, q2,..., q l is also the loading matrix at this time, respectively representing the regression relationship of T to X and Y, both obtained through the internal model; Is the decomposed principal space (PCS), which contains the information of Y in X; Is the residual space (RS) after eliminating the influence of Y from X; In contrast, Is the RS after eliminating the influence of X from Y.
[0049] Since MSPC is often used for fault detection in chemical processes, Because it contains the quality information in the process data, usually using T 2 Statistic monitoring, Then SPE is usually used for monitoring, which makes the development of Y relatively less. And combined with the analysis of the fan operation principle, RS is the stationary signal after eliminating the influence of the system latent variables on the generator characteristics. Based on the above analysis, the present invention will conduct research on RS, and further construct the MSLVR model considering the non-stationarity and dynamic characteristics of the fan temperature signal.
[0050] (2) Construction of the multi-scale latent variable regression model
[0051] a. Data preprocessing method based on multivariate variational mode decomposition
[0052] As a mature non-stationary signal processing technology, mode decomposition is widely used in the multi-time scale feature extraction of time series data by decomposing complex signals into a series of intrinsic mode functions (IMFs). The multivariate variational mode decomposition (MVMD) technology can effectively suppress the deficiencies such as mode mixing, endpoint effect, and uncontrollable mode components compared with the common empirical mode decomposition (EMD). And compared with the variational mode decomposition (VMD), it can jointly process the multi-dimensional input X and take into account the correlation between variables. Therefore, the present invention will adopt the MVMD technology to perform multi-scale expansion on the coupled variable X to make the decomposition result more consistent and physically meaningful. For the input X(t) = {x1(t), x2(t),..., x7(t)} of the present invention, extract the predefined K multivariate modulated oscillations z k (t) from the input data including 7 data channels, and two conditions need to be satisfied: the sum of the bandwidths of the extracted modes is the smallest; the sum of the extracted modes can accurately recover the original signal. The goal of MVMD is to simultaneously optimize the mode components of multiple signals, so that the decomposition maintains the consistency of the modes between channels, which can be expressed as the following minimization optimization goal
[0053]
[0054] Among them, {z k,i} is the set of different mode components; {ω k} is the set of center frequencies of {z k,i}; z k,i is the k-th mode component of the i-th channel; ωk is the central frequency of the k-th mode; is the analytic representation of the vector signal z k,i (t) obtained by using the Hilbert transform operator; x i (t) is the input data of the i-th channel.
[0055] Based on the above problems, an augmented Lagrangian function is constructed to remove the constraints of the multivariate variational optimization problem. The variational optimization problem is solved by the alternating direction method of multipliers. Finally, after multiple iterations, the multivariate original signal X can be extended into a dataset with different time scales.
[0056] b. Serial structure latent variable regression model
[0057] After the above MVMD decomposition, X will become a multi-scale dataset Z is a three-dimensional data of N×K×7, where N is the number of samples and K is the mode function. There are relatively few existing studies on the multi-scale extension of the latent variable model. If the constructed multi-scale dataset is flattened and directly input into LVR for latent variable extraction, the influence of each mode locally on Y will be ignored.
[0058] However, the PCA decomposition method with a serial structure can be used for reference. First, PCA is applied to extract PCs as linear features, and then in the RS of the PCA decomposition, KPCA is used for non-linear feature extraction. Combining this idea with the scenario of the present invention, if the input data consists of multiple time-scale subsets, and if the first mode is used as the LVR input, then the T extracted by the model is the latent variable of the system coupling variable at this time scale. Then, in the RS of Y after eliminating this influence, there should still be information at other time scales. Assuming that there is a linear relationship between the subsequent mode subsets and the corresponding RS, then a serial structure LVR is constructed, which can sequentially extract the latent variables of different modes and eliminate them one by one. Finally, an RS that eliminates all time-scale influences of X can be obtained. In addition to being simple and efficient, this architecture can retain the multi-scale information in the data and avoid the problem of local information loss caused by mixing the features of all modes together.
[0059] Based on the above, the MVMD decomposition and the introduction of a serial structure can be combined to construct an MSLVR model. The principle of MSLVR is as Figure 2 shown. Before that, Z needs to be reconstructed. The specific operation is to convert the data grouped by variables originally into data grouped by modes and arrange them in order from low frequency to high frequency. At this time After recombination, the data of mode 1, Z 1 is used as the input to extract and eliminate the latent variable of Z 1 contained in Y. The outer model in the formula can be converted into
[0060]
[0061] Among them, t 1 is the latent variable of Z 1 ; w 1 and q 1 are the weight vectors in the first mode respectively, and γ1 is the regularization coefficient at the first calculation.
[0062] After solving, Y is decomposed into the following form
[0063]
[0064] Among them, is the extracted latent variable matrix about Z 1 ; the explanations of the remaining variables are similar; l1 is the number of latent variables of the model at the first calculation; T 1 (Q 1 ) is the PCS, which contains the information about Z 1 in Y; is the RS after eliminating the influence of Z 1 through the regression relationship.
[0065] At this time, still contains the information about the remaining time scales of X. Therefore, according to the relationships given in equations (7) to (8), for the remaining components in , they are eliminated in order from high frequency to low frequency, and equation (7) can be transformed into
[0066]
[0067] Among them, j takes values from 2 to K in sequence, and β takes values from 1 to K in sequence.
[0068] The subsequent decomposition relationship can be expressed as follows
[0069]
[0070] After the above operations, Y is gradually obtained after eliminating the influences of multiple time scales That is, RS.
[0071] In the past, when using the latent variable model for fault detection, P and Q were obtained in the training stage, and since T changes with the data, T can be equivalently T = XR, R = W(P T W) -1 . Therefore, in the online stage when new data Z new is input, the latent variables in different modes can be calculated by the following relationships
[0072]
[0073] R j = Wj [(P j ) T W j -1 (12)
[0074] Among them, j takes values in sequence from 1 to K.
[0075] Then when there are Y new inputs, the new RS can be expressed by the following formula
[0076]
[0077] Obtained through the above calculations and are respectively the RSs of the DFIG temperature signal after eliminating the influence of the system coupling latent variables in the offline stage and the online stage. Thus, based on the RS, the monitored quantity is constructed to achieve online detection.
[0078] (3) Development of the monitored quantity based on the residual subspace
[0079] Hotelling's T 2 statistic's main objective is to test whether there is a significant difference between two multivariate mean vectors, and it has strong advantages in system and global monitoring, and is more suitable for the detection of various types of generator faults. Thus, the present invention will be based on construct T 2 statistic. Since the decomposed RS still contains a large variance, and although the influence of Z is eliminated in the RS, there is still a correlation between the internal variables, which is not conducive to the construction of T 2 . Therefore, PCA still needs to be used to orthogonally decompose the RS to remove noise and correlation ε e is noise, contains most of the residual information and can be obtained through the following calculations
[0080]
[0081] Among them, T e and P e are respectively the score vector and the loading matrix obtained by PCA decomposition.
[0082] The calculation of T 2 is
[0083]
[0084] Among them is a diagonal matrix.
[0085] Regarding the control limit of T 2 of Tl 2 It is obtained by the kernel density estimation method in the training stage, and the confidence level of the present invention is taken as 0.01.
[0086] When there is a new sample input, the new T 2 is
[0087]
[0088]
[0089] When online detection is performed, exceeding the control limit T l 2 it can be determined that the generator equipment has a fault.
[0090] (4) DFIG fault detection method based on MSLVR
[0091] In summary, the DFIG fault detection method for wind turbine based on the MSLVR model proposed by the present invention consists of Figure 3 as shown, and is divided into two parts: offline and online, which can be summarized as:
[0092] a. Offline stage
[0093] Based on the mutual information technology, determine the system coupling variable matrix X that affects the overall temperature characteristic Y of the DFIG. And obtain the historical data of X and Y from the SCADA system.
[0094] Perform MVMD decomposition on X and reconstruct it into a multi-scale data set Z divided by time scale. Input the standardized Z and Y into the LVR with a serial structure for cross-validation training to determine the key parameters of the model. Extract T under different time scale components in sequence to obtain the relationship matrix R that can reflect the relationship between X and T, and the matrix Q of the regression relationship between T and Y. Eliminate the influence of different modal Ts in sequence to obtain the final RS.
[0095] For the RS after eliminating the influence of multi-scale latent variables on the DFIG temperature characteristic, combine it with PCA for orthogonal decomposition to construct the T 2 statistic based on RS, obtain P e and Λ, and determine the control limit T l 2 .
[0096] b. Online stage
[0097] According to the constructed model and method, perform online detection to obtain real-time data X new and Y new , perform multi-scale expansion on X new to obtain a new multi-scale data set Znew 。
[0098] For Z new and Y new Perform normalization processing to obtain R and Q calculated in the offline stage, and obtain a new RS.
[0099] Obtain P calculated in the offline stage e and Λ, calculate the new T according to equations (16) and (17) 2 , when exceeding the control limit T l 2 , determine that the DFIG has a fault, otherwise it is normal.
[0100] 3. Experiments and Results
[0101] (1) Dataset Division and Experimental Setup
[0102] a. Fault Description and Dataset Division
[0103] To verify the effectiveness and advancement of the detection method of this application, three typical fault cases of actual generators will be used for analysis. The case data were all collected and recorded from the wind farm site described above. According to the statistical information of this wind farm, in addition to the bearing faults that have been widely studied, the DFIG faults that cause long downtime and high maintenance costs also include stator winding inter-turn short circuit faults and rotor slip ring wear faults. The downtime and losses caused by the remaining faults are much smaller than these three types of faults. When an inter-turn short circuit occurs in the stator winding, the local winding current will increase significantly, resulting in electromagnetic imbalance and local overheating. Since the short circuit loop will form a low-impedance path, the short circuit area will bear too much current and generate a large amount of heat. The slip ring is used to transfer current to the rotor winding. As the slip ring and brush wear, the contact resistance increases, resulting in heat generation at the contact point. This heat will gradually spread to the rotor system, causing the temperature to rise.
[0104] According to the on-site maintenance records, the No. 25 fan was shut down for maintenance due to bearing faults at 7:40 on April 25, 2019. The No. 18 fan was shut down due to stator winding faults at 23:40 on November 18, 2022. The generator rotor slip ring fault occurred at 23:50 on July 5, 2020 in the No. 47 fan, and it was immediately shut down for replacement. This invention will conduct analysis and research based on the above three types of faults, and the dataset division is shown in Table 1.
[0105] Table 1 Dataset Division
[0106]
[0107] b. Experimental Setup
[0108] This application will demonstrate the detection results of MSLVR, MVMD-LVR, and LVR through Case 1 to verify the effectiveness of the method of the present invention. Case 2 will show the ablation experiment results to verify the necessity of the improvement. Since machine learning and neural networks are used to construct the NBM analysis to predict residuals, it is the current mainstream method for fault detection of generators based on temperature signals. Therefore, Case 3 will show the detection results of MSLVR and the classical machine learning algorithm SVM and the current popular neural network model Transformer to verify the advancement of the method.
[0109] Before testing, training data is required to determine the model parameters. To avoid the method being overly complex, the present invention sets the number of modes decomposed by MVMD to 3, α = 600, and the model convergence error threshold to 1e-6. This setting applies to all models of the present invention involving MVMD. The determined X in the present invention is decomposed into three scales of low, medium, and high frequencies after being processed by MVMD. Since they are all temperature variables, the signals of the three scales can respectively represent the long-term trend, periodic changes, and short-term fluctuation components of temperature, and effectively improve the non-stationarity of the original data. After training, the model parameters are shown in Table 2.
[0110] Table 2 Model Parameters
[0111]
[0112] (2) Case Analysis
[0113] To quantify the experimental results and further verify the effectiveness and advancement of the method of the present invention. The detection performance of the method will be evaluated by calculating the false positive rate (FPR) and the false negative rate (FNR), and the calculation formulas are
[0114]
[0115] where FP is the number of data detected as faulty under normal conditions; TN represents the number detected as normal under normal working conditions; FN is the number of data detected as normal under abnormal conditions; TP is the number of slices detected as abnormal under abnormal working conditions; the smaller the FNR and FPR, the better the performance of the detection method.
[0116] At the same time, to compare the early warning performance of the method, the early warning time margin Δt is defined, and the calculation formula is
[0117] Δt = t f -t s (19)
[0118] where t f is the fault time in the record; t s is the alarm time when a fault is confirmed to occur (when the monitored values of 6 consecutive points, that is, 1 hour, exceed the threshold, a fault is confirmed to occur).
[0119] a. Case 1: In this case, the detection results of bearing faults are as Figure 4 shown in and Table 3, where the left figure is the monitoring result control chart and the right is the alarm result chart. Values exceeding the threshold are marked as 1, and those not exceeding are marked as 0. By observing Figure 4 the display results, it can be seen that all three methods can detect early anomalies caused by faults 13.3 days earlier than the maintenance records, and the alarm frequency increases significantly after the fault is confirmed. It can be clearly seen from the alarm chart and the table that the false alarm points of the MSLVR detection method and MVMD-LVR based on the present invention are significantly fewer than those of LVR on the eve of fault confirmation. This shows that integrating the MVMD technology can expand the data scale, improve the data quality, effectively enhance the ability of LVR to extract latent variables, and thus improve the robustness of the detection method. By comparison, the MSLVR method constructed in the present invention further suppresses the false alarm phenomenon compared with the LVR method that simply combines MVMD, indicating that the serial structure can pay more comprehensive attention to the latent variables of each modality than the direct combination method. And the FPR and FNR of MSLVR are both smaller than those of the former, indicating better detection performance.
[0120] Table 3 Results of Case 1
[0121]
[0122] b. Case 2: In this case, a model that only eliminates the influence of the first modal component is defined as MSLVR (1) , and the component that eliminates the influence of the first two modal components is MSLVR (2) . The detection results of the three methods for generator winding faults are as Figure 5 shown in and Table 4. The description of Figure 5 is the same as that in Case 1. It is not difficult to find that all three cases can detect generator anomalies 29.6 days in advance, and the number of alarms also increases significantly after the fault is confirmed. It should be noted that in the initial stage of the fault, the monitored quantity will not continuously exceed the threshold, which is due to the adjustment effect of the wind turbine cooling system. From the comparison chart, it can be seen that MSLVR after eliminating the influence of the three modal latent variables has relatively fewer false alarm phenomena and a lower miss rate compared with MSLVR (1) and MSLVR (2) . This shows that although IMF1 contains most of the information of X, the rapid fluctuations will still affect the residual subspace, which is not conducive to detection. By comparing Figure 5 it can be seen that although the detection results of MSLVR (2) and MSLVR are not very different, eliminating the influence of low-frequency noise is still beneficial to online detection.
[0123] Table 4 Results of Case 2
[0124]
[0125] c. Case 3: To enable the Transformer and SVM models to comprehensively detect generator faults, according to the conventional physical explanations based on such models, the SPE is constructed as the monitoring variable for the predicted error of the output, and the threshold is determined using KDE. For the generator carbon brush slip ring fault, the detection results are as Figure 6 shown in Table 5. Since the results of SVM and Transformer after the full deterioration of the fault are the same as those of LVR, for ease of analysis, Figure 6 only the local detection effects of SVM and Transformer are shown. It can be found that the fault has deteriorated at Figure 6 (a) the 4,667th sampling point. Therefore, in this case, the full deterioration time is taken as the time when the fault is recorded. It is not difficult to find that both the methods based on SVM and Transformer can detect the occurrence of the fault, and the FPR and FNR are not very different. However, the time when the MSLVR method first detects a large number of abnormal points is 4.1 days earlier than the full deterioration time, and the fault trend is obvious, and the FPR is also relatively low. The possible reason is that SVM and Transformer only combine multiple individual features instead of comprehensively considering the complex relationship between two sets of multivariate variables. In addition, the quality of the training data will affect the performance of the constructed NBM, and a large amount of data and the data cleaning process are often required. However, MSLVR can effectively reduce the interference of noise and irrelevant information through the principle of projection dimensionality reduction. And the deficiency of SPE in terms of globality may also be one of the reasons why the model has a higher FPR in the research of the present invention and fails to capture the early fault characteristics of the generator earlier. The FNR of the method of the present invention is higher than that of SVM and Transformer. The main reason is that the detection time of SVM and Transformer is close to the time of full fault occurrence, so a relatively low false negative rate is shown.
[0126] Table 5 Results of Case 3
[0127]
[0128] (3) Comprehensive analysis
[0129] To further analyze the online detection ability of the method proposed in the present invention, a comprehensive analysis will be conducted on four models, namely MSLVR, LVR combined with MVMD, SVM, and Transformer, based on the above three cases. In addition to the average Δt, in order to comprehensively evaluate the detection performance of each method, the daily false positive rate (DFPR), daily false negative rate (DFNR), and daily accuracy (DA) are added. Specifically, the test data of the three types of faults are divided into 210 time slices according to each day (144 points). Among them, 153 normal slices are marked (Case 1: 76, Case 2: 49, Case 3: 28), and 57 abnormal slices (Case 1: 14, Case 2: 31, Case 3: 12). It is stipulated that a slice with abnormal points continuously exceeding 1 hour (6 sampling points) in the normal slices is abnormal. The principles of DFPR and DFNR are consistent with formula (18), and the principle of DA is
[0130]
[0131] Figure 7 The confusion matrix showing the calculation results is presented. In the coordinates, 0 represents normal and 1 represents abnormal. The abscissa is the actual classification of the three cases, and the ordinate represents the prediction situations of the four methods. At the same time, the complexity of the model will also affect the online detection ability of the method. Therefore, the model training time and test time are added as indicators of computational efficiency (excluding the MVMD stage). A computer with a 2.3 GHz processor and 16 GB of RAM is used to obtain the results. The calculation results of the average indicators of the three cases for the methods involved are shown in Table 6. It can be found that the method of the present invention has the highest accuracy and can detect early anomalies caused by faults in a timely manner while maintaining a low false positive rate. By comparing each indicator, it is found that the detection performance of the method of this application is also superior to that of similar methods and mainstream methods. Although the training and test times are longer than those of LVR and SVM, they are much shorter than those of the Transformer model.
[0132] Table 6 Comprehensive analysis results
[0133]
Claims
1. A method for early fault detection of double-fed induction generators in wind power systems based on multi-scale latent variable regression, characterized by: It first analyzes the recorded data with a 10-minute sampling interval in the fan SCADA system, determines the generator temperature variable as the monitoring variable, and calculates the information interaction intensity between the monitoring variable and other system variables through comprehensive mutual information technology, thereby determining the coupling variables that mainly affect the generator temperature characteristics; then, the multivariate variational mode decomposition (MVMD) is used to decompose the coupling variable matrix, and the original non-stationary signal is expanded into modal components of multiple time scales, where the low-frequency component reflects the long-term trend, the medium-frequency component represents the periodic change, and the high-frequency component reflects the short-term fluctuation; the above decomposition results are reorganized into a multi-scale data set divided by time scale, and input into the serial structural latent variable regression (LVR) model. By extracting and eliminating the potential influence of each modal component on the generator temperature signal in turn, the residual subspace (RS) with the interference of the coupling variables eliminated is finally obtained; RS is combined with principal component analysis (PCA) for orthogonal decomposition to remove noise and redundant information, and a Hotelling's T based on RS is constructed. 2 The monitoring statistics are used, and the control limits are determined by the kernel density estimation method. In practical applications, when the monitoring statistics exceed the control limits, it can be determined that the generator has failed. The effectiveness and accuracy of the method are verified through three typical fault cases, including bearing fault, stator winding fault and rotor slip ring fault, and sensitive detection and early warning of faults are achieved in each case.
2. According to the method for early fault detection of a doubly-fed induction generator in a wind power system based on multi-scale latent variable regression according to claim 1, the present application comprises the following steps: (1) Offline stage a. Extract the operation data with a sampling interval of 10 minutes within five years from the wind turbine SCADA system. According to the temperature characteristics of the generator, the mutual information technology is used to calculate the coupling variable matrix related to it, and the coupling variables are selected as input. b. The coupled variable matrix is processed using multivariate variational mode decomposition (MVMD) to decompose the variables into low-frequency, medium-frequency and high-frequency modal components to form a multi-scale data set Z with physical significance, where the low-frequency component reflects the long-term trend, the medium-frequency component represents the periodic change, and the high-frequency component reflects the short-term fluctuation; The serial structural latent variable regression model is used to extract the potential influence of each modal component on the temperature signal in turn, and the interference of the mode is eliminated one by one until the residual subspace that eliminates all time scale interferences of the coupling variables is obtained; thus, the construction of the multi-scale latent variable regression (MSLVR) model is completed. c. Perform orthogonal decomposition of the residual subspace combined with principal component analysis (PCA) to remove noise and correlation and construct Hotelling's T 2 Monitoring statistics; d. Use the kernel density estimation method to determine the control limits of the monitoring statistics, with a confidence level of 0.01 as the standard; (2) Online stage a. Perform online detection based on the constructed model and method to obtain real-time data X new and Y new , for X new Also use MVMD for multi-scale expansion to obtain a new multi-scale dataset Z new . b. For Z new and Y new The normalization process is performed to obtain the new generator temperature main space (PCS) calculated in the offline phase, and the new RS is obtained. c. Obtain the statistical correlation matrix calculated in the offline phase and calculate the new T 2 Statistics, when it exceeds the control limit, it is determined that the DFIG is faulty, otherwise it is normal.
3. The method for early fault detection of a doubly-fed induction generator in a wind power system based on multi-scale latent variable regression according to claim 1 is characterized by: Considering the influence of wind turbine system, information measurement technology is used to analyze the coupling between DFIG and the system, and supervised MSPC technology is introduced to obtain potential relationships, aiming to keenly and robustly capture the early anomalies caused by DFIG faults.
4. The method for early fault detection of a doubly-fed induction generator in a wind power system based on multi-scale latent variable regression according to claim 1 is characterized by: developing The MSLVR model combines the multivariate variational mode decomposition (MVMD) technology to decompose the coupling variables of the original non-stationary signal to ensure the physical meaning of the decomposition results and the correlation between variables. A serial structural latent variable regression (LVR) model is established to gradually extract the latent variables of multi-scale modal components and eliminate the interference of modal components on the target signal. The weight vector is determined by minimizing the error between latent variables in each step of the latent variable extraction process, and a regularization term is combined to avoid collinearity problems.
5. The method for early fault detection of a doubly-fed induction generator in a wind power system based on multi-scale latent variable regression according to claim 1 is characterized by: The Hotelling's T2 monitoring statistic constructed by combining the residual subspace (RS) with PCA is used to quantify the abnormality of the generator temperature signal. The control limit is determined by combining the kernel density estimation method. When the monitoring statistic value exceeds the control limit, it is determined that a fault has occurred. This method can comprehensively detect different types of DFIG faults.
Citation Information
Cited By
Wind power gear box intelligent fault early warning method and system based on machine learning
CN120998009A