Abnormal data detection and cleaning method for wind turbines based on LSTM-AE integrated sharing framework
By adopting the LSTM-AE integrated sharing framework and adaptive threshold method in the detection and cleaning of wind turbine abnormal data, the problems of overfitting and misjudgment are solved, the detection accuracy and cleaning accuracy are improved, and the accuracy of wind power monitoring and prediction models are improved.
Patent Information
- Application Number
- CN202111560423.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-12-20
AI Technical Summary
The prior art has problems of overfitting and misjudgment in the detection and cleaning of abnormal data of wind turbines, especially when abnormal data accounts for a large proportion, resulting in low detection accuracy and cleaning accuracy.
The method based on the LSTM-AE integrated sharing framework is adopted, and the impact ratio of each unit data is optimized during the model training process by hidden state sharing module, and combined with the cleaning method of adaptive thresholds, the accuracy of abnormal data detection and cleaning is improved.
It effectively improves the abnormal data detection accuracy and cleaning accuracy of the wind turbine unit model, reduces overfitting and misjudgment, and improves the accuracy of the wind power monitoring and prediction model.
Smart Images

Figure CN114357866B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of abnormal data detection and cleaning of wind turbines, and relates to an abnormal data detection and cleaning method for wind turbines based on an LSTM-AE integrated sharing framework. Background Art
[0002] Energy shortage is the main problem that restricts the international community from maintaining rapid development. Wind energy, as a renewable, sustainable and clean energy, is rapidly becoming an important part of the carbon neutral energy strategy, and the number and scale of wind farms are constantly expanding. However, the uncertainty of wind energy itself leads to large-scale power fluctuations in the power grid system after wind power is connected to the grid, which affects the safe operation of the system. In order to make wind energy a reliable source of energy, it is very important to establish an efficient and accurate wind power monitoring and prediction model through high-quality wind turbine operation data. However, wind turbines are affected by factors such as equipment quality, working environment, and operating status. There will be a large number of abnormal data in the collected wind turbine operation data that do not conform to the normal output characteristics of wind turbines. The existence of these "dirty data" will affect further predictive modeling analysis, cause information misjudgment, and even have an adverse impact on the safe operation of the power grid system. Therefore, effective detection and cleaning of abnormal data in wind turbine operation data is a necessary prerequisite for improving the accuracy of predictive modeling analysis. Finding a reliable method for detecting and cleaning abnormal data of wind turbines is of great significance.
[0003] It is a common method to collect wind turbine operation data through the wind farm's data acquisition and supervision (SCADA) system. Abnormal data in wind turbine operation data is mainly due to large deviations between the actual operating conditions of the wind turbine and the designed operating conditions, as well as sensor failures and transmission noise. At present, there are many wind turbine abnormal data detection and cleaning methods based on SCADA system data, which can be mainly divided into two categories. One is a method based on mathematical statistics, which mainly detects abnormal data through mathematical statistics methods based on the distribution characteristics of the data combined with prior theoretical knowledge. This type of method ignores the time series characteristics of wind turbine operation data during modeling. Therefore, it is difficult for this type of method to accurately describe the abnormalities caused by the changing trend of the wind turbine operation status. The other is an artificial intelligence-based method that mainly starts from the characteristics and structure of the data itself, and mines the deep information hidden in the data through deep learning. It can effectively process high-dimensional variables, and can effectively learn the changing trend of the wind turbine operation status through time series-related models, and can effectively identify abnormalities in the changing trend. For example, the wind turbine data anomaly detection method based on the LSTM-AE network model, such as Figure 3The method combines the characteristics of LSTM in effectively processing time series with the advantages of AE in effectively extracting the essential characteristics of data, learns and reconstructs the input sequence, and uses the reconstruction error for anomaly detection. However, in actual application, some units in the original data may have abnormal data that accounts for too large a proportion of the overall data, resulting in overfitting of the data of the unit in the training phase when modeling based on the operating data of a single unit, and the inability to clearly differentiate between normal data and abnormal data in the detection phase, resulting in low accuracy. The various performance parameters of wind turbine operating data vary greatly under different working conditions, and the criteria for judging reconstruction errors under different working conditions are different. The traditional method of using fixed thresholds to clean up abnormalities in reconstruction errors is prone to misjudgment and omission. Summary of the invention
[0004] In order to solve the deficiencies in the prior art, this application proposes a method for detecting and cleaning abnormal data of wind turbines based on an LSTM-AE integrated sharing framework to improve the detection accuracy and cleaning accuracy of abnormal data of wind turbine models in the overall wind farm. In view of the fact that some units in the original data may have abnormal data that accounts for too large a proportion of the overall data during actual application, which leads to overfitting of the data of the unit during the training phase when modeling based on the operating data of a single unit, and the inability to clearly differentiate between normal data and abnormal data during the detection phase, resulting in a low accuracy rate, this invention proposes a hidden state sharing module that can optimize parameters and adjust the influence proportion of each unit data during model training by combining the data of adjacent units to help with training during the training process, and combines the LSTM-AE network model structure to design an integrated sharing framework that can effectively conduct joint training of multi-unit models. In view of the problem that the traditional method of cleaning the reconstruction error by fixing the threshold is prone to misjudgment and missed judgment, the present invention constructs the expected function between the probability density of the error in the multivariate Gaussian distribution and the corresponding reconstruction value through the AE network, and proposes an adaptive threshold cleaning method. By comparing the expected error probability density of the reconstructed value with the probability density of the actual error, the corresponding error threshold is adaptively adjusted under different working conditions to improve the accuracy of abnormal data cleaning.
[0005] The technical solution adopted by the present invention is as follows:
[0006] The abnormal data detection and cleaning method of wind turbines based on the LSTM-AE integrated sharing framework includes the following steps:
[0007] Step 1: preprocess the SCADA data of n neighboring wind turbines in the same wind farm to obtain n original time series with the same timestamp;
[0008] Step 2: construct an integrated sharing framework based on LSTM-AE, which includes a sliding window amplification layer, an encoding layer, a hidden state sharing layer and a decoding layer; the input of the sliding window amplification layer is the original time series, and the sliding window amplification layer performs feature engineering on the aligned data to obtain an amplified time series sequence;
[0009] The encoder in the encoding layer learns the amplified time series to obtain the hidden state and inputs it into the hidden state sharing layer;
[0010] The hidden state sharing layer optimizes and adjusts the influence proportion of the hidden state of each unit during the model training process to obtain the shared hidden state, and outputs the shared hidden state to the decoding layer;
[0011] The decoding layer reconstructs the amplified time series sequence through the input shared hidden state and outputs the corresponding amplified time series sequence;
[0012] The integrated shared framework loss function is obtained by superimposing the reconstruction error and penalty term of the wind turbines. The loss function is introduced into the shared hidden state as a penalty term for joint training of multi-unit models.
[0013] Step 3: Based on the trained integrated shared framework, the encoder E in the trained integrated shared framework is i With decoder D i Corresponding to the split, by the re-encoder E i With decoder D i Build an LSTM-AE model for a single wind turbine and perform abnormal data detection;
[0014] Step 4: Set the cleaning index ξ, and perform abnormality judgment and cleaning by comparing the expected error probability density of the reconstructed value with the probability density of the actual error.
[0015] Furthermore, the method of constructing the hidden state sharing layer is:
[0016] Add a hidden state sharing module in the encoding layer and the decoding layer, which is the encoding layer E = (E1, E2...E n ) designs a linear weight matrix for each encoder Shared hidden state is by superimposing each hidden state The corresponding linear weight matrix The product of is expressed as follows:
[0017]
[0018] Among them, the function f(·) is a linear superposition function, which shares the hidden state As the decoding layer D=(D1,D2 . . . D n) is the input to each decoder in .
[0019] Furthermore, multi-unit model joint training:
[0020] Based on the reconstruction error loss of each wind turbine i And the penalty term obtains the integrated shared framework loss function, which is expressed as follows:
[0021]
[0022] Among them, loss i represents the reconstruction error of unit i, loss represents the integrated shared framework loss function, is the penalty term, λ is the control The importance of the penalty effect in the loss function; are the jth vector in the amplified sequence and the reconstructed sequence in unit i, respectively.
[0023] Further, abnormal data detection:
[0024] Use the amplified time series of unit i after sliding window amplification As the input of the LSTM-AE model of unit i, the reconstructed sequence of the model output is obtained through the decoder is the xth reconstruction value in the reconstruction sequence of the i-th unit; calculate the error between the input amplification timing sequence and the reconstruction sequence to obtain the error vector sequence It is expressed as follows:
[0025] e j =|h j -h″ j |
[0026] E′ i =(e1, e2...)
[0027] Among them, e j is the jth vector parameter in the error vector; |·| is the absolute value function, h j and h″ j are the jth parameters in the amplification vector and the reconstruction vector, For E i The x-th error vector in .
[0028] Further, the adaptive threshold cleaning process:
[0029] Step 1: Establish a multivariate Gaussian distribution model: For the error vector sequence E i Standardization, then establish E i Multivariate Gaussian distribution model E i~N(μ,∑), the estimates of the parameters μ and ∑ of the multivariate Gaussian distribution model are given by the maximum likelihood method;
[0030] Step 2: Fit the error vector The probability density and reconstruction value of Nonlinear expectation function between: probability density of error vector in multivariate Gaussian distribution as follows:
[0031]
[0032] Using AE network to fit error probability density With reconstruction value The nonlinear expected function of , that is, the expected error probability density estimator f: as follows:
[0033]
[0034] Among them, W P is the weight coefficient matrix of the fitting AE network, b P is its offset, and f is the expected error probability density estimation function;
[0035] Step 3: Set the cleaning index ξ, and perform abnormality determination and cleaning by comparing the expected error probability density of the reconstructed value with the probability density of the actual error, as follows:
[0036]
[0037] Among them, f(·) is the expected error probability density estimator, and η is the set error offset. By setting η, the error threshold is adaptively adjusted. If ξ is positive, it means that the input vector corresponding to the error vector is abnormal data and needs to be cleaned.
[0038] Furthermore, the latest start time to the earliest end time is used as the time interval of the SCADA data of the entire wind turbine group. According to the time interval, the SCADA data of each wind turbine group is filtered and sorted to obtain the original time series with the same timestamp corresponding to each group, thus completing the preprocessing of the SCADA data.
[0039] Furthermore, the sliding window amplification layer generates an amplified time series: in the original time series T of each unit input, T = (S1, S2...S C ) uses a window with a span of 2n and an interval of n to slide, and calculates the statistical features and derived features of each vector parameter in the vector set within each window interval, and uses the statistical features and derived features as new amplification vectors, thereby obtaining a new short-term correlation amplification time series sequence T′=(H1, H2...H k ).
[0040] Beneficial effects of the present invention:
[0041] The present invention is applied to the field of abnormal data detection and cleaning of wind turbines, and can effectively detect and clean abnormal data in the operating data of wind turbines, which is of great significance to the establishment of efficient and accurate wind power monitoring and prediction models. In view of the fact that some units in the original data may have abnormal data that accounts for too large a proportion of the overall data during actual application, resulting in overfitting of the data of the unit during the training phase when modeling based on the operating data of a single unit, and the inability to clearly differentiate between normal data and abnormal data during the detection phase, resulting in a low accuracy rate, this invention proposes a hidden state sharing module that can optimize parameters and adjust the influence proportion of each unit data during model training by combining the data of other units to help with training during the training process, and designs an integrated sharing framework that can effectively carry out joint training of multi-unit models in combination with the LSTM-AE network model structure. In view of the problem that the traditional method of cleaning the reconstruction error by fixing the threshold is prone to misjudgment and missed judgment, the present invention constructs the expected function between the probability density of the error in the multivariate Gaussian distribution and the corresponding reconstruction value through the AE network, and proposes an adaptive threshold cleaning method. By comparing the expected error probability density of the reconstructed value with the probability density of the actual error, the corresponding error threshold is adaptively adjusted under different working conditions to improve the accuracy of abnormal data cleaning. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic diagram of the overall process of the abnormal data detection and cleaning method of wind turbines based on the LSTM-AE integrated sharing framework of the present invention;
[0043] Figure 2 This is a structural diagram of the LSTM-AE integrated sharing framework according to the present invention;
[0044] Figure 3 This is a structural diagram of the basic LSTM-AE model described in the present invention. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0046] like Figure 1As shown, according to an embodiment of the present invention, the abnormal data detection and cleaning method of wind turbines based on the LSTM-AE integrated sharing framework includes four basic steps: SCADA data preprocessing; building and training based on the LSTM-AE integrated sharing framework; reorganizing the basic LSTM-AE model corresponding to the unit and performing abnormal data detection; and adaptive threshold cleaning. Step 1: SCADA data preprocessing
[0047] In this embodiment, the SCADA data of 9 neighboring wind turbines in the same wind farm in the same year is taken as an example. The collection period of the operating data of the wind turbines with timestamps in the SCADA data used this time is from November 2017 to October 2018, and the collection time interval is 10 minutes. The monitored performance parameters include wind speed, power and wind rotor speed; the data set contains a total of 497,837 time series data.
[0048] From all the SCADA data of the nine wind turbines, the latest start time to the earliest end time is used as the time interval of the SCADA data of the entire wind turbine. According to the time interval, the SCADA data of each wind turbine is filtered and sorted to obtain 9 original time series with the same timestamp. Each vector in the original time series contains wind speed, power and rotor speed information.
[0049] In the same wind farm and in the same time period, the operating status and change trend of the wind turbine group should be the same. Therefore, the hidden state information in the LSTM-AE isomorphic model is the same. The original SCADA data is aligned. In the subsequent overall joint training, the shared hidden state is the hidden state of the data at the same time.
[0050] Step 2: Build an integrated shared framework based on LSTM-AE and train it
[0051] The integrated sharing framework includes a sliding window amplification layer, an encoding layer, a hidden state sharing layer and a decoding layer; wherein the aligned data (i.e., the original time series) is used as the input of the sliding window amplification layer, and the sliding window amplification layer performs feature engineering on the aligned data to obtain an amplified time series sequence.
[0052] The encoder in the encoding layer learns the amplified time series to obtain the hidden state and inputs it into the hidden state sharing layer;
[0053] The hidden state sharing layer optimizes and adjusts the influence ratio of each unit's hidden state during model training to obtain the shared hidden state, and outputs the shared hidden state to the decoding layer.
[0054] The decoding layer reconstructs the amplified time series sequence through the input shared hidden state and outputs the corresponding amplified time series sequence. During training, multi-model joint training is performed by superimposing the loss function of the multi-unit model and introducing the shared hidden state as a penalty term.
[0055] In this embodiment, the specific process of building an integrated sharing framework based on LSTM-AE is as follows:
[0056] Step 1: Generate an augmented time series through a sliding window augmentation layer: In the original time series T of each unit input, T = (S1, S2...S C ) uses a window with a span of 2n and an interval of n to slide, and calculates the statistical features and derived features of each vector parameter in the vector set within each window interval, and uses the statistical features and derived features as new amplification vectors, thereby obtaining a new short-term correlation amplification time series sequence T′=(H1, H2...H k ), the statistical features and derived features include mean (MEA), maximum (MAX), minimum (MIN), standard deviation (STD), peak distance (P2P) and three quartiles. The feature space of the amplified time series is much larger than that of the original time series, which helps the automatic encoder to identify the most representative features during the model training process; among them, S i is the i-th vector of the original time series; i = 1, 2, .., C, where C is the length of the original time series. j is the expansion vector generated for the jth window interval. j = 1, 2, .., k, where k is the number of windows;
[0057] Step 2: Construct the encoding layer and decoding layer: The encoding layer has 9 encoders E with the same structure i , the decoding layer has 9 decoders D with the same structure i , denoted as E = (E1, E2...E9) and D = (D1, D2...D9), the amplified time series of unit i after sliding window amplification The corresponding encoder E in the input encoding layer E i In the example, the encoder learns the input amplified time series to obtain the hidden state, and the decoder D i After receiving the hidden state, the amplified time series is reconstructed to obtain
[0058] Step 3: Construct a hidden state sharing layer: Add a hidden state sharing module in the encoding layer and the decoding layer, and design a linear weight matrix for each encoder in the encoding layer E = (E1, E2...E9) Shared hidden state is by superimposing each hidden state And the corresponding linear weight matrix The specific formula is as follows:
[0059]
[0060] in, Denotes the encoder E of unit i in the coding layer E i The output hidden state is Represents its corresponding linear weight matrix, which can be learned during training and adjusted by weight Reduce the influence of the hidden state of the abnormal unit in the shared hidden state, so that the unit can train its decoder through the shared hidden state with a large proportion of normal data. The function f(·) is a linear superposition function. Represents the integrated output of the hidden state sharing layer, sharing the hidden state As the decoding layer D=(D1,D2 . . . D n ) is the input of each decoder in;
[0061] Step 4: Joint training of multi-unit models: based on the reconstruction error loss of each wind turbine i And the penalty term is used to obtain the integrated shared framework loss function. The specific formula is as follows:
[0062]
[0063]
[0064] Among them, loss i represents the reconstruction error of unit i, loss represents the integrated shared framework loss function, is the penalty term, λ is the control The importance of the penalty effect in the loss function; are the jth vector in the amplified sequence and the reconstructed sequence in unit i, respectively.
[0065] Step 3: Based on the trained integrated sharing framework, build the LSTM-AE model of a single wind turbine and perform abnormal data detection
[0066] Step 1: Build the LSTM-AE model of a single wind turbine: The encoder E in the encoding layer E = (E1, E2...E9) and the decoding layer D = (D1, D2...D9) in the trained integrated sharing framework are i With decoder D i The corresponding split is performed by the re-encoder E i With decoder D i Construct the LSTM-AE model of a single wind turbine as follows Figure 3 shown.
[0067] Step 2, abnormal data detection: use the amplified time series of unit i after sliding window amplification As the input of the LSTM-AE model of unit i, the reconstructed sequence of the model output is obtained through the decoder is the xth reconstruction value in the reconstruction sequence of the i-th unit; calculate the error between the input amplification timing sequence and the reconstruction sequence to obtain the error vector sequence The specific formula is as follows:
[0068] e j =|h j -h″ j | (3.1)
[0069] E′ i =(e1, e2...) (3.2)
[0070] Among them, e j is the jth vector parameter in the error vector; |·| is the absolute value function, h j and h″ j are the jth parameters in the amplification vector and the reconstruction vector, For E i The x-th error vector in .
[0071] Step 4: Adaptive Threshold Cleaning
[0072] Step 1: Establish a multivariate Gaussian distribution model: For the error vector sequence E i Standardization, then establish E i Multivariate Gaussian distribution model E i ~N(μ,∑), the estimates of the parameters μ and ∑ of the multivariate Gaussian distribution model are given by the maximum likelihood method;
[0073] Step 2: Fit the error vector The probability density and reconstruction value of Nonlinear expectation function between: probability density of error vector in multivariate Gaussian distribution The specific formula is as follows:
[0074]
[0075] Using AE network to fit error probability density With reconstruction value The nonlinear expected function of , that is, the expected error probability density estimator f: The formula is as follows:
[0076]
[0077] Among them, WP is the weight coefficient matrix of the fitting AE network, b P is its offset, and f is the expected error probability density estimation function;
[0078] Step 3: Set the cleaning index ξ, and perform abnormality judgment and cleaning by comparing the expected error probability density of the reconstructed value with the probability density of the actual error. The specific formula is as follows:
[0079]
[0080] Among them, f(·) is the expected error probability density estimator, and η is the set error offset. By setting η, the error threshold is adaptively adjusted. If ξ is positive, it means that the input vector corresponding to the error vector is abnormal data and needs to be cleaned.
[0081] Comparative analysis with other popular wind turbine abnormal data detection and cleaning models: In order to verify the effectiveness of the integrated sharing framework and the adaptive threshold, the present invention conducted a large number of experiments, comparing the cleaning effects of the integrated sharing framework based on LSTM-AE with the RNN-AE model and the LSTM-AE model under fixed thresholds and adaptive thresholds, and using the average of the performance indicators of 9 wind turbines as the final result. The results are shown in Table 1. The F value is a performance indicator for comprehensively evaluating the recall rate and precision rate. The larger the better. The formula is as follows:
[0082]
[0083]
[0084]
[0085] TP is the number of samples that are judged to be normal and are actually normal, FP is the number of samples that are judged to be abnormal but are actually normal, and FN is the number of samples that are judged to be abnormal but are actually normal.
[0086] Table 1 Comparison of cleaning results
[0087]
[0088]
[0089] In summary, the present invention discloses a method for detecting and cleaning abnormal data of a wind turbine based on an LSTM-AE integrated sharing framework, which belongs to the field of abnormal data detection and cleaning of wind turbines. The specific steps are as follows: SCADA system data is grouped according to units and then data is aligned by timestamp; an integrated sharing framework model is constructed, wherein the sliding window amplification layer amplifies the original data by feature engineering, the encoder in the encoding layer learns the input amplification sequence, the hidden state sharing layer optimizes and adjusts the influence proportion of the hidden state of each unit during the model training process and integrates the output to the decoding layer, the decoding layer reconstructs the input sequence by the input shared hidden state, superimposes the loss function of the multi-unit model and introduces the shared hidden state as a penalty term for multi-model joint training; the integrated sharing framework is split and the basic LSTM-AE model corresponding to the unit is reorganized, and the error sequence is calculated according to the reconstructed sequence and the input sequence; the error sequence is modeled by multivariate Gaussian distribution, and the probability density of the error and the expected function of the reconstructed value are constructed by the AE network, and the expected probability density of the reconstructed value is compared with the actual error probability density to perform data cleaning. The present invention can be applied to abnormal data detection and cleaning of wind turbines, and can effectively identify and clean abnormal data in the operation data of wind turbines.
[0090] The above embodiments are only used to illustrate the design ideas and features of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, any equivalent changes or modifications made based on the principles and design ideas disclosed by the present invention are within the protection scope of the present invention.
Claims
1. A wind turbine abnormal data detection and cleaning method based on LSTM-AE integrated sharing framework, characterized in that: The steps include: Step 1: preprocess the SCADA data of n neighboring wind turbines in the same wind farm to obtain n original time series with the same timestamp; Step 2: construct an integrated sharing framework based on LSTM-AE, which includes a sliding window amplification layer, an encoding layer, a hidden state sharing layer and a decoding layer; the input of the sliding window amplification layer is the original time series, and the sliding window amplification layer performs feature engineering on the aligned data to obtain an amplified time series sequence; The encoder in the encoding layer learns the amplified time series to obtain the hidden state and inputs it into the hidden state sharing layer; The hidden state sharing layer optimizes and adjusts the influence proportion of the hidden state of each unit during the model training process to obtain the shared hidden state, and outputs the shared hidden state to the decoding layer; The decoding layer reconstructs the amplified time series sequence through the input shared hidden state and outputs the corresponding amplified time series sequence; The integrated shared framework loss function is obtained by superimposing the reconstruction error and penalty term of the wind turbines, and the loss function is introduced into the shared hidden state as a penalty term for joint training of multi-unit models; Joint training of multi-unit models: Based on the reconstruction error loss of each wind turbine i And the penalty term obtains the integrated shared framework loss function, which is expressed as follows: Among them, loss i represents the reconstruction error of unit i, loss represents the integrated shared framework loss function, is the penalty term, λ is the control The importance of the penalty effect in the loss function; are the jth vector in the amplified sequence and the reconstructed sequence in unit i, respectively; Step 3: Based on the trained integrated shared framework, the encoder E in the trained integrated shared framework is i With decoder D i Corresponding to the split, by the re-encoder E i With decoder D i Build an LSTM-AE model for a single wind turbine and perform abnormal data detection; Step 4: Set the cleaning index ξ, and perform abnormality determination and cleaning by comparing the expected error probability density of the reconstructed value with the probability density of the actual error; Adaptive threshold cleaning process: Step 1: Establish a multivariate Gaussian distribution model: For the error vector sequence E i Standardization, then establish E i Multivariate Gaussian distribution model E i ~N(μ,∑), the estimates of the parameters μ and ∑ of the multivariate Gaussian distribution model are given by the maximum likelihood method; Step 2: Fit the error vector The probability density and reconstruction value of Nonlinear expectation function between: probability density of error vector in multivariate Gaussian distribution as follows: Using AE network to fit error probability density With reconstruction value The nonlinear expected function, that is, the expected error probability density estimator f: as follows: Among them, W P is the weight coefficient matrix of the fitting AE network, b P is its offset, f is the expected error probability density estimation function; Step 3: Set the cleaning index ξ, and perform abnormality determination and cleaning by comparing the expected error probability density of the reconstructed value with the probability density of the actual error, as follows: Among them, f(·) is the expected error probability density estimator, and η is the set error offset. By setting η, the error threshold is adaptively adjusted. If ξ is positive, it means that the input vector corresponding to the error vector is abnormal data and needs to be cleaned.
2. According to claim 1, the method for detecting and cleaning abnormal data of wind turbines based on the LSTM-AE integrated sharing framework is characterized in that: The method to construct the hidden state sharing layer is: Add a hidden state sharing module in the encoding layer and the decoding layer, and the encoding layer E = (E1, E2…E n ) designs a linear weight matrix for each encoder Shared hidden state is by superimposing each hidden state And the corresponding linear weight matrix The product of is expressed as follows: Among them, the function f(·) is a linear superposition function, which shares the hidden state As the decoding layer D=(D1,D2...D n ) is the input to each decoder in .
3. The abnormal data detection and cleaning method for wind turbines based on the LSTM-AE integrated sharing framework according to claim 1 is characterized in that: Abnormal data detection: Use the amplified time series of unit i after sliding window amplification As the input of the LSTM-AE model of unit i, the reconstructed sequence of the model output is obtained through the decoder is the xth reconstruction value in the reconstruction sequence of the i-th unit; calculate the error between the input amplification timing sequence and the reconstruction sequence to obtain the error vector sequence It is expressed as follows: e j =|h j -h j ”| THE' i \(e1,e2…) Among them, e j is the jth vector parameter in the error vector; |·| is the absolute value function, h j and h″ j are the jth parameters in the amplification vector and the reconstruction vector, For E i The x-th error vector in .
4. The abnormal data detection and cleaning method for wind turbines based on the LSTM-AE integrated sharing framework according to claim 1 is characterized in that: The latest start time to the earliest end time is taken as the time interval of the SCADA data of the entire wind turbine group. According to the time interval, the SCADA data of each wind turbine group is filtered and sorted to obtain the original time series with the same timestamp corresponding to each unit, thus completing the preprocessing of the SCADA data.
5. The abnormal data detection and cleaning method for wind turbines based on the LSTM-AE integrated sharing framework according to claim 1 is characterized in that: The sliding window amplification layer generates an amplified time series: the original time series T of each unit is input as (S1, S2…S C ) uses a span of 2n and a window with an interval of n to slide, and calculates the statistical characteristics and derived characteristics of each vector parameter in the vector set within each window interval, and uses the statistical characteristics and derived characteristics as new amplification vectors, thereby obtaining a new short-term correlation amplification time series sequence T'=(H1,H2…H k ).
Citation Information
Patent Citations
Power time series data anomaly detection method based on long-term and short-term memory network
CN112308402A
Large-scale multivariate time series data anomaly detection method oriented to cloud environment
CN112784965A