A method and device for filling missing state variables of a wind turbine

Through the random forest algorithm and Block Recurrent Transformer model, the problem of incomplete state variables caused by missing wind turbine data was solved, accurate state variable filling was achieved, and the research and analysis quality of wind turbines was improved.

CN116680567BActive Publication Date: 2025-09-12CSIC HAIZHUANG WINDPOWER CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310684853.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-09
Publication Date
2025-09-12
Estimated Expiration
2043-06-09

AI Technical Summary

Technical Problem

During the operation of wind turbines, the state variables are incomplete due to missing data, which affects the quality and results of research and analysis.

Method used

The random forest algorithm is used to determine the feature input variables from the set of candidate feature variables. The Block Recurrent Transformer model is used to train the missing data filling model, and the filling accuracy is evaluated using a test sample set. If the standard accuracy is met, state variable filling is performed. Otherwise, the model is adjusted using the Akaike information criterion and finite difference regression vector clustering.

Benefits of technology

It achieves accurate filling of missing state variables of wind turbines, ensures the completeness and accuracy of research and analysis, and improves the state detection capability of wind turbines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116680567B_ABST
    Figure CN116680567B_ABST
Patent Text Reader

Abstract

The present application provides a method and apparatus for filling missing state variables of a wind turbine, wherein a random forest algorithm is used to determine at least one feature input variable from a set of candidate feature variables; a first initial model for filling missing data of a wind power generation state is trained using an overall training sample set to obtain a first target model; the first target model is tested using an overall test sample set to obtain the missing data filling accuracy of the first target model; a determination is made as to whether the missing data filling accuracy of the first target model exceeds a standard accuracy; if so, a reference state variable of a target wind turbine is input into the first target model to obtain a predicted value of the state variable to be filled; and the predicted value of the state variable to be filled is filled into the state variable to obtain complete predicted state information of the target wind turbine. The above method is used to fill missing state variables of a wind turbine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wind turbine state management, and in particular to a method and device for filling missing state variables of a wind turbine. Background Art

[0002] The wind power industry has developed rapidly in recent years, with its scale growing annually. To improve power quality and wind turbine utilization efficiency, various research and analysis methods are required for wind turbine state variables and data collected by data acquisition and monitoring control systems, such as wind power forecasting and wind turbine performance evaluation. However, during actual wind turbine operation, recorded data is inevitably missing due to various human factors, extreme weather conditions, and instrument failures.

[0003] However, the research found that the missing state variables of wind turbines undermine the integrity of the wind turbine state data, directly affecting the quality and results of various research and analysis of wind turbine state variables and data, making it impossible to properly detect the wind turbine status. Therefore, how to fill the missing state variables of wind turbines has become an urgent problem to be solved. Summary of the Invention

[0004] In view of this, an object of the present invention is to provide a method and apparatus for filling missing state variables of a wind turbine, so as to fill missing state variables of the wind turbine.

[0005] In a first aspect, an embodiment of the present application provides a method for filling missing state variables of a wind turbine, the method comprising:

[0006] Determine at least one feature input variable from a set of candidate feature variables using a random forest algorithm based on pre-configured feature output variables, wherein the set of candidate feature variables includes at least one state feature variable of a target wind turbine, and at least one of the feature input variables is a strongly correlated variable of the feature output variable;

[0007] A first target model is obtained by performing model training on a first initial model for filling missing data of a wind power generation state using the overall training sample set, wherein the training output samples in the overall training sample set are feature output variables that meet a preset number of at least one feature output variable, and the training input samples in the overall training sample set are strongly correlated variables of the training output samples in the overall training sample set;

[0008] The first target model is tested using the entire test sample set to obtain a missing data filling accuracy rate of the first target model, wherein the test output samples in the entire test sample set are the remaining feature output variables of at least one feature output variable except the training output samples, and the test input samples in the entire test sample set are strongly correlated variables of the test output samples in the entire test sample set;

[0009] Determining whether the missing data filling accuracy of the first target model exceeds the standard accuracy;

[0010] If the missing data filling accuracy of the first target model exceeds the standard accuracy, the reference state variable of the target wind turbine is input into the first target model to obtain a predicted value of the state variable to be filled in the target wind turbine, wherein the reference state variable is a characteristic variable that is not missing in the state characteristic variables of the target wind turbine, and the state variable to be filled is a characteristic variable that is missing in the state characteristic variables of the target wind turbine;

[0011] The predicted value of the state variable to be filled is filled into the state variable to be filled to obtain complete predicted state information of the target wind turbine.

[0012] Optionally, after determining whether the missing data filling accuracy of the first target model exceeds a standard accuracy, the method further includes:

[0013] If the missing data filling accuracy of the first target model does not exceed the standard accuracy, determining the delay order between at least one of the feature input variables and the feature output variable using the Akaike information criterion;

[0014] Dividing the scope of at least one of the characteristic input variables according to the delay order to obtain at least one group of characteristic input variable groups, wherein the characteristic input variables in each group of the characteristic input variable groups all fall within the same variable interval;

[0015] For each set of the feature input variable groups, each feature input variable in the feature input variable group is used as a model input of a second initial model, a strongly correlated variable of each feature input variable in the feature input variable group is used as a model output of the second initial model, and the second initial model is trained to obtain a second target model obtained by training the feature input variable group;

[0016] The to-be-filled state variables of the target wind turbine are filled using a third target model, wherein the third target model is a second target model trained using a target feature input variable group, the target feature input variable group is a feature input variable group whose included feature input variables fall within a target interval, and the target interval is a variable interval within which strongly correlated variables of the missing state data of the target wind turbine fall.

[0017] Optionally, the step of scoping at least one of the characteristic input variables according to the delay order to obtain at least one group of characteristic input variable groups includes:

[0018] Establishing a finite difference regression vector of at least one of the characteristic input variables according to the delay order;

[0019] At least one of the feature input variables is scoped according to clustering of a finite difference regression vector of at least one of the feature input variables to obtain at least one group of feature input variable groups.

[0020] Optionally, after filling the predicted value of the to-be-filled state variable into the to-be-filled state variable to obtain complete predicted state information of the target wind turbine, the method further includes:

[0021] For each of the reference state variables, calculating a first autocorrelation coefficient between the reference state variable and a predicted value of the state variable to be filled for filling the reference state variable;

[0022] Calculating a second autocorrelation coefficient between the reference state variable and the observed value of the state variable to be filled in for filling in the reference state variable;

[0023] Determining whether a difference between the first autocorrelation coefficient and the second autocorrelation coefficient is within a preset threshold range;

[0024] If the difference between the first autocorrelation coefficient and the second autocorrelation coefficient is within a preset threshold range, the complete predicted state information obtained after filling the state variable to be filled is marked as filling successful.

[0025] Optionally, after determining whether the difference between the first autocorrelation coefficient and the second autocorrelation coefficient is within a preset threshold range, the method further includes:

[0026] If the difference between the first autocorrelation coefficient and the second autocorrelation coefficient is not within a preset threshold range, the complete predicted state information obtained after filling the state variable to be filled is marked as filling failure.

[0027] Optionally, after marking the complete predicted state information obtained after filling the state variable to be filled as successfully filled, the method further includes:

[0028] The model prediction effect of the first target model is evaluated according to the filling rate of the first target model, the filling accuracy of the first target model and the filling time of the first target model, wherein the filling rate of the first target model is the ratio of the number of complete prediction state information marked as successfully filled after filling using the first target model to the number of all complete prediction state information, the filling accuracy of the first target model is the difference between the complete prediction state information marked as successfully filled after filling using the first target model and the complete simulation state variables, the complete simulation state variables are the complete state variables of the target wind turbine obtained by the simulation algorithm, and the filling time of the first target model is the filling duration of a single complete prediction state information when filling using the first target model.

[0029] Optionally, the evaluating the model prediction effect of the first target model according to the filling rate of the first target model, the filling accuracy of the first target model, and the filling time of the first target model includes:

[0030] Performing density clustering using the DBSCAN clustering algorithm based on the filling rate of the first target model, the filling accuracy of the first target model, and the filling time of the first target model to determine outliers in the complete prediction state information marked as successfully filled after filling using the first target model;

[0031] Determining whether the number of abnormal values ​​exceeds a preset number;

[0032] If the number of the abnormal values ​​does not exceed the preset number, the model prediction effect of the first target model is evaluated as qualified;

[0033] If the number of the outliers exceeds a preset number, the model prediction effect of the first target model is evaluated as unqualified.

[0034] In a second aspect, an embodiment of the present application provides a device for filling missing state variables of a wind turbine, the device comprising:

[0035] a feature input variable determination module, configured to determine at least one feature input variable from a set of candidate feature variables using a random forest algorithm based on pre-configured feature output variables, wherein the set of candidate feature variables includes at least one state feature variable of a target wind turbine, and at least one of the feature input variables is a strongly correlated variable of the feature output variable;

[0036] a first target model determination module, configured to perform model training on a first initial model for filling missing data of a wind power generation state using the overall training sample set to obtain a first target model, wherein the training output samples in the overall training sample set are feature output variables that meet a preset number of at least one feature output variable, and the training input samples in the overall training sample set are strongly correlated variables of the training output samples in the overall training sample set;

[0037] a missing data filling accuracy determination module, configured to test the first target model using the entire test sample set to obtain the missing data filling accuracy of the first target model, wherein the test output samples in the entire test sample set are the remaining feature output variables of at least one feature output variable other than the training output samples, and the test input samples in the entire test sample set are strongly correlated variables of the test output samples in the entire test sample set;

[0038] A first judgment module is used to judge whether the missing data filling accuracy of the first target model exceeds the standard accuracy;

[0039] a predicted value determination module, configured to input the reference state variables of the target wind turbine into the first target model to obtain predicted values ​​of the to-be-filled state variables of the target wind turbine if the missing data filling accuracy of the first target model exceeds a standard accuracy, wherein the reference state variables are characteristic variables that are not missing in the state characteristic variables of the target wind turbine, and the to-be-filled state variables are characteristic variables that are missing in the state characteristic variables of the target wind turbine;

[0040] The complete predicted state information determination module is configured to fill the predicted value of the state variable to be filled into the state variable to be filled to obtain the complete predicted state information of the target wind turbine.

[0041] Optionally, the device further comprises:

[0042] a delay order determination module, configured to, after the first determination module determines whether the missing data filling accuracy of the first target model exceeds a standard accuracy rate, determine a delay order between at least one of the feature input variables and the feature output variable using the Akaike information criterion if the missing data filling accuracy of the first target model does not exceed the standard accuracy rate;

[0043] a characteristic input variable group determination module, configured to divide the scope of at least one of the characteristic input variables according to the delay order to obtain at least one group of characteristic input variable groups, wherein the characteristic input variables in each group of the characteristic input variable groups all fall within the same variable interval;

[0044] a second target model determination module configured to, for each set of the feature input variable groups, use each feature input variable in the feature input variable group as a model input of a second initial model, use a strongly correlated variable of each feature input variable in the feature input variable group as a model output of the second initial model, and perform model training on the second initial model to obtain a second target model obtained by training the feature input variable group;

[0045] The module for filling in the state variables to be filled is used to fill in the state variables to be filled in of the target wind turbine using a third target model, wherein the third target model is a second target model trained using a target feature input variable group, the target feature input variable group is a feature input variable group whose feature input variables fall within a target interval, and the target interval is a variable interval within which the strongly correlated variables of the missing state data of the target wind turbine fall.

[0046] Optionally, when the characteristic input variable group determination module is used to scope-divide at least one characteristic input variable according to the delay order to obtain at least one group of characteristic input variable groups, it is specifically used to:

[0047] Establishing a finite difference regression vector of at least one of the characteristic input variables according to the delay order;

[0048] At least one of the feature input variables is scoped according to clustering of a finite difference regression vector of at least one of the feature input variables to obtain at least one group of feature input variable groups.

[0049] Optionally, the device further comprises:

[0050] a first autocorrelation coefficient determining module configured to calculate, for each reference state variable, a first autocorrelation coefficient between the reference state variable and the predicted value of the state variable to be filled in, after the complete predicted state information determining module fills the predicted value of the state variable to be filled in to obtain the complete predicted state information of the target wind turbine;

[0051] A second autocorrelation coefficient determination module is used to calculate a second autocorrelation coefficient between the reference state variable and the observed value of the state variable to be filled for filling the reference state variable;

[0052] A second judging module is configured to judge whether a difference between the first autocorrelation coefficient and the second autocorrelation coefficient is within a preset threshold range;

[0053] The first marking module is configured to mark the complete predicted state information obtained after filling the state variable to be filled as successfully filled if the difference between the first autocorrelation coefficient and the second autocorrelation coefficient is within a preset threshold range.

[0054] Optionally, the device further comprises:

[0055] A second marking module is used to mark the complete prediction state information obtained after filling the state variable to be filled as a filling failure if the difference between the first autocorrelation coefficient and the second autocorrelation coefficient is not within the preset threshold range after the second judgment module determines whether the difference between the first autocorrelation coefficient and the second autocorrelation coefficient is within the preset threshold range.

[0056] Optionally, the device further comprises:

[0057] A model prediction effect evaluation module is used to evaluate the model prediction effect of the first target model according to the filling rate of the first target model, the filling accuracy of the first target model and the filling time of the first target model after the first marking module marks the complete prediction state information obtained after filling the state variable to be filled as successfully filled, wherein the filling rate of the first target model is the ratio of the number of complete prediction state information marked as successfully filled after filling using the first target model to the number of all complete prediction state information, the filling accuracy of the first target model is the difference between the complete prediction state information marked as successfully filled after filling using the first target model and the complete simulation state variable, the complete simulation state variable is the complete state variable of the target wind turbine obtained by the simulation algorithm, and the filling time of the first target model is the filling time of a single complete prediction state information when filling using the first target model.

[0058] Optionally, when the model prediction effect evaluation module is used to evaluate the model prediction effect of the first target model based on the filling rate of the first target model, the filling accuracy of the first target model, and the filling time of the first target model, it is specifically used to:

[0059] Performing density clustering using the DBSCAN clustering algorithm based on the filling rate of the first target model, the filling accuracy of the first target model, and the filling time of the first target model to determine outliers in the complete prediction state information marked as successfully filled after filling using the first target model;

[0060] Determining whether the number of abnormal values ​​exceeds a preset number;

[0061] If the number of the abnormal values ​​does not exceed the preset number, the model prediction effect of the first target model is evaluated as qualified;

[0062] If the number of the outliers exceeds a preset number, the model prediction effect of the first target model is evaluated as unqualified.

[0063] The technical solutions provided by this application include but are not limited to the following beneficial effects:

[0064] Based on the pre-configured characteristic output variables, a random forest algorithm is used to determine at least one characteristic input variable from a set of candidate characteristic variables, wherein the set of candidate characteristic variables includes at least one state characteristic variable of the target wind turbine, and at least one of the characteristic input variables is a strongly correlated variable of the characteristic output variable. Through the above steps, a characteristic input variable having a strong correlation with the characteristic output variable can be obtained to provide training samples for subsequent model training.

[0065] The first initial model for filling missing data of wind power generation status is trained using the overall training sample set to obtain a first target model, wherein the training output samples in the overall training sample set are feature output variables that meet a preset number of at least one feature output variable, and the training input samples in the overall training sample set are strongly correlated variables of the training output samples in the overall training sample set; through the above steps, the first initial model can be trained according to the feature output variables and the strongly correlated variables of the feature output variables to obtain the first target model for filling missing data of wind power generation status.

[0066] The first target model is tested using the overall test sample set to obtain the missing data filling accuracy of the first target model, wherein the test output samples in the overall test sample set are the remaining feature output variables in at least one of the feature output variables except the training output samples, and the test input samples in the overall test sample set are strongly correlated variables of the test output samples in the overall test sample set; it is determined whether the missing data filling accuracy of the first target model exceeds the standard accuracy; through the above steps, the prediction accuracy of the trained first target model can be tested to determine whether the first target model can be used to perform actual prediction value prediction.

[0067] If the missing data filling accuracy of the first target model exceeds the standard accuracy, the reference state variable of the target wind turbine is input into the first target model to obtain the predicted value of the state variable to be filled of the target wind turbine, wherein the reference state variable is a characteristic variable that is not missing in the state characteristic variables of the target wind turbine, and the state variable to be filled is a characteristic variable that is missing in the state characteristic variables of the target wind turbine; through the above steps, the predicted value of the state variable to be filled can be determined using the first target model whose prediction accuracy meets the requirements.

[0068] The predicted value of the state variable to be filled is filled into the state variable to be filled to obtain the complete predicted state information of the target wind turbine; through the above steps, the predicted value of the state variable to be filled can be used to fill the state variable to be filled to obtain the complete predicted state information of the target wind turbine.

[0069] Using the above method, after using the random forest algorithm to determine the training input samples and training output samples for model training, the first initial model is trained using the training sample set consisting of the training input samples and the training output samples to obtain a first target model, and then the first target model is tested for prediction accuracy using the test sample set consisting of the test input samples and the test output samples. When the test passes, the first target model is used to determine the predicted values ​​of the to-be-filled state variables of the target wind turbine, and then the predicted values ​​of the to-be-filled state variables are used to fill in the to-be-filled state variables to obtain complete predicted state information of the target wind turbine, so as to fill in the missing state variables of the wind turbine.

[0070] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0072] Figure 1 A flow chart of a method for filling missing state variables of a wind turbine provided by the first embodiment of the present invention is shown;

[0073] Figure 2 A flow chart showing a second method for filling missing state variables of a wind turbine provided by the first embodiment of the present invention is shown;

[0074] Figure 3 A flowchart of a method for determining a feature input variable group provided by the first embodiment of the present invention is shown;

[0075] Figure 4 A flowchart of a complete prediction state information marking method provided by the first embodiment of the present invention is shown;

[0076] Figure 5 A flowchart of a first target model prediction effect evaluation method provided by the first embodiment of the present invention is shown;

[0077] Figure 6 A device for filling missing state variables of a wind turbine provided by a second embodiment of the present invention is shown. DETAILED DESCRIPTION

[0078] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.

[0079] Example 1

[0080] To facilitate understanding of this application, Figure 1 The flowchart of a method for filling missing state variables of a wind turbine provided by the first embodiment of the present invention is shown to describe the content of the first embodiment of the present application in detail.

[0081] See also Figure 1 As shown, Figure 1 A flowchart of a method for filling missing state variables of a wind turbine provided by a first embodiment of the present invention is shown, wherein the method comprises steps S101 to S106:

[0082] S101: Based on pre-configured characteristic output variables, use a random forest algorithm to determine at least one characteristic input variable from a set of candidate characteristic variables, wherein the set of candidate characteristic variables includes at least one state characteristic variable of a target wind turbine, and at least one of the characteristic input variables is a strongly correlated variable of the characteristic output variable.

[0083] Specifically, the state characteristic variables are the state characteristic data of the target wind turbine, including but not limited to output power, wind turbine speed, nacelle temperature, main bearing temperature, generator bearing temperature, gearbox temperature and other operating parameters of the wind turbine during operation, as well as its inherent equipment properties.

[0084] A random forest algorithm is used to select at least one characteristic input variable that is strongly correlated with the pre-configured characteristic output variable from the set of candidate characteristic variables (the state characteristic variables include environmental quantities: tower base temperature, ambient temperature, instantaneous wind speed, instantaneous wind direction; mechanical quantities: wind wheel speed, generator speed, blade pitch angle, gearbox oil pump outlet pressure, gearbox inlet oil pressure; electrical quantities: grid three-phase voltage, grid three-phase current, active power; temperature quantities: gearbox inlet oil temperature, gearbox oil pool temperature, hub temperature, tower base control cabinet temperature).

[0085] Random forest is a set of decision tree classifiers {h(X,θ k ),k=1,2,3,…,K}, where θ k is a random vector that obeys independent and identical distribution, K represents the number of decision trees in the random forest, and under a given independent variable X, each decision tree classifier determines the optimal classification result by voting.

[0086] Given a set of classifiers h1(X),h2(X),…,h k (X), the training set of each classifier is randomly sampled from the original dataset (,X) that follows a random distribution, and the residual function mg(,Y) is defined as

[0087]

[0088] Where I(·) is the characteristic function, av k (·) represents the average, and j is the upper bound of the generalization error. The margin function measures the extent to which the average number of correct classifications exceeds the average number of incorrect classifications. The larger the margin value, the more reliable the classification prediction.

[0089] Generalization Error PE * Defined as:

[0090] PE * = X,Y ((X,Y)<0)

[0091] Where, P X,Y is the probability of covering the X,Y space, and the subscript X,Y indicates the probability P covers the X,Y space.

[0092] The classification accuracy of the random forest algorithm is defined as:

[0093]

[0094] In the formula, TP (true positive) represents true positive; TN (true negative) represents true negative; FP (false positive) represents false positive; FN (false negative) represents false negative.

[0095] The random forest-based wrapper feature selection method (RFFS) is adopted. The variable importance measure of the random forest algorithm is used to sort the features. Then, the sequential backward search method is used to remove the least important feature (the one with the smallest importance score) from the feature set each time. The iterations are performed successively and the classification accuracy is calculated. Finally, the feature set with the least number of variables and the highest classification accuracy is obtained as the feature selection result.

[0096] Setting the number of spanning trees to 300 and varying the number of feature parameters, we conducted comparative analysis for 33, 23, and 13 feature parameters, respectively. The results are shown in the table below. The following six features—wind speed, power, generator speed, ambient temperature, gearbox oil temperature, and generator stator coil temperature—have the highest overall evaluation of nacelle temperature, main bearing temperature, generator bearing temperature, and gearbox temperature. Their importance shows no sign of diminishing as the features are filtered. While the importance ranking of the top six features shifts slightly, the overall trend remains unchanged. Wind speed, in particular, has the greatest impact.

[0097] The importance scores of state characteristic parameters with high correlation with dominant characteristic variables are shown in the following table:

[0098] State parameters Importance Rating wind speed 0.876651 power 0.801122 Generator speed 0.695477 Ambient temperature 0.446921 Gearbox oil temperature 0.412265 Generator stator coil temperature 0.398651

[0099] S102: Using the overall training sample set, a first initial model for filling missing data of wind power generation status is trained to obtain a first target model, wherein the training output samples in the overall training sample set are feature output variables that meet a preset number of at least one feature output variable, and the training input samples in the overall training sample set are strongly correlated variables of the training output samples in the overall training sample set.

[0100] Specifically, the first initial model is the Block Recurrent Transformer multi-input multi-output model (BRT). BRT is a deep learning model that combines the Transformer and RNN structures. It can capture both long-term and local dependencies, and utilizes the Transformer's self-attention mechanism, with excellent modeling capabilities and computational efficiency. In BRT, the Transformer Encoder is mainly responsible for capturing global dependencies in the input sequence, while the RNN is mainly responsible for capturing local dependencies in the input sequence. By combining the Transformer and RNN structures, BRT can capture both long-term and local dependencies, thereby better modeling sequence data. In addition, BRT also utilizes the Transformer's self-attention mechanism, which can calculate the representation of each element in the sequence without traversing the entire sequence, thereby improving the computational efficiency of the model.

[0101] In the actual modeling process, data preprocessing is first required. The feature output variables and feature input variables obtained in step S101 are normalized in a high-dimensional space. A preset number (or a preset percentage, for example, 70%) of the feature output variables are randomly selected as training output samples (the remaining are used for subsequent model validation). The first initial model is trained using variables that are strongly correlated with the training output samples obtained using the random forest variables in step S101 as training input samples from the overall training sample set. The training process and parameters are designed based on the structure of the BRT neural network. Specifically, the BRT consists of multiple blocks, each containing a Transformer Encoder and a RNN. In each block, the input sequence first passes through the Transformer Encoder to obtain a new representation. This new representation is then input into the RNN, which maintains a state vector internally and combines the new representation with the previous state vector to obtain a new state vector. This new state vector is then passed to the next block to process the next input sequence. Therefore, it is necessary to adjust the number of blocks and embedding dimensions based on the neural network parameter settings.

[0102] S103: Use the overall test sample set to test the first target model to obtain the missing data filling accuracy of the first target model, wherein the test output samples in the overall test sample set are the remaining feature output variables of at least one of the feature output variables except the training output samples, and the test input samples in the overall test sample set are strongly correlated variables of the test output samples in the overall test sample set.

[0103] Specifically, the feature input variables and feature output variables not selected in step S102 are used as samples for model testing to form the overall test sample set. The overall test sample set is used to evaluate the performance of the model. A different test sample from the overall test sample set is selected for model testing each time. It is determined whether the result obtained after each test input sample is input to the model is the same as the test output sample. If they are the same, the test is marked as successful; if they are different, the test is marked as unsuccessful.

[0104] The number of successful tests is counted, and the ratio of the number of successful tests to the total number of tests is calculated, and the ratio is determined as the missing data filling accuracy of the first target model.

[0105] S104: Determine whether the missing data filling accuracy of the first target model exceeds the standard accuracy.

[0106] Specifically, a standard accuracy rate is pre-configured according to user needs and is used to evaluate the model prediction accuracy rate of the first target model.

[0107] S105: If the missing data filling accuracy of the first target model exceeds the standard accuracy, the reference state variable of the target wind turbine is input into the first target model to obtain the predicted value of the state variable to be filled of the target wind turbine, wherein the reference state variable is the characteristic variable that is not missing in the state characteristic variables of the target wind turbine, and the state variable to be filled is the filling value of the missing characteristic variable in the state characteristic variables of the target wind turbine.

[0108] Specifically, if the missing data filling accuracy of the first target model exceeds the standard accuracy, it means that the prediction effect of the first target model meets the standard, meets user needs, and can be used for actual missing data filling. The first target model is then used to fill in the actual status data of the target wind turbine.

[0109] The specific usage process is: inputting the reference state variable of the target wind turbine into the first target model to obtain the predicted value of the state variable to be filled of the target wind turbine, wherein the reference state variable is the characteristic variable that is not missing in the state characteristic variables of the target wind turbine, and the state variable to be filled is the filling value of the missing characteristic variable in the state characteristic variables of the target wind turbine.

[0110] For example, the radiation amount and temperature value are non-missing characteristic input variables, the active power is a missing characteristic output variable, and the radiation amount and temperature value are strongly correlated variables with the active power. Then, the radiation amount and temperature value are input into the first target model to obtain the predicted value of the active power of the target wind turbine.

[0111] S106: Filling the predicted value of the state variable to be filled into the state variable to be filled to obtain complete predicted state information of the target wind turbine.

[0112] Specifically, filling the missing part with the predicted value can obtain complete prediction status information. For example, filling the predicted value of active power into the active power part can obtain the complete prediction status information of "radiation-temperature value-active power" of the target wind turbine that is not missing.

[0113] In one possible embodiment, see Figure 2 As shown, Figure 2 A flowchart of a second method for filling missing state variables of a wind turbine provided in the first embodiment of the present invention is shown. After determining whether the missing data filling accuracy of the first target model exceeds the standard accuracy, the method includes steps S201 to S206:

[0114] S201: If the missing data filling accuracy of the first target model does not exceed the standard accuracy, the delay order between at least one of the feature input variables and the feature output variable is determined using the Akaike information criterion.

[0115] S202: Dividing the scope of at least one of the characteristic input variables according to the delay order to obtain at least one group of characteristic input variable groups, wherein the characteristic input variables in each group of the characteristic input variable groups all fall into the same variable interval.

[0116] Specifically, if the missing data filling accuracy of the first target model does not exceed the standard accuracy, it means that the first target model cannot complete the missing data filling as required, so it is necessary to use differential dynamic regression vectors in each working domain to establish a multi-input and multi-output neural network model.

[0117] The Akaike Information Criterion (AIC) is used to determine the model delay order using a system identification method. AIC, based on the KL distance (relative entropy), provides a compromise between the complexity of the estimated model and the quality of the model fit to the data. The AIC value is defined as follows:

[0118]

[0119] Where AIC represents the model delay order, N is the number of values ​​in the estimation data set; ε(t, θ N ) is the prediction error vector; t represents the number of selected variable sequences; θ N represents the estimated parameter; n p is the number of estimated parameters; n y is the number of model outputs. The order n is determined by selecting the model with the smallest AIC value a and n b .

[0120] Considering the delay characteristics of input and output variables, a finite difference regression vector is established:

[0121] x(k)=[y T (k-1)y T (k-2)...y T (kn a )u T (k-1)…u T (kn b )]

[0122] In the formula, x(k) is the established finite difference regression vector, k is the independent variable, y(k) is the output, u(k) is the input, and n a and n b are the delay orders of input and output respectively, and T is the transposed sign. Considering the spatial distribution characteristics of the data points (x(k), y(k)) and the similarity between the parameter vectors, a feature vector is formed in each local data set, and the finite difference action space is divided according to the clustering of the feature vectors. Each data point (x(k), y(k)) is taken as a data center and a local data set C is established. k , C k Contains the data center (x(k), y(k)) and its adjacent (k-1) data points (x Ck (j), y Ck (j))(j=1,2,...,k-1), where (x Ck (j), y Ck (j)) represents the data points adjacent to the data center in the data set. Finally, according to C kCalculate the empirical covariance P for the data points in k The covariance matrix V k . And calculate the divergence matrix Q k To measure C k The intra-class dispersion of the data points in is as follows:

[0123]

[0124] Among them, M k C k The mean of the input vector for each data point in .

[0125] The eigenvector can be regarded as a random vector that obeys the Gaussian distribution. According to the characteristics of the Gaussian distribution, its variance R k It can be expressed as R k =[V k 0;0 Q k ], the eigenvector is taken as the mean value M k Confidence It can be measured by the formula:

[0126]

[0127] Where n = n a n y +n b n u , n a and n b are the input and output delay orders, n y and n u is the dimension of the output and input vectors.

[0128] The K-Means algorithm is used to cluster the eigenvectors of the local data subset and divide the data points corresponding to the local eigenvectors into S groups, representing S finite difference domains. Clustering belongs to unsupervised learning, and K-means clustering is the most basic and commonly used clustering algorithm. Its basic idea is to iteratively find a partitioning scheme for K clusters that minimizes the loss function corresponding to the clustering results. The loss function can be defined as the sum of squared errors J(c, μ) between each sample and the center point of the cluster to which it belongs:

[0129]

[0130] Where x i represents the i-th sample, c i is x i The cluster to which it belongs, Represents c i The center point of the cluster, M is the total number of samples.

[0131] The core goal of K-Means is to divide a given dataset into K clusters and give the center point corresponding to each sample data. The specific steps can be divided into 4 steps:

[0132] Data preprocessing, mainly standardization and outlier filtering;

[0133] Randomly select K centers and record them as

[0134] Define the loss function:

[0135]

[0136] Let t = 0, 1, 2, ... be the number of iterations, and repeat the following process until J(c, μ) converges;

[0137] For each sample x i , assign it to the nearest center

[0138]

[0139] in, is x i The cluster to which it belongs, k is the sequence number of the dataset, Represents the t-th iteration of the cluster centers of the k-th dataset.

[0140] For each class center k, recalculate the center of the class

[0141]

[0142] in, represents the t+1th iteration of the cluster center of the kth data set, and μ is the selected cluster center.

[0143] To clearly represent each scope, we studied hyperplane estimation between scopes and used support vector machines to classify the coefficients of each hyperplane equation. Because it's uncertain whether the data is completely linearly separable, we used a soft-margin support vector machine (SVM) to classify the coefficients of each hyperplane equation. This method exhibits better robustness and generalization capabilities than hard-margin SVMs.

[0144]

[0145] Where J is the hyperplane coefficient, x k is the independent variable, ζ k is the slack variable of each equation reflecting the degree to which the data does not meet the hard interval constraint, φ is the normal vector of the switching surface of different scopes, Tis the transposed matrix of φ; d is the offset; ζ is the slack variable that reflects the degree to which the data does not meet the hard interval constraint; γ represents the penalty coefficient that can be adjusted from 0 to 1; y k is the data classification label with values ​​of 1 and -1, defined as y k (x k )=sgn(φ T x k +d); m is the total amount of data; st represents the constraint condition.

[0146] S203: For each group of the feature input variable groups, each feature input variable in the feature input variable group is used as the model input of the second initial model, and the strongly correlated variables of each feature input variable in the feature input variable group are used as the model output of the second initial model. The second initial model is trained to obtain a second target model obtained through the training of the feature input variable group.

[0147] Specifically, since the feature input variables in each set of the feature input variable groups all fall within the same variable interval, for each set of feature input variable groups containing feature input variables that fall within the same variable interval, each feature input variable in the feature input variable group is used as a model input of the second initial model, and the strongly correlated variables of each feature input variable in the feature input variable group are used as the model output of the second initial model. The second initial model is then trained to obtain a second target model obtained by training with the feature input variable group. In other words, as many second target models can be obtained by training as there are feature input variable groups.

[0148] S204: Use the third target model to fill in the to-be-filled state variables of the target wind turbine, wherein the third target model is the second target model trained using the target feature input variable group, the target feature input variable group is a feature input variable group whose feature input variables fall within a target interval, and the target interval is the variable interval into which the strongly correlated variables of the missing state data of the target wind turbine fall.

[0149] Specifically, since the variable intervals of the feature input variables in the feature input variable groups used to train different second target models are different, different second target models are used to fill in the predicted values ​​of the to-be-filled state variables that fall into different variable intervals, that is, to predict the predicted values ​​of the to-be-filled state variables that fall into the variable intervals into which the feature input variables in the feature input variable groups used to train them fall.

[0150] In one possible embodiment, see Figure 3 As shown, Figure 3A flowchart of a method for determining a characteristic input variable group provided by a first embodiment of the present invention is shown, wherein the method of dividing the scope of at least one characteristic input variable according to the delay order to obtain at least one characteristic input variable group includes steps S301 to S302:

[0151] S301: Establishing a finite difference regression vector of at least one of the characteristic input variables according to the delay order.

[0152] S302: Scope-dividing at least one of the feature input variables according to clustering of a finite difference regression vector of at least one of the feature input variables to obtain at least one group of feature input variable groups.

[0153] In one possible embodiment, see Figure 4 As shown, Figure 4 A flowchart of a method for marking complete predicted state information provided by the first embodiment of the present invention is shown. After the predicted value of the to-be-filled state variable is filled into the to-be-filled state variable to obtain the complete predicted state information of the target wind turbine, the method further includes steps S401 to S404:

[0154] S401: For each reference state variable, calculate a first autocorrelation coefficient between the reference state variable and a predicted value of a state variable to be filled for filling the reference state variable.

[0155] Specifically, step S105 can obtain the predicted value of the state variable to be filled of the target wind turbine, and for each reference state variable, calculate the first autocorrelation coefficient between the reference state variable and the predicted value of the state variable to be filled.

[0156] S402: Calculate a second autocorrelation coefficient between the reference state variable and the observed value of the state variable to be filled for filling the reference state variable.

[0157] Specifically, the observation values ​​of the reference state variable and the state variable to be filled are obtained through the sensor or data acquisition system on the wind turbine, and then the second autocorrelation coefficient between the reference state variable and the observation values ​​of the state variable to be filled is calculated.

[0158] Among them, the calculation formula of the autocorrelation coefficient acf(x) is as follows

[0159]

[0160] Among them, N is the sequence length, k is the sequence interval, t is the number of selected variable sequences, and x t For the selected variables, is the mean of the complete series, x t-k is the reference variable to be substituted into the calculation.

[0161] S403: Determine whether the difference between the first autocorrelation coefficient and the second autocorrelation coefficient is within a preset threshold range.

[0162] Specifically, it is determined whether the first autocorrelation coefficient is higher than the second autocorrelation coefficient.

[0163] S404: If the difference between the first autocorrelation coefficient and the second autocorrelation coefficient is within a preset threshold range, the complete predicted state information obtained after filling the state variable to be filled is marked as successfully filled.

[0164] Specifically, if the first autocorrelation coefficient increases compared to the second autocorrelation coefficient, the complete predicted state information obtained after filling the state variable to be filled is marked as successful filling.

[0165] In a feasible embodiment, after determining whether the difference between the first autocorrelation coefficient and the second autocorrelation coefficient is within a preset threshold range, the method further includes:

[0166] If the difference between the first autocorrelation coefficient and the second autocorrelation coefficient is not within a preset threshold range, the complete predicted state information obtained after filling the state variable to be filled is marked as filling failure.

[0167] Specifically, if the first autocorrelation coefficient decreases compared to the second autocorrelation coefficient, the complete predicted state information obtained after filling the state variable to be filled is marked as a filling failure.

[0168] In a feasible embodiment, after marking the complete predicted state information obtained after filling the state variable to be filled as filling successful, the method further includes:

[0169] The model prediction effect of the first target model is evaluated according to the filling rate of the first target model, the filling accuracy of the first target model and the filling time of the first target model, wherein the filling rate of the first target model is the ratio of the number of complete prediction state information marked as successfully filled after filling using the first target model to the number of all complete prediction state information, the filling accuracy of the first target model is the difference between the complete prediction state information marked as successfully filled after filling using the first target model and the complete simulation state variables, the complete simulation state variables are the complete state variables of the target wind turbine obtained by the simulation algorithm, and the filling time of the first target model is the filling duration of a single complete prediction state information when filling using the first target model.

[0170] Specifically, the filling rate is defined as the ratio of the number of complete prediction state information marked as successfully filled to the total number of complete prediction state information. This indicator evaluates the applicability of the filling algorithm and is calculated according to the following formula:

[0171]

[0172] Where, PCE (%) is the filling rate, N SI The number of complete prediction status information marked as filled successfully, N I Indicates the total number of complete prediction status information.

[0173] The evaluation methods of filling accuracy are divided into: direct evaluation method and classification performance evaluation method.

[0174] Direct evaluation methods evaluate the accuracy of missing data by calculating the difference between the original values ​​of the virtual missing data in the simulated dataset and the estimated values ​​obtained by the missing data filling algorithm. For discrete data, the correct prediction ratio is often used as an evaluation metric, calculated using the following formula.

[0175]

[0176] Where, PCE (%) is the filling rate, N CE N is the number of predicted values ​​of the state variables to be filled used by the complete predicted state information marked as successfully filled, E Indicates the total number of predicted values ​​of the state variables to be filled.

[0177] For continuous data, the root mean square error is usually used as the evaluation indicator and is calculated using the following formula.

[0178]

[0179] Where RMSE is the root mean square error, NMV Represents the total number of state variables to be filled, V A V represents the observed value of the state variable to be filled. E Represents the predicted value of the state variable to be filled.

[0180] The mean absolute error ratio can also be used as an evaluation indicator and calculated using the following formula.

[0181]

[0182] Where, MAPE is the mean absolute error rate, N MV Represents the total number of state variables to be filled, V A V represents the observed value of the state variable to be filled. E Represents the predicted value of the state variable to be filled.

[0183] The classification performance evaluation method is a filling algorithm performance evaluation method combined with subsequent data application, which is suitable for supporting

[0184] A missing data imputation algorithm for datasets used in classification applications. First, the original dataset is virtually imputed with missing data to create an incomplete dataset. The resulting complete dataset, processed by the missing data imputation algorithm, is then divided into a training set and a test set. The training set is used to train the classifier, while the test set is used to test the classifier's performance. By evaluating the classification performance of different missing data imputation algorithms using the same classifier and comparing the performance of the classifier on the original complete dataset, a performance metric for the missing data imputation algorithm can be obtained.

[0185] Filling time is the time it takes to fill a single missing data point in a sensor network dataset, from the start to the end. The specific calculation method for this time varies depending on the filling algorithm. In actual measurements, since the filling time for a single missing data point is often relatively short, the sum of the filling times for multiple missing data points is typically measured, and the average time is used to evaluate time performance.

[0186] In one possible embodiment, see Figure 5 As shown, Figure 5 A flowchart of a method for evaluating the prediction effect of a first target model provided by the first embodiment of the present invention is shown, wherein the method for evaluating the model prediction effect of the first target model based on the filling rate of the first target model, the filling accuracy of the first target model, and the filling time of the first target model includes steps S501 to S504:

[0187] S501: Density clustering is performed using the DBSCAN clustering algorithm based on the filling rate of the first target model, the filling accuracy of the first target model, and the filling time of the first target model to determine the outliers in the complete prediction status information marked as successfully filled after filling using the first target model.

[0188] Specifically, the three indicators of filling rate, filling accuracy, and filling time are density clustered in three-dimensional space using DBSCAN. The DBSCAN algorithm is an unsupervised machine learning method based on density clustering. It does not require the number of clusters to be set in advance. It uses two important parameters, radius ε and neighborhood density threshold Z, to reflect the compactness of data distribution and identify clusters with irregular shapes. When determining a cluster, the DBSCAN algorithm searches the neighborhood of each test object in the dataset. If the number of objects contained in the neighborhood exceeds the neighborhood density threshold Z, a new cluster with the object as the core is created. The algorithm then starts from the core object, finds all density-reachable objects, and merges them into a cluster. The process terminates when no points in the dataset can be added to any cluster. Points that do not fall into any cluster are considered outliers.

[0189] S502: Determine whether the number of abnormal values ​​exceeds a preset number.

[0190] S503: If the number of the abnormal values ​​does not exceed the preset number, the model prediction effect of the first target model is evaluated as qualified.

[0191] S504: If the number of the abnormal values ​​exceeds a preset number, the model prediction effect of the first target model is evaluated as unqualified.

[0192] Specifically, when the outliers do not reach a certain proportion (do not exceed a preset number), it proves that the BRT neural network is better at filling the multivariate heterogeneous abnormal data, and the model prediction effect of the first target model is evaluated as qualified; conversely, if the number of the outliers reaches a certain proportion (exceeds a preset number), the model prediction effect of the first target model is evaluated as unqualified, and the BRT neural network parameters need to be re-determined.

[0193] Furthermore, our proposed method for filling missing state variables in wind turbines can be applied to edge-based intelligent sensing devices. Edge-based intelligent sensing devices process and analyze data close to the data source, sending the processed data to the cloud or other systems for further processing. In offshore wind power projects, applying this data filling method to edge-based sensing devices can improve system efficiency and reliability.

[0194] In offshore wind power projects, edge intelligent sensing devices can be installed on wind turbines to collect various sensor data, including temperature, wind speed, vibration, and other information. The proposed missing value filling algorithm is integrated into the modules of the edge intelligent sensing device, including an outlier identification module, a random forest feature selection module, a finite difference regression vector workspace partitioning module, a BRT neural network multi-input and multi-output modeling module, and a filling effect evaluation module. This covers the main functions, including data analysis and processing, and missing value filling. This can significantly reduce the interaction content between edge servers and central servers, and reduce computing and transmission time while ensuring data quality.

[0195] The present invention proposes a method for filling missing state variables in a wind turbine. This method targets abnormal data from a wind turbine generator set. After selecting characteristic variables and dividing similar operating conditions, it uses normal data to train a neural network model. The missing values ​​are effectively predicted and filled using a neural network algorithm, effectively filling in abnormal data in wind power generation. This method effectively fills in abnormal data in wind power generation, making wind power generation operation control more precise. The proposed BRT neural network missing value filling method can reduce filling errors caused by data correlation. Compared to general data filling methods, the multi-input and multi-output model established by the BRT neural network can process multiple variables simultaneously and use the relationships between different variables to fill in missing values. This enables BRT to perform better in data filling tasks. The adaptive DBSCAN density clustering method is used to evaluate the filling effect of the model. When the model filling effect is poor, the model parameters can be automatically changed to achieve a better filling effect.

[0196] Example 2

[0197] See also Figure 6 As shown, Figure 6 The present invention shows a device for filling missing state variables of a wind turbine provided by the second embodiment of the present invention, wherein the device includes

[0198] A feature input variable determination module 601 is configured to determine at least one feature input variable from a set of candidate feature variables using a random forest algorithm based on pre-configured feature output variables, wherein the set of candidate feature variables includes at least one state feature variable of a target wind turbine, and at least one of the feature input variables is a strongly correlated variable of the feature output variable;

[0199] A first target model determination module 602 is configured to perform model training on a first initial model for filling missing data of a wind power generation state using the entire training sample set to obtain a first target model, wherein the training output samples in the entire training sample set are feature output variables that meet a preset number of at least one feature output variable, and the training input samples in the entire training sample set are strongly correlated variables of the training output samples in the entire training sample set;

[0200] a missing data filling accuracy determination module 603, configured to test the first target model using the entire test sample set to obtain the missing data filling accuracy of the first target model, wherein the test output samples in the entire test sample set are the remaining feature output variables of at least one feature output variable except the training output samples, and the test input samples in the entire test sample set are strongly correlated variables of the test output samples in the entire test sample set;

[0201] A first judgment module 604 is configured to judge whether the missing data filling accuracy of the first target model exceeds a standard accuracy;

[0202] a predicted value determination module 605 configured to input the reference state variables of the target wind turbine into the first target model to obtain predicted values ​​of the to-be-filled state variables of the target wind turbine if the missing data filling accuracy of the first target model exceeds a standard accuracy, wherein the reference state variables are characteristic variables that are not missing in the state characteristic variables of the target wind turbine, and the to-be-filled state variables are filled values ​​of the missing characteristic variables in the state characteristic variables of the target wind turbine;

[0203] The complete predicted state information determining module 606 is configured to fill the predicted value of the state variable to be filled into the state variable to be filled to obtain the complete predicted state information of the target wind turbine.

[0204] In one feasible embodiment, the device further comprises:

[0205] a delay order determining module, configured to, after determining whether the missing data filling accuracy of the first target model exceeds a standard accuracy rate, determine a delay order between at least one of the feature input variables and the feature output variable using an Akaike information criterion if the missing data filling accuracy of the first target model does not exceed the standard accuracy rate;

[0206] a characteristic input variable group determination module, configured to divide the scope of at least one of the characteristic input variables according to the delay order to obtain at least one group of characteristic input variable groups, wherein the characteristic input variables in each group of the characteristic input variable groups all fall within the same variable interval;

[0207] a second target model determination module configured to, for each set of the feature input variable groups, use each feature input variable in the feature input variable group as a model input of a second initial model, use a strongly correlated variable of each feature input variable in the feature input variable group as a model output of the second initial model, and perform model training on the second initial model to obtain a second target model obtained by training the feature input variable group;

[0208] The module for filling in the state variables to be filled is used to fill in the state variables to be filled in of the target wind turbine using a third target model, wherein the third target model is a second target model trained using a target feature input variable group, the target feature input variable group is a feature input variable group whose feature input variables fall within a target interval, and the target interval is a variable interval within which the strongly correlated variables of the missing state data of the target wind turbine fall.

[0209] In a feasible implementation manner, the characteristic input variable group determination module, when used to scope at least one of the characteristic input variables according to the delay order to obtain at least one group of characteristic input variable groups, is specifically used to:

[0210] Establishing a finite difference regression vector of at least one of the characteristic input variables according to the delay order;

[0211] At least one of the feature input variables is scoped according to clustering of a finite difference regression vector of at least one of the feature input variables to obtain at least one group of feature input variable groups.

[0212] In one feasible embodiment, the device further comprises:

[0213] a first autocorrelation coefficient determining module configured to calculate, for each reference state variable, a first autocorrelation coefficient between the reference state variable and the predicted value of the state variable to be filled in, after the complete predicted state information determining module fills the predicted value of the state variable to be filled in to obtain the complete predicted state information of the target wind turbine;

[0214] A second autocorrelation coefficient determination module is used to calculate a second autocorrelation coefficient between the reference state variable and the observed value of the state variable to be filled for filling the reference state variable;

[0215] A second judging module is configured to judge whether a difference between the first autocorrelation coefficient and the second autocorrelation coefficient is within a preset threshold range;

[0216] The first marking module is configured to mark the complete predicted state information obtained after filling the state variable to be filled as successfully filled if the difference between the first autocorrelation coefficient and the second autocorrelation coefficient is within a preset threshold range.

[0217] In one feasible embodiment, the device further comprises:

[0218] A second marking module is used to mark the complete prediction state information obtained after filling the state variable to be filled as a filling failure if the difference between the first autocorrelation coefficient and the second autocorrelation coefficient is not within the preset threshold range after the second judgment module determines whether the difference between the first autocorrelation coefficient and the second autocorrelation coefficient is within the preset threshold range.

[0219] In one feasible embodiment, the device further comprises:

[0220] A model prediction effect evaluation module is used to evaluate the model prediction effect of the first target model according to the filling rate of the first target model, the filling accuracy of the first target model and the filling time of the first target model after the first marking module marks the complete prediction state information obtained after filling the state variable to be filled as successfully filled, wherein the filling rate of the first target model is the ratio of the number of complete prediction state information marked as successfully filled after filling using the first target model to the number of all complete prediction state information, the filling accuracy of the first target model is the difference between the complete prediction state information marked as successfully filled after filling using the first target model and the complete simulation state variable, the complete simulation state variable is the complete state variable of the target wind turbine obtained by the simulation algorithm, and the filling time of the first target model is the filling time of a single complete prediction state information when filling using the first target model.

[0221] In one feasible implementation, when the model prediction effect evaluation module is used to evaluate the model prediction effect of the first target model based on the filling rate of the first target model, the filling accuracy of the first target model, and the filling time of the first target model, it is specifically used to:

[0222] Performing density clustering using the DBSCAN clustering algorithm based on the filling rate of the first target model, the filling accuracy of the first target model, and the filling time of the first target model to determine outliers in the complete prediction state information marked as successfully filled after filling using the first target model;

[0223] Determining whether the number of abnormal values ​​exceeds a preset number;

[0224] If the number of the abnormal values ​​does not exceed the preset number, the model prediction effect of the first target model is evaluated as qualified;

[0225] If the number of the outliers exceeds a preset number, the model prediction effect of the first target model is evaluated as unqualified.

[0226] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0227] The missing state variable filling device for the wind turbine provided in the embodiment of the present invention can be specific hardware on the device or software or firmware installed on the device. The implementation principle and technical effects of the device provided in the embodiment of the present invention are the same as those of the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference can be made to the corresponding content in the aforementioned method embodiment. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can all refer to the corresponding processes in the aforementioned method embodiment, and will not be repeated here.

[0228] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some communication interface, the indirect coupling or communication connection of the device or unit may be electrical, mechanical or other forms.

[0229] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0230] In addition, each functional unit in the embodiment provided by the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0231] It should be noted that similar numbers and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are only used to distinguish the description and are not to be understood as indicating or implying relative importance.

[0232] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. However, such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. They should all be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A method for filling missing state variables of a wind turbine, characterized in that: The method comprises: Determine at least one feature input variable from a set of candidate feature variables using a random forest algorithm based on pre-configured feature output variables, wherein the set of candidate feature variables includes at least one state feature variable of a target wind turbine, and at least one of the feature input variables is a strongly correlated variable of the feature output variable; A first target model is obtained by performing model training on a first initial model for filling missing data of a wind power generation state using the overall training sample set, wherein the training output samples in the overall training sample set are feature output variables that meet a preset number of at least one feature output variable, and the training input samples in the overall training sample set are strongly correlated variables of the training output samples in the overall training sample set; The first target model is tested using the entire test sample set to obtain a missing data filling accuracy rate of the first target model, wherein the test output samples in the entire test sample set are the remaining feature output variables of at least one feature output variable except the training output samples, and the test input samples in the entire test sample set are strongly correlated variables of the test output samples in the entire test sample set; Determining whether the missing data filling accuracy of the first target model exceeds the standard accuracy; If the missing data filling accuracy of the first target model exceeds the standard accuracy, the reference state variable of the target wind turbine is input into the first target model to obtain a predicted value of the state variable to be filled in the target wind turbine, wherein the reference state variable is a characteristic variable that is not missing in the state characteristic variables of the target wind turbine, and the state variable to be filled is a characteristic variable that is missing in the state characteristic variables of the target wind turbine; Filling the predicted value of the state variable to be filled into the state variable to be filled to obtain complete predicted state information of the target wind turbine; After determining whether the missing data filling accuracy of the first target model exceeds the standard accuracy, the method further includes: If the missing data filling accuracy of the first target model does not exceed the standard accuracy, determining the delay order between at least one of the feature input variables and the feature output variable using the Akaike information criterion; Dividing the scope of at least one of the characteristic input variables according to the delay order to obtain at least one group of characteristic input variable groups, wherein the characteristic input variables in each group of the characteristic input variable groups all fall within the same variable interval; For each set of the feature input variable groups, each feature input variable in the feature input variable group is used as a model input of a second initial model, a strongly correlated variable of each feature input variable in the feature input variable group is used as a model output of the second initial model, and the second initial model is trained to obtain a second target model obtained by training the feature input variable group; The to-be-filled state variables of the target wind turbine are filled using a third target model, wherein the third target model is a second target model trained using a target feature input variable group, the target feature input variable group is a feature input variable group whose included feature input variables fall within a target interval, and the target interval is a variable interval within which strongly correlated variables of the missing state data of the target wind turbine fall.

2. The method according to claim 1, characterized in that The step of dividing the scope of at least one of the characteristic input variables according to the delay order to obtain at least one group of characteristic input variable groups includes: Establishing a finite difference regression vector of at least one of the characteristic input variables according to the delay order; At least one of the feature input variables is scoped according to clustering of a finite difference regression vector of at least one of the feature input variables to obtain at least one group of feature input variable groups.

3. The method according to claim 1, characterized in that After filling the predicted value of the to-be-filled state variable into the to-be-filled state variable to obtain complete predicted state information of the target wind turbine, the method further includes: For each of the reference state variables, calculating a first autocorrelation coefficient between the reference state variable and a predicted value of the state variable to be filled for filling the reference state variable; Calculating a second autocorrelation coefficient between the reference state variable and the observed value of the state variable to be filled in for filling in the reference state variable; Determining whether a difference between the first autocorrelation coefficient and the second autocorrelation coefficient is within a preset threshold range; If the difference between the first autocorrelation coefficient and the second autocorrelation coefficient is within a preset threshold range, the complete predicted state information obtained after filling the state variable to be filled is marked as filling successful.

4. The method according to claim 3, characterized in that After determining whether the difference between the first autocorrelation coefficient and the second autocorrelation coefficient is within a preset threshold range, the method further includes: If the difference between the first autocorrelation coefficient and the second autocorrelation coefficient is not within a preset threshold range, the complete predicted state information obtained after filling the state variable to be filled is marked as filling failure.

5. The method according to claim 3, characterized in that After marking the complete predicted state information obtained after filling the state variable to be filled as successfully filled, the method further includes: The model prediction effect of the first target model is evaluated according to the filling rate of the first target model, the filling accuracy of the first target model and the filling time of the first target model, wherein the filling rate of the first target model is the ratio of the number of complete prediction state information marked as successfully filled after filling using the first target model to the number of all complete prediction state information, the filling accuracy of the first target model is the difference between the complete prediction state information marked as successfully filled after filling using the first target model and the complete simulation state variables, the complete simulation state variables are the complete state variables of the target wind turbine obtained by the simulation algorithm, and the filling time of the first target model is the filling duration of a single complete prediction state information when filling using the first target model.

6. The method according to claim 5, characterized in that The evaluating the model prediction effect of the first target model according to the filling rate of the first target model, the filling accuracy of the first target model, and the filling time of the first target model includes: Performing density clustering using the DBSCAN clustering algorithm based on the filling rate of the first target model, the filling accuracy of the first target model, and the filling time of the first target model to determine outliers in the complete prediction state information marked as successfully filled after filling using the first target model; Determining whether the number of abnormal values ​​exceeds a preset number; If the number of the abnormal values ​​does not exceed the preset number, the model prediction effect of the first target model is evaluated as qualified; If the number of the outliers exceeds a preset number, the model prediction effect of the first target model is evaluated as unqualified.

7. A device for filling missing state variables of a wind turbine, characterized in that: The device comprises: a feature input variable determination module, configured to determine at least one feature input variable from a set of candidate feature variables using a random forest algorithm based on pre-configured feature output variables, wherein the set of candidate feature variables includes at least one state feature variable of a target wind turbine, and at least one of the feature input variables is a strongly correlated variable of the feature output variable; a first target model determination module, configured to perform model training on a first initial model for filling missing data of a wind power generation state using the overall training sample set to obtain a first target model, wherein the training output samples in the overall training sample set are feature output variables that meet a preset number of at least one feature output variable, and the training input samples in the overall training sample set are strongly correlated variables of the training output samples in the overall training sample set; a missing data filling accuracy determination module, configured to test the first target model using the entire test sample set to obtain the missing data filling accuracy of the first target model, wherein the test output samples in the entire test sample set are the remaining feature output variables of at least one feature output variable other than the training output samples, and the test input samples in the entire test sample set are strongly correlated variables of the test output samples in the entire test sample set; A first judgment module is used to judge whether the missing data filling accuracy of the first target model exceeds the standard accuracy; a predicted value determination module, configured to input the reference state variables of the target wind turbine into the first target model to obtain predicted values ​​of the to-be-filled state variables of the target wind turbine if the missing data filling accuracy of the first target model exceeds a standard accuracy, wherein the reference state variables are characteristic variables that are not missing in the state characteristic variables of the target wind turbine, and the to-be-filled state variables are characteristic variables that are missing in the state characteristic variables of the target wind turbine; a complete predicted state information determination module, configured to fill the predicted value of the state variable to be filled into the state variable to be filled to obtain the complete predicted state information of the target wind turbine; a delay order determination module, configured to, after the first determination module determines whether the missing data filling accuracy of the first target model exceeds a standard accuracy rate, determine a delay order between at least one of the feature input variables and the feature output variable using the Akaike information criterion if the missing data filling accuracy of the first target model does not exceed the standard accuracy rate; a characteristic input variable group determination module, configured to divide the scope of at least one of the characteristic input variables according to the delay order to obtain at least one group of characteristic input variable groups, wherein the characteristic input variables in each group of the characteristic input variable groups all fall within the same variable interval; a second target model determination module configured to, for each set of the feature input variable groups, use each feature input variable in the feature input variable group as a model input of a second initial model, use a strongly correlated variable of each feature input variable in the feature input variable group as a model output of the second initial model, and perform model training on the second initial model to obtain a second target model obtained by training the feature input variable group; The module for filling in the state variables to be filled is used to fill in the state variables to be filled in of the target wind turbine using a third target model, wherein the third target model is a second target model trained using a target feature input variable group, the target feature input variable group is a feature input variable group whose feature input variables fall within a target interval, and the target interval is a variable interval within which the strongly correlated variables of the missing state data of the target wind turbine fall.

8. The device according to claim 7, characterized in that The characteristic input variable group determination module is configured to divide the scope of at least one characteristic input variable according to the delay order to obtain at least one characteristic input variable group, specifically configured to: Establishing a finite difference regression vector of at least one of the characteristic input variables according to the delay order; At least one of the feature input variables is scoped according to clustering of a finite difference regression vector of at least one of the feature input variables to obtain at least one group of feature input variable groups.

Citation Information

Patent Citations

  • Voltage missing data identification method based on improved random forest algorithm

    CN113468796A

  • Method for predicting heavy metal migration in waste incineration process based on random forest algorithm

    CN115017980A