Wind generating set fatigue load estimation method based on combination of I-RAE and CatBoost

By combining I-RAE and CatBoost methods, a low-dimensional characteristic model of fatigue loads for wind turbines is constructed, which solves the problems of low accuracy and poor generalization ability of fatigue load estimation in the prior art, and achieves higher estimation accuracy and applicability.

CN119989926APending Publication Date: 2025-05-13JINAN UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510405807.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing fatigue load estimation method of wind turbine units has problems such as difficulty in measuring, low accuracy and poor generalization ability, and it is difficult to effectively apply in real-time scenarios.

Method used

The fatigue load estimation method of wind turbines based on the combination of I-RAE and CatBoost is adopted. By constructing a statistical feature data set of key parameters, the low-dimensional features are extracted using the I-RAE dimensionality reduction method, and the CatBoost model is optimized through grid search and self-service sampling to perform real-time estimation of fatigue load.

Benefits of technology

The accuracy and generalization ability of estimating fatigue loads of wind turbine towers and transmission systems is significantly improved, providing reliable data support, and optimizing control strategies for wind turbine operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989926A_ABST
    Figure CN119989926A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of wind power generation, and discloses a wind generating set fatigue load estimation method based on the combination of I-RAE and CatBoost. Comprising the following steps: step 1, determining key parameters related to fatigue loads of a tower and a transmission system of the wind generating set; step 2, constructing a wind generating set tower and transmission system fatigue load estimation model taking the key parameters as input; and step 3, estimating the fatigue load through the estimation model. According to the fatigue load estimation method combining I-RAE and CatBoost, the problems of difficult measurement, low precision, poor generalization ability and the like in a traditional fatigue load estimation method are effectively solved by combining statistical analysis and an improved dimension reduction technology, the estimation accuracy of the fatigue load of the tower and the transmission system of the wind generating set is remarkably improved, and the fatigue load estimation accuracy of the tower and the transmission system of the wind generating set is improved. And reliable data support is provided for operation and maintenance of the wind generating set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of wind power generation, and in particular to a method for estimating fatigue load of a wind power generator set based on a combination of I-RAE and CatBoost. Background Art

[0002] With the rapid development of wind power generation, the safe operation and reliability of wind turbines have become the focus of attention. During the operation of wind turbines, the transmission system and tower are subjected to complex fatigue loads. Tower fatigue load (DEL Mt ) and transmission system fatigue load (DEL Ts ) is not only affected by the frequent changes in wind speed and direction, but is also closely related to the operating status of the unit, the temperature and humidity of the external environment, and other factors. Therefore, accurate estimation of fatigue load is crucial to optimizing control strategies and extending the service life of the unit.

[0003] At present, the estimation of fatigue load mainly relies on physical models and data-driven methods. Physical models calculate fatigue loads through complex mechanical simulations. Some researchers have proposed using damage equivalent loads (DEL) to effectively calculate fatigue loads, but this process requires a large amount of stress measurement data and relies on the rain flow counting method and Palmgren-Miner method for calculation. Not only are the measurement and calculation processes complicated, but the real-time performance is poor and it is difficult to apply to real-time scenarios. Data-driven methods are evaluated through machine learning models. There are deficiencies in existing methods in terms of modeling input selection, data dimensionality reduction, and model generalization capabilities. There is little attention paid to the generalization performance of data-driven models, and the model is not very transferable. Summary of the invention

[0004] In order to solve the problems in the background technology, the present invention provides a method for estimating fatigue load of a wind turbine generator set based on the combination of I-RAE and CatBoost.

[0005] The present invention is implemented by the following technical solution: a method for estimating fatigue load of a wind turbine generator set based on the combination of I-RAE and CatBoost, comprising the following steps:

[0006] Step 1: Determine the key parameters related to the fatigue load of the wind turbine;

[0007] Step 2: constructing a wind turbine generator fatigue load estimation model using the key parameters as input;

[0008] Step 3, estimating fatigue load by using the estimation model;

[0009] The step 2 comprises:

[0010] Step 2.1, constructing a simulation data set, each set of data in the simulation data set includes a time series of each of the key parameters;

[0011] Step 2.2: constructing a statistical feature data set of each key parameter based on the simulation data set;

[0012] Step 2.3, using the I-RAE dimensionality reduction method to reduce the dimensionality of the statistical feature data set to extract low-dimensional features closely related to fatigue loads, thereby constructing a low-dimensional feature set;

[0013] Step 2.4: Using the low-dimensional feature set as training data and test data, the CatBoost model is optimized through grid search and bootstrap sampling methods, and finally used for fatigue load estimation.

[0014] Preferably, in step 1, the fatigue load of the wind turbine generator set includes the tower fatigue load (DEL Mt ) and transmission system fatigue load (DEL Ts ), correspondingly, the key parameters include wind speed, active power, generator rotor speed and pitch angle.

[0015] Preferably, in step 2.1, the simulation data set is generated based on a high-precision wind turbine simulation tool.

[0016] Preferably, in step 2.2, the statistical characteristics of each key parameter include mean characteristics, range characteristics, quantile characteristics, discrete characteristics, morphological characteristics, and percentage characteristics. The statistical characteristics cover the overall distribution and variation patterns of time series data, providing more valuable input for subsequent dimensionality reduction and evaluation.

[0017] Preferably, the statistical features of the various parameters also include but are not limited to the following features: third-order central moment, third-order moment, mean absolute deviation divided by mean, mean absolute deviation divided by standard deviation, median absolute deviation divided by mean, median absolute deviation divided by standard deviation, square of mean absolute deviation, square of median absolute deviation, cube of mean absolute deviation.

[0018] Preferably, in order to improve the robustness of the model to noisy data, in step 2.3, the regression autoencoder (RAE) dimensionality reduction method introduces noise processing and a random seed mechanism, wherein: the noise processing is used to improve the robustness of the model to noisy data, and the random seed mechanism is used to ensure that the initial parameters of the model are consistent each time the training is carried out to avoid the instability of the dimensionality reduction result due to randomness;

[0019] Gaussian noise is added before the training data is input into RAE to simulate the actual data noise interference. The formula is:

[0020] X raw_noisy ij=X rawi,j+noise factory×ò(i,j)

[0021] Among them, X raw_noisy is the data after adding noise; X raw is the original data; noise_factory is the noise coefficient; ò(i,j) is a random value that follows a normal distribution with a mean of 0 and a standard deviation of 1.

[0022] The random seed mechanism is performed using the following formula;

[0023] r n+1 =(a×r n +c)modξ

[0024] Wherein, mod represents the modulo operation; a, c, ξ are preset constants; r n It is the pseudo-random number generated for the nth time. The initial r0 is determined by the given random seed, and based on the starting value r0, the subsequent pseudo-random number sequence is generated according to the linear congruential formula.

[0025] Specifically, when initializing the weights of the neural network layer, the fixed sequence numbers generated by the above method are used to assign values, so that the initial parameters of the model are the same each time the code is run. For data partitioning and noise addition operations, the generated random number sequence is fixed by setting parameters and initial seeds. As a result, model training and hyperparameter adjustment will be more controllable, making the experimental results more credible.

[0026] Preferably, in step 2.4, CatBoost model construction includes:

[0027] Load the training set X_ formed after dimensionality reduction train and the test set X_ test , and two target variable sets Y_ formed by theoretical calculation of fatigue loads of tower and transmission system DELMt and Y_ DELTs , where X_ train Used for model validation and hyperparameter tuning; X_ test For performance testing;

[0028] Use grid search to replace the complexity of manual parameter adjustment and find the target variable Y_ DELMt Set the parameter grid param_grid_DEL Mt , the target variable Y_ DELTs Set the parameter grid param_grid_DEL Ts , as a parameter adjustment framework;

[0029] Determine the number of bootstrap samplings and perform multiple bootstrap samplings, randomly select the training set index, and generate the bootstrap sample set X_train_bootstrap_DEL Mt and Y_train_bootstrap_DEL Mt , for DEL Ts Similarly;

[0030] Create a CatBoost regressor using the parameters in the parameter grid, train the CatBoost regressor with the bootstrap sample set, estimate the test set, output the average evaluation metric, and add it to the corresponding list, and plot DEL respectively. Mt and DEL Ts Comparison chart of the true value and the estimated value.

[0031] The step 3 estimates the fatigue load by using the estimation model, including real-time acquisition of key parameter data, data preprocessing, data input into the fatigue load estimation model and data output, i.e., output DEL Mt and DEL Ts The estimated value of .

[0032] The present invention provides an electronic device, comprising a processor and a memory;

[0033] The memory is used to store programs;

[0034] The processor executes the program to implement the above-mentioned evaluation method.

[0035] The present invention also proposes a computer-readable storage medium, wherein the storage medium stores a program, and the program is executed by a processor to implement the above-mentioned evaluation method.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] This paper proposes a fatigue load estimation method combining I-RAE with CatBoost. By combining statistical analysis and improved dimensionality reduction technology, it effectively solves the problems of difficult measurement, low accuracy and poor generalization ability in traditional fatigue load estimation methods, and significantly improves the DEL Mt and DEL Ts The estimation accuracy provides reliable data support for the operation and maintenance of wind turbines.

[0038] The present invention extracts statistical features from the original data and constructs a low-dimensional feature set, thus avoiding the problem of high computational complexity and susceptibility to noise interference caused by directly using high-dimensional time series data for modeling. In addition, these statistical features reflect data fluctuations from different angles, effectively reducing the data dimension while retaining key information, providing more valuable input for subsequent dimensionality reduction and evaluation.

[0039] The present invention introduces noise processing and random seed mechanisms in the RAE training process, which can ensure that the initial parameters of the model are consistent each time it is trained, avoid the instability of the dimensionality reduction result due to randomness, and improve the robustness of the model to noisy data. At the same time, the present invention optimizes the hyperparameters of CatBoost through grid search and self-service sampling methods, and can evaluate the model performance under different hyperparameter combinations, so as to select the best hyperparameters. Multiple sub-sample sets are generated through the self-service sampling method, and each sub-sample set is used to train a CatBoost model. Finally, the estimation accuracy is improved through ensemble learning, and the generalization ability of the model is enhanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 A flowchart for constructing the estimation model proposed by the present invention;

[0041] Figure 2 It is a time series of a group of data in the simulation data set used in the method proposed by the present invention;

[0042] Figure 3 Comparison chart of using LightGBM model estimation on validation set under different data processing methods;

[0043] Figure 4 Comparison chart of using CatBoost model estimation on validation set under different data processing methods;

[0044] Figure 5 Comparison chart of using LightGBM model estimation on the test set under different dimensionality reduction methods;

[0045] Figure 6 Comparison chart of using CatBoost model estimation on the test set under different dimensionality reduction methods;

[0046] Figure 7 This is a workflow diagram of the estimation method proposed in the present invention. DETAILED DESCRIPTION

[0047] The present invention is further described below in conjunction with the accompanying drawings and specific implementation methods. It should be noted that, under the premise of no conflict, the various embodiments or technical features described below can be arbitrarily combined to form a new embodiment.

[0048] Example:

[0049] Aiming at the difficulty of estimating fatigue load of wind turbine generator set, the present invention proposes a method for estimating fatigue load of wind turbine generator set tower and transmission system based on the combination of I-RAE and CatBoost. The model data of wind turbine generator set comes from the National Renewable Energy Laboratory (NREL) of the United States, and the rated power of this type of wind turbine generator set is 5MW. Based on the above settings, the specific implementation steps of the method proposed in this paper will be further described in detail below:

[0050] Reference Figure 1 - Figure 7 The fatigue load estimation method of a wind turbine generator set based on the combination of I-RAE and CatBoost proposed in the present invention has the following specific implementation steps:

[0051] Step 1: Combine the fatigue load mechanism analysis of the wind turbine tower and transmission system to clearly estimate the key parameters required.

[0052] It should be noted that during the operation of wind turbines, the transmission system and tower, as key components, are subjected to complex fatigue loads for a long time. Their fatigue conditions are directly related to the safe and stable operation of the unit, so they are studied as modeling objects.

[0053] For the transmission system, the torque imbalance between the wind wheel and the generator is the key cause of damage; the dynamic modeling of the transmission system is as follows:

[0054]

[0055] Among them, T m and T e are the torques on the rotor side and the generator side respectively; T s is the shaft torque of the transmission system, i.e. the main shaft torque; Expressed as the angular acceleration of the wind rotor; Expressed as the angular acceleration of the generator; J m is the rotor inertia; J e is the generator inertia; N gear K is the gear ratio of the gearbox, which indicates the speed ratio between the input shaft and the output shaft of the gearbox; sp is the spring constant; K vi is the viscous friction coefficient; ψ is the torsion angle difference between the two ends of the main shaft, It is the torsional angular velocity difference between the two ends of the main shaft.

[0056] In formula (1), T m and T e The value of can be calculated by the following model:

[0057]

[0058] Where ρ is the air density; R is the blade radius; β is the pitch angle; ω r is the wind wheel speed; ω e is the generator rotor speed, and ω e =N gear ω r ; C p (ω r ,v,β) is the power coefficient, which is used to measure the efficiency of the wind wheel in converting wind energy into mechanical energy; v is the wind speed; P e is the active power of the generator; μ is the generator efficiency.

[0059] For the tower, the thrust generated by the wind field makes the tower form an inverted cantilever structure, and the root of the tower bears a huge load. The dynamic modeling of the tower is as follows:

[0060] F t =0.5πρR 2 C t (λ,β)V 2 (4)

[0061] M t =hF t (5)

[0062] Among them, F t is the thrust generated by the wind field on the tower; λ is the tip speed ratio; C t (λ,β) is the thrust coefficient; M t is the bending moment borne by the tower root; h represents the height of the tower.

[0063] According to material damage theory, the fatigue degree of a material does not depend entirely on the instantaneous stress at a certain moment, but on the stress fluctuations over a historical period of time. Based on this theory, the damage equivalent load (DEL) is proposed, which is defined as the amplitude of the sinusoidal stress that produces the same damage as the original signal within a time T at a constant frequency f. The calculation formula is:

[0064]

[0065] Among them, σ j is the stress amplitude; n j is the number of load cycles; M is the total number of load conditions, j∈{1,2,…,M}; m is Coefficient parameters.

[0066] According to the above-mentioned analysis of the stress mechanism of the transmission system and the tower structure, the present invention selects wind speed, active power, generator rotor speed and pitch angle, four parameters that are easy to measure and closely related to fatigue load, as the basis for research, and performs fatigue load data-driven modeling based on these four parameters.

[0067] Step 2: Generate a data set based on the SimWindFarm toolbox.

[0068] The data set in this implementation scheme is based on the SimWindFarm toolbox and is generated by the Monte Carlo experiment method. According to the analysis of the DEL calculation and related parameters in step 1, it can be seen that the wind speed, active power, generator rotor speed and pitch angle of the wind turbine are crucial to the fatigue load, so the experiment is used to simulate the different operating conditions of the wind turbine and record the above four parameters. A total of 48,000 different operating conditions are set in the Monte Carlo experiment, and 300 seconds of time series data are collected under each operating condition, and data is collected every 1 second. Since four parameters are recorded, a total of 1,200 seconds of time series data are recorded for each operating condition. Therefore, the result can be regarded as a data set of 48,000 sets of data, each of which contains 1,200 features, and this data set is used as a training set.

[0069] In subsequent experiments, 0.2% of the 48,000 sets of data were used as validation sets to evaluate the performance of the model. In addition, in order to further verify the generalization performance of the model, SimWindFarm was used to generate 10 additional sets of data with completely different working conditions from the training scenarios, each set also containing 1,200 features, and used as test sets.

[0070] Step 3: Perform statistical analysis on the collected time series data and construct a statistical feature set.

[0071] Directly using the high-dimensional time series data obtained in step 2 to model the model will result in high computational complexity and be susceptible to uncertainty noise interference. The present invention extracts statistical features from the original time series data and constructs a low-dimensional feature set. Therefore, 7 types of statistical indicators are calculated for the collected time series data, including:

[0072] Mean characteristics: mean, median, mode, geometric mean;

[0073] Range characteristics: range, interquartile range, maximum value, minimum value;

[0074] Quantile characteristics: upper quartile, lower quartile, upper third, lower third;

[0075] Discrete characteristics: standard deviation, variance, coefficient of variation, range coefficient, square root of coefficient of variation, square of coefficient of variance, median absolute deviation, mean absolute deviation;

[0076] Morphological characteristics: skewness, kurtosis, adjusted skewness, adjusted kurtosis, squared kurtosis, cubed kurtosis, squared skewness times variance, kurtosis times variance;

[0077] Percentile features: 10% to 90% quantiles;

[0078] Other features: third central moment, third moment, mean absolute deviation divided by mean, mean absolute deviation divided by standard deviation, median absolute deviation divided by mean, median absolute deviation divided by standard deviation, mean absolute deviation squared, median absolute deviation squared, mean absolute deviation cubed.

[0079] Therefore, 46 statistical indicators are extracted from the time series data of each parameter, totaling 184 dimensional features. These statistical features reflect data fluctuations from different angles, effectively reducing the data dimension while retaining key information, providing more valuable input for subsequent dimensionality reduction and evaluation.

[0080] Step 4: Use the I-RAE dimensionality reduction method to further extract low-dimensional feature representation from the statistical feature set.

[0081] The present invention uses an improved regression autoencoder (I-RAE) method to further reduce the dimension of the statistical feature set obtained in step three. It should be noted that the regression autoencoder (RAE) is a type of autoencoder, including an encoder and a decoder, and the network structure of RAE includes an input layer, a hidden layer, a regression layer and an output layer. It is through the above-mentioned multi-layer structure that RAE can effectively capture the nonlinear relationship in high-dimensional data and extract key features related to fatigue loads.

[0082] And on the basis of considering the reconstruction loss, an additional target loss is introduced for the regression task, so that the features after dimensionality reduction are closely related to the fatigue load. re , measured using the mean square error (MSE) loss function, whose mathematical formula is:

[0083]

[0084] Where n is the total amount of data, satisfying i∈{1,2,…,n}; x i Represents the i-th input data; Represents the data reconstructed by the decoder, and the reconstruction loss is minimized through continuous iteration until convergence.

[0085] The target loss L measures the difference between the estimated value and the actual value. tg , can be calculated using the SmoothL1 loss function. For example, if the target variable is Y = {y1,y2,...,y n}, the estimated value is The target loss can be expressed as:

[0086]

[0087] Therefore, the joint loss function of RAE is:

[0088] L=L re +ηLtg (9)

[0089] Among them, η is a weighting parameter used to balance the importance of reconstruction loss and target loss.

[0090] In addition, in order to improve the robustness of the model to noisy data, noise processing and random seed mechanisms are introduced in the RAE training process. Specifically, Gaussian noise is added before the training data is input into RAE to simulate the noise interference of actual data. The formula is:

[0091] X raw_noisy ij=X raw i,j+noise factory×ò(i,j) (10)

[0092] Among them, X raw_noisy is the data after adding noise; X raw is the original data; noise_factory is the noise coefficient; ò(i,j) is a random value that follows a normal distribution with a mean of 0 and a standard deviation of 1.

[0093] At the same time, the random seed is set to ensure that the initial parameters of the model are consistent each time it is trained, so as to avoid the instability of the dimensionality reduction results due to randomness. The random seed mechanism uses the following formula:

[0094] r n+1 =(a×r n +c)modξ (11)

[0095] Wherein, mod represents the modulo operation; a, c, ξ are preset constants; r n It is the pseudo-random number generated last time. The initial r0 is determined by the given random seed. Based on the starting value r0, the subsequent pseudo-random number sequence is generated according to the linear congruential formula.

[0096] The value of ξ is crucial, and a larger prime number is usually chosen. This is because a larger prime number can ensure that the generated pseudo-random number sequence has a longer period, which means that before repeating the next round, the sequence can generate more different random numbers, thereby reducing the repetition and correlation of random numbers. On this basis, different values ​​of a and c are tried to observe the performance of the generated random numbers in the experiment.

[0097] The above-mentioned RAE principle and improvement scheme are the I-RAE scheme proposed in the present invention. Its training process includes setting random seeds, data reading, data standardization, adding noise, defining encoders and decoders, training models and evaluating performance. The generated model is used to reduce the dimension of 48,000 sets of data sets and 10 sets of test sets to obtain low-dimensional feature sets.

[0098] Step 5: Estimate fatigue load based on CatBoost model and optimize model parameters.

[0099] Based on the low-dimensional feature set obtained by the I-RAE dimensionality reduction method, this scheme uses the CatBoost model for fatigue load estimation. The CatBoost model is a machine learning model based on the gradient boosting algorithm that can effectively process high-dimensional data.

[0100] Specifically, CatBoost constructs multiple decision trees, each of which attempts to correct the estimation error of the previous tree, and finally weights the evaluation results of all trees to obtain the estimated value of fatigue load.

[0101] In order to ensure the estimation accuracy of the model, the present invention optimizes the hyperparameters of CatBoost through grid search and bootstrap sampling to evaluate the model performance under different hyperparameter combinations, thereby selecting the best hyperparameters. The model training process includes data loading, parameter setting, bootstrap sampling, model training and evaluation.

[0102] First, the 48,000 feature sets after I-RAE dimension reduction are divided into training sets and validation sets. The training set is used for model training, the validation set is used to evaluate model performance, and 10 test sets are used to evaluate the generalization ability of the model. During the training process, the self-service sampling method is used to generate multiple sub-sample sets, each of which is used to train a CatBoost model. Finally, the estimation accuracy is improved through ensemble learning, and the generalization ability of the model is enhanced. The specific algorithm flow is as follows:

[0103] Load the training set X_ formed after dimensionality reduction train and the test set X_ test , and two target variable sets Y_ formed by theoretical calculation of fatigue loads of tower and transmission system DELMt and Y_ DELTs , where X_ train Used for model validation and hyperparameter tuning; X_ test For performance testing;

[0104] Use grid search to replace the complexity of manual parameter adjustment and find the target variable Y_ DELMt Set the parameter grid param_grid_DEL Mt , the target variable Y_ DELTs Set the parameter grid param_grid_DEL Ts , as a parameter adjustment framework;

[0105] Determine the number of bootstrap samplings and perform multiple bootstrap samplings, randomly select the training set index, and generate the bootstrap sample set X_train_bootstrap_DEL Mtand Y_train_bootstrap_DEL Mt , for DEL Ts Similarly; create a CatBoost regressor using the parameters in the parameter grid, train the CatBoost regressor with the bootstrap sample set, and then estimate the test set, output the average evaluation index, and add it to the corresponding list, and plot DEL respectively. Mt and DEL Ts Comparison chart of the true value and the estimated value.

[0106] In order to comprehensively consider the performance of the model under multiple different training data and reduce the influence of accidental factors, the present invention trains multiple models through the self-service sampling method, and each model is tested on the test set. Finally, the average value of these test results is taken. The final evaluation result can better represent the true level of this model.

[0107] For the present invention Figure 1-Figure 5 , explained in detail as follows:

[0108] Figure 1 The overall framework of the fatigue load estimation method proposed in the present invention is demonstrated. In order to verify the effectiveness of the method proposed in this scheme, in addition to the method based on statistical indicators, a data set based on the original time series is also used for comparison. In addition, the following three methods are used to reduce the dimension of the data set: recursive feature elimination (RFE), AE, and RAE, and then compared with the dimension reduction results of the I-RAE scheme proposed in the present invention, and a LightGBM model for comparison is constructed in the estimation modeling, which is compared with the Catboost model proposed in the present invention.

[0109] For convenience, the processing methods based on the original time series data and statistical data can be simplified to Otd and Sd respectively. Therefore, the I-RAE dimensionality reduction method based on Sd and modeled with Catboost can be expressed as Sd-I-RAE-C. If it is modeled with LightGBM, it can be expressed as Sd-I-RAE-L. Similarly, other methods can be expressed.

[0110] Figure 2 300 time series values ​​of wind speed, active power, generator rotor speed and pitch angle required to calculate the target variables for one of 48,000 random operating conditions are shown, as well as the tower bending moment M used to calculate the theoretical value of fatigue load. t and spindle torque T s .

[0111] Figure 3 , Figure 4The figure shows the comparison between the fatigue load estimation results and the true value of the validation set under different schemes. The 48,000 groups of data are divided into training set and validation set, and the validation set is divided at a ratio of 0.2%, that is, 96 samples are used as the validation set. Under the premise of the I-RAE dimensionality reduction method, the two data processing methods of Otd and Sd are compared, and the estimation results based on the LightGBM model used for comparison and the Catboost model proposed in the present invention are respectively shown.

[0112] Reference Figure 3 , Figure 4 As a result, the following conclusions can be drawn:

[0113] The results of dimensionality reduction estimation after data processing using the Sd method are much better than those of Otd in terms of fit with the true value. Therefore, in order to further verify the generalization ability of the model in subsequent experiments, experiments will only be conducted based on the Sd scheme.

[0114] Based on the above comparative experiments, Table 1 shows the evaluation index results obtained by different dimensionality reduction methods and estimation models. Four evaluation indicators are used in Table 1, namely: mean absolute error (MAE), mean square error (MSE), root mean square error (RMSE) and determination coefficient (R 2 ).

[0115] Table 1 - Evaluation results of various methods based on Sd in the test set

[0116]

[0117] Figure 5 This is a comparison chart between the fatigue load estimation results using different dimensionality reduction methods under the LightGBM model used for comparison and the actual values ​​of the test set.

[0118] Figure 6 The figure is a comparison chart of fatigue load estimation results using different dimensionality reduction methods under the Catboost model proposed in the present invention and the actual value of the test set.

[0119] The results show that the Sd-I-RAE-C scheme integrating the proposed method has better estimation effect on the test set of the two target variables. Mt R 2 is 0.961, DEL Ts R 2 The MAE value is 0.982, and the Sd-I-RAE-L method has the lowest DEL Mt The best result R 2Only 0.88, compared with the proposed Sd-I-RAE-C scheme improved by about 8%. At the same time, the estimation error of the Sd-I-RAE-C scheme for different groups of data in the test set is only between 1% and 4%, indicating that this method not only has high estimation accuracy, but also has stronger generalization ability.

[0120] In summary, the estimation method proposed in this scheme has the following improvements:

[0121] The fatigue load estimation method combining I-RAE and CatBoost proposed in this paper effectively solves the problems of difficult measurement, low accuracy and poor generalization ability in traditional fatigue load estimation methods by combining statistical analysis with improved dimensionality reduction technology and hyperparameter setting technology, and significantly improves the DEL Mt and DEL Ts The estimation accuracy and generalization ability provide reliable data support for the operation and maintenance of wind turbines.

[0122] Once the above model is built, real-time estimation can be performed as follows:

[0123] Real-time data collection: The wind speed sensor, power sensor, speed sensor and pitch angle sensor installed on the wind turbine generator set monitor the operating status of the wind turbine generator set in real time. With a cycle of 5 minutes, data is collected every 1 second, and each collection lasts for 300 seconds to obtain the time series data of wind speed, active power, generator rotor speed and pitch angle during the time period.

[0124] Statistical calculation: The collected 300-second time series data is calculated according to the established statistical indicator set, which comprehensively describes the distribution and change patterns of the data from different angles and generates a 184-dimensional statistical feature data set.

[0125] Data dimensionality reduction: Use the trained I-RAE model to reduce the dimensionality of the statistical feature data set. The I-RAE model has determined the network structure and optimal hyperparameters during training, and the model is directly loaded in actual applications. The statistical feature data is input into the model to map the high-dimensional data to a low-dimensional space, output a low-dimensional feature set closely related to the fatigue load, remove redundant information, and improve the efficiency and accuracy of subsequent model calculations.

[0126] Fatigue load estimation: The low-dimensional feature set after dimensionality reduction is used as input, and the trained CatBoost model is used to estimate fatigue load. The CatBoost model has optimized hyperparameters through grid search and bootstrap sampling during the training phase, and can accurately fit the complex relationship between low-dimensional features and fatigue loads. Model output DEL Mt and DEL TsThe estimated value reflects the fatigue load on the tower and transmission system of the wind turbine generator set within 5 minutes.

[0127] Application and feedback of results: The fatigue load prediction results can be directly applied to the operation and maintenance decisions of wind turbines. If the predicted fatigue load approaches or exceeds the set threshold, the system automatically issues an early warning to prompt the operation and maintenance personnel to promptly inspect and maintain the relevant components to prevent potential failures. At the same time, these data can also be used to optimize the control strategy of wind turbines, adjust the pitch angle according to the fatigue load conditions, reduce fatigue damage, and improve the stability and reliability of wind turbine operation. In addition, each prediction result is recorded, and the accumulated data is used for subsequent model optimization and performance evaluation to continuously improve the accuracy and adaptability of the model.

[0128] In actual operation, the entire process is implemented through an automated system to ensure the efficiency and accuracy of data collection, processing and prediction. At the same time, sensors need to be calibrated and maintained regularly to ensure data quality, and the model needs to be updated and optimized to adapt to changes in the wind turbine operating environment and working conditions.

[0129] The above-mentioned embodiments are only preferred embodiments of the present invention and cannot be used to limit the scope of protection of the present invention. Any non-substantial changes and substitutions made by technicians in this field on the basis of the present invention shall fall within the scope of protection required by the present invention.

Claims

1. A wind turbine fatigue load estimation method based on the combination of I-RAE and CatBoost is characterized by: The steps include: Step 1: Determine the key parameters related to the fatigue load of the wind turbine tower and transmission system; Step 2: construct a wind turbine tower and transmission system fatigue load estimation model using the key parameters as input; Step 3, estimating fatigue load by using the estimation model; The step 2 comprises: Step 2.1, constructing a simulation data set, each set of data in the simulation data set includes a time series of each of the key parameters; Step 2.2: constructing a statistical feature data set of each key parameter based on the simulation data set; Step 2.3, using the I-RAE dimensionality reduction method to reduce the dimensionality of the statistical feature data set to extract low-dimensional features closely related to fatigue loads, thereby constructing a low-dimensional feature set; Step 2.4: Using the low-dimensional feature set as training data and test data, the CatBoost model is optimized through grid search and bootstrap sampling methods, and finally used for fatigue load estimation.

2. The method for estimating fatigue load of a wind turbine generator set based on the combination of I-RAE and CatBoost as claimed in claim 1, characterized in that: In the step 1, the fatigue load of the wind turbine generator set includes the tower fatigue load and the transmission system fatigue load. Correspondingly, the key parameters include wind speed, active power, generator rotor speed and pitch angle.

3. The method for estimating fatigue load of a wind turbine generator set based on the combination of I-RAE and CatBoost as claimed in claim 1, characterized in that: In the step 2.1, the simulation data set is generated based on a high-precision wind turbine simulation tool.

4. The method for estimating fatigue load of a wind turbine generator set based on the combination of I-RAE and CatBoost as claimed in claim 1, characterized in that: In step 2.2, the statistical characteristics of each key parameter include mean characteristics, range characteristics, quantile characteristics, discrete characteristics, morphological characteristics and percentage characteristics.

5. The method for estimating fatigue load of a wind turbine generator set based on the combination of I-RAE and CatBoost as claimed in claim 4, characterized in that: The statistical characteristics of each key parameter also include but are not limited to: third-order central moment, third-order moment, mean absolute deviation divided by mean, mean absolute deviation divided by standard deviation, median absolute deviation divided by mean, median absolute deviation divided by standard deviation, square of mean absolute deviation, square of median absolute deviation, and cube of mean absolute deviation.

6. The method for estimating fatigue load of a wind turbine generator set based on the combination of I-RAE and CatBoost as claimed in claim 1, characterized in that: In step 2.3, the I-RAE dimensionality reduction method introduces noise processing and random seed mechanisms, wherein: noise processing is used to improve the robustness of the model to noise data, and the random seed mechanism is used to ensure that the initial parameters of the model are consistent each time it is trained, so as to avoid instability of the dimensionality reduction results due to randomness; Gaussian noise is added before the training data is input into RAE to simulate the actual data noise interference. The formula is: X raw_noisy i j=X raw i,j+noise factory×ò(i,j) Among them, X raw_noisy is the data after adding noise; X raw is the original data; noise_factory is the noise coefficient; ò(i,j) is a random value that follows a normal distribution with a mean of 0 and a standard deviation of 1; The random seed mechanism is performed using the following formula; r n+1 =(a×r n +c)modξ Wherein, mod represents the modulo operation; a, c, ξ are preset constants; r n is the pseudo-random number generated for the nth time.

7. The method for estimating fatigue load of a wind turbine generator set based on the combination of I-RAE and CatBoost as claimed in claim 1, characterized in that: In step 2.4, the CatBoost model is constructed as follows: Load the training set X_ formed after dimensionality reduction train and the test set X_ test , and two target variable sets Y_ formed by theoretical calculation of fatigue loads of tower and transmission system DELMt and Y_ DELTs , where X_ train Used for model validation and hyperparameter tuning; X_ test For performance testing; Use grid search to replace the complexity of manual parameter adjustment and find the target variable Y_ DELMt Set the parameter grid param_grid_DEL Mt , the target variable Y_ DELTs Set the parameter grid param_grid_DEL Ts , as a parameter adjustment framework; Determine the number of bootstrap samplings and perform multiple bootstrap samplings, randomly select the training set index, and generate the bootstrap sample set X_train_bootstrap_DEL Mt and Y_train_bootstrap_DEL Mt , for DEL Ts Similarly; Create a CatBoost regressor using the parameters in the parameter grid, train the CatBoost regressor with the bootstrap sample set, estimate the test set, output the average evaluation metric, and add it to the corresponding list, and plot DEL respectively. Mt and DEL Ts Comparison chart of the true value and the estimated value.

8. The method for estimating fatigue load of a wind turbine generator set based on the combination of I-RAE and CatBoost as claimed in claim 1, characterized in that: In step 3, fatigue load is estimated by the estimation model, including real-time acquisition of key parameter data, data preprocessing, data input into the fatigue load estimation model and data output, and finally DEL is obtained. Mt and DEL Ts The estimated value of .

9. An electronic device, characterized in that: including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Chemical process safety risk early warning system based on edge calculation

    CN121599478A