Wind power blade icing state detection method

Through multimodal data fusion and Transformer model processing of wind power blade sensor data, the problems of low detection accuracy and poor robustness in the existing technology are solved, and high-precision identification and timely warning of the icy state of wind power blades are achieved.

CN120296658APending Publication Date: 2025-07-11JINTIANHONG ENERGY TECH (BEIJING) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510361916.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art has problems with low detection accuracy and poor robustness in wind power blade icing detection, especially in complex timing dependencies and nonlinear relationships of long-term data in complex dynamic environments.

Method used

Multimodal data fusion, maximum-minimum normalization and sliding window methods are used to process sensor data, feature compression and timing dependency modeling are combined with encoder and Transformer models, data is processed through self-attention mechanism and occlusion multi-head attention mechanism, and icing probability prediction is optimized using binary cross entropy loss function.

Benefits of technology

It improves the accuracy and robustness of the detection of the icy state of wind power blades, can better capture the dynamic changes of environmental factors, reduce false detection rates and missed detection rates, and ensure the safe and efficient operation of the wind turbine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296658A_ABST
    Figure CN120296658A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of detection methods, and particularly relates to a wind power blade icing state detection method. The method for detecting the icing state of the wind power blade is good in accuracy and robustness. The method comprises the following steps: step 1, collecting multi-modal data at a wind power blade; step 2, fusing the multi-modal data; step 3, a Min-Max Normalization (Min-Max Normalization) method is adopted to map the data to a uniform numerical value range, and the data are mapped to the uniform numerical value range through the Min-Max Normalization (Min-Max Normalization) method; dividing the data by adopting a sliding window method to generate a plurality of time sequence subsequences; 4, performing feature compression on the data through an encoder; 5, the decoder carries out data processing through a multi-head attention masking mechanism; and step 6, calculating the icing probability of the processed data through a binary cross entropy loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of detection methods, and particularly relates to a method for detecting the icing state of wind turbine blades. Background Art

[0002] With the gradual popularization of global wind energy utilization, the wind power industry has developed rapidly. As an important energy equipment, the operation stability and safety of wind turbines are particularly important. During the operation of wind turbines, blade icing is a major hidden danger affecting the normal operation and power generation efficiency of the turbines. Especially in cold regions, when the temperature drops sharply, the surface of wind turbine blades is prone to icing and form an ice layer, which will not only increase the weight of the blades, change the aerodynamic characteristics, but also may cause the turbine to stall, malfunction or even be damaged. Therefore, accurately and timely monitoring the icing state of wind turbine blades is the key to ensuring the safe and efficient operation of wind turbines.

[0003] Traditional icing detection methods mainly rely on meteorological data (such as temperature, humidity, wind speed, etc.) and physical models. These methods often ignore the complex dynamic environment faced by wind turbine blades during actual operation. For example, meteorological data cannot accurately reflect the actual icing state of wind turbine blades, and the influence of environmental changes is often difficult to capture, which makes the traditional methods have a high false detection rate and missed detection rate. In addition, since icing often occurs as a dynamic process, involving real-time monitoring of multiple sensor data, how to effectively fuse multiple data sources and improve the detection accuracy has become an urgent problem in the current technology.

[0004] In recent years, machine learning and deep learning technologies have been widely used in the abnormal detection of wind turbine blade icing. By analyzing multi-dimensional sensor data and using data-driven methods for abnormal detection, this provides a new idea for improving the detection accuracy and reliability. However, the existing methods based on models such as encoders still have some limitations, especially unable to effectively capture the complex temporal dependencies in long-term time series data, and the robustness and generalization ability of the models are poor when facing high noise and variable environments. Therefore, improving the temporal data modeling ability and multi-modal data fusion ability of the models has become an important direction for the current technological development.

[0005] The prior art 1, "Method and System for Blade Icing Recognition Based on Conditional Variational Autoencoder", proposes a method based on the conditional variational autoencoder (CVAE). This method trains the conditional variational autoencoder to identify the abnormal state of blade icing. This technology relies on the Gaussian mixture model to cluster data and locates abnormal parameters through the reconstruction probability, and then conducts anomaly detection and diagnostic reasoning (Inventors: Zhu Junjie, etc., Application Publication No.: CN117469105A). Prior art 1 can relatively accurately identify the anomalies of blade icing, but its disadvantage is that although the conditional variational autoencoder can model data, there is still a problem of insufficient ability to model long-distance dependencies in complex time-series data. The process of wind turbine blade icing is a dynamically changing time-series process, in which there are complex time-series dependencies among multiple environmental factors. The conditional variational autoencoder model is difficult to effectively capture these long-term dependencies, thus affecting the detection accuracy and robustness of the model in complex environments.

[0006] The prior art 2, "Method and System for Monitoring Wind Turbine Blade Icing", uses the Extreme Gradient Boosting (XGBoost) model and the Isolation Forest model, and combines the real-time operation data of the wind turbine and the environmental temperature to monitor the wind turbine blade icing. This technology improves the accuracy of monitoring and the generalizability of the model by establishing a dynamic regression model and an error monitoring mechanism (Inventors: Liu Yu, etc., Authorization Publication No.: CN114753980B). Although prior art 2 improves the monitoring accuracy to a certain extent, due to excessive reliance on manual feature engineering and traditional machine learning models (such as XGBoost), and the lack of the ability to model long-term dependencies in time-series data, it is easy to be unable to effectively process the non-linear relationships and long-term time-series dependencies in the data in a complex dynamic environment. Therefore, the adaptability and generalization ability of the model may be insufficient when facing changing environmental conditions. Summary of the Invention

[0007] The present invention aims at the above problems and provides a method for detecting the icing state of wind turbine blades with good accuracy and robustness.

[0008] To achieve the above object, the present invention adopts the following technical solutions. The present invention includes the following steps:

[0009] Step 1: Collect multi-modal data at the wind turbine blade.

[0010] Step 2: Fuse the multi-modal data.

[0011] Step 3: Use the Min-Max Normalization method to map the data to a unified numerical range; and use the sliding window method to divide the data to generate multiple time-series subsequences.

[0012] Step 4: Feature compression of data is performed by an encoder;

[0013] Step 5: The decoder processes data through a masked multi-head attention mechanism;

[0014] Step 6: The processed data calculates the icing probability through a binary cross-entropy loss function.

[0015] As a preferred solution, in step 1 of the present invention, the multi-modal data includes temperature, humidity, air density, and air pressure data.

[0016] The collected data includes temperature change data, humidity fluctuation data, air density signal data, and air pressure change data during the operation of the wind turbine. The temperature change data, humidity fluctuation data, air density signal data, and air pressure change data can be collected by an anemometer tower near the wind turbine.

[0017] As another preferred solution, in step 2 of the present invention, let the data of the temperature, humidity, air density, and air pressure sensors (the sensors mentioned here and subsequently refer to the sensors on the anemometer tower) be T(t), H(t), V(t), and P(t) respectively, where t represents the time step; if there is missing data at some time steps t k , the second-order Lagrange interpolation formula is used to process the missing data; the second-order Lagrange interpolation formula is as follows:

[0018]

[0019] where f(t k ) is the missing sensor data value, t k is the time step to be interpolated, t1, t2, t3 are known adjacent time steps, and f(t1), f(t2), f(t3) are the sensor data values of the known time steps respectively;

[0020] Using this interpolation formula, the sensor data T(t k ), H(t k ), V(t k ), P(t k ) of the missing time step are calculated.

[0021] As another preferred solution, in step 2 of the present invention, the Pearson correlation coefficient is used to analyze the correlation between the temperature, humidity, air density, air pressure sensor data and the icing state of the wind turbine blade; the calculation formula of the Pearson correlation coefficient is:

[0022]

[0023] where cov(X,Y) is the covariance of variables X and Y, σ X and σY is the standard deviation of X and Y; The Pearson correlation coefficient is used to measure the linear correlation between two variables, with a value range of [-1, 1]. When the value is 1 or -1, it indicates a perfect positive or negative correlation, and when the value is 0 or close to 0, it indicates no correlation;

[0024] Perform a correlation analysis on the sensor data T(t), H(t), V(t), P(t) at each time step to obtain the correlation coefficients between them and the icing state variables; According to the results of the Pearson correlation coefficient, select the features with a stronger correlation with wind power as the input features for the final fusion.

[0025] As another preferred solution, in step 2 of the present invention, the sensor data are fused by weighted summation, and the weights are adjusted according to the results of the correlation analysis to generate the comprehensive feature F(t):

[0026] F(t) = w T ·T(t) + w H ·H(t) + w V ·V(t) + w P ·P(t)

[0027] where, w T , w H , w V and w P are the weight coefficients adjusted according to the Pearson correlation coefficient, reflecting the influence degree of each sensor data on wind power; The feature with a larger weight coefficient indicates that the sensor data has a greater influence on wind power, and vice versa.

[0028] As another preferred solution, in step 3 of the present invention, the unified numerical range is [0, 1]. For the original data x(t) of a certain sensor at time step t, its normalization processing formula is:

[0029]

[0030] where, min(x) and max(x) respectively represent the minimum and maximum values of the sensor data in the entire dataset, and x norm (t) is the normalized data value, ensuring that each type of sensor data is compressed into the interval of [0, 1].

[0031] As another preferred solution, in step 3 of the present invention, by setting a window w with a fixed size and sliding it in the data according to time steps, a data subset with time series dependence is obtained; Setting the window size as w, the subsequence within the i-th time window is:

[0032] S i = {x norm (i), xnorm (i + 1), …, x norm (i + w - 1)}

[0033] Each time window S i corresponds to the sensor data of w consecutive time steps; the subsequence S generated by the sliding window i is used as input data for subsequent processing by the Autoencoder and Transformer modules.

[0034] As another preferred solution, the step size of the sliding window described in the present invention is usually set to 1.

[0035] As another preferred solution, in step 4 of the present invention, the input data of the encoder is the normalized multi-modal data S i ∈ R 30×4 , where the window w = 30 and the feature dimension d = 4 represent the sensor data of four sensors, namely temperature, humidity, air density, and air pressure; R represents the set of real numbers, that is, each element in the normalized sensor data is a real number.

[0036] The encoder structure is as follows: in the embedding input layer, the input features are extended to the dimension matching the multi-head attention through a linear layer. Starting from the input data S i ∈ R 30×4 is converted to S i ′ ∈ R 30×64 ; each subsequence contains the sensor data of 30 consecutive time steps, and each time step contains 64 feature dimensions; then in the fully connected layer, the encoder part uses multiple fully connected layers for step-by-step dimensionality reduction. The specific structure is: 30×64 → 256 → 128 → 64; the input data is compressed to 256 dimensions and 128 dimensions through the fully connected layer, and finally outputs 64-dimensional low-dimensional features h(t); in the activation function, Leaky ReLU is used as the activation function, and the negative slope coefficient is 0.01 to avoid gradient disappearance and retain negative value information;

[0037] The compressed low-dimensional features h(t) are passed to the Transformer model to capture the dynamic changes and long-term dependencies in the data;

[0038] In the Transformer model, the multi-head self-attention mechanism (Multi-HeadAttention) is adopted to enhance the learning ability of the model; the outputs of multiple attention heads are combined, and each attention head captures the features of the input data from different perspectives; the self-attention calculation formula is as follows:

[0039]

[0040] where Q, K, and V are linearly mapped from h(t).

[0041] As another preferred solution, the number of attention heads H = 8 in the present invention, and the dimension of each attention head is Each attention head captures different features in the input data; the dimension of the self-attention calculation formula is 8×8.

[0042] As another preferred solution, the Transformer model in the present invention adopts residual connection and layer normalization, and enhances the non-linear expression ability of the model through a per-bit feed-forward network (FFN).

[0043] As another preferred solution, the structure of the feed-forward network in the present invention is 64→16→64. The input features are compressed from 64 dimensions to 16 dimensions through a fully connected layer and then expanded back to 16 dimensions. The activation function adopts Leaky ReLU; 6 layers of Transformer blocks are stacked, and each layer adopts a combination of multi-head self-attention and feed-forward network for deep learning and time series dependence modeling;

[0044] The Transformer model constructs the time series feature z(t) through the self-attention mechanism and the number of stacked layers. The time series feature z(t) expresses the state of the wind turbine blade at each time step and contains the complex time series dependence information of the original input data; the time series feature z(t) is represented by the following formula:

[0045]

[0046] Among them, h(t) is the feature compressed by the Autoencoder, and α i (t) is the weight coefficient calculated through the self-attention mechanism, which is used to dynamically weight the feature importance of different time steps.

[0047] As another preferred solution, in step 5 of the present invention, the input of the decoder is the time series feature z(t) ∈ R output by the Transformer model 64 , in the masked multi-head attention calculation, the number of heads is set to H = 8, and the dimension of each attention head is A mask matrix M is introduced in the attention calculation to mask the correlation weights of future time steps:

[0048]

[0049] Among them, the mask matrix M is a lower triangular matrix, the positions of future time steps are set to -∞, and the current and historical time steps are 0;

[0050] After calculation, the output result is processed through residual connection and layer normalization, and the formula is:

[0051] z masked(t) = LayerNorm(z(t) + MaskedAttention(z(t))

[0052] Next, after the position-wise feed-forward network (Position-wise FFN) processes z_masked(t) ∈ R^64, it outputs a new feature z_weighted(t), which also goes through residual connection and layer normalization. The position-wise feed-forward network maps the 64-dimensional feature to 16 dimensions through linear projection, then expands it back to 64 dimensions, and finally processes it through the Leaky ReLU activation function with a negative slope coefficient of 0.01; the output of the feed-forward network also goes through residual connection and layer normalization, and the formula is:

[0053] z weighted (t) = LayerNorm(z masked (t) + FFN(z masked (t))

[0054] After completing feature weighting, the original data size is reconstructed through the decoder; the decoder structure uses fully connected layers with gradually increasing dimensions and finally restores to the same dimension as the input data.

[0055] As another preferred solution, the decoder in the present invention expands the feature z weighted (t) ∈ R 64 to 128 dimensions and 256 dimensions through fully connected layers. The last fully connected layer in the decoder maps the processed data back to the size of the original input window, which is 30×64, that is, reshapes the 256-dimensional feature into a 30×64 matrix to ensure that the reconstructed data is consistent with the size of the input window.

[0056] Secondly, the position-wise feed-forward network in the decoder uses the Leaky ReLU activation function, and the output layer uses the Sigmoid function to limit the reconstructed value within the range of [0,1], which is consistent with the range of the normalized input data; the mathematical representation of the reconstruction process is:

[0057]

[0058] where W i and b d are decoder parameters, is the reconstructed sensor data window.

[0059] In addition, in step 6 of the present invention, by calculating the difference between the predicted icing probability output by the model and the actual label, the binary cross-entropy loss function is used to quantify the prediction error; the larger the cross-entropy value, the less effectively the model can predict the icing probability, and the worse the prediction quality, indicating that there are anomalies in the input data; on the contrary, a smaller cross-entropy value indicates better model prediction and that the input data is within the normal range;

[0060] Let the feature after being processed by the Autoencoder and the Transformer be z weighted (t), and the reconstructed data obtained through the decoder is The original input data is F(t); for the prediction of the icing probability, the reconstructed data output by the decoder is converted into the probability value of icing through the Sigmoid activation function; the output probability of the decoder is:

[0061]

[0062] where, W i and b d are the parameters of the decoder, is the predicted icing probability generated by the decoder, represents the icing occurrence probability corresponding to each time step t;

[0063] Using the binary cross-entropy loss function to optimize the generation of the icing probability, the calculation formula is as follows:

[0064]

[0065] where w = 30 is the length of the time series, y i is the actual label, indicating whether it is iced at time step t, 1 means iced, and 0 means not iced; is the predicted icing probability of the model.

[0066] Advantages of the present invention.

[0067] The present invention fuses multi-modal data at the wind turbine blade, ensuring the collaborative analysis of sensor data and improving the accuracy and robustness of icing state recognition. When dealing with the problem of wind turbine blade icing, the present invention can more comprehensively consider environmental factors, avoiding the limitations of over-reliance on manual feature engineering or the inability of the model to handle non-linear time series dependencies in traditional methods.

[0068] The present invention uses the maximum-minimum normalization method to standardize various types of sensor data, eliminating the scale differences of different sensor data; at the same time, the sliding window method is used to process time series data to capture time-dependent relationships.

[0069] Icing is a process, not instantaneously iced, and reaching the icing standard is a threshold. What is generated in step 3 of the present invention is a time series, and the trend is seen through time dependence, and the icing probability is calculated through the binary cross-entropy loss function, rather than a binary classification problem (iced or not iced), which is helpful for early warning. Description of the drawings

[0070] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. The protection scope of the present invention is not limited only to the description of the following content.

[0071] Figure 1 is the overall architecture diagram of the present invention.

[0072] Figure 2 is the system flow chart of the present invention.

[0073] Figure 3 is the comparison of RMSE and MAPE performance of different models in the wind turbine blade icing detection task. Specific Embodiments

[0074] As shown in the figure, the present invention collects and fuses multi-dimensional raw data.

[0075] During the inspection process of wind turbine blades, multi-modal data such as temperature, humidity, air density, and air pressure are collected first. The collected data includes the temperature change, humidity fluctuation, air density signal, and air pressure change during the operation of the wind turbine. These data provide information in multiple dimensions for the operation state of the wind turbine blade, and thus provide a rich reference basis for subsequent anomaly detection.

[0076] To improve the accuracy and robustness of anomaly detection, the present invention unifies the processing of data from different sensors through a multi-modal data fusion method. Specifically, the raw data of temperature, humidity, air density, and air pressure are preprocessed and then fused to generate a comprehensive feature vector as the input for the subsequent model. Through this data fusion, it can be ensured that the data from different sensors are analyzed collaboratively at the same time step, maximizing the accuracy of ice state recognition.

[0077] In the data fusion stage, let the data of the temperature, humidity, air density, and air pressure sensors be T(t), H(t), V(t), and P(t) respectively, where t represents the time step. If there is missing data at some time steps t k the second-order Lagrange interpolation formula is used to process the missing data to ensure the continuity of the data. The second-order Lagrange interpolation formula is as follows:

[0078]

[0079] where f(t k ) is the missing sensor data value, t k is the time step to be interpolated, t1, t2, t3 are the known adjacent time steps, and f(t1), f(t2), f(t3) are the sensor data values of the known time steps respectively.

[0080] Using this interpolation formula, the sensor data T(t k), H(t k ), V(t k ), P(t k ), to ensure the data integrity at each time step.

[0081] To improve the accuracy and robustness of data fusion, the Pearson correlation coefficient is used to analyze the correlation between the data of temperature, humidity, air density, barometric pressure sensors and the icing state of wind turbine blades. The calculation formula of the Pearson correlation coefficient is:

[0082]

[0083] where cov(X,Y) is the covariance of variables X and Y, and σ X and σ Y are the standard deviations of X and Y. The Pearson correlation coefficient is used to measure the linear correlation between two variables, with a value range of [-1,1]. When the value is 1 or -1, it indicates a perfect positive or negative correlation, and when the value is 0 or close to 0, it indicates no correlation.

[0084] Perform correlation analysis on the sensor data T(t), H(t), V(t), P(t) at each time step to obtain the correlation coefficients between them and the icing state variable. According to the results of the Pearson correlation coefficient, select the features with stronger correlation with wind power as the input features for the final fusion.

[0085] On this basis, fuse the sensor data by weighted summation and adjust the weights according to the results of the correlation analysis to generate the comprehensive feature F(t):

[0086] F(t) = w T ·T(t) + w H ·H(t) + w V ·V(t) + w P ·P(t)

[0087] where w T , w H , w V , and w P are the weight coefficients adjusted according to the Pearson correlation coefficient, reflecting the influence degree of each sensor data on wind power. A feature with a larger weight coefficient indicates that the sensor data has a greater influence on wind power, and vice versa.

[0088] Through this weighted fusion method, it can ensure that the sensor data highly correlated with wind power and icing state is given priority, and an optimized time series feature F(t) is generated as the input for the subsequent model.

[0089] The present invention is about maximum - minimum normalization and sliding window generation.

[0090] For all kinds of sensor data collected, before inputting it into the subsequent model, it is first normalized to eliminate the influence of the scale differences of different sensor data. Specifically, the Min-Max Normalization method is used to map the data of each sensor to a unified numerical range [0, 1], so as to ensure that the data of each sensor can be compared on the same scale. For the original data x(t) of a certain sensor at time step t, its normalization formula is:

[0091]

[0092] where min(x) and max(x) respectively represent the minimum and maximum values of the sensor data in the entire dataset, and x norm (t) is the normalized data value, ensuring that the data of each sensor is compressed into the interval [0, 1].

[0093] Secondly, considering that the wind turbine blade operation data has strong temporal characteristics, the sliding window method is used to divide the data to generate multiple temporal subsequences. The sliding window technique obtains a subset of data with temporal dependence by setting a window w of a fixed size and sliding it in the data according to time steps. Setting the window size to w, the subsequence within the i-th time window is:

[0094] S i ={x norm (i), x norm (i + 1), …, x norm (i + w - 1)}

[0095] Each time window S i corresponds to the sensor data of consecutive w time steps. In this process, the step size of the sliding window is usually set to 1 to ensure that each time step in the temporal data can be fully learned. The subsequence S i generated by the sliding window will be used as input data for subsequent Autoencoder and Transformer modules for processing, and these subsequences can better capture the temporal dependence relationship of the data.

[0096] The present invention features feature compression and temporal dependence modeling.

[0097] In the present invention, an encoder model is used to perform feature compression on multimodal data. The input data is the normalized multimodal data S i ∈R 30×4 , where the window size w = 30 and the feature dimension d = 4 represent the data of four sensors: temperature, humidity, air density, and air pressure.

[0098] The encoder structure is as follows: In the embedded input layer, the input features are extended to the dimension matching the multi-head attention through a linear layer, that is, from the input data S i ∈R 30×3 is converted to S i ′∈R 30×64 . Each subsequence contains sensor data of 30 consecutive time steps, and each time step contains 64 feature dimensions. Immediately after that, in the fully connected layer, the encoder part uses multiple fully connected layers for step-by-step dimensionality reduction. The specific structure is: 30×64→256→128→64. The input data is compressed to 256 dimensions and 128 dimensions through the fully connected layer, and finally the 64-dimensional low-dimensional feature h(t) is output. In the activation function, Leaky ReLU is used as the activation function, and the negative slope coefficient is 0.01 to avoid gradient vanishing and retain negative value information.

[0099] Next, the compressed low-dimensional feature h(t) is passed to the Transformer model to further capture the dynamic changes and long-term dependencies in the data.

[0100] In the Transformer model, the multi-head self-attention mechanism (Multi-HeadAttention) is adopted to enhance the learning ability of the model. By combining the outputs of multiple attention heads, each head can capture the features of the input data from different perspectives. Specifically, the number of heads H = 8 is selected, and the dimension of each attention head is so as to ensure that each head can capture different features in the input data. The self-attention calculation method is as follows:

[0101]

[0102] Among them, Q, K, and V are linearly mapped and generated by h(t), and the dimension is 8×8.

[0103] To prevent gradient vanishing and optimize the training process, the Transformer model adopts residual connection and layer normalization, and these techniques effectively improve the training efficiency of the model. Next, through the pointwise feed-forward network (FFN), the non-linear expression ability of the model is further enhanced. The structure of the feed-forward network is 64→16→64. The input features are compressed from 64 dimensions to 16 dimensions through the fully connected layer and then expanded back to 16 dimensions. The activation function also uses Leaky ReLU. To improve the expression ability of the model, the present invention stacks 6 layers of Transformer blocks, and each layer adopts a combination of multi-head self-attention and feed-forward network, so as to realize deep learning and time series dependency modeling.

[0104] The present invention utilizes the self-attention mechanism and the encoder to capture long-term dependencies in multi-modal data, and improves the adaptability of the model to complex dynamic environments through the fusion of multi-modal data, thereby realizing more accurate and robust ice accretion anomaly detection for wind turbine blades and solving problems such as poor model accuracy and weak generalization ability in the prior art.

[0105] Based on the internal structure and hierarchy of the Transformer model with self-attention mechanism, the present invention is used to take multi-modal data as input to check the ice accretion condition of the blade. The encoder is used to compress data features. The Transformer enhances the modeling ability of long-time series dependencies through the multi-head (such as eight-head) self-attention mechanism. Especially in complex dynamic environments, it improves the accuracy of ice accretion detection and can accurately identify the risk of ice accretion during the inspection of wind turbine blades.

[0106] The self-attention mechanism enables the model to pay more attention to the key time points of abnormal changes, thereby improving the accuracy and robustness of anomaly detection. Especially in the complex dynamic environment of wind turbine blades, it can effectively capture the long-term dependencies and time series patterns in the data. The present invention can solve the problems of insufficient modeling ability for time series data and poor adaptability to dynamic environments in the prior art.

[0107] The encoder of the present invention adopts a deep fully-connected network activated by Leaky ReLU (30×4→256→128→64) to realize feature dimensionality reduction.

[0108] The present invention integrates the eight-head self-attention mechanism and six-layer stacked Transformer modules to capture long-time series dependencies.

[0109] The present invention adjusts the feature weights through the masked multi-head attention mechanism and the position-wise feed-forward network, focuses on the changes at key time points, thereby enhancing the anomaly detection ability of the model; effectively captures the non-linear relationships and dynamic changes in time series data.

[0110] Finally, the Transformer model constructs the time series feature z(t) through the self-attention mechanism and the number of stacked layers, which expresses the state of the wind turbine blade at each time step and contains the complex time series dependency information of the original input data. The time series feature z(t) can be expressed by the following formula:

[0111]

[0112] where h(t) is the feature compressed by the Autoencoder, and α i (t) is the weight coefficient calculated by the self-attention mechanism, dynamically weighting the importance of features at different time steps.

[0113] Self-attention weighting and feature restoration of the present invention.

[0114] After being processed by the Transformer module, the obtained temporal feature z(t) ∈ R 64 Combines the long-term dependencies and dynamic changes in the temporal data. To further enhance the focusing ability on key time steps, the present invention dynamically adjusts the feature weights through the Masked Multi-Head Attention mechanism to capture the dependencies between time steps and ensure the autoregressive property. Specifically, the input to the decoder is the temporal feature z(t) ∈ R 64 output by the Transformer model, and this feature will be further processed through the Masked Multi-Head Attention mechanism. In the Masked Multi-Head Attention calculation, the number of heads is set to H = 8, and the dimension of each attention head is A mask matrix M is introduced in the attention calculation to mask the correlation weights of future time steps:

[0115]

[0116] where the mask matrix M is a lower triangular matrix, the positions of future time steps are set to -∞, and the current and historical time steps are 0.

[0117] The decoder of the present invention introduces Masked Multi-Head Attention and a lower triangular mask matrix to strengthen the feature focusing ability on key time steps.

[0118] After the calculation, the output result is processed through a residual connection and layer normalization, and the formula is:

[0119] z masked (t) = LatyerNorm(z(t) + MaskedAttention(z(t))

[0120] Next, the Position-wise FFN outputs z masked (t) ∈ R 64 as the input to be passed to the Position-wise FFN. This network maps the 64-dimensional feature to 16 dimensions through a linear projection, then expands it back to 64 dimensions, and finally processes it through the Leaky ReLU activation function with the negative slope coefficient set to 0.01. This process ensures consistency with the aforementioned framework. The output of the feed-forward network will also go through a residual connection and layer normalization, and the formula is:

[0121] z weighted (t) = LayerNorm(z masked (t) + FFN(z masked (t))

[0122] After completing feature weighting, the model reconstructs the original data size through the decoder. The decoder structure uses fully connected layers with gradually increasing dimensions and finally restores to the same dimension as the input data. Specifically, the decoder expands the feature z weighted (t) ∈ R 64 to 128 dimensions and 256 dimensions through fully connected layers, and finally the output layer reshapes the 256-dimensional feature into a 30×64 matrix to ensure that the reconstructed data is consistent with the input window size.

[0123] To introduce non-linear transformation, the Leaky ReLU activation function is used in the hidden layer, while the Sigmoid function is used in the output layer to limit the reconstructed value within the range of [0,1], which is consistent with the range of the normalized input data. The mathematical representation of the reconstruction process is:

[0124]

[0125] where W i and b d are the decoder parameters, is the reconstructed sensor data window.

[0126] What is generated in step 3 of the present invention is a time series, and the transformer itself has an attention mechanism. It can identify trends through time dependence and calculate the icing probability through the binary cross-entropy loss function, rather than a binary classification problem (icing or not icing), which helps with early warning.

[0127] Reconstruction error calculation and improvement of anomaly detection accuracy in the present invention.

[0128] The present invention is trained by providing training data and labels (icing or not icing), and an anomaly detection is performed using the trained model. The model calculates the icing probability based on the sensor data and determines whether there is abnormal icing through a set threshold.

[0129] After being processed by the encoder and Transformer, the generated icing probability is compared with the original data to evaluate the icing prediction quality of the data. Specifically, by calculating the difference between the icing probability output by the model and the actual label, the binary cross-entropy loss function is used to quantify the prediction error. The larger the cross-entropy value, the less effectively the model can predict the icing probability, indicating poor prediction quality and potentially indicating anomalies in the input data; conversely, a smaller cross-entropy value indicates better model prediction performance and that the input data is within the normal range.

[0130] Let the feature after being processed by the Auto encoder and Transformer be z weighted (t), and the reconstructed data obtained through the decoder is The original input data is F(t). For the prediction of icing probability, the reconstructed data output by the decoder is converted into the probability value of icing through the Sigmoid activation function. Specifically, the output probability of the decoder is:

[0131]

[0132] where W i and b d are the parameters of the decoder, is the predicted icing probability generated by the decoder, representing the probability of icing occurrence corresponding to each time step t.

[0133] To optimize the generation of icing probability, the binary cross-entropy loss function is used, and its calculation formula is as follows:

[0134]

[0135] where w = 30 is the length of the time series, y i is the actual label, indicating whether icing occurs at time step t (1 means icing, 0 means no icing), is the predicted icing probability of the model.

[0136] After calculating the icing probability, the icing situation can be judged according to the threshold. A standard with a threshold of 50% can be set. The reason for choosing 50% as the threshold is that the icing phenomenon is usually a gradual accumulation process, and the probability of its occurrence is not determined instantaneously but increases gradually over time. When the predicted icing probability by the model exceeds 50%, it indicates that the risk of icing within this time step is relatively high enough to trigger an alarm and take relevant measures. This threshold can effectively prevent premature alarms and can also respond in a timely manner when the icing risk is obvious, thus ensuring the safety of wind turbine blades under icing conditions. Therefore, choosing 50% as the threshold is a reasonable choice that balances the accuracy of icing prediction and the timeliness of early warning.

[0137] This loss function optimizes the model's ability to generate icing probability by minimizing the difference between the predicted icing probability and the actual label, and ensures that the model can accurately identify the risk of wind turbine blade icing, issue alarms in a timely manner, and improve the anomaly detection ability.

[0138] To verify the effectiveness of the Transformer model proposed in the present invention in wind turbine blade icing detection, the present invention has conducted a large number of experiments and comparative analyses. The experimental environment used the Windows 10 operating system, and the hardware configuration was an 8th generation Intel Core i7 processor, equipped with an NVIDIA RTX 2060 CPU. The experimental framework was based on Python 3.12 and the deep learning frameworks of PyTorch and Keras. During the training process, appropriate hyperparameters and optimization strategies were adopted to ensure the stability and accuracy of the algorithm.

[0139] The dataset used in the present invention is from a wind farm, specifically including multi-source time series data and icing conditions provided by the wind measurement towers in 3 areas of the wind farm. The data includes information such as temperature, humidity, and air pressure, which are collected by sensors at different heights at different time steps. The data collection time span was from 0:00:00 on January 1, 2024 to 23:55:00 on March 31, 2024, and the sampling frequency was once every 15 minutes. The dataset records the air temperature (unit: °C), air pressure (unit: kPa), humidity (unit: %), air density (kg / m 3 ) and whether icing occurred.

[0140] The total data volume was 77,760 records, but due to situations such as shutdowns and faults, some data was invalid. After inspection, the final valid data volume was 75,563 records. For model training, samples with normal operation and no icing were labeled as 0, and samples with icing were labeled as 1 and used as training labels. The dataset was large, with normal samples accounting for 63.6% and icing samples only accounting for 36.4%. Since the proportion of normal samples in the dataset was relatively high, in order to balance the proportion of icing and normal samples, the present invention adopted the method of random undersampling and discarded some normal samples. After adjustment, the proportion of icing samples to normal samples reached 1:1, that is, there were 27,504 positive and negative samples respectively. It was specifically divided into 3 datasets, and each dataset represented the data collected from one area of the wind farm, aiming to ensure the representativeness of the data from different wind farms and further improve the generalization ability and robustness of the model through cross-wind farm area data training and verification.

[0141] In the experiment, 80% of the dataset was used for training, 10% for validation, and 10% for testing. During the model evaluation process, RMSE (Root Mean Square Error) and MAPE (Mean Absolute Percentage Error) were used as the main performance indicators to measure the prediction error and prediction accuracy of the model respectively.

[0142] RMSE (Root Mean Squared Error) is a commonly used metric to measure the gap between predicted values and actual values. For each time step t, the true value is y t , and the model's predicted value is Then the calculation formula for RMSE is:

[0143]

[0144] where n is the number of samples in the dataset, y t is the true label at the t-th time step, is the predicted value at the t-th time step. The smaller the value of RMSE, the smaller the prediction error of the model, and the better the prediction effect of the model.

[0145] MAPE (Mean Absolute Percentage Error) is a commonly used metric to measure the percentage error of predicted values relative to actual values. Specifically, MAPE evaluates the performance of the model by calculating the error percentage for each time step and averaging the errors for all time steps. Its calculation formula is:

[0146]

[0147] where y t and represent the true value and the predicted value at the t-th time step respectively, and n is the number of samples in the dataset. The smaller MAPE is, the higher the degree of fitting of the model to the data and the stronger the prediction accuracy.

[0148] By using these two metrics, RMSE and MAPE, the present invention can comprehensively evaluate the performance of the proposed Transformer model in the task of wind turbine blade icing detection. RMSE is mainly used to measure the error of the model in numerical prediction, while MAPE helps to evaluate the performance of the model in terms of relative error. The two together reflect the prediction accuracy and fitting ability of the model.

[0149] To verify the effectiveness of the Transformer model in wind turbine blade icing detection, several common models of the same type in the training environment of the present invention were used for performance comparison on the dataset. All models were trained on the same dataset and evaluated for performance using the same evaluation metrics (RMSE and MAPE).

[0150] (1) GCN Model: This model only uses the Graph Convolutional Network (GCN) to model the spatial dependencies between multi-sensor signals, ignoring the temporal information. This model mainly focuses on capturing spatial dependencies and is suitable for tasks with obvious static data features. Through the training and testing of the GCN model, the performance difference between its effect in the icing detection task and that of the GCN-Transformer model can be compared.

[0151] (2) LSTM Model: This model uses the Long Short-Term Memory network (LSTM) to process temporal data. LSTM can capture long-term dependency information well. When constructing the model, 128 LSTM neurons were selected, and the learning rate and batch size were adjusted through experiments. Finally, the best LSTM model was trained on the dataset. The LSTM model was compared with other models based on temporal data to evaluate its accuracy in the wind turbine blade icing detection task.

[0152] (3) CNN Model: This model uses the Convolutional Neural Network (CNN) to process the wind turbine blade icing detection task. CNN extracts local features of data through convolutional layers and is especially good at processing image-like data, but it can also achieve good results on temporal data. By constructing a CNN network and training the data, its performance in the detection task was evaluated and further compared with other models.

[0153] The above comparison models represent different deep learning architectures respectively: temporal modeling, spatial modeling, and spatio-temporal combination, aiming to evaluate the advantages and disadvantages of the models from different perspectives.

[0154] Table 1 shows the performance comparison results after experiments on the data of three wind farm areas, which lists the experimental results of the Transformer, SVM, LSTM, and XGBoost models. During the experiment, the dataset was divided into three parts, each corresponding to the data of one wind farm area, and the performance of each model on these datasets was evaluated. The evaluation metrics were RMSE and MAPE, which were used to measure the prediction error and accuracy of the model in the icing detection task.

[0155] Table 1

[0156]

[0157]

[0158] According to the experimental results in Table 1, it can be seen that the Transformer model shows good performance on the datasets of the three wind farm areas. In terms of the two indicators of RMSE and MAPE, Transformer outperforms other comparison models in all the data of the wind farm areas, showing lower prediction errors and higher accuracy. According to the average values, the average RMSE of Transformer is 0.57 and the MAPE is 3.60%, which are significantly lower than those of GCN (average RMSE = 0.72, MAPE = 3.98%) and other comparison models.

[0159] Figure 3 The performance comparison of different models under the two evaluation indicators of RMSE and MAPE is shown. Compared with other models, Transformer has achieved an overall performance improvement on all datasets.

[0160] It can be understood that the above specific description of the present invention is only for explaining the present invention and is not limited to the technical solutions described in the embodiments of the present invention. Those of ordinary skill in the art should understand that the present invention can still be modified or equivalently replaced to achieve the same technical effects; as long as the usage requirements are met, they are all within the protection scope of the present invention.

Claims

1. Method for detecting icing state of wind turbine blade, characterized in that It includes the following steps: Step 1: Collect multi-modal data at the wind turbine blade. Step 2: Fuse the multi-modal data. Step 3: Use the max-min normalization method to map the data to a unified numerical range; and use the sliding window method to divide the data to generate multiple time series subsequences. Step 4: Compress the features of the data through an encoder. Step 5: The decoder processes the data through the masked multi-head attention mechanism. Step 6: Calculate the icing probability of the processed data through the binary cross-entropy loss function.

2. The method for detecting the icing state of a wind turbine blade according to claim 1, wherein In the above Step 1, the multi-modal data includes temperature, humidity, air density, and air pressure data. In step 2, let the data of the temperature, humidity, air density, and air pressure sensors be T(t), H(t), V(t), and P(t), respectively, where t represents the time step; if there are missing data at some time steps t k with missing data, use the second-order Lagrange interpolation formula to handle the missing data; the second-order Lagrange interpolation formula is as follows: where f(t k ) is the missing sensor data value, t k is the time step for which interpolation is required, t1, t2, and t3 are known adjacent time steps, and f(t1), f(t2), and f(t3) are the sensor data values at the known time steps respectively; Using this interpolation formula, calculate the sensor data T(t k )、H(t k )、V(t k )、P(t k ) at the missing time step.

3. The wind turbine blade icing state detection method according to claim 1, characterized in that In the above Step 2, the Pearson correlation coefficient is used to analyze the correlation between the temperature, humidity, air density, air pressure sensor data and the icing state of the wind turbine blade; the calculation formula of the Pearson correlation coefficient is: Among them, cov(X,Y) is the covariance of variables X and Y, and σ X and σ Y are the standard deviations of X and Y; the Pearson correlation coefficient is used to measure the linear correlation between two variables, with a value range of [-1,1]. When the value is 1 or -1, it indicates a perfect positive or negative correlation, and when the value is 0 or close to 0, it indicates no correlation; Perform correlation analysis on the sensor data T(t), H(t), V(t), P(t) at each time step to obtain their correlation coefficients with the icing state variable; according to the results of the Pearson correlation coefficient, select the features with stronger correlation with the wind power as the input features for the final fusion.

4. The wind turbine blade icing state detection method according to claim 1, characterized in that In the above Step 2, the sensor data are fused by weighted summation, and the weights are adjusted according to the results of the correlation analysis to generate the comprehensive feature F(t): F(t) = w T ·T(t) + w H ·H(t) + W V ·V(t) + w P ·P(t) Among them, w T , w H , w V and w P are weight coefficients adjusted according to the Pearson correlation coefficient, respectively, reflecting the influence degree of each sensor data on wind power; a feature with a larger weight coefficient indicates that the sensor data has a greater influence on wind power, and vice versa.

5. The wind turbine blade icing state detection method according to claim 1, characterized in that In Step 3, the unified numerical range is [0,1]. For the original data x(t) of a certain sensor at time step t, its normalization formula is: where min(x) and max(x) respectively represent the minimum and maximum values of the sensor data in the entire dataset, and x norm (t) is the normalized data value, ensuring that each type of sensor data is compressed into the interval [0, 1]; In the above Step 3, by setting a window w of a fixed size and sliding it in the data according to the time step, a data subset with time series dependence is obtained; assuming the window size is w, the subsequence within the i-th time window is: S i = {x norm (i), x norm (i + 1), …, x norm (i + w - 1)} Each time window S i corresponds to the sensor data of consecutive w time steps; the subsequence S generated by the sliding window i is used as input data for subsequent processing by the Autoencoder and Transformer modules.

6. The wind turbine blade icing state detection method according to claim 1, wherein In step 4, the input data of the encoder is the normalized multimodal data S i ∈R 30×4 , where the window w = 30 and the feature dimension d = 4 represent the data of four sensors, namely temperature, humidity, air density, and air pressure; R represents the set of real numbers, that is, each element in the normalized sensor data is a real number; The encoder structure is as follows: In the embedded input layer, the input features are extended to the dimension matching the multi-head attention through a linear layer, starting from the input data \(S\) i \(\in \mathbb{R}\) 30×4 and converted to \(S'\) i \(\in \mathbb{R}\) 30×64 ; each subsequence contains sensor data for 30 consecutive time steps, and each time step contains 64 feature dimensions; immediately afterwards, in the fully connected layer, the encoder part uses multiple fully connected layers for step-by-step dimensionality reduction, and the specific structure is: \(30\times64\rightarrow256\rightarrow128\rightarrow64\); the input data is compressed to 256 dimensions and 128 dimensions through the fully connected layer, and finally the 64-dimensional low-dimensional feature \(h(t)\) is output; in the activation function, Leaky ReLU is used as the activation function with a negative slope coefficient of 0.01 to avoid gradient vanishing and retain negative value information; The compressed low-dimensional feature h(t) is passed to the Transformer model to capture the dynamic changes and long-term dependence relationships in the data. In the Transformer model, the multi-head self-attention mechanism is used to enhance the learning ability of the model; the outputs of multiple attention heads are combined, and each attention head captures the features of the input data from different perspectives; the self-attention calculation formula is as follows: Among them, Q, K, and V are linearly mapped from h(t).

7. The wind turbine blade icing state detection method according to claim 5, wherein The above Transformer model adopts residual connection and layer normalization, and enhances the nonlinear expression ability of the model through a pointwise feed-forward network. The structure of the above feed-forward network is 64→16→64. The input features are compressed from 64 dimensions to 16 dimensions through a fully connected layer and then expanded back to 16 dimensions. The activation function uses Leaky ReLU; 6 layers of Transformer blocks are stacked, and each layer adopts a combination of multi-head self-attention and a feed-forward network for deep learning and time series dependence modeling. The Transformer model constructs the time series feature z(t) through the self-attention mechanism and the number of stacked layers. The time series feature z(t) represents the state of the wind turbine blade at each time step and contains the complex time series dependence information of the original input data. The time series feature z(t) is represented by the following formula: Among them, h(t) is the feature compressed by the Autoencoder, and α i (t) is the weight coefficient calculated by the self-attention mechanism, which is used to dynamically weight the feature importance at different time steps.

8. The wind turbine blade icing state detection method according to claim 1, characterized in that In the above step 5, the input of the decoder is the temporal feature z(t) ∈ R output by the Transformer model 64 , in the masked multi-head attention calculation, the number of heads is set to H = 8, and the dimension of each attention head is A mask matrix m is introduced in the attention calculation to mask the correlation weights of future time steps: Among them, the mask matrix m is a lower triangular matrix, the future time step positions are set to -∞, and the current and historical time steps are 0. After calculation, the output result is processed through residual connection and layer normalization, and the formula is: z masked (t) = LayerNorm(z(t) + MaskedAttention(z(t)) Next, after the per-bit feed-forward network processes z_masked(t) ∈ R^64, a new feature z_weighted(t) is output, and it goes through residual connection and layer normalization; the per-bit feed-forward network maps the 64-dimensional feature to 16 dimensions through linear projection, then expands it back to 64 dimensions, and finally is processed through the Leaky ReLU activation function with a negative slope coefficient of 0.01; the output of the feed-forward network also goes through residual connection and layer normalization, and the formula is: z weighted (t) = LayerNorm(z masked (t) + FFN(z masked (t)) After feature weighting is completed, the original data size is reconstructed through the decoder; the decoder structure is restored to the same dimension as the input data through fully connected layers with gradually increasing dimensions.

9. The method for detecting the icing state of a wind turbine blade according to claim 8, wherein The per-bit feed-forward network in the decoder uses the Leaky ReLU activation function, and the output layer uses the Sigmoid function to limit the reconstructed value within the range of [0,1], which is consistent with the range of the normalized input data; the mathematical representation of the reconstruction process is: where W i and b d are decoder parameters, is the reconstructed sensor data window.

10. The wind turbine blade icing state detection method according to claim 1, wherein In step 6, by calculating the difference between the icing probability output by the model and the actual label, the binary cross-entropy loss function is used to quantify the prediction error; the larger the cross-entropy value, the less effectively the model predicts the icing probability, and the worse the prediction quality, indicating that there are anomalies in the input data; Conversely, a smaller cross-entropy value indicates that the model has a better prediction effect and the input data is within the normal range; Let the feature after being processed by the Autoencoder and the Transformer be z weighted (t), and the reconstructed data obtained by passing it through the decoder is The original input data is F(t); for the prediction of the icing probability, the reconstructed data output by the decoder is converted into the probability value of icing through the Sigmoid activation function; The output probability of the decoder is: Among them, W i and b d are parameters of the decoder, is the icing probability prediction generated by the decoder, representing the icing occurrence probability corresponding to each time step t; Using the binary cross-entropy loss function, optimize the generation of the icing probability, and the calculation formula is as follows: where w = 30 is the length of the time series, y i is the actual label indicating whether it freezes at time step t, with 1 indicating freezing and 0 indicating non-freezing; is the predicted freezing probability of the model.

Citation Information

Patent Citations

  • Wind turbine blade icing detection method based on deep learning and storage medium

    CN112832960A

  • Power line icing detection method, device and equipment and storage medium

    CN114387246A

  • Battery voltage estimation method and device and storage medium

    CN116660757A

  • Wind turbine generator blade icing degree prediction method, device, equipment and medium

    CN119442031A