Wind turbine gearbox state monitoring method and system based on self-supervised contrast residual map network

By applying a self-supervised and comparative residual graph network in wind turbines, problems such as long data time intervals and large subjectivity of feature selection in gearbox status monitoring are solved, efficient and accurate status monitoring and reliable early warning are achieved, and the operating efficiency and economic benefits of the wind farm are improved.

CN120087406AActive Publication Date: 2025-06-03HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510574232.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-06-03
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

The existing wind turbine gearbox status monitoring methods have problems such as long data time intervals, large subjectivity of feature selection, weak model generalization ability, large uncertainty in different calculations and low early warning reliability, resulting in low monitoring efficiency, low accuracy and high false alarm rate.

Method used

The method based on self-supervised comparison residual graph network is adopted to achieve accurate and efficient monitoring of the gearbox status of the wind turbine unit through steps such as data preprocessing, adaptive feature selection, graph data sample construction, model training, health guidance calculation and fault threshold determination.

Benefits of technology

Accurate monitoring of the gearbox status of wind turbine units is achieved, reliable warning is issued 30 to 40 hours in advance, and the accuracy of abnormal identification exceeds 90%, reducing operation and maintenance costs and reducing downtime caused by gearbox failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087406A_ABST
    Figure CN120087406A_ABST
Patent Text Reader

Abstract

The invention discloses a wind turbine gearbox state monitoring method and system based on a self-supervised contrast residual map network, and the method comprises the steps: obtaining original SCADA data, and carrying out the preprocessing; based on a self-adaptive feature selection method, selecting data features which are highly related to the wind turbine gearbox; constructing the data features into a graph data sample; based on the graph neural network and a comparative learning method, constructing a comparative residual graph neural network; loading a specific residual image neural network model trained by health data, inputting an unknown state sample into a pre-training model, performing distance calculation in a multi-dimensional space on an output result of the pre-training model and a prediction result of a health sample, and constructing gearbox health guidance based on an exponentially weighted moving average and a normalized average distance; determining a fault threshold based on a statistical process control technique; and comparing the relative size of the gearbox health guidance and the fault threshold value to realize online state monitoring. According to the invention, accurate and efficient monitoring of the operation state of the wind turbine generator gearbox is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of wind power operation and maintenance, and particularly to a method and system for monitoring the state of a wind turbine gearbox based on a self-supervised contrast residual graph network. Background Art

[0002] Wind turbines are often in harsh environments and frequently malfunction, resulting in an increasing demand for operation and maintenance. Traditional operation and maintenance methods such as shutdown maintenance, patrol inspection maintenance, and regular maintenance have many drawbacks. For example, shutdown maintenance affects the efficiency of wind farms, patrol inspection maintenance is costly and inefficient, and regular maintenance faces problems of cost, efficiency, and ineffective maintenance investment in normal units. Therefore, predictive maintenance can be applied in real industrial scenarios because it can accurately locate abnormal units, accurately predict faults, and achieve efficient and low-cost maintenance. Accurate and effective state monitoring is a core task of predictive maintenance.

[0003] Predictive maintenance for each component of a wind turbine is costly and, in many cases, redundant work. Therefore, monitoring the state of wind turbine subsystems has become a reasonable option. Among the multiple subsystems inside a wind turbine, the gearbox has a high failure rate and expensive repair costs, and has the greatest impact on the wind turbine after a failure. Therefore, the gearbox is the preferred monitoring target for the state monitoring task of subsystems.

[0004] Regarding the method for monitoring the state of the gearbox, traditional model-based and signal processing methods have problems such as difficult modeling and high costs. Data-driven state monitoring methods can avoid the investment in additional sensing devices and, at the same time, define and model different states based on an end-to-end method. In the prior art, although the monitoring method based on a normal behavior model can detect anomalies by comparing healthy data with real-time data, it has problems such as a large randomness in feature selection and weak model expression ability, resulting in a short warning time (usually less than 24 hours) and low accuracy (usually less than 80%). Some prominent problems are as follows: 1) Long data time interval: The data sampling interval of the SCADA system is long (usually 10 minutes), making it difficult to accurately extract early fault signals from it.

[0005] 2) Strong subjectivity in feature selection: Mainly relying on manual experience to select data features easily introduces subjective factors, resulting in information omission or redundancy; 3) Insufficient model generalization ability: Traditional machine learning models ignore the correlation between different data features and are difficult to capture dynamic fault features; 4) Uncertainty in difference calculation: When calculating the difference between normal samples and samples to be monitored, uncertainty will be generated due to the application of a single distance metric; 5) Low warning reliability: A single distance metric cannot comprehensively reflect the health status. At the same time, the setting of the fault threshold relies heavily on historical experience, resulting in a high false alarm rate.

[0006] As described above, although the data-driven method can effectively extract valuable information from the wind turbine data, its operation process still faces many challenges. Therefore, a new method and system for monitoring the state of the wind turbine gearbox are needed to solve the above problems. Summary of the Invention

[0007] The present invention provides a method and system for monitoring the state of a wind turbine gearbox based on a self-supervised contrast residual graph network, aiming to solve the defects existing in the actual operation of the existing methods for monitoring the state of the wind turbine gearbox. Through innovative technical means, accurate and efficient monitoring of the operating state of the wind turbine gearbox is realized, potential fault hazards are detected in advance, reliable warnings are issued, thereby reducing the operation and maintenance costs of the wind turbine, reducing the downtime caused by gearbox failures, and improving the overall operation efficiency and economic benefits of the wind farm.

[0008] The present invention adopts the following technical solutions to solve the technical problems: A method for monitoring the state of a wind turbine gearbox based on a self-supervised contrast residual graph network includes the following steps: Step S1, data preprocessing: Obtain the original SCADA data from the SCADA system installed in the wind farm and preprocess the original SCADA data. Step S2, adaptive feature selection: Based on the adaptive feature selection method, select the data features with strong correlation with the wind turbine gearbox. Step S3, graph data sample construction: Construct the data features after adaptive feature selection into graph data samples. Step S4, model training: Based on the graph neural network and the contrast learning method, construct a contrast residual graph neural network for information mining and model training. Step S5, health guidance calculation: Load the contrast residual graph neural network model trained with healthy data, input the unknown state samples into the pre-trained model, calculate the distance in the multi-dimensional space between the output result of the pre-trained model and the prediction result of the healthy samples, and construct the gearbox health guidance based on the exponentially weighted moving average and the normalized average distance. Step S6, fault threshold determination: Determine the fault threshold based on the statistical process control method. Step S7, state monitoring and warning: Compare the relative magnitudes of the gearbox health guidance and the fault threshold to realize the online state monitoring of the wind turbine gearbox.

[0009] Further, in step S1, the original SCADA data may contain null values, error values, and abnormal data caused by sensor failures; the preprocessing method includes: using a data cleaning algorithm to automatically identify and delete null value records; for error values, they are corrected according to the upper and lower limit ranges and logical relationships of the data, and deleted if they cannot be corrected; for abnormally fluctuating data, it is judged in combination with historical data and physical laws, and if it is caused by a sensor failure, data interpolation or filtering algorithms are used for repair.

[0010] Further, in step S2, the method for selecting data features strongly correlated with the wind turbine gearbox based on the adaptive feature selection method is as follows: For the original SCADA data, Spearman coefficients are selected for correlation analysis, and the formula is:

[0011] where is the Spearman correlation coefficient, is the rank difference of each pair of observed values, is the total number of observed values; Based on the operation and maintenance reports from the wind farm, five data features, namely the average oil temperature of the gearbox, the average oil temperature at the inlet of the gearbox, the average value of the oil inlet pressure of the gearbox, the average value of the impeller speed, and the average value of the generator speed, are selected as core features, and the average correlation coefficient between other features and the core features is calculated. The formula is:

[0012] where is the average correlation coefficient of other features relative to the core features, is the th other feature and the th core feature's correlation coefficient, is the number of core features; A quantization index is set. When the average correlation coefficient exceeds this quantization index, the corresponding data feature is selected; the variables with correlation coefficients exceeding this index are determined as strongly correlated variables and jointly form the model training samples with the core features; Finally, the feature data after feature selection is processed by maximum - minimum normalization, and the formula is:

[0013] where is the normalized data, is the original data, and are the minimum and maximum values in the original data respectively.

[0014] Further, in step S3, the method of constructing the data features after adaptive feature selection into graph data samples is as follows: Construct the features after adaptive feature selection into graph data samples. In this graph data sample, each feature serves as a node, and the correlation between features serves as an edge. The correlation uses the Spearman coefficient, and the adjacency matrix is constructed according to this process. The adjacency matrix is shown as follows:

[0015] In the formula, n is the number of nodes, ρ i represents the Spearman correlation coefficient between the i-th node and the j-th node, where, i , j = 1, 2,..., n.

[0016] Further, in step S4, the ratio residual graph neural network model adopts a self-supervised learning mode, sets two encoders for contrastive learning, the two encoders share weight parameters, each encoder contains a graph convolutional layer, a max pooling layer, and a residual connection, and at the same time uses the Leaky–ReLU function as the activation function, with the negative half-axis coefficient set to 0.1; the residual connection in the model is mainly used to ensure information integrity; During the model training process, the original samples are masked with Gaussian random noise with a standard deviation of 0.1, and the original samples and the masked samples are respectively input into the two encoders; the objective function of the model training is based on the contrastive loss, and at the same time L2-regularization is added during the training process; the calculation formula of the objective function is:

[0017] where and are the positive and negative sample feature representations respectively, sim represents the cosine similarity, is the temperature parameter; N is the number of samples, λ is the L2-regularization coefficient, w is the model parameter; the Stochastic Gradient Descent optimizer SGD with Nesterov momentum is used for model training.

[0018] Further, in step S5, the specific method for calculating the health guidance is as follows: Construct the graph data sample of the gearbox to be monitored, load the ratio residual graph neural network model trained with health data into the monitoring system, input the unknown state sample into the pre-trained model, and obtain a prediction value; Then adopt the Manhattan distance , Euclidean distance Cosine similarity 、 Chebyshev distance Calculate the average distance between healthy samples and samples to be monitored using the mean of these 4 distance metrics , and the calculation formula is:

[0019]

[0020]

[0021]

[0022]

[0023] Among them, a , b are two vectors for which the distance is to be calculated, a i 、 b i are respectively the elements in the two vectors, n is the dimension of the vector, Normed indicates normalizing the calculation result; on this basis, construct a gearbox health guideline based on exponentially weighted moving average HI , and the formula is:

[0024] where is the time health guideline, is the penalty factor, is the time health guideline, is the time normalized average distance; Next, perform median filtering denoising on the calculation result, and set the filter window size to 100.

[0025] Furthermore, in step S6, the method for determining the fault threshold based on the statistical process control method is as follows: Use the sample mean and the standard deviation to represent the mean and variance of the normal distribution, define the range exceeding the preset value as the fault state, and calculate the threshold according to the calculated state indication parameter according to the principle of statistical process control; considering the definition of the gearbox health guideline HI , when selecting the regional boundary value here, only consider the value corresponding to the upper boundary, and this upper boundary is the failure threshold Th defined here, and the calculation formula is:

[0026] Among them, is the mean of the health guidance sequence, is the standard deviation of the health guidance sequence.

[0027] A wind turbine gearbox condition monitoring system based on a self-supervised contrast residual graph network, comprising: a data preprocessing module, an adaptive feature selection module, a graph data sample construction module, a model training module, a health guidance calculation module, a fault threshold determination module, and a condition monitoring and early warning module; The data preprocessing module is used to obtain the original SCADA data from the SCADA system installed in the wind farm and preprocess the original SCADA data; The adaptive feature selection module selects data features with strong correlation with the wind turbine gearbox based on the adaptive feature selection method; The graph data sample construction module is used for graph data sample construction, and constructs the data features after adaptive feature selection into graph data samples; The model training module constructs a contrast residual graph neural network based on the graph neural network and the contrast learning method for information mining and model training; The health guidance calculation module loads the contrast residual graph neural network model trained with health data, inputs the unknown state samples into the pre-trained model, calculates the distance in the multi-dimensional space between the output result of the pre-trained model and the prediction result of the health samples, and constructs the gearbox health guidance based on the exponentially weighted moving average and the normalized average distance; The fault threshold determination module determines the fault threshold based on the statistical process control method; The condition monitoring and early warning module is used to compare the relative magnitudes of the gearbox health guidance and the fault threshold to achieve online condition monitoring of the wind turbine gearbox.

[0028] The beneficial effects of the present invention: (1) Precise and efficient monitoring: Adaptive feature selection effectively removes redundant information, retains key features, improves the monitoring efficiency and reliability, enables the model to focus on the core data, and enhances the monitoring accuracy.

[0029] (2) Powerful model performance: The contrast residual graph neural network model combines the advantages of multiple technologies. Through contrast learning and residual connection, it combines the correlation relationships between data features, enhances the model's expression ability and information extraction ability, accurately mines data features, and accurately identifies the gearbox condition.

[0030] (3) Reliable early warning mechanism: Construct health guidelines based on multiple distance metrics, construct gearbox health guidelines by combining exponentially weighted moving average, and determine the fault threshold by combining statistical process control to improve the reliability of the method. It can achieve abnormal early warning 30 - 40 hours in advance, and the abnormal recognition accuracy rate exceeds 90%, which buys time for the operation and maintenance process and reduces losses.

[0031] (4) Wide applicability: The method of the present invention realizes an end-to-end state monitoring of the gearbox for the SCADA data of wind turbines, provides methods and ideas, and can realize this function for different wind turbines and supporting SCADA systems, with generalization, and can meet the state monitoring requirements of different wind turbine units. Description of the Drawings

[0032] Figure 1 It is the flowchart of the method of the present invention.

[0033] Figure 2 It is the flowchart of the adaptive feature selection and graph sample construction of the present invention.

[0034] Figure 3 It is the architecture diagram of the CRGN model of the present invention. Specific Embodiments

[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0036] Referring to the attached Figure 1 drawings, the present invention provides a method for monitoring the state of a wind turbine gearbox based on a self-supervised contrast residual graph network, including the following steps: Step S1, data preprocessing, obtaining the original SCADA data from the SCADA system installed in the wind farm, and preprocessing the original SCADA data; The original SCADA data may have null values, error values, and abnormal data caused by sensor failures; the preprocessing methods include: using a data cleaning algorithm to automatically identify and delete null value records; for error values, correcting them according to the upper and lower limit ranges and logical relationships of the data, and deleting them if they cannot be corrected; for abnormally fluctuating data, judging in combination with historical data and physical laws, and if it is caused by a sensor failure, using data interpolation or filtering algorithms for repair.

[0037] Step S2, Adaptive Feature Selection: Based on the adaptive feature selection method, select data features that are highly correlated with the wind turbine gearbox. For the original SCADA data, first use the Anderson-Darling test to evaluate the normal distribution of each feature. Since in most cases the original data features do not follow a normal distribution, the Spearman coefficient is selected to conduct a correlation analysis. The formula is:

[0038] where is the Spearman correlation coefficient, is the rank difference of each pair of observed values, is the total number of observed values; Based on the operation and maintenance reports from the wind farm, select 5 data features, namely the average oil temperature of the gearbox, the average oil temperature at the inlet of the gearbox, the average oil inlet pressure of the gearbox, the average impeller speed, and the average generator speed, as the core features. Calculate the average correlation coefficient between other features and the core features. The formula is:

[0039] where is the average correlation coefficient of other features relative to the core features, is the th correlation coefficient between the th other feature and the th core feature, Set a quantization index. When the average correlation coefficient exceeds this quantization index, the corresponding data feature is selected. According to the information in the literature, this quantization index generally takes values in , and in this invention, it is selected as 0.6 based on the specific experimental process. Determine the variables with correlation coefficients exceeding this index as strongly correlated variables, which together with the core features form the model training samples; Finally, perform min-max normalization on the feature data after feature selection. The formula is:

[0040] where is the normalized data, is the original data, and are the minimum and maximum values in the original data respectively.

[0041] Step S3, Graph Data Sample Construction: Construct the data features after adaptive feature selection into graph data samples; Construct the features after adaptive feature selection into graph data samples. In these graph data samples, each feature serves as a node, and the correlation between features serves as an edge. The correlation is calculated using the Spearman coefficient, and the adjacency matrix is constructed according to this process. The adjacency matrix is shown as follows:

[0042] In the formula, n is the number of nodes, ρ ij represents the i th node and the j th node, and the Spearman correlation coefficient between them. Among them, i , j = 1, 2, …, n. The feature selection and graph sample construction process based on correlation analysis is as shown in Figure 2 .

[0043] Step S4, model training. Based on the graph neural network and the contrastive learning method, construct a contrastive residual graph neural network (CRGN) for information mining and model training; The composition form of the contrastive residual graph neural network model can be seen in Appendix Figure 3 . It adopts a self-supervised learning mode, sets two encoders for contrastive learning, and the two encoders share weight parameters. Each encoder contains a graph convolutional layer, a max pooling layer, and a residual connection. At the same time, the Leaky–ReLU function is used as the activation function, and the negative half-axis coefficient is set to 0.1; the residual connection in the model is mainly used to ensure information integrity; During the model training process, the original samples are masked with Gaussian random noise with a standard deviation of 0.1, and the original samples and masked samples are respectively input into the two encoders; the objective function of the model training is based on the contrastive loss, and L2-regularization is added during the training process; the calculation formula of the objective function is:

[0044] where and are the positive and negative sample feature representations respectively, sim represents the cosine similarity, is the temperature parameter (set to 0.7); N is the number of samples, λ is the L2-regularization coefficient, w is the model parameter; the Stochastic Gradient Descent optimizer SGD with Nesterov momentum is used for model training.

[0045] Step S5, Health guidance calculation: Load the pre-trained contrast residual graph neural network model trained with health data. Input the unknown state samples into the pre-trained model, calculate the distance in the multi-dimensional space between the output result of the pre-trained model and the prediction result of the health samples, and construct the gearbox health guidance based on the exponentially weighted moving average and the normalized average distance. The specific method for health guidance calculation is as follows: Construct the graph data samples of the gearbox to be monitored based on the same data preprocessing and sampling mode as above. Load the pre-trained contrast residual graph neural network model trained with health data into the monitoring system, and input the unknown state samples into the pre-trained model to obtain a predicted value. Then use the Manhattan distance , Euclidean distance , cosine similarity , Chebyshev distance to calculate the average distance between the health samples and the samples to be monitored using the mean of these 4 distance metrics , and the calculation formula is:

[0046]

[0047]

[0048]

[0049]

[0050] Among them, a , b are two vectors for which the distance is to be calculated, a i , b i are the respective elements in the two vectors, n is the vector dimension, Normed indicates normalizing the calculation result; on this basis, construct the gearbox health guidance based on the exponentially weighted moving average (EWMA) HI , and the formula is:

[0051] Among them is the health guidance at time, is the penalty factor, is the health guidance at is the normalized average distance at Next, median filtering denoising is performed on the calculation results, and the filter window size is set to 100.

[0052] Step S6, fault threshold determination, determining the fault threshold based on the Statistical Process Control (SPC) method; From a statistical perspective, the SPC method can use the relevant principles of statistical analysis to monitor the production process in real time and scientifically distinguish between the common causes and special causes of product quality fluctuations in the production process. Assume that the quality variable of a certain product follows a normal distribution during the production process, and its quality characteristic X follows a normal distribution with a mean of and a standard deviation of , then, the probability P of the quality characteristic within is calculated by the formula:

[0053] In practical applications, the sample mean and the standard deviation are used to represent the mean and variance of the normal distribution. Then, we define the range exceeding as the fault state. Therefore, according to the calculated state indication parameters, the threshold calculation is performed according to the SPC principle. Considering the HI definition of the value, when selecting the regional boundary value here, only the value corresponding to the upper boundary is considered, and this upper boundary is the failure threshold Th defined here, and the calculation formula is:

[0054] Among them, is the mean of the health guidance sequence, is the standard deviation of the health guidance sequence.

[0055] Step S7, status monitoring and early warning, comparing the relative magnitudes of the gearbox health guidance and the fault threshold to achieve online status monitoring of the wind turbine gearbox. During the status monitoring process, when the calculated HI exceeds the corresponding failure threshold Th , it can be considered that the operating state of the unit at the corresponding moment is abnormal, and the existing problems need to be investigated before continuing to work.

[0056] The present invention also provides a wind turbine gearbox status monitoring system based on a self-supervised contrast residual graph network, including: a data preprocessing module, an adaptive feature selection module, a graph data sample construction module, a model training module, a health guidance calculation module, a fault threshold determination module, and a status monitoring and early warning module; The data preprocessing module is used to obtain the original SCADA data from the SCADA system installed in the wind farm and preprocess the original SCADA data; The adaptive feature selection module selects data features with strong correlation with the wind turbine gearbox based on the adaptive feature selection method; The graph data sample construction module is used for graph data sample construction, and constructs the data features after adaptive feature selection into graph data samples; The model training module constructs a contrast residual graph neural network based on the graph neural network and the contrast learning method for information mining and model training; The health guidance calculation module loads the contrast residual graph neural network model trained with health data, inputs the unknown state samples into the pre-trained model, calculates the distance in the multi-dimensional space between the output result of the pre-trained model and the prediction result of the health samples, and constructs the gearbox health guidance based on the exponentially weighted moving average and the normalized average distance; The fault threshold determination module determines the fault threshold based on the statistical process control method; The state monitoring and warning module is used to compare the relative sizes of the gearbox health guidance and the fault threshold to realize the online state monitoring of the wind turbine gearbox.

[0057] After the original SCADA data is preprocessed such as removing null values, based on the adaptive feature selection method, data features with strong correlation with the wind turbine gearbox are selected, and data samples are established based on the correlation. The process of feature selection is adaptive, so the features selected for different wind turbines are dynamically changing. On this basis, based on the Gaussian random process, Gaussian random noise is added to the data samples to construct a mask for model training based on the contrast learning method. The standard deviation of the Gaussian random noise is set to 。The offline model training process is based on contrast learning, constructs 2 encoders, and shares the initial weights. Each encoder is a graph neural network containing residual connections. The model training process is based on the backpropagation algorithm, and is trained using the original data samples in the healthy state. After training, the pre-trained model is loaded into the online monitoring system, the graph samples of the online data in the unknown state are constructed, and after calculation by the pre-trained model, the distance in the multi-dimensional space between the obtained prediction result and the prediction result of the health samples is calculated. Based on this distance metric, combined with the exponentially weighted moving average (EWMA) method, a health guidance is constructed HI ,and then the failure threshold is established based on the statistical process control (SPC) method Th 。Compare HI and Th 's relative sizes to realize the online state monitoring of the wind turbine gearbox.

[0058] The present invention can achieve precise and efficient monitoring: Adaptive feature selection effectively removes redundant information, retains key features, improves the monitoring efficiency and reliability, enables the model to focus on core data, and enhances the monitoring accuracy. It has powerful model performance: By combining the advantages of multiple technologies with the residual graph neural network model, through contrastive learning and residual connections, and combining the correlation relationships between data features, it enhances the model's expression ability and information extraction ability, accurately mines data features, and accurately identifies the state of the gearbox. It has a reliable early warning mechanism: Based on multiple distance metrics, a health guideline is constructed, combined with the exponentially weighted moving average to construct a gearbox health guideline, and combined with statistical process control to determine the fault threshold, improving the reliability of the method. It can achieve an abnormal early warning 30 - 40 hours in advance, with an abnormal recognition accuracy rate exceeding 90%, saving time for the operation and maintenance process and reducing losses. It has wide applicability: The method of the present invention realizes an end-to-end state monitoring of the gearbox for the SCADA data of wind turbines, provides methods and ideas, and can achieve this function for different wind turbines and their supporting SCADA systems, with generalization, and can meet the state monitoring requirements of different wind turbine units.

[0059] Embodiment This embodiment provides a method for monitoring the state of a wind turbine gearbox based on a self-supervised contrastive residual graph network, including the following steps: (1) Data acquisition and preprocessing Data is acquired from the SCADA system installed in the wind farm. These data contain several features, covering various information of the operation of the wind turbine unit, such as power, wind speed, temperature, pressure, etc. The original data may have null values, error values, and abnormal data caused by sensor failures. The data cleaning algorithm is used to automatically identify and delete null value records; for error values, they are corrected according to the upper and lower limit ranges and logical relationships of the data, and deleted if they cannot be corrected; for abnormally fluctuating data, it is judged in combination with historical data and physical laws. If it is caused by a sensor failure, data interpolation or filtering algorithms are used for repair.

[0060] (2) Adaptive feature selection Data distribution test: Perform the Anderson-Darling test on each feature of the original SCADA data to judge whether it follows a normal distribution. This test calculates a specific statistic and compares it with the critical value to draw a conclusion. If the feature data does not follow a normal distribution, it indicates that its distribution form is relatively complex. At this time, using the Pearson correlation coefficient based on the normal distribution assumption will produce biases, so the Spearman correlation coefficient is selected for subsequent correlation analysis.

[0061] Core feature determination: Combining past maintenance experience of wind turbines, the working principle of the gearbox, and the maintenance reports provided by the wind farm, five features, namely the average oil temperature of the gearbox, the average oil temperature at the inlet of the gearbox, the average value of the oil inlet pressure of the gearbox, the average value of the impeller speed, and the average value of the generator speed, are determined as core features. The average oil temperature of the gearbox reflects the internal thermal state of the gearbox. Excessive oil temperature may indicate increased gear wear or poor lubrication; the average oil temperature at the inlet of the gearbox affects the initial performance of the lubricating oil; the oil inlet pressure of the gearbox is related to the lubrication effect; the impeller speed and the generator speed are closely related to the transmission efficiency of the gearbox, and their abnormal changes can directly reflect the operating state of the gearbox.

[0062] Calculation of correlation coefficients and feature screening: Taking each core feature as a reference, calculate the Spearman correlation coefficient between other features and it. For example, when calculating the correlation coefficient between feature and the core feature F1 (average oil temperature of the gearbox), calculate for all data samples according to the Spearman correlation coefficient formula. After the calculation is completed, calculate the average value of the correlation coefficients between each non-core feature and the five core features to obtain the average correlation coefficient . Set the quantization index to 0.6, and combine the non-core features with an average correlation coefficient greater than 0.6 with the five core features to form the original data samples for model training.

[0063] Data normalization: Use the maximum-minimum normalization method to process the selected feature data. Through normalization, map the data of different features to the interval [0,1], eliminate the influence of dimensions, make different features equally important in model training, and improve the model training effect.

[0064] (3) Model training Construction of graph data samples: Construct the features after adaptive feature selection into graph data samples. Regard each feature as a node in the graph, and determine the edges according to the Spearman correlation coefficient between features. Different edges are used to construct the adjacency matrix. In this way, each feature and its relationship with other features are presented in the form of a graph, providing structured data for the training of graph neural networks.

[0065] Construction of Contrast Residual Graph Neural Network (CRGN): CRGN adopts a self-supervised learning mode and sets two encoders for contrast learning. The two encoders share weight parameters. Each encoder contains a graph convolutional layer (GCN), a max pooling layer, and a residual connection. At the same time, the Leaky–ReLU function is used as the activation function, and the negative half-axis coefficient is set to 0.1. The residual connection in the model is mainly used to ensure information integrity. Composition. GCN extracts features from graph data through convolutional operations, combining the processing ability of graph neural networks for graph-structured data and the feature extraction advantages of convolutional neural networks, which is suitable for processing graph data samples of wind turbine gearbox condition monitoring.

[0066] Contrast learning training: Gaussian random noise with a standard deviation of 0.1 is used to mask the original data to generate masked data. The original data and the masked data are respectively input into the two encoders, and the model is allowed to learn the key features of the data through contrast learning. The objective function based on the contrast loss is:

[0067] where and are the feature representations of the positive and negative samples respectively, sim represents the cosine similarity, is the temperature parameter. The temperature parameter controls the smoothness of the contrast loss function and affects the training effect of the model. The Stochastic Gradient Descent (SGD) optimizer with Nesterov momentum is used for model training, and the momentum is set to 0.9. Nesterov momentum enables the optimizer to consider the change trend of the gradient when updating parameters, accelerating the convergence of the model; at the same time, L2 regularization is used to prevent the model from overfitting and improve the generalization ability of the model. The hyperparameter settings for model training are shown in the following table: Number Hyperparameter Value 1 Learning rate 1e-4 2 Training epochs 3000 3 Momentum 0.9 4 Batch size 256 5 Number of GCN layers in each encoder 4 6 L2 - regularization coefficient 1e-5 7 Temperature coefficient for controlling shape 0.7 8 Input channels of GCN 32 9 Channels of GCN hidden layer 64 10 Output layer channels 1280 During the training process, the loss value of the model on the validation set is recorded every certain number of training rounds to observe the training trend of the model. If the loss value on the validation set no longer decreases after multiple rounds of training, it is considered that the model has reached the convergence state and the training is stopped.

[0068] (4) Health guidance calculation and fault threshold determination Health guidance calculation: Based on the same data preprocessing and sampling mode, graph samples of the gearbox to be monitored are constructed. The trained CRGN model with healthy data is loaded, and the unknown state samples are input into the pre-trained model to obtain the model output results. Multiple distance metrics such as Manhattan distance, Euclidean distance, cosine similarity, and Chebyshev distance are used to calculate the average distance between the healthy samples and the unknown samples .

[0069] Construct the health index HI of the gearbox based on the exponentially weighted moving average (EWMA) and the normalized average distance. The calculation formula is

[0070] where is the health index at time is the penalty factor is the health index at time is the normalized average distance at time

[0071] Apply median filtering to the calculated health index results to remove noise. The size of the filtering window is set to 100. Median filtering sorts the data within the window and takes the middle value as the filtering output, which can effectively remove sudden noise and make the health index more accurately reflect the actual state of the gearbox.

[0072] Fault threshold determination: Determine the fault threshold based on the statistical process control (SPC) method. In practical applications, use the sample mean and the standard deviation to represent the mean and variance of the normal distribution. Then we define the range exceeding as the fault state. Therefore, according to the calculated state indicator parameters, perform threshold calculation according to the SPC principle. Considering the HI definition of the value, when selecting the boundary value of the region here, only consider the value corresponding to the upper boundary, and this upper boundary is the failure threshold Th defined here. The calculation formula is:[[]]

[0073] During the state monitoring process, when the calculated HI exceeds the corresponding failure threshold Th , it can be considered that the operating state of the unit at the corresponding time is abnormal. After troubleshooting the existing problems, the work can be continued.

[0074] (5) State monitoring and early warning During the operation of the wind turbine, collect SCADA data in real time. According to the above processes of data preprocessing, adaptive feature selection, graph data sample construction, model calculation, health index calculation, and fault threshold comparison, continuously monitor the operating state of the gearbox. When the health index HI exceeds the fault threshold ThWhen it is determined that the operating state of the wind turbine gearbox is abnormal, the system immediately issues a warning signal. The warning signal is notified to the operation and maintenance personnel by means of text messages, emails, or pop-up windows in the wind farm monitoring system. At the same time, information such as the time of the abnormality occurrence, the health guidance value, and relevant characteristic data is recorded in the system, which is convenient for the operation and maintenance personnel to analyze the cause of the fault later. After receiving the warning, the operation and maintenance personnel check and maintain the gearbox according to the detailed information provided by the system, combined with the historical operation data of the wind turbine and the actual on-site situation, such as checking the gear wear condition, the quality of the lubricating oil, the bearing state, etc., and timely eliminate potential fault hazards to ensure the safe and stable operation of the wind turbine. After the fault is processed, the maintenance records and processing results are fed back into the system to update the health status information of the gearbox and start a new round of condition monitoring.

[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A wind turbine gearbox condition monitoring method based on self-supervised contrast residual graph network, characterized in that: The steps include: Step S1, data preprocessing, obtaining original SCADA data from the SCADA system installed in the wind farm, and preprocessing the original SCADA data; Step S2, adaptive feature selection, selecting data features with strong correlation with the wind turbine gearbox based on an adaptive feature selection method; Step S3, graph data sample construction, constructing the data features after adaptive feature selection into graph data samples; Step S4, model training, based on graph neural network and contrastive learning method, construct a contrastive residual graph neural network for information mining and model training; Step S5, health guide calculation, load the contrast residual graph neural network model trained with health data, input the unknown state sample into the pre-trained model, obtain the distance calculation between the output result of the pre-trained model and the prediction result of the healthy sample in multi-dimensional space, and construct the gearbox health guide based on the exponentially weighted moving average and normalized average distance; Step S6, determining the fault threshold, determining the fault threshold based on a statistical process control method; Step S7, status monitoring and early warning, compares the relative size of the gearbox health guide and the fault threshold to achieve online status monitoring of the wind turbine gearbox.

2. A wind turbine gearbox condition monitoring method based on a self-supervised contrast residual graph network according to claim 1, characterized in that: In step S1, the original SCADA data may contain null values, error values, and abnormal data caused by sensor failure; Preprocessing methods include: using data cleaning algorithms to automatically identify and delete records with null values; For error values, corrections are made based on the upper and lower limits of the data and the logical relationship. If correction is not possible, they are deleted. For abnormal fluctuating data, judgments are made based on historical data and physical laws. If it is caused by sensor failure, data interpolation or filtering algorithms are used to repair it.

3. A wind turbine gearbox condition monitoring method based on self-supervised contrast residual graph network according to claim 2, characterized in that: In step S2, based on the adaptive feature selection method, the method for selecting data features with strong correlation with the wind turbine gearbox is as follows: For raw SCADA data, select Spearman The coefficients are used for correlation analysis, and the formula is: in is the Spearman correlation coefficient, is the rank difference of each pair of observations, is the total number of observations; based on the operation and maintenance report from the wind farm, five data features, namely, the average oil temperature of the gearbox, the average oil temperature at the gearbox inlet, the average oil inlet pressure of the gearbox, the average impeller speed, and the average generator speed, are selected as the core features, and the correlation coefficients of other features and the core features are calculated. The formula is: in is the average correlation coefficient of other features relative to the core features, It is Other features and The correlation coefficient of the core features is is the number of core features; A quantitative index is set. When the average correlation coefficient exceeds the quantitative index, the corresponding data feature is selected. The variables whose correlation coefficient exceeds the index are determined as strongly correlated variables, which together with the core features constitute the model training samples. Finally, the feature data after feature selection is processed by maximum and minimum normalization. The formula is: in is the normalized data, is the original data, and are the minimum and maximum values ​​in the original data, respectively.

4. The method for monitoring the state of a wind turbine gearbox based on a self-supervised contrast residual graph network according to claim 3, characterized in that: In step S3, the method of constructing the data features after adaptive feature selection into a graph data sample is as follows: construct the features after adaptive feature selection into a graph data sample, in which each feature is used as a node, the correlation between features is used as an edge, the correlation uses the Spearman coefficient, and constructs an adjacency matrix according to the process. The adjacency matrix is ​​shown in the following formula: In the formula, n is the number of nodes, ρ ij Indicates i Nodes and j The Spearman correlation coefficient between nodes, where i , j =1, 2,…, n.

5. The method for monitoring the state of a wind turbine gearbox based on a self-supervised contrast residual graph network according to claim 4, characterized in that: In step S4, the residual graph neural network model adopts a self-supervised learning mode, sets two encoders for comparative learning, and the two encoders share weight parameters. Each encoder contains a graph convolution layer, a maximum pooling layer, and a residual connection. At the same time, the Leaky-ReLU function is used as the activation function, and its negative semi-axis coefficient is set to 0.1; the residual connection is mainly used in the model to ensure information integrity; During the model training process, the original samples were masked with Gaussian random noise with a standard deviation of 0.1, and the original samples and masked samples were input into two encoders respectively; the objective function of the model training was based on contrast loss, and L2-regularization was added during the training process; The objective function calculation formula is: in and They are positive and negative sample feature representations, sim represents cosine similarity, is the temperature parameter; N is the number of samples, λ is the L2-regularization coefficient, w are model parameters; the stochastic gradient descent optimizer SGD with Nesterov momentum is used for model training.

6. A method for monitoring the state of a wind turbine gearbox based on a self-supervised contrast residual graph network according to claim 5, characterized in that: In step S5, the specific method of calculating the health guidance is as follows: construct a graph data sample of the gearbox to be monitored, load the residual graph neural network model trained with health data into the monitoring system, input the unknown state sample into the pre-trained model, and obtain a predicted value; then use the Manhattan distance , Euclidean distance , cosine similarity , Chebyshev distance The average distance between healthy samples and samples to be monitored is calculated by taking the mean of these four distance metrics. , the calculation formula is: in, a , b are the two vectors whose distance is to be calculated, a i , b i are the elements in the two vectors respectively, n is the dimension of the vector, Normed Indicates that the calculation results are normalized; on this basis, the gearbox health guideline is constructed based on the exponentially weighted moving average HI , the formula is: in for Always healthy guidance, is the penalty factor, for Always healthy guidance, for The normalized average distance at each moment is calculated; then, the calculated result is subjected to median filtering for denoising, and the filter window size is set to 100.

7. A method for monitoring the state of a wind turbine gearbox based on a self-supervised contrast residual graph network according to claim 6, characterized in that: In step S6, the method for determining the fault threshold based on the statistical process control method is as follows: using the sample mean and standard deviation Represents the mean and variance of the normal distribution, defines the range exceeding the preset value as a fault state, and calculates the threshold value according to the statistical process control principle based on the calculated state indication parameters; considering the gearbox health guide HI When selecting the region boundary value here, only the value corresponding to the upper boundary is considered. The upper boundary is the failure threshold Th defined here, and the calculation formula is: in, is the mean of the health guide series, is the standard deviation of the health guide series.

8. A wind turbine gearbox condition monitoring system based on a self-supervised contrast residual graph network, characterized in that: include: Data preprocessing module, adaptive feature selection module, graph data sample construction module, model training module, health guidance calculation module, fault threshold determination module and status monitoring and early warning module; The data preprocessing module is used to obtain original SCADA data from the SCADA system installed in the wind farm and preprocess the original SCADA data; The adaptive feature selection module selects data features with strong correlation with the wind turbine gearbox based on an adaptive feature selection method; The graph data sample construction module is used for graph data sample construction, and constructs the data features after adaptive feature selection into graph data samples; The model training module constructs a contrastive residual graph neural network based on a graph neural network and a contrastive learning method for information mining and model training; The health guidance calculation module loads the comparative residual graph neural network model trained with health data, inputs the unknown state sample into the pre-trained model, obtains the distance calculation between the output result of the pre-trained model and the prediction result of the health sample in the multi-dimensional space, and constructs the gearbox health guidance based on the exponentially weighted moving average and the normalized average distance; The fault threshold determination module determines the fault threshold based on a statistical process control method; The state monitoring and early warning module is used to compare the relative sizes of the gearbox health guide and the fault threshold to achieve online state monitoring of the wind turbine gearbox.

Citation Information

Patent Citations

  • Distribution network fault positioning method based on short-time matrix pencil method and graph neural network

    CN118191510A