A Wind Turbine Gearbox Condition Monitoring Method and System Based on Self-Supervised Contrastive Residual Map Network

Through the self-supervised and comparative residual graph network, the problems of long data time intervals, strong subjectivity of feature selection, insufficient generalization capability of model and low early warning reliability in the gearbox status monitoring of wind turbines are solved, and efficient and accurate gearbox status monitoring and early warning are achieved, which improves the operating efficiency and economic benefits of wind farms.

CN120087406BActive Publication Date: 2025-07-22HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510574232.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-07-22
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

The existing wind turbine gearbox status monitoring methods have problems such as long data time intervals, strong subjectivity of feature selection, insufficient model generalization ability, uncertainty in difference calculations and low warning reliability, resulting in short warning time and low accuracy.

Method used

A self-supervised comparison residual graph network is adopted, and precise monitoring of gearbox status is achieved through adaptive feature selection, graph data sample construction, comparison residual graph neural network training and multi-dimensional distance calculation, combined with exponential weighted moving average and statistical process control.

Benefits of technology

Abnormal warning is achieved 30 to 40 hours in advance, with an accuracy of abnormal identification exceeding 90%, reducing operation and maintenance costs, reducing downtime caused by gearbox failure, and improving wind farm operation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087406B_ABST
    Figure CN120087406B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for monitoring the state of a wind turbine gearbox based on a self-supervised contrast residual graph network, including: obtaining original SCADA data and performing preprocessing; selecting data features with strong correlation with the wind turbine gearbox based on an adaptive feature selection method; constructing the data features into graph data samples; constructing a contrast residual graph neural network based on a graph neural network and a contrast learning method; loading a contrast residual graph neural network model trained with healthy data, inputting unknown state samples into the pre-trained model, calculating the distance in a multi-dimensional space between the output result of the pre-trained model and the prediction result of the healthy samples, constructing a gearbox health guideline based on exponential weighted moving average and normalized average distance; determining a fault threshold based on statistical process control technology; comparing the relative magnitudes of the gearbox health guideline and the fault threshold to achieve online state monitoring. The present invention realizes accurate and efficient monitoring of the operating state of the wind turbine gearbox.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of wind power operation and maintenance, and particularly to a method and system for monitoring the state of a wind turbine gearbox based on a self-supervised contrast residual graph network. Background Art

[0002] Wind turbines are often in harsh environments and have frequent failures, resulting in an increasing demand for operation and maintenance. Traditional operation and maintenance methods such as shutdown maintenance, inspection and maintenance, and regular maintenance have many drawbacks. For example, shutdown maintenance affects the efficiency of wind farms, inspection and maintenance is costly and inefficient, and regular maintenance faces problems of cost, efficiency, and ineffective maintenance investment in normal units. Therefore, predictive maintenance can be applied in real industrial scenarios because it can accurately locate abnormal units, accurately predict faults, and achieve efficient and low-cost maintenance. Accurate and effective condition monitoring is a core task of predictive maintenance.

[0003] Predictive maintenance of each component of a wind turbine is costly and, in many cases, redundant work. Therefore, condition monitoring of wind turbine subsystems has become a reasonable option. Among the multiple subsystems inside a wind turbine, the gearbox has a high failure rate and expensive repair costs, and has the greatest impact on the wind turbine after a failure. Therefore, the gearbox is the first choice for the condition monitoring task of subsystems.

[0004] Regarding the method for gearbox condition monitoring, traditional model-based and signal processing methods have problems such as difficult modeling and high costs. Data-driven condition monitoring methods can avoid additional investment in sensing devices and, at the same time, define and model different states based on end-to-end methods. In the prior art, although the monitoring method based on a normal behavior model can detect anomalies by comparing healthy data with real-time data, it has problems such as a large randomness in feature selection and weak model expression ability, resulting in a short warning time (usually less than 24 hours) and low accuracy (usually less than 80%). Some prominent problems are as follows:

[0005] 1) Long data time interval: The data sampling interval of the SCADA system is long (usually 10 minutes), making it difficult to accurately extract early fault signals from it.

[0006] 2) Strong subjectivity in feature selection: Mainly relying on manual experience to select data features easily introduces subjective factors, resulting in information omission or redundancy;

[0007] 3) Insufficient model generalization ability: Traditional machine learning models ignore the correlation between different data features and are difficult to capture dynamic fault features;

[0008] 4) Uncertainty in difference calculation: When calculating the difference between normal samples and samples to be monitored, uncertainty will be generated due to the application of a single distance metric;

[0009] 5) Low warning reliability: A single distance metric cannot comprehensively reflect the health status. At the same time, the setting of the fault threshold relies heavily on historical experience, resulting in a high false alarm rate.

[0010] As described above, although the data-driven method can effectively extract valuable information from the wind turbine data, its operation process still faces many challenges. Therefore, a new method and system for monitoring the state of the wind turbine gearbox are needed to solve the above problems. Summary of the Invention

[0011] The present invention provides a method and system for monitoring the state of a wind turbine gearbox based on a self-supervised contrast residual graph network, aiming to solve the defects existing in the actual operation of the existing methods for monitoring the state of the wind turbine gearbox. Through innovative technical means, accurate and efficient monitoring of the operating state of the wind turbine gearbox is achieved, potential fault hazards are detected in advance, reliable warnings are issued, thereby reducing the operation and maintenance costs of the wind turbine, reducing the downtime caused by gearbox failures, and improving the overall operation efficiency and economic benefits of the wind farm.

[0012] The present invention adopts the following technical solutions to solve the technical problems:

[0013] A method for monitoring the state of a wind turbine gearbox based on a self-supervised contrast residual graph network includes the following steps:

[0014] Step S1, data preprocessing: Obtain the original SCADA data from the SCADA system installed in the wind farm and preprocess the original SCADA data;

[0015] Step S2, adaptive feature selection: Based on the adaptive feature selection method, select the data features with strong correlation with the wind turbine gearbox;

[0016] Step S3, graph data sample construction: Construct the data features after adaptive feature selection into graph data samples;

[0017] Step S4, model training: Based on the graph neural network and the contrast learning method, construct a contrast residual graph neural network for information mining and model training;

[0018] Step S5, health guidance calculation: Load the contrast residual graph neural network model trained with healthy data, input the unknown state samples into the pre-trained model, calculate the distance in the multi-dimensional space between the output result of the pre-trained model and the prediction result of the healthy samples, and construct the gearbox health guidance based on the exponentially weighted moving average and the normalized average distance;

[0019] Step S6, fault threshold determination: Determine the fault threshold based on the statistical process control method;

[0020] Step S7, status monitoring and warning: Compare the relative magnitudes of the gearbox health guidelines and the fault thresholds to achieve online status monitoring of the wind turbine gearbox.

[0021] Furthermore, in step S1, the original SCADA data may contain null values, error values, and abnormal data caused by sensor failures. The preprocessing methods include: using a data cleaning algorithm to automatically identify and delete null value records; for error values, correct them according to the upper and lower limit ranges and logical relationships of the data, and delete them if they cannot be corrected; for abnormally fluctuating data, make a judgment in combination with historical data and physical laws. If it is caused by a sensor failure, use data interpolation or a filtering algorithm for repair.

[0022] Furthermore, in step S2, based on the adaptive feature selection method, the method for selecting data features strongly correlated with the wind turbine gearbox is as follows:

[0023] For the original SCADA data, select Spearman coefficient to conduct a correlation analysis. The formula is:

[0024]

[0025] where is the Spearman correlation coefficient, is the rank difference of each pair of observed values, is the total number of observed values;

[0026] Based on the operation and maintenance reports from the wind farm, select 5 data features including the average oil temperature of the gearbox, the average oil temperature at the inlet of the gearbox, the average oil inlet pressure of the gearbox, the average impeller speed, and the average generator speed as the core features, and calculate the average correlation coefficient of other features with the core features. The formula is:

[0027]

[0028] where is the average correlation coefficient of other features relative to the core features, is the th other feature and the th core feature's correlation coefficient, is the number of core features;

[0029] Set a quantization index. When the average correlation coefficient exceeds this quantization index, the corresponding data features are selected; determine the variables with correlation coefficients exceeding this index as strongly correlated variables, and jointly form the model training samples with the core features;

[0030] Finally, perform maximum-minimum normalization on the feature data after feature selection. The formula is:

[0031]

[0032] wherein is the normalized data, is the original data, and are the minimum and maximum values in the original data, respectively.

[0033] Furthermore, in step S3, the method of constructing the data features after adaptive feature selection into graph data samples is as follows:

[0034] Construct the features after adaptive feature selection into graph data samples. In this graph data sample, each feature is used as a node, and the correlation between features is used as an edge. The correlation is calculated using the Spearman coefficient, and the adjacency matrix is constructed according to this process. The adjacency matrix is shown as follows:

[0035]

[0036] In the formula, n is the number of nodes, ρ i represents the Spearman correlation coefficient between the i-th node and the j-th node, where i , j = 1, 2,..., n.

[0037] Furthermore, in step S4, the residual graph neural network model adopts a self-supervised learning mode. Two encoders are set for contrastive learning. The two encoders share weight parameters. Each encoder contains a graph convolutional layer, a max pooling layer, and a residual connection. At the same time, the Leaky–ReLU function is used as the activation function, and the negative half-axis coefficient is set to 0.1; the residual connection in the model is mainly used to ensure information integrity;

[0038] During the model training process, the original samples are masked with Gaussian random noise with a standard deviation of 0.1, and the original samples and the masked samples are respectively input into the two encoders; the objective function of the model training is based on the contrastive loss, and L2-regularization is added during the training process; the calculation formula of the objective function is:

[0039]

[0040] where and are the positive and negative sample feature representations, respectively, sim represents the cosine similarity, is the temperature parameter; N is the number of samples, λ is the L2-regularization coefficient, wThey are model parameters; the model is trained using the Stochastic Gradient Descent (SGD) optimizer with Nesterov momentum.

[0041] Further, in step S5, the specific method for calculating the health guidance is as follows:

[0042] Construct a graph data sample of the gearbox to be monitored, load the trained residual graph neural network model with health data into the monitoring system, input the unknown state sample into the pre-trained model, and obtain a predicted value;

[0043] Then, use the Manhattan distance , Euclidean distance , cosine similarity , and Chebyshev distance The mean of these four distance metrics is used to calculate the average distance between the health sample and the sample to be monitored , and the calculation formula is:

[0044]

[0045]

[0046]

[0047]

[0048]

[0049] Among them, a , b are two vectors for which the distance is to be calculated, a i , b i are the respective elements in the two vectors, n is the dimension of the vector, Normed indicates normalizing the calculation result; on this basis, a gearbox health guidance is constructed based on the exponentially weighted moving average HI , and the formula is:

[0050]

[0051] Among them is the health guidance at time, is the penalty factor, is the health guidance at is the normalized average distance at

[0052] Next, median filtering denoising is performed on the calculation results, and the filter window size is set to 100.

[0053] Furthermore, in step S6, the method for determining the fault threshold based on the statistical process control method is as follows:

[0054] Use the sample mean and the standard deviation to represent the mean and variance of the normal distribution. Define the range exceeding the preset value as the fault state, and calculate the threshold according to the calculated state indication parameters based on the statistical process control principle; Considering the definition of the gearbox health guideline HI here, when selecting the region boundary value, only consider the value corresponding to the upper boundary, and this upper boundary is the failure threshold defined here Th , and the calculation formula is:

[0055]

[0056] where, is the mean of the health guideline sequence, is the standard deviation of the health guideline sequence.

[0057] A wind turbine gearbox condition monitoring system based on a self-supervised contrast residual graph network, comprising: a data preprocessing module, an adaptive feature selection module, a graph data sample construction module, a model training module, a health guideline calculation module, a fault threshold determination module, and a condition monitoring and warning module;

[0058] The data preprocessing module is used to obtain the original SCADA data from the SCADA system installed in the wind farm and preprocess the original SCADA data;

[0059] The adaptive feature selection module selects data features with strong correlation with the wind turbine gearbox based on the adaptive feature selection method;

[0060] The graph data sample construction module is used for graph data sample construction, and constructs the data features after adaptive feature selection into graph data samples;

[0061] The model training module constructs a contrast residual graph neural network based on the graph neural network and the contrast learning method for information mining and model training;

[0062] The health guideline calculation module loads the contrast residual graph neural network model trained with health data, inputs the unknown state samples into the pre-trained model, calculates the distance in the multi-dimensional space between the output result of the pre-trained model and the prediction result of the health samples, and constructs the gearbox health guideline based on the exponentially weighted moving average and the normalized average distance;

[0063] The fault threshold determination module determines the fault threshold based on the statistical process control method;

[0064] The state monitoring and early warning module is used to compare the relative magnitudes of the gearbox health guideline and the fault threshold to achieve online state monitoring of the wind turbine gearbox.

[0065] Advantages of the present invention:

[0066] (1) Precise and efficient monitoring: Adaptive feature selection effectively removes redundant information, retains key features, improves monitoring efficiency and reliability, enables the model to focus on core data, and enhances monitoring accuracy.

[0067] (2) Powerful model performance: The contrast residual graph neural network model combines the advantages of multiple technologies. Through contrastive learning and residual connections, it combines the correlation relationships between data features, enhances the model's expression ability and information extraction ability, accurately mines data features, and accurately identifies the gearbox state.

[0068] (3) Reliable early warning mechanism: A health guideline is constructed based on multiple distance metrics, a gearbox health guideline is constructed by combining the exponential weighted moving average, and a fault threshold is determined by combining statistical process control to improve the reliability of the method. It can achieve abnormal early warning 30 - 40 hours in advance, with an abnormal recognition accuracy rate exceeding 90%, saving time for the operation and maintenance process and reducing losses.

[0069] (4) Wide applicability: The method of the present invention realizes end-to-end state monitoring of the gearbox for wind turbine SCADA data, provides methods and ideas, and can achieve this function for different wind turbines and their supporting SCADA systems, with generalization and can meet the state monitoring requirements of different wind turbine units. Description of the Drawings

[0070] Figure 1 It is the flowchart of the method of the present invention.

[0071] Figure 2 It is the flowchart of the adaptive feature selection and graph sample construction of the present invention.

[0072] Figure 3 It is the architecture diagram of the CRGN model of the present invention. Detailed Embodiments

[0073] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0074] Reference appendix Figure 1 , the present invention provides a method for monitoring the state of a wind turbine gearbox based on a self-supervised contrast residual graph network, including the following steps:

[0075] Step S1, data preprocessing, obtaining the original SCADA data from the SCADA system installed in the wind farm, and preprocessing the original SCADA data;

[0076] The original SCADA data may have null values, error values, and abnormal data caused by sensor failures; the preprocessing methods include: using a data cleaning algorithm to automatically identify and delete null value records; for error values, correcting them according to the upper and lower limit ranges and logical relationships of the data, and deleting them if they cannot be corrected; for abnormally fluctuating data, judging in combination with historical data and physical laws, and if it is caused by a sensor failure, using data interpolation or a filtering algorithm for repair.

[0077] Step S2, adaptive feature selection, selecting data features with strong correlation with the wind turbine gearbox based on the adaptive feature selection method;

[0078] For the original SCADA data, first use the Anderson-Darling test to evaluate the normal distribution of each feature. Since the original data features do not follow a normal distribution in most cases, the Spearman coefficient is selected to carry out the correlation analysis, and the formula is:

[0079]

[0080] where is Spearman the correlation coefficient, is the rank difference of each pair of observed values, is the total number of observed values;

[0081] Based on the operation and maintenance reports from the wind farm, select 5 data features, namely the average oil temperature of the gearbox, the average oil temperature at the inlet of the gearbox, the average value of the oil inlet pressure of the gearbox, the average value of the impeller speed, and the average value of the generator speed, as the core features, and calculate the average value of the correlation coefficients between other features and the core features. The formula is:

[0082]

[0083] where is the average value of the correlation coefficients of other features relative to the core features, is the th other feature and the th core feature's correlation coefficient, is the number of core features;

[0084] Set a quantization index. When the average correlation coefficient exceeds this quantization index, the corresponding data feature is selected; according to the information in the literature, the general value range of this quantization index is , and in the present invention, it is selected as 0.6 based on the specific experimental process. Determine the variables with correlation coefficients exceeding this index as strongly correlated variables, which together with the core features constitute the model training samples;

[0085] Finally, perform min-max normalization on the feature data after feature selection. The formula is:

[0086]

[0087] where is the data after normalization, is the original data, and are the minimum and maximum values in the original data respectively.

[0088] Step S3, graph data sample construction. Construct the data features after adaptive feature selection into graph data samples;

[0089] Construct the features after adaptive feature selection into graph data samples. In this graph data sample, each feature serves as a node, and the correlation between features serves as an edge. The correlation uses the Spearman coefficient, and the adjacency matrix is constructed according to this process. The adjacency matrix is shown as follows:

[0090]

[0091] In the formula, n is the number of nodes, ρ ij represents the i th node and the j th node. The Spearman correlation coefficient between them, where i , j = 1, 2,..., n. The process of feature selection and graph sample construction based on correlation analysis is as Figure 2 shown.

[0092] Step S4, model training. Based on the graph neural network and the contrastive learning method, construct a contrastive residual graph neural network (CRGN) for information mining and model training;

[0093] The composition form of the contrastive residual graph neural network model can be seen in Appendix Figure 3As shown, it adopts a self-supervised learning mode, sets two encoders for contrastive learning, and the two encoders share weight parameters. Each encoder contains a graph convolutional layer, a max pooling layer, and a residual connection. At the same time, the Leaky–ReLU function is used as the activation function, and the negative half-axis coefficient is set to 0.1; the residual connection in the model is mainly used to ensure information integrity;

[0094] During the model training process, the original samples are masked with Gaussian random noise with a standard deviation of 0.1, and the original samples and masked samples are respectively input into the two encoders; the objective function of the model training is based on the contrastive loss, and L2-regularization is added during the training process; the calculation formula of the objective function is:

[0095]

[0096] where and are the positive and negative sample feature representations respectively, sim represents the cosine similarity, is the temperature parameter (set to 0.7); N is the number of samples, λ is the L2-regularization coefficient, w are the model parameters; the Stochastic Gradient Descent optimizer SGD with Nesterov momentum is used for model training.

[0097] Step S5, health guidance calculation. Load the contrastive residual graph neural network model trained with health data, input the unknown state samples into the pre-trained model, calculate the distance in the multi-dimensional space between the output result of the pre-trained model and the prediction result of the health samples, and construct the gearbox health guidance based on the exponentially weighted moving average and the normalized average distance;

[0098] The specific method for health guidance calculation is as follows:

[0099] Construct the graph data samples of the gearbox to be monitored based on the same data preprocessing and sampling mode as above, load the trained contrastive residual graph neural network model with health data into the monitoring system, input the unknown state samples into the pre-trained model, and obtain a predicted value;

[0100] Then, use the Manhattan distance , Euclidean distance , cosine similarity , Chebyshev distance The mean of these 4 distance metrics is used to calculate the average distance between the health samples and the samples to be monitored. The calculation formula is:

[0101]

[0102]

[0103]

[0104]

[0105]

[0106] Among them, a , b are two vectors for which the distance is to be calculated, a i , b i are respectively the elements in the two vectors, n is the vector dimension, Normed indicates normalizing the calculation result; on this basis, a gearbox health guideline is constructed based on the exponentially weighted moving average (EWMA). HI , and the formula is:

[0107]

[0108] Among them is the time health guideline, is the penalty factor, is the time health guideline, is the time normalized average distance;

[0109] Next, median filtering denoising is performed on the calculation result, and the filtering window size is set to 100.

[0110] Step S6, fault threshold determination, determine the fault threshold based on the Statistical Process Control (SPC) method;

[0111] From a statistical perspective, the SPC method can use the relevant principles of statistical analysis to monitor the production process in real time and scientifically distinguish the common causes and special causes of product quality fluctuations in the production process. Assume that the quality variable of a certain product follows a normal distribution during the production process, and its quality characteristic X follows a normal distribution with a mean of , and a standard deviation of . Then, the calculation formula for the probability P of the quality characteristic within is:

[0112]

[0113] In practical applications, use the sample mean and the standard deviation represent the mean and variance of the normal distribution, then we define the range exceeding as the fault state. Therefore, according to the calculated state indication parameters, the threshold is calculated according to the SPC principle. Considering the HI definition of the Th value, when selecting the boundary value of the region here, only the value corresponding to the upper boundary is considered, and this upper boundary is the failure threshold

[0114]

[0115] defined here, and the calculation formula is: is the mean of the health guidance sequence, is the standard deviation of the health guidance sequence.

[0116] Step S7, state monitoring and early warning. Compare the relative magnitudes of the gearbox health guidance and the fault threshold to achieve online state monitoring of the wind turbine gearbox. During the state monitoring process, when the calculated HI exceeds the corresponding failure threshold Th , it can be considered that the operating state of the unit at the corresponding moment is abnormal. After troubleshooting the existing problems, the work can continue.

[0117] The present invention also provides a wind turbine gearbox state monitoring system based on a self-supervised contrast residual graph network, including: a data preprocessing module, an adaptive feature selection module, a graph data sample construction module, a model training module, a health guidance calculation module, a fault threshold determination module, and a state monitoring and early warning module;

[0118] The data preprocessing module is used to obtain the original SCADA data from the SCADA system installed in the wind farm and preprocess the original SCADA data;

[0119] The adaptive feature selection module selects data features with strong correlation with the wind turbine gearbox based on the adaptive feature selection method;

[0120] The graph data sample construction module is used for graph data sample construction, and constructs the data features after adaptive feature selection into graph data samples;

[0121] The model training module constructs a contrast residual graph neural network based on the graph neural network and the contrast learning method for information mining and model training;

[0122] The health guidance calculation module loads the contrast residual graph neural network model trained with health data, inputs the unknown state samples into the pre-trained model, calculates the distance in the multi-dimensional space between the output result of the pre-trained model and the prediction result of the health samples, and constructs the gearbox health guidance based on the exponentially weighted moving average and the normalized average distance;

[0123] The fault threshold determination module determines the fault threshold based on the statistical process control method;

[0124] The state monitoring and warning module is used to compare the relative magnitudes of the gearbox health guidance and the fault threshold to achieve online state monitoring of the wind turbine gearbox.

[0125] After the original SCADA data undergoes preprocessing such as data null value removal, based on the adaptive feature selection method, data features with strong correlation with the wind turbine gearbox are selected, and data samples are established based on the correlation. The process of feature selection is adaptive, so the features selected for different wind turbines are dynamically changing. On this basis, based on the Gaussian random process, Gaussian random noise is added to the data samples to construct a mask for model training based on the contrast learning method. The standard deviation of the Gaussian random noise is set to During the offline model training process based on contrast learning, 2 encoders are constructed and share the initial weights. Each encoder is a graph neural network containing residual connections. The model training process is based on the backpropagation algorithm and is trained using the original data samples of the healthy state. After training is completed, the pre-trained model is loaded into the online monitoring system to construct graph samples for the online data of the unknown state, and through the calculation of the pre-trained model, the distance in the multi-dimensional space between the obtained prediction result and the prediction result of the health samples is calculated. Based on this distance metric and combined with the exponentially weighted moving average (EWMA) method, the health guidance is constructed HI Then, the failure threshold is established based on the statistical process control (SPC) method Th . Compare HI and Th to achieve online state monitoring of the wind turbine gearbox.

[0126] The present invention can achieve precise and efficient monitoring: Adaptive feature selection effectively removes redundant information, retains key features, improves monitoring efficiency and reliability, enables the model to focus on core data, and enhances monitoring accuracy. It has powerful model performance: By combining the advantages of multiple technologies in the contrast residual graph neural network model, through contrastive learning and residual connections, and by combining the correlation relationships between data features, it enhances the model's expressive ability and information extraction ability, accurately mines data features, and accurately identifies the state of the gearbox. It has a reliable early warning mechanism: A health guidance is constructed based on multiple distance metrics, a gearbox health guidance is constructed by combining exponential weighted moving average, and a fault threshold is determined by combining statistical process control to improve the reliability of the method. It can achieve an abnormal early warning 30 - 40 hours in advance, with an abnormal recognition accuracy exceeding 90%, which buys time for the operation and maintenance process and reduces losses. It has wide applicability: The method of the present invention realizes an end-to-end state monitoring of the gearbox for the SCADA data of wind turbines, provides methods and ideas, and can implement this function for different wind turbines and their supporting SCADA systems, with generalization, and can meet the state monitoring requirements of different wind turbine units.

[0127] Embodiment

[0128] This embodiment provides a method for monitoring the state of a wind turbine gearbox based on a self-supervised contrast residual graph network, including the following steps:

[0129] (1) Data acquisition and preprocessing

[0130] Data is acquired from the SCADA system installed in the wind farm. These data contain several features, covering various information of the operation of the wind turbine unit, such as power, wind speed, temperature, pressure, etc. The original data may have null values, error values, and abnormal data caused by sensor failures. Data cleaning algorithms are used to automatically identify and delete null value records; for error values, they are corrected according to the upper and lower limit ranges and logical relationships of the data, and are deleted if they cannot be corrected; for abnormally fluctuating data, it is judged by combining historical data and physical laws. If it is caused by a sensor failure, data interpolation or filtering algorithms are used for repair.

[0131] (2) Adaptive feature selection

[0132] Data distribution test: Perform an Anderson-Darling test on each feature of the original SCADA data to determine whether it follows a normal distribution. This test calculates a specific statistic and compares it with a critical value to draw a conclusion. If the feature data does not follow a normal distribution, it indicates that its distribution form is relatively complex. At this time, using the Pearson correlation coefficient based on the normal distribution assumption will produce biases, so the Spearman correlation coefficient is selected for subsequent correlation analysis.

[0133] Core feature determination: Combining past maintenance experience of wind turbines, the working principle of the gearbox, and the maintenance reports provided by the wind farm, five features, namely the average oil temperature of the gearbox, the average oil temperature at the inlet of the gearbox, the average value of the oil inlet pressure of the gearbox, the average value of the impeller speed, and the average value of the generator speed, are determined as core features. The average oil temperature of the gearbox reflects the internal thermal state of the gearbox. Excessive oil temperature may imply increased gear wear or poor lubrication; the average oil temperature at the inlet of the gearbox affects the initial performance of the lubricating oil; the oil inlet pressure of the gearbox is related to the lubrication effect; the impeller speed and the generator speed are closely related to the transmission efficiency of the gearbox, and their abnormal changes can directly reflect the operating state of the gearbox.

[0134] Calculation of correlation coefficients and feature screening: Taking each core feature as a reference, calculate the Spearman correlation coefficient between other features and it. For example, when calculating the correlation coefficient between feature and the core feature F1 (average oil temperature of the gearbox), calculate for all data samples according to the Spearman correlation coefficient formula. After the calculation is completed, calculate the average value of the correlation coefficients between each non-core feature and the five core features to obtain the average correlation coefficient . Set the quantization index to 0.6, and combine the non-core features with an average correlation coefficient greater than 0.6 with the five core features to form the original data samples for model training.

[0135] Data normalization: Use the maximum-minimum normalization method to process the selected feature data. Through normalization, map the data of different features to the interval [0, 1], eliminate the influence of dimensions, make different features equally important in model training, and improve the model training effect.

[0136] (3) Model training

[0137] Construction of graph data samples: Construct the features after adaptive feature selection into graph data samples. Consider each feature as a node in the graph, and determine the edges according to the Spearman correlation coefficients between features. Different edges are used to construct the adjacency matrix. In this way, each feature and its relationship with other features are presented in the form of a graph, providing structured data for the training of graph neural networks.

[0138] Construction of Contrast Residual Graph Neural Network (CRGN): CRGN adopts a self-supervised learning mode and sets two encoders for contrastive learning. The two encoders share weight parameters. Each encoder contains a graph convolutional layer (GCN), a max pooling layer, and a residual connection. At the same time, the Leaky–ReLU function is used as the activation function, and the negative half-axis coefficient is set to 0.1. The residual connection in the model is mainly used to ensure information integrity. Composition. GCN extracts features from graph data through convolutional operations, combines the processing ability of graph neural networks for graph-structured data and the feature extraction advantages of convolutional neural networks, and is suitable for processing graph data samples for the condition monitoring of wind turbine gearboxes.

[0139] Contrastive learning training: The original data is masked with Gaussian random noise with a standard deviation of 0.1 to generate masked data. The original data and the masked data are respectively input into the two encoders, and the model is allowed to learn the key features of the data through contrastive learning. The objective function based on the contrastive loss is:

[0140]

[0141] where and are the feature representations of the positive and negative samples respectively, sim represents the cosine similarity, is the temperature parameter. The temperature parameter controls the smoothness of the contrastive loss function and affects the training effect of the model. The model is trained using a stochastic gradient descent (SGD) optimizer with Nesterov momentum, and the momentum is set to 0.9. Nesterov momentum enables the optimizer to consider the change trend of the gradient when updating parameters, accelerating the convergence of the model; at the same time, L2 regularization is used to prevent the model from overfitting and improve the generalization ability of the model. The hyperparameter settings for model training are shown in the following table:

[0142] Number Hyperparameter Value 1 Learning rate 1e-4 2 Training epochs 3000 3 Momentum 0.9 4 Batch size 256 5 Number of GCN layers in each encoder 4 6 L2 - regularization coefficient 1e-5 7 Temperature coefficient for controlling shape 0.7 8 Input channels of GCN 32 9 Channels of GCN hidden layer 64 10 Output layer channels 1280

[0143] During the training process, the loss value of the model on the validation set is recorded every certain number of training rounds to observe the training trend of the model. If the loss value on the validation set no longer decreases after multiple rounds of training, it is considered that the model has reached the convergence state and the training is stopped.

[0144] (4) Health guidance calculation and fault threshold determination

[0145] Health guidance calculation: Based on the same data preprocessing and sampling pattern, construct graph samples of the gearbox to be monitored. Load the CRGN model trained with healthy data, input the unknown state samples into the pre-trained model, and obtain the model output results. Calculate the average distance between the healthy samples and the unknown samples using various distance metrics such as Manhattan distance, Euclidean distance, cosine similarity, and Chebyshev distance .

[0146] Construct the gearbox health index HI based on the exponentially weighted moving average (EWMA) and the normalized average distance. The calculation formula is

[0147]

[0148] where is the health index at time , is the penalty factor , is the normalized average distance at time

[0149] Apply median filtering to the calculated health index results to remove noise. The size of the filtering window is set to 100. Median filtering sorts the data within the window and takes the median value as the filtering output, which can effectively remove sudden noise and make the health index more accurately reflect the actual state of the gearbox

[0150] Fault threshold determination: Determine the fault threshold based on the statistical process control (SPC) method. In practical applications, use the sample mean and the standard deviation to represent the mean and variance of the normal distribution. Then we define the range exceeding as the fault state. Therefore, according to the calculated state indicator parameters, calculate the threshold according to the SPC principle. Considering the HI value definition, when selecting the regional boundary value here, only consider the value corresponding to the upper boundary, and this upper boundary is the failure threshold Th defined here. The calculation formula is

[0151]

[0152] During the state monitoring process, when the calculated HI exceeds the corresponding failure threshold Th , it can be considered that the operating state of the unit at the corresponding time is abnormal. After troubleshooting the existing problems, work can continue

[0153] (5) State monitoring and early warning

[0154] During the operation of the wind turbine, the SCADA data is collected in real time. According to the above processes of data preprocessing, adaptive feature selection, graph data sample construction, model calculation, health guidance calculation, and fault threshold comparison, the operation status of the gearbox is continuously monitored. When the health guidance HI exceeds the fault threshold Th , it is determined that the operation status of the gearbox of the wind turbine is abnormal, and the system immediately issues a warning signal. The warning signal is notified to the operation and maintenance personnel by means of text messages, emails, or pop-up windows in the wind farm monitoring system. At the same time, the time of the abnormality occurrence, the health guidance value, relevant feature data, and other information are recorded in the system to facilitate the operation and maintenance personnel to analyze the cause of the fault later. After receiving the warning, the operation and maintenance personnel check and maintain the gearbox according to the detailed information provided by the system, combined with the historical operation data of the wind turbine and the actual situation on site, such as checking the gear wear condition, the quality of the lubricating oil, the bearing status, etc., and timely eliminate the potential fault hazards to ensure the safe and stable operation of the wind turbine. After the fault is processed, the maintenance records and processing results are fed back into the system to update the health status information of the gearbox and start a new round of status monitoring.

[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for monitoring the state of a wind turbine gearbox based on a self-supervised contrastive residual graph network, characterized in that, It includes the following steps: Step S1, data preprocessing: Obtain the original SCADA data from the SCADA system installed in the wind farm, and preprocess the original SCADA data; Step S2, adaptive feature selection: Based on the adaptive feature selection method, select the data features with strong correlation with the wind turbine gearbox; Step S3, graph data sample construction: Construct the data features after adaptive feature selection into graph data samples; Step S4, model training: Based on the graph neural network and the contrastive learning method, construct a contrastive residual graph neural network for information mining and model training; The contrastive residual graph neural network model adopts a self-supervised learning mode, sets two encoders for contrastive learning, the two encoders share weight parameters, each encoder contains a graph convolutional layer, a max pooling layer, and a residual connection, and at the same time uses the Leaky–ReLU function as the activation function, and the negative half-axis coefficient is set to 0.1; the residual connection is used in the model to ensure information integrity; During the model training process, the original samples are masked with Gaussian random noise with a standard deviation of 0.1, and the original samples and the masked samples are respectively input into the two encoders; the objective function of the model training is based on the contrastive loss, and L2-regularization is added during the training process; The calculation formula of the objective function is: where is the feature representation, and are the positive and negative samples respectively, sim represents the cosine similarity, is the temperature parameter; N is the number of samples, K is the number of negative samples, λ is the L2 - regularization coefficient, w are the model parameters; Stochastic Gradient Descent optimizer SGD with Nesterov momentum is used for model training; Step S5, health guidance calculation: Load the contrastive residual graph neural network model trained with health data, input the unknown state samples into the pre-trained model, calculate the distance in the multi-dimensional space between the output result of the pre-trained model and the prediction result of the health samples, and construct the gearbox health guidance based on the exponentially weighted moving average and the normalized average distance; Step S6, fault threshold determination: Determine the fault threshold based on the statistical process control method; Step S7, status monitoring and warning: Compare the relative sizes of the gearbox health guidance and the fault threshold to achieve online status monitoring of the wind turbine gearbox.

2. The method for monitoring the state of a wind turbine gearbox based on a self-supervised contrastive residual graph network according to claim 1, wherein, In step S1, the preprocessing method includes: using a data cleaning algorithm to automatically identify and delete null value records; for error values, correct them according to the upper and lower limit ranges and logical relationships of the data, and delete them if they cannot be corrected; for abnormally fluctuating data, judge in combination with historical data and physical laws, and if it is caused by a sensor failure, use data interpolation or filtering algorithms for repair.

3. A method for monitoring the state of a wind turbine gearbox based on a self-supervised contrastive residual graph network according to claim 2, characterized in that, In step S2, based on the adaptive feature selection method, the method for selecting the data features with strong correlation with the wind turbine gearbox is as follows: For the original SCADA data, select Spearman the coefficient to conduct a correlation analysis. The formula is: where is Spearman the correlation coefficient, is the rank difference of each pair of observed values, b is the total number of observed values; Based on the operation and maintenance reports from the wind farm, select the five data features of the average oil temperature of the gearbox, the average oil temperature at the inlet of the gearbox, the average value of the oil inlet pressure of the gearbox, the average value of the impeller speed, and the average value of the generator speed as the core features, and calculate the average correlation coefficient of other features with the core features. The formula is: wherein is the average value of the correlation coefficients of other features relative to the core feature, is the s correlation coefficient between the r th other feature and the th core feature, and is the number of core features; Set a quantization index. When the average correlation coefficient exceeds this quantization index, the corresponding data feature is selected; determine the variables with correlation coefficients exceeding this index as strongly correlated variables, and jointly form the model training samples with the core features; Finally, perform maximum-minimum normalization on the feature data after feature selection. The formula is: wherein is the normalized data, is the original data, and are the minimum and maximum values in the original data, respectively.

4. A method for monitoring the state of a wind turbine gearbox based on a self-supervised contrastive residual graph network according to claim 3, characterized in that, In step S3, the method for constructing the data features after adaptive feature selection into graph data samples is as follows: Construct the features after adaptive feature selection into graph data samples. In this graph data sample, each feature serves as a node, and the correlation between features serves as an edge. The correlation uses the Spearman coefficient, and the adjacency matrix is constructed according to this process. The adjacency matrix is shown in the following formula: In the formula, f is the number of nodes, is the correlation coefficient between the s th other feature and the r th core feature, where r , t = 1, 2, …, f .

5. A method for monitoring the state of a wind turbine gearbox based on a self-supervised contrastive residual graph network according to claim 4, characterized in that, In step S5, the specific method for calculating the health guidance is as follows: Construct the graph data sample of the gearbox to be monitored, load the trained contrast residual graph neural network model with health data into the monitoring system, input the unknown state sample into the pre-trained model, and obtain a predicted value; Then, the Manhattan distance , Euclidean distance , cosine similarity , Chebyshev distance is used to calculate the average distance between the healthy samples and the samples to be monitored by taking the mean of these four distance metrics , and the calculation formula is: Among them, a , b are two vectors for which the distance is to be calculated, , are respectively the individual elements in the two vectors, n is the dimension of the vector, Normed indicates normalization of the calculation result; on this basis, a gearbox health guideline is constructed based on exponential weighted moving average HI , and the formula is: Among them is the health guidance at a certain moment, is the penalty factor, is the health guidance at a certain moment, is the normalized average distance at a certain moment; Next, perform median filtering denoising on the calculation result, and set the filter window size to 100.

6. The method for monitoring the state of a wind turbine gearbox based on a self-supervised contrast residual map network according to claim 5, wherein In step S6, the method for determining the fault threshold based on the statistical process control method is as follows: Using the sample mean and the standard deviation to represent the mean and variance of the normal distribution, defining the range exceeding the preset value as the fault state, calculating the threshold according to the calculated state indication parameter according to the principle of statistical process control; considering the gearbox health guideline HI definition, when selecting the regional boundary value here, only consider the value corresponding to the upper boundary, and this upper boundary is the failure threshold defined here Th , and the calculation formula is: Among them, is the mean of the health guidance sequence, is the standard deviation of the health guidance sequence.

7. A monitoring system for implementing the wind turbine gearbox condition monitoring method based on the self-supervised contrast residual graph network according to any one of claims 1-6, characterized in that, Including: A data preprocessing module, an adaptive feature selection module, a graph data sample construction module, a model training module, a health guidance calculation module, a fault threshold determination module, and a status monitoring and warning module; The data preprocessing module is used to obtain the original SCADA data from the SCADA system installed in the wind farm and preprocess the original SCADA data; The adaptive feature selection module selects the data features with strong correlation with the wind turbine gearbox based on the adaptive feature selection method; The graph data sample construction module is used for graph data sample construction, and constructs the data features after adaptive feature selection into graph data samples; The model training module constructs a contrast residual graph neural network based on the graph neural network and the contrast learning method for information mining and model training; The health guidance calculation module loads the trained contrast residual graph neural network model with health data, inputs the unknown state sample into the pre-trained model, calculates the distance in the multi-dimensional space between the output result of the pre-trained model and the prediction result of the health sample, and constructs the gearbox health guidance based on the exponentially weighted moving average and the normalized average distance; The fault threshold determination module determines the fault threshold based on the statistical process control method; The status monitoring and warning module is used to compare the relative sizes of the gearbox health guidance and the fault threshold to achieve online status monitoring of the wind turbine gearbox.

Citation Information

Patent Citations

  • Distribution network fault positioning method based on short-time matrix pencil method and graph neural network

    CN118191510A