Wind power gear box intelligent fault early warning method and system based on machine learning

By using machine learning methods to perform time-frequency joint analysis and adaptive feature selection on multi-source monitoring data of wind turbine gearboxes, a dynamic fault feature weight matrix is ​​constructed. Combined with deep neural networks and time-series pattern matching, the accuracy and adaptability issues of early fault identification in wind turbine gearboxes are solved, and efficient fault early warning is achieved.

CN120998009AActive Publication Date: 2025-11-21华电重庆新能源有限公司

Patent Information

Application Number
CN202511518049.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2025-11-21
Estimated Expiration
2045-10-23

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively identify early-stage faults in wind turbine gearboxes. Traditional single-parameter monitoring is easily masked by noise, and multi-parameter fusion technology cannot fully uncover correlated features, resulting in poor accuracy and reliability of early warnings. Static models also lack adaptability.

Method used

A machine learning-based approach is adopted, which uses time-frequency joint analysis of multi-source monitoring data, adaptive feature selection and dynamic fault feature weight matrix, combined with deep neural network to construct fault evolution feature space, and uses time-series pattern matching algorithm to identify fault development patterns and generate graded early warning signals.

Benefits of technology

It enables accurate identification of early faults in wind turbine gearboxes, improves the accuracy and reliability of early warning, adapts to different fault types and operating conditions, reduces misjudgments and omissions, and provides more comprehensive data support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998009A_ABST
    Figure CN120998009A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of wind power equipment monitoring, and discloses a wind power gear box intelligent fault early warning method and system based on machine learning. The method comprises the steps that multi-source monitoring data such as vibration signals, temperature data and oil analysis data of the wind power gear box are acquired, and multi-scale operation characteristics are extracted through time-frequency conjoint analysis; key fault sensitive features are determined through an adaptive feature selection algorithm, and a dynamic fault feature weight matrix is constructed in combination with a historical fault case library; multi-modal data fusion is adopted to generate an enhanced fault feature set, and modal decomposition is carried out on the enhanced fault feature set to obtain a trend component and a fluctuation component; a fault evolution feature space is constructed by using a deep neural network based on two components, then a fault development mode is identified by using a time sequence mode matching algorithm, and finally a graded early warning signal is generated according to a matching degree with a preset mode, so that fault features can be comprehensively captured, and safe operation of a wind power gear box is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wind power equipment monitoring technology, specifically to a method and system for intelligent fault early warning of wind turbine gearboxes based on machine learning. Background Technology

[0002] During the rapid development of the wind power industry, the wind turbine gearbox, as the core transmission component of the wind turbine generator set, directly affects the stable operation and power generation efficiency of the entire unit. Because wind turbine gearboxes operate in complex environments, they must withstand frequent fluctuations in wind loads and cope with harsh natural conditions such as low temperatures, high humidity, and sandstorms. This makes them highly susceptible to faults such as gear wear, bearing spalling, and shaft deformation. If these faults are not detected and addressed promptly, they can lead to unit shutdowns for maintenance, resulting in considerable economic losses, or even complete gearbox failure and major safety accidents such as turbine collapse. Currently, the technical means for early warning of wind turbine gearbox faults in the industry are mainly divided into two categories: single-parameter monitoring and traditional multi-parameter fusion. Single-parameter monitoring technology usually relies on only one of the vibration signals or temperature data for fault diagnosis. For example, vibration sensors are used to collect vibration signals from the gearbox, and traditional signal processing methods such as Fourier transform are used to extract features and identify faults. However, this approach has significant limitations. When the gearbox is in the early stage of a fault, the fault characteristics are very weak in a single signal and are easily masked by the background noise generated by the normal operation of the equipment. It is difficult to effectively capture early faults, and the problem is often only discovered when the fault characteristics become significant, by which time the best maintenance opportunity has been missed. While traditional multi-parameter fusion technologies attempt to combine various monitoring data such as vibration, temperature, and oil levels for fault early warning, they suffer from shortcomings in data processing and feature utilization. These technologies mostly employ simple feature splicing or weighted summation methods for data fusion, failing to fully consider the inherent correlations and dynamic changes between different monitoring parameters. For example, during gearbox operation, a slow increase in temperature may be synergistically related to the gradual increase in gear wear. Traditional multi-parameter fusion technologies cannot effectively uncover this correlation, resulting in low feature recognition after fusion and an inability to accurately reflect the actual development state of the fault. Furthermore, traditional technologies often use fixed feature selection criteria, failing to adaptively adjust feature selection strategies based on different operating conditions and fault types of the gearbox. This easily introduces redundant features or misses key fault-sensitive features, thus affecting the accuracy and reliability of fault early warning. In addition, fault early warning models built using traditional technologies are mostly static models, unable to dynamically update model parameters with the accumulation of historical fault cases and changes in equipment operating conditions. This leads to poor adaptability of the model to new faults or faults under complex operating conditions, resulting in a gradual decline in early warning effectiveness. Summary of the Invention

[0003] The purpose of this invention is to provide a machine learning-based intelligent fault early warning method and system for wind turbine gearboxes to solve the problems mentioned in the background art.

[0004] To achieve the above objectives, this invention provides a machine learning-based intelligent fault early warning method for wind turbine gearboxes, the method comprising: Acquire multi-source monitoring data during the operation of the wind turbine gearbox, including vibration signals, temperature data, and oil analysis data; The multi-source monitoring data are subjected to time-frequency joint analysis to extract multi-scale operational features; Based on the aforementioned multi-scale operational characteristics, key fault-sensitive features are determined using an adaptive feature selection algorithm. Based on the historical failure case library and the key failure sensitivity features, a dynamic failure feature weight matrix is ​​constructed. The historical failure case library includes the characteristic patterns and weight assignments of past failure events; A multimodal data fusion method is used to enhance the dynamic fault feature weight matrix to generate an enhanced fault feature set. Modal decomposition is performed on the enhanced fault feature set to separate the trend component characterizing the fault development process and the fluctuation component characterizing the fault suddenness. Based on the trend component and fluctuation component, a fault evolution feature space is constructed using a deep neural network. The deep neural network has powerful nonlinear fitting and feature learning capabilities, and can further mine the deep features of fault evolution from trend components and fluctuation components, and construct a feature space that better reflects the essence of fault development. In the fault evolution feature space, a time-series pattern matching algorithm is used to identify fault development patterns; Based on the degree of matching between the fault development mode and the preset fault mode, a graded early warning signal is generated; The preset fault mode is a preset mode used for comparison and matching in the fault evolution feature space, and is composed of typical fault data in the historical fault case library.

[0005] Preferably, the time-frequency joint analysis of the multi-source monitoring data includes: Wavelet packet transform is performed on the vibration signal to extract the energy distribution characteristics of different frequency bands; Establish a time series model for temperature data and extract temperature change trend features; Spectral feature extraction was performed on oil analysis data to obtain the distribution characteristics of wear particles; The energy distribution characteristics, temperature change trend characteristics, and wear particle distribution characteristics are fused at the feature level.

[0006] Preferably, the determination of key fault-sensitive features using an adaptive feature selection algorithm includes: Calculate the correlation coefficients between each feature dimension and historical fault records; A ranking of feature importance is constructed based on the correlation coefficients; A sliding window mechanism is used to dynamically adjust the importance ranking of the features; Features with a predetermined proportion before sorting are selected as key fault-sensitive features.

[0007] Preferably, the construction of the dynamic fault feature weight matrix includes: The critical fault sensitivity features are classified according to the equipment's operating conditions. Establish feature weight allocation rules for each type of working condition; The feature weight allocation rule is dynamically adjusted based on real-time operating condition data. Perform matrix operations between the adjusted feature weights and the critical fault-sensitive features.

[0008] Preferably, the feature enhancement of the dynamic fault feature weight matrix using a multimodal data fusion method includes: Principal component analysis is performed on the dynamic fault feature weight matrix to extract the main feature vectors after principal component analysis. The main feature vectors are then nonlinearly combined with the original features, and new combined features are generated through feature cross-operation.

[0009] Preferably, the modal decomposition of the enhanced fault feature set includes: The enhanced fault feature set is processed using the empirical mode decomposition method; The first three intrinsic mode components after decomposition are extracted as wave components; The remaining components are superimposed and reconstructed to obtain the trend components; Calculate the energy ratio of the fluctuation component to the trend component.

[0010] Preferably, the construction of the fault evolution feature space includes: The trend component is input into a long short-term memory network to extract long-term evolutionary features; The fluctuation components are input into a convolutional neural network to extract local anomaly features; The long-term evolutionary features and local anomaly features are then concatenated. The spliced ​​features are dimensionality reduced using an autoencoder.

[0011] Preferably, the step of using a time-series pattern matching algorithm to identify fault development patterns includes: Establish a fault mode template library in the reduced feature space; Calculate the dynamic time warp distance between the real-time feature sequence and each fault mode template; The most suitable fault mode is determined based on the dynamic time warping distance; Evaluate the similarity between the current feature sequence and the best-matching fault mode.

[0012] Preferably, the generation of graded early warning signals includes: Warning level thresholds are determined based on the degree of similarity. A level one warning signal is generated when the similarity exceeds the first threshold. A secondary warning signal is generated when the similarity exceeds the second threshold but is lower than the first threshold. A level 3 warning signal is generated when the similarity exceeds the third threshold but is lower than the second threshold. The first threshold is the lowest quantile of the similarity score between when the fault is about to occur or when it is already in the early stages of occurrence; The second threshold is the lowest quantile of the similarity score when the device status has shown a strong similarity to a certain failure mode. The third threshold is the lowest quantile of the similarity score of the device when it is in a stage where there may be early signs of abnormality or its operating state begins to deviate from the healthy baseline state.

[0013] Preferably, the present invention also includes a machine learning-based intelligent fault early warning system for wind turbine gearboxes. The system includes a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of the machine learning-based intelligent fault early warning method for wind turbine gearboxes described above.

[0014] Compared with the prior art, the beneficial effects of the present invention are: This method first acquires multi-source monitoring data, including vibration signals, temperature data, and oil analysis data, during the operation of the wind turbine gearbox. Compared to traditional single-parameter monitoring techniques, multi-source monitoring data can reflect the gearbox's operating status from different perspectives. Vibration signals can directly reflect the mechanical state changes of moving parts such as gears and bearings; temperature data can reflect the thermal operating characteristics of the equipment; and oil analysis data can indirectly determine the wear condition of internal components through information such as metal particles and contaminant content in the oil. The introduction of multi-source data enriches the sources of fault characteristics, effectively avoiding the problem of fault characteristics being masked by noise in single-parameter monitoring, and providing more comprehensive data support for early fault identification.

[0015] In the data processing stage, this method uses time-frequency joint analysis to extract multi-scale operating features. Time-frequency joint analysis technology can analyze monitoring signals from both the time and frequency domains simultaneously. Compared with the limitation of traditional Fourier transform, which can only analyze signals in the frequency domain, it can more accurately capture the characteristic changes of signals in different time and frequency ranges. In particular, for non-stationary and nonlinear signals generated by early faults, it can extract richer and more subtle multi-scale operating features. These features can more clearly reflect the nascent state of early gearbox faults, laying the foundation for subsequent fault early warning.

[0016] Based on multi-scale operational characteristics, an adaptive feature selection algorithm is used to determine key fault-sensitive features. This algorithm dynamically adjusts the criteria and weights for feature selection according to changes in the actual operating conditions and fault types of the gearbox, automatically filtering out features that contribute significantly to fault identification and eliminating redundant and interfering features. This adaptive feature selection method avoids the shortcomings of traditional fixed feature selection criteria that cannot adapt to changes in operating conditions, ensuring that the selected features accurately point to the essence of the fault and improving the efficiency and accuracy of subsequent fault identification.

[0017] A dynamic fault feature weight matrix is ​​constructed based on a historical fault case database and key fault sensitivity features. This matrix is ​​not fixed but continuously updated as historical fault cases accumulate and new fault types emerge, reflecting in real time the changing importance of different fault features in fault identification. This dynamic update mechanism makes the setting of fault feature weights more closely match actual fault situations, avoiding the problem of poor adaptability to new faults caused by fixed parameters in traditional static models, and improving the model's ability to identify different fault types.

[0018] A multimodal data fusion method is used to enhance the dynamic fault feature weight matrix and generate an enhanced fault feature set. This fusion method is not a simple feature superposition, but rather a deep exploration of the intrinsic correlation and synergistic change patterns between different modal data such as vibration, temperature, and oil. Through the complementary fusion of multimodal data, the recognizability of fault features is enhanced, so that the enhanced fault feature set can more comprehensively and accurately reflect the fault state of the gearbox and reduce misjudgments and omissions caused by insufficient feature information.

[0019] Modal decomposition of the enhanced fault feature set separates the trend component characterizing the fault development process and the fluctuation component characterizing the fault's sudden onset. This decomposition method clearly distinguishes the long-term trend and short-term sudden characteristics of fault development, facilitating a more detailed analysis of the fault's evolution. The trend component can be used to track the gradual process of a fault from its inception to its development, while the fluctuation component can capture abnormal changes during a sudden fault event. By analyzing and comprehensively utilizing these two components separately, the development stage and severity of the fault can be grasped more accurately, providing a clearer basis for subsequent identification of fault development patterns.

[0020] Based on trend and fluctuation components, a fault evolution feature space is constructed using deep neural networks. Deep neural networks possess powerful nonlinear fitting and feature learning capabilities, enabling them to further extract deeper features of fault evolution from the trend and fluctuation components, thus constructing a feature space that better reflects the essence of fault development. Within this feature space, fault features of different fault types and development stages can form a highly discriminative feature distribution, creating favorable conditions for the accurate identification of fault development patterns.

[0021] In the fault evolution feature space, a time-series pattern matching algorithm is used to identify fault development patterns. This algorithm can accurately match the time-series features in the fault evolution process. By comparing the current fault evolution features with known fault development patterns, it can quickly and accurately identify the specific development pattern of the current fault, including fault type, development speed, possible deterioration path and other information. Attached Figure Description

[0022] Figure 1 This is a schematic diagram illustrating the working principle of the intelligent fault early warning method for wind turbine gearboxes based on machine learning as described in this invention. Figure 2 This is a flowchart of the time-frequency joint analysis; Figure 3 A flowchart for constructing the dynamic fault feature weight matrix. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] Please see Figure 1This invention provides a machine learning-based intelligent fault early warning method for wind turbine gearboxes. The method includes: acquiring multi-source monitoring data of the wind turbine gearbox during operation, including vibration signals, temperature data, and oil analysis data. This data originates from sensors and monitoring systems installed on the gearbox and comprehensively reflects the equipment's operating status. After acquiring the data, time-frequency joint analysis is performed on the multi-source monitoring data. This step aims to extract the equipment's operating characteristics from both the time and frequency domains. Time-frequency joint analysis can capture transient changes and periodic patterns in the signals, thereby extracting multi-scale operating characteristics. Based on the extracted multi-scale operating characteristics, an adaptive feature selection algorithm is used to determine key fault-sensitive features. This adaptive feature selection algorithm can dynamically adjust the importance of features according to data distribution and fault history, selecting the feature subset with the highest correlation to the fault.

[0025] Based on a historical fault case database and key fault sensitivity features, a dynamic fault feature weight matrix is ​​constructed. The historical fault case database contains feature patterns and weight assignments of past fault events, and the dynamic matrix can update feature weights in real time according to changes in operating conditions. A multimodal data fusion method is used to enhance the features of the dynamic fault feature weight matrix. Multimodal fusion integrates information from different data sources and generates an enhanced fault feature set through nonlinear transformation, improving the discriminative ability of the features. Modal decomposition is performed on the enhanced fault feature set to separate the trend component representing the fault development process and the fluctuation component representing the sudden characteristics of the fault. Modal decomposition decomposes the signal into different frequency components; the trend component reflects the gradual evolution of the fault, and the fluctuation component captures sudden anomalies. Based on the trend and fluctuation components, a fault evolution feature space is constructed using a deep neural network. The deep neural network can learn a high-level representation of fault evolution, forming a feature space containing spatiotemporal information. In the fault evolution feature space, a temporal pattern matching algorithm is used to identify fault development patterns. Temporal pattern matching compares real-time data with preset patterns to evaluate the similarity of fault development. Based on the degree of matching between the fault development mode and the preset fault mode, a graded early warning signal is generated. The higher the degree of matching, the more urgent the early warning level, thus realizing intelligent fault early warning.

[0026] Example 1: See Figure 2Time-frequency joint analysis processes vibration signals, temperature data, and oil analysis data generated during the operation of wind turbine gearboxes. Vibration signals are typically acquired by accelerometers mounted on the gearbox bearing housing or gearbox body. The sampling frequency needs to be appropriately set according to the gear meshing frequency and its harmonic components to capture high-frequency resonance information. Temperature data comes from temperature sensors embedded in the bearings or oil sump, recording their time-varying sequence values. Oil analysis data is obtained through online particle counters or periodic laboratory analysis of oil samples, including information on the concentration and size distribution of wear metal particles. Wavelet packet transform analysis is performed on the vibration signals. Wavelet packet transform is a more refined time-frequency analysis method than wavelet transform, capable of further decomposing the high-frequency components of the signal. Choosing appropriate wavelet basis functions and the number of decomposition levels is crucial; for example, using the Db4 wavelet for 4-level decomposition decomposes the original vibration signal into multiple sub-bands with different frequency ranges. The energy value of each sub-band signal is calculated as a feature. These energy distribution features reflect the distribution of vibration energy of different components (such as gears and bearings) inside the gearbox at different frequency bands. When local damage occurs, the energy in a specific frequency band will be significantly enhanced. A time series model is established for temperature data. An autoregressive integral moving average (ARIMA) model is used to model the temperature time series, and the model parameters are determined through training with historical data. Temperature change trend features are extracted from the fitted model, including the long-term trend term of the temperature series, the seasonal fluctuation amplitude, and the statistical characteristics of the residual series. These features help identify slowly developing faults such as overheating trends in the gearbox or decreased cooling system efficiency. Spectral features are extracted from oil analysis data. Using oil spectral analysis data, the concentration change trends of wear metal elements such as iron and copper are focused on. At the same time, the size distribution of wear particles is analyzed, and the proportion of particles in different size ranges is calculated. The wear particle distribution characteristics can directly characterize the wear state of gear meshing surfaces or bearing rolling elements and are an important basis for judging wear-related faults.

[0027] The energy distribution features, temperature change trend features, and wear particle distribution features extracted from different data sources are fused at the feature level. The fusion operation is performed at the feature level, aligning and concatenating feature vectors with different physical meanings and dimensions. Before concatenation, various features are typically normalized to eliminate the influence of dimensions; for example, Z-score normalization is used to transform feature values ​​to a similar numerical range. The fused feature vector forms a comprehensive multi-scale operational feature vector, which simultaneously contains key information from vibration, temperature, and oil data sources in both the time and frequency domains, providing comprehensive input for subsequent feature selection. Based on the fused multi-scale operational features, an adaptive feature selection process is initiated to filter out key fault-sensitive features. This process first calculates the correlation coefficient between each dimension of the feature vector and historical fault record labels to quantify the linear correlation strength between each feature and the fault state. When calculating the correlation coefficient, accumulated historical datasets are used, which contain feature samples under normal conditions and various known fault states. The Pearson correlation coefficient is a commonly used metric, with values ​​ranging from -1 to +1; a larger absolute value indicates a stronger correlation between the feature and the fault. Based on the absolute values ​​of the correlation coefficients of each feature obtained from the calculation, a feature importance ranking list is constructed, and the features are arranged in descending order according to the strength of their correlation with the fault.

[0028] A sliding window mechanism is employed to dynamically adjust the feature importance ranking, adapting to slow changes in gearbox operating conditions and external environmental conditions. The sliding window covers feature data from a recent period, such as the past 30 days. The system periodically (e.g., daily) slides the window forward, recalculating the correlation coefficient between each feature and the current operating state (which may include newly emerging anomalies) using the new data within the window, and updating the feature importance ranking. This mechanism allows feature selection to dynamically respond to the latest changes in equipment status; for example, when a new wear pattern begins to emerge, the importance ranking of associated oil features gradually increases.

[0029] From the updated feature importance ranking list, features ranking at the top of a predetermined proportion are selected as critical fault-sensitive features. This predetermined proportion is a parameter that needs to be set according to the specific application scenario; for example, selecting the top 20% of features. These selected features constitute a concise and efficient feature subset, which is considered to be the most indicative of the current gearbox health status. This feature subset will be used to subsequently construct a dynamic fault feature weight matrix, thereby concentrating computational resources on monitoring and analyzing the most critical indicators, improving the accuracy and timeliness of fault early warning. The entire process, through continuous time-frequency analysis and adaptive feature selection, ensures that the fault monitoring model can keep up with changes in equipment status, maintaining its sensitivity and reliability.

[0030] Example 2: See Figure 3 The historical fault case database is a structured database that systematically archives records of various fault events that have occurred in wind turbine gearboxes in the past. Each record typically includes a complete set of characteristic data at the time of the fault, the finally confirmed fault type, the specific time of the fault, and the equipment's operating parameters at that time. This case database needs to be maintained and updated regularly to continuously absorb newly occurring fault case data to maintain its timeliness and coverage. The first step in constructing the matrix is ​​to classify and organize the key fault-sensitive features based on the real-time operating conditions of the equipment. Operating conditions are a series of external and internal condition parameters that affect the equipment's state, mainly including the load on the gearbox, the range of output shaft speed, ambient temperature, and lubricating oil temperature and pressure. The operating condition classification operation is usually completed using unsupervised clustering algorithms, such as the K-means clustering algorithm. This algorithm can automatically divide a large number of historical operating data points into several different clusters based on the characteristics of the operating condition parameters. Each cluster represents a typical operating condition category, such as high-load high-speed operating condition, low-load variable speed operating condition, and low-temperature start-up operating condition. This classification allows complex continuous operating condition variables to be transformed into finite discrete operating condition states, facilitating targeted management in the future.

[0031] For each identified typical operating condition, corresponding feature weighting rules need to be established. The formulation of these rules relies on the knowledge contained in the historical failure case database. Initial weight values ​​are determined by analyzing the strength of the correlation between each key failure-sensitive feature and different failure types within a specific operating condition category. This quantitative analysis of the correlation strength can be achieved using various statistical or machine learning methods. For example, a logistic regression model can be used to calculate the contribution of each feature to the occurrence of the failure (i.e., the coefficient), or prior weights can be assigned by domain experts based on long-term practical experience. Generally, features showing a strong positive or negative correlation with a certain failure mode are assigned higher weights, indicating that the feature has a more significant indicative meaning for failure early warning under that operating condition. The initial weighting rules constitute a weight vector with the same dimension as the number of key failure-sensitive features.

[0032] Because the operating status, performance degradation, and external environment of equipment are not static, fixed weight allocation rules are insufficient to maintain optimal early warning performance. A dynamic adjustment mechanism is needed to adaptively update feature weights based on real-time collected operating data and feature values. The core of this dynamic adjustment mechanism is to continuously monitor current operating parameters and categorize them into a pre-established operating condition category. The system uses the initial weight rules corresponding to that category as a basis. Further fine-tuning is calculated based on the deviation of real-time feature values ​​from historical normal or fault-initiating patterns. For example, a sliding time window statistical method can be used to calculate the mean and standard deviation of each feature value within the current window in real time and compare them with the historical baseline. If the current value of a feature shows a persistent deviation, its weight may be appropriately increased to enhance the system's sensitivity to this abnormal change. This adjustment is a continuous iterative process, ensuring that weight allocation can sensitively respond to the latest evolution trends of equipment status.

[0033] The final feature weight vector, determined after dynamic adjustment, is combined with the real-time acquired key fault-sensitive feature vectors through matrix operations to generate a dynamic fault feature weight matrix. Matrix operations typically employ Hadamard product (element-wise multiplication), ensuring that each feature value is associated with its importance weight. The weighted feature value not only contains the absolute value of the physical quantity but also incorporates its relative importance for fault diagnosis. The resulting dynamic fault feature weight matrix is ​​a two-dimensional data structure. Its rows typically represent consecutive timestamps, recording the sequence of equipment status changes over time, while its columns represent the weighted key fault-sensitive features. Essentially, this matrix is ​​a weighted feature sequence that highlights the most noteworthy fault indications at specific times and under specific operating conditions. This matrix is ​​not static but continuously updated with the influx of new monitoring data and changes in operating conditions, becoming a digital image that dynamically reflects changes in equipment health status. This provides more focused and targeted input data for subsequent feature enhancement and fault pattern recognition. The entire construction process involves four stages: working condition classification, rule formulation, dynamic adjustment, and matrix operation. It closely integrates static historical knowledge, dynamic real-time data, and changing operating environment, enabling the fault early warning system to have stronger context awareness and adaptability.

[0034] Example 3: The dynamic fault feature weight matrix is ​​a two-dimensional data structure containing weighted critical fault-sensitive features. Its rows correspond to time-series samples, and its columns correspond to weighted feature values ​​from different sensor sources. This matrix reflects the relative importance of different features under specific operating conditions. The first step in feature enhancement is to perform principal component analysis (PCA) on this matrix. PCA is a classic linear transformation technique whose core objective is to uncover the main variation patterns within the data. The analysis process begins by calculating the covariance matrix of the dynamic fault feature weight matrix, which characterizes the degree of linear correlation between different weighted feature dimensions. Then, the eigenvalues ​​and corresponding eigenvectors of this covariance matrix are solved. The magnitude of the eigenvalues ​​represents the variance of the data distribution along the corresponding eigenvector direction; a larger variance indicates richer information contained in that direction. The eigenvalues ​​are sorted from largest to smallest, and the eigenvectors corresponding to the k largest eigenvalues ​​are selected to form a projection matrix. The value of k is usually determined by the cumulative variance contribution rate; for example, the smallest k value that makes the cumulative contribution rate exceed 95% is retained. Multiplying the original dynamic fault feature weight matrix with this projection matrix maps the original data from a high-dimensional feature space to a new, lower-dimensional orthogonal subspace, resulting in the principal eigenvectors after principal component analysis. These principal component vectors are linear recombinations of the original weighted features, which are uncorrelated with each other and capture the most significant trends in the data.

[0035] The next crucial step is to nonlinearly combine the obtained principal feature vectors with the original dynamic fault feature weight matrix. This nonlinear combination aims to overcome the limitations of linear transformations in principal component analysis and uncover more complex interactions between features. This combination process can be implemented using a shallow neural network structure that takes the concatenated vectors as input, containing both the principal feature vectors and the original weighted feature values. The network typically has one or more hidden layers, which use nonlinear activation functions, such as rectified linear units or hyperbolic tangent functions. Through the nonlinear transformation of the neural network, a set of intermediate feature representations is generated. These intermediate feature representations are no longer simple linear combinations of the original features but rather new feature abstractions containing higher-order interaction information between features.

[0036] Based on the generated intermediate feature representations, feature cross operations are performed to produce new combined features. Feature cross operations focus on exploring the interaction effects between features from different data modes (such as vibration, temperature, and oil fluid), which may have a stronger indicative power for fault modes. A specific cross operation can be described as follows: given two specific elements in the intermediate feature representation vector, each derived from a different data mode, calculate the interaction term between them. This operation can be expressed as:

[0037] in: These represent feature elements in the intermediate feature representation that originate from a specific data mode (e.g., vibration mode). These represent feature elements derived from another data mode (e.g., temperature mode). and This is an optional transformation function, which in some designs can be an identity function, a small affine transformation, or a nonlinear mapping, used for preprocessing input elements. (Symbol) This represents an interactive operation, which can be a multiplication operation or a more complex bilinear form. The result of the operation... This generates a new combined feature that attempts to explicitly capture the joint variation pattern between vibration and temperature features. By systematically selecting feature pairs from different modes and performing the aforementioned cross-operation, a large number of potentially valuable combined features can be generated.

[0038] The enhanced fault feature set consists of three parts: features from the original dynamic fault feature weight matrix, main eigenvectors extracted through principal component analysis, and a series of new combined features generated through nonlinear combination and feature cross-operations. This enhanced feature set has higher dimensionality and richer feature representation. It not only retains the main statistical properties of the original data but also incorporates deep pattern information discovered through nonlinear transformations and cross-operations. This multimodal data fusion and feature enhancement strategy aims to overcome the limitations of information from a single data source, improve the representation ability of complex fault modes through fusion and interaction, and lay a more solid data foundation for subsequent modal decomposition and fault evolution analysis.

[0039] Example 4: Suppose an enhanced fault feature set over a period of time is obtained from a wind turbine gearbox monitoring system exhibiting early signs of bearing wear. This feature set contains multiple feature sequences after fusion and enhancement, such as features reflecting high-frequency vibration, features characterizing temperature trends, and features describing oil particle concentration. To clearly illustrate the data changes before and after modal decomposition, consider a simplified example. Referring to Table 1, it shows the original values ​​of a vibration-related enhanced feature sequence at five consecutive time points, the values ​​of the two main components obtained after empirical modal decomposition, and the calculated energy ratio.

[0040] Table 1: Empirical Mode Decomposition of Enhanced Feature Sequences Time point Original eigenvalues Fluctuation component (IMF1) Trend component (residual) Wave component energy Trend component energy Energy ratio t1 1.52 0.25 1.27 0.0625 1.6129 0.0387 t2 1.61 0.31 1.30 0.0961 1.6900 0.0569 t3 1.45 -0.18 1.63 0.0324 2.6569 0.0122 t4 1.78 0.22 1.56 0.0484 2.4336 0.0199 t5 1.69 0.15 1.54 0.0225 2.3716 0.0095 Empirical Mode Decomposition (EMD) is used to perform mode decomposition on the enhanced fault feature set. EMD is an adaptive signal processing method suitable for analyzing nonlinear and non-stationary signals. It can decompose complex signal sequences into a finite number of intrinsic mode functions (EMFs). The processing object is each feature sequence in the enhanced fault feature set. Taking the "original feature value" sequence in the table above as an example, EMD achieves decomposition through an iterative sieving process. First, all local extrema (maxima and minima) in the signal sequence are identified. Then, cubic spline interpolation is used to connect all maxima and minima to form the upper and lower envelopes of the signal. The mean of the upper and lower envelopes is calculated to obtain the first mean envelope. This mean envelope is subtracted from the original signal sequence to obtain an intermediate sequence. This intermediate sequence is checked to see if it satisfies the two conditions of the EMFs: the number of zero-crossing points in the entire sequence is equal to or at most differs from the number of extrema by one; and at any time, the mean of the upper envelope defined by the local maxima and the lower envelope defined by the local minima are both zero. If the intermediate sequence does not meet the conditions, it is treated as a new "original signal," and the above screening process is repeated until the conditions are met. The resulting sequence is the first intrinsic mode function (IMF1), which represents the highest frequency component in the signal. The first IMF is separated from the original signal, and the resulting residual signal is treated as a new signal. The entire screening process is repeated to extract the second IMF (IMF2), the third IMF (IMF3), and so on, until the residual signal becomes a monotonic sequence or a constant sequence, and no more IMFs can be extracted.

[0041] The first three intrinsic mode components (IMFs) after decomposition are extracted as fluctuation components. In gearbox fault analysis, the first few IMFs typically contain high-frequency components, impulsive components, and noise in the signal. These components are closely related to the instantaneous impact response caused by faults such as gear pitting and bearing cracks. The "Fluctuation Component (IMF1)" in Table 1 simulates the value of the first extracted IMF, fluctuating around zero and capturing high-frequency details and sudden changes in the original signal. The remaining components are superimposed to reconstruct the trend component. The remaining component includes all subsequent IMFs except the first three, as well as the final residual term. These components represent low-frequency trends and long-term drift components in the signal. Adding all these residual components at their corresponding time points reconstructs the signal's trend term. The "Trend Component (Residual)" in Table 1 is a simplified representation, assuming that after EMD decomposition, the residual itself directly represents the trend term, showing a relatively smooth and slow change, reflecting the gradual degradation process of equipment performance. The energy ratio of the fluctuation component to the trend component is calculated. Energy is an indicator of signal strength, and the sum of squares of the components is usually calculated as an energy estimate. As shown in Table 1, the square of the fluctuation component value at each time point is calculated as the "fluctuation component energy," and the square of the trend component value is calculated as the "trend component energy." Then, the ratio of the fluctuation component energy to the trend component energy is calculated to obtain the "energy ratio." This energy ratio provides a quantitative indicator to measure the significance of sudden activities relative to the slow evolution trend during the development of a fault. When the energy ratio continues to increase, it may indicate that the fault is transitioning from a slow accumulation stage to an accelerated development stage.

[0042] Based on the separated trend and fluctuation components, a fault evolution feature space is constructed using a deep neural network. The trend component is input into a Long Short-Term Memory (LSTM) network, a special type of recurrent neural network with cell states and gating mechanisms (input gate, forget gate, output gate), which can effectively learn and remember long-term dependencies in time series. The trend component sequence is sequentially input into an LSTM network. At each time step, the LSTM unit updates its cell state based on the current input and the hidden state of the previous time step. After sequence processing, the hidden state of the final time step, or the vector obtained by aggregating the hidden states of all time steps (e.g., average pooling), can be regarded as the long-term evolution feature extracted from the trend component. This feature vector encodes the long-term change pattern of the device's health status, such as the cumulative trend of wear.

[0043] The fluctuation components are input into a convolutional neural network (CNN). The CNN, using one-dimensional convolutional kernels to perform a sliding window operation on the time series, effectively extracts pattern features at local time scales. Each convolutional kernel acts as a feature detector, extracting local anomalous patterns from the fluctuation component sequence, such as short-duration impact pulses or the envelope shape of periodic impacts. Pooling layers (such as max pooling) are typically used after the convolutional layers to reduce data dimensionality and enhance the invariance of features to small time shifts. After several layers of convolution and pooling operations, the extracted local feature map is flattened into a feature vector. This vector serves as the local anomalous feature extracted from the fluctuation components, capturing transient details in the signal related to the suddenness of the fault.

[0044] Long-term evolutionary features and local anomaly features are concatenated. This concatenation operation is typically performed along the feature dimension, connecting the long-term evolutionary feature vector obtained from LSTM and the local anomaly feature vector obtained from CNN end-to-end to form a longer composite feature vector. This composite feature vector simultaneously contains macroscopic long-term trend information and microscopic local anomaly information of fault evolution, theoretically providing a more comprehensive description of the fault's development state. Dimensionality reduction is then performed on the concatenated features using an autoencoder. An autoencoder is an unsupervised neural network model consisting of an encoder and a decoder. The encoder maps the high-dimensional input features (i.e., the concatenated composite feature vector) to a low-dimensional latent space representation; this low-dimensional representation is the dimensionality-reduced fault evolution feature. The decoder then attempts to reconstruct the original high-dimensional input from this low-dimensional feature. The goal of training the autoencoder is to minimize the reconstruction error, thereby forcing the encoder to learn the most important and essential structural information from the input data. The dimensionality-reduced fault evolution feature space not only eliminates redundant information in the original composite features, but may also uncover the intrinsic laws hidden behind the original data that are related to the physical evolution of faults, providing a more concise and effective feature representation for subsequent time-series pattern matching.

[0045] Example 5: Assume the fault evolution feature space is a low-dimensional space after dimensionality reduction by an autoencoder. The equipment state at each time point can be represented by a point in this space, and the state points over a continuous period constitute a trajectory characterizing the evolution of the equipment's health status. The first step in time-series pattern matching is to establish a fault mode template library in this dimensionality-reduced feature space. The construction of the template library relies on historically accumulated fault data. For each known typical fault type, such as rolling bearing outer ring damage, gear tooth surface spalling, or shaft misalignment, the state evolution sequence in this feature space within a certain period before the fault occurs is extracted from historical cases. Each sequence undergoes data cleaning and alignment to ensure they have the same time length or are standardized using interpolation methods. These sequences constitute the template for the corresponding fault mode. A complete template library contains multiple such template sequences, each template being labeled with the fault type and information such as the typical development rate of the fault.

[0046] Calculating the dynamic time warping distance between the real-time feature sequence and each fault mode template is the core step in pattern matching. Dynamic time warping is an algorithm used to measure the similarity between two time series, effectively handling nonlinear scaling and phase differences along the time axis. Assuming a recent state evolution sequence of a device is monitored in real time, it's necessary to determine which fault mode in the template library it most closely resembles. The dynamic time warping algorithm works by constructing a cumulative distance matrix, where rows correspond to each point in the real-time sequence and columns correspond to each point in a specific fault template sequence. The algorithm calculates the local distance (usually using Euclidean distance) between each pair of points in the real-time and template sequences, then finds a curved path from the top left to the bottom right corner of the matrix that minimizes the sum of all local distances along this path. This minimum path sum is the dynamic time warping distance. The smaller this distance value, the more similar the real-time sequence is to the currently compared fault template in shape, even if their rates of change differ. The algorithm iterates through each fault template in the template library, calculating its dynamic time warping distance to the real-time sequence for each template. Based on a set of calculated dynamic time warped distance values, the fault mode that best matches the real-time sequence is determined. This process is comparative; the system selects the smallest distance from all calculated distance values, and the fault template corresponding to this distance is determined to be the best-matching mode for the current real-time sequence. For example, assuming the real-time sequence's distance to the "bearing outer ring damage" template is 5.2, its distance to the "gear wear" template is 8.7, and its distance to the "misalignment" template is 12.1, then "bearing outer ring damage" is determined to be the best-matching fault mode. This smallest distance value is recorded because it directly reflects the degree of matching.

[0047] Assessing the similarity between the current feature sequence and the best-matching fault mode is fundamental for early warning classification. Similarity assessment requires transforming the dynamic time-warped distance (DTW) into a more intuitive, fixed-range metric. Since DTW is an absolute value, its magnitude is affected by factors such as sequence length and feature scale, thus requiring normalization. A common approach is to use historical data to statistically determine a distance distribution range, mapping the minimum distance calculated in real-time to a similarity score between 0 and 1. The mapping function can be designed such that the smaller the distance, the closer the similarity score is to 1, indicating greater similarity; the larger the distance, the closer the similarity score is to 0, indicating greater difference. For example, an exponential decay function can be used for the transformation: similarity score = exp(-λ * DTW distance), where λ is a scaling parameter used to adjust the sensitivity of similarity to distance changes. This transformation yields a quantified similarity index that clearly indicates the degree of proximity between the current equipment operating state and a known fault mode precursor.

[0048] The system classifies warning levels based on calculated similarity scores, generating tiered warning signals. The threshold values ​​reflect the application scenario and risk tolerance. Multiple warning levels are typically used to differentiate the urgency of the fault risk; for example, three thresholds might correspond to three warning levels. The specific values ​​of these thresholds are determined through analysis of numerous historical fault cases, focusing on the changing patterns of similarity scores before a fault occurs. For instance, if analysis shows that a similarity score consistently above 0.85 indicates a very high probability of a corresponding fault occurring in the near future, the threshold for Level 1 warning is set at 0.85. When the real-time calculated similarity score exceeds this first threshold, the system generates a Level 1 warning signal. A Level 1 warning typically indicates an impending or early-stage fault, requiring immediate shutdown for inspection or repair. This signal may trigger the highest-level alarm notification. The threshold for Level 2 warnings is set as an intermediate value, such as 0.70. When the similarity score exceeds the second threshold of 0.70 but has not yet reached the first threshold of 0.85, the system generates a Level 2 warning signal. Level 2 warnings indicate that the equipment status shows a strong similarity to a certain failure mode, significantly increasing the risk of failure. This requires close attention from maintenance personnel, increased monitoring frequency, and the preparation of necessary maintenance plans. Level 3 warnings have a relatively low threshold, such as 0.50. When the similarity score exceeds the third threshold of 0.50 but is below the second threshold of 0.70, the system generates a Level 3 warning signal. Level 3 warnings indicate that the equipment may have early signs of anomalies, or that its operating status is beginning to deviate from a healthy baseline. More detailed data analysis and trend tracking are recommended, and the equipment should be included as a key focus of regular inspections. This tiered warning mechanism allows maintenance strategies to be matched with the level of failure risk, avoiding overreaction to minor anomalies while enabling rapid response to high-risk situations. This achieves optimized allocation of maintenance resources and effective early intervention in failures. The entire time-series pattern matching and warning generation process links abstract failure evolution characteristic sequences with specific failure modes and historical experience. Through quantitative similarity assessment and tiered warnings, it provides intuitive and actionable decision support for equipment condition-based maintenance.

[0049] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0050] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A machine learning-based intelligent fault early warning method for wind turbine gearboxes, characterized in that, include: Acquire multi-source monitoring data during the operation of the wind turbine gearbox, including vibration signals, temperature data, and oil analysis data; The multi-source monitoring data are subjected to time-frequency joint analysis to extract multi-scale operational features; Based on the aforementioned multi-scale operational characteristics, key fault-sensitive features are determined using an adaptive feature selection algorithm. Based on the historical failure case library and the key failure sensitivity features, a dynamic failure feature weight matrix is ​​constructed. The historical failure case library includes the characteristic patterns and weight assignments of past failure events; The construction of the dynamic fault feature weight matrix includes: The critical fault sensitivity features are classified according to the equipment's operating conditions. Establish feature weight allocation rules for each type of working condition; The feature weight allocation rule is dynamically adjusted based on real-time operating condition data. Perform matrix operations between the adjusted feature weights and the key fault-sensitive features. A multimodal data fusion method is used to enhance the dynamic fault feature weight matrix to generate an enhanced fault feature set. Modal decomposition is performed on the enhanced fault feature set to separate the trend component characterizing the fault development process and the fluctuation component characterizing the fault suddenness. Based on the trend component and fluctuation component, a fault evolution feature space is constructed using a deep neural network. The deep neural network has powerful nonlinear fitting and feature learning capabilities, and can further mine the deep features of fault evolution from trend components and fluctuation components, and construct a feature space that better reflects the essence of fault development. In the fault evolution feature space, a time-series pattern matching algorithm is used to identify fault development patterns; Based on the degree of matching between the fault development mode and the preset fault mode, a graded early warning signal is generated; The preset fault mode is a preset mode used for comparison and matching in the fault evolution feature space, and is composed of typical fault data in the historical fault case library.

2. The intelligent fault early warning method for wind turbine gearboxes based on machine learning according to claim 1, characterized in that, The time-frequency joint analysis of the multi-source monitoring data includes: Wavelet packet transform is performed on the vibration signal to extract the energy distribution characteristics of different frequency bands; Establish a time series model for temperature data and extract temperature change trend features; Spectral feature extraction was performed on oil analysis data to obtain the distribution characteristics of wear particles; The energy distribution characteristics, temperature change trend characteristics, and wear particle distribution characteristics are fused at the feature level.

3. The intelligent fault early warning method for wind turbine gearboxes based on machine learning according to claim 2, characterized in that, The determination of key fault-sensitive features through the adaptive feature selection algorithm includes: Calculate the correlation coefficients between each feature dimension and historical fault records; A ranking of feature importance is constructed based on the correlation coefficients; A sliding window mechanism is used to dynamically adjust the importance ranking of the features; Features with a predetermined proportion before sorting are selected as key fault-sensitive features.

4. The intelligent fault early warning method for wind turbine gearboxes based on machine learning according to claim 1, characterized in that, The feature enhancement of the dynamic fault feature weight matrix using a multimodal data fusion method includes: Principal component analysis is performed on the dynamic fault feature weight matrix to extract the main feature vectors after principal component analysis. The main feature vectors are then nonlinearly combined with the original features, and new combined features are generated through feature cross-operation.

5. The intelligent fault early warning method for wind turbine gearboxes based on machine learning according to claim 4, characterized in that, The modal decomposition of the enhanced fault feature set includes: The enhanced fault feature set is processed using the empirical mode decomposition method; The first three intrinsic mode components after decomposition are extracted as wave components; The remaining components are superimposed and reconstructed to obtain the trend components; Calculate the energy ratio of the fluctuation component to the trend component.

6. The intelligent fault early warning method for wind turbine gearboxes based on machine learning according to claim 5, characterized in that, The construction of the fault evolution feature space includes: The trend component is input into a long short-term memory network to extract long-term evolutionary features; The fluctuation components are input into a convolutional neural network to extract local anomaly features; The long-term evolutionary features and local anomaly features are then concatenated. The spliced ​​features are dimensionality reduced using an autoencoder.

7. The intelligent fault early warning method for wind turbine gearboxes based on machine learning according to claim 6, characterized in that, The method of using time-series pattern matching to identify fault development patterns includes: Establish a fault mode template library in the dimensionality-reduced feature space; Calculate the dynamic time warp distance between the real-time feature sequence and each fault mode template; The most suitable fault mode is determined based on the dynamic time warping distance; Evaluate the similarity between the current feature sequence and the best-matching fault mode.

8. The intelligent fault early warning method for wind turbine gearboxes based on machine learning according to claim 7, characterized in that, The generation of graded early warning signals includes: Warning level thresholds are determined based on the degree of similarity. A level one warning signal is generated when the similarity exceeds the first threshold. A secondary warning signal is generated when the similarity exceeds the second threshold but is lower than the first threshold. A level 3 warning signal is generated when the similarity exceeds the third threshold but is lower than the second threshold. The first threshold is the lowest quantile of the similarity score between when the fault is about to occur or when it is already in the early stages of occurrence; The second threshold is the lowest quantile of the similarity score when the device status has shown a strong similarity to a certain failure mode. The third threshold is the lowest quantile of the similarity score of the device when it is in a stage where there may be early signs of abnormality or its operating state begins to deviate from the healthy baseline state.

9. A machine learning-based intelligent fault early warning system for wind turbine gearboxes, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the intelligent fault early warning method for wind turbine gearboxes based on machine learning as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Gearbox fault diagnosis method and system

    CN111855202A

  • Hydroelectric generating set vibration trend prediction method and system

    CN112651290A

  • Distributed photovoltaic power generation power prediction method and device, equipment and storage medium

    CN117613850A

  • Photovoltaic power station regulation and operation optimization method, device, equipment and medium

    CN118748538A

  • Power distribution system reliability analysis method and system

    CN119813203A

Cited By

  • Gearbox health state stage identification method and system based on oil characteristics

    CN121350892A

  • Special vehicle fault diagnosis method and system based on wide near infrared spectrum analysis

    CN121364170A

  • Gear measurement center measurement process automation and intelligent scheduling system

    CN121535599A

  • Fan gear box wear evaluation method, device and equipment and storage medium

    CN121595197A

  • Fault detection method and system for concrete mixing device

    CN121677846A