Wind turbine generator health index construction and state monitoring method based on sparse dictionary distance

By constructing a health index for wind turbines using a sparse dictionary distance method, the problems of insufficient robustness and physical mechanism in existing technologies are solved, enabling sensitive monitoring and accurate early warning of early faults in wind turbines.

CN121897528APending Publication Date: 2026-04-21SOUTH CHINA UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTH CHINA UNIV OF TECH
Filing Date
2026-02-02
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing methods for constructing health indicators for wind turbines lack robustness and generalization ability when facing complex operating conditions and noise interference, and lack clear physical mechanism support, making it difficult to identify faults in the early stages.

Method used

The sparse dictionary distance method is adopted to construct health indicators through the theory of sparse signal representation. It combines sparse coding and dictionary learning algorithms to learn the baseline dictionary of the device health status and calculate the sparse dictionary distance to monitor the device status. It has good physical mechanism support and robustness.

Benefits of technology

It achieves keen detection of early faults in complex operating conditions and noisy environments, has earlier fault warning capabilities, is applicable to different equipment and operating conditions, and has good interpretability and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121897528A_ABST
    Figure CN121897528A_ABST
Patent Text Reader

Abstract

The invention discloses a wind turbine generator health index construction and state monitoring method based on sparse dictionary distance. The method comprises the following steps: (1) collecting an original vibration acceleration signal of a wind turbine generator in a health state; (2) collecting a signal training sample, and combining a sparse coding algorithm and a dictionary learning algorithm to construct a reference health dictionary representing a healthy operation mode of the equipment; (3) preprocessing the real-time vibration signal, collecting a signal training sample, and obtaining a real-time feature dictionary in combination with a sparse coding algorithm and a dictionary learning algorithm; (4) calculating a sparse dictionary distance of the reference health dictionary and the real-time feature dictionary based on Euclidean distance measurement; (5) obtaining all sparse dictionary distance calculation results and constructing health indexes; and (6) a dynamic fault threshold value is set based on the statistical characteristics of the health stage, and the operation state of the wind turbine generator is evaluated. According to the invention, early faults in the equipment can be found earlier, accurate state monitoring of the wind turbine generator is realized, and earlier fault early warning capability is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of intelligent fault diagnosis and health monitoring of wind turbines, specifically involving a method for constructing health indicators and monitoring the status of wind turbines based on sparse dictionary distance. Background Technology

[0002] With the development of modern industry and the growing aspiration for environmental protection, people's demand for green energy and sustainable development is increasing. Wind power generation is an important green energy source, and wind turbines play a crucial role in wind power generation. However, wind turbine equipment operates under complex and variable conditions and loads for extended periods, making it highly susceptible to wear, fatigue fractures, and other failures. Failure to detect and address these failures promptly can lead to shutdowns affecting production, or even serious safety accidents. Maintaining large mechanical equipment like wind turbines typically requires significant time and financial resources. Therefore, effective health monitoring and fault diagnosis of wind turbine equipment in advance, along with the development of maintenance plans, are of significant engineering importance.

[0003] Main drivetrain failure is one of the most frequent types of failure in wind turbines, and these failures primarily occur in critical components such as gearboxes within the main drivetrain. Currently, condition monitoring of wind turbine equipment mainly relies on two methods: physical model-based and data-driven approaches. Physical model-based methods, based on the principles of solid mechanics, materials fatigue, and aerodynamics, utilize load simulation and the construction of cumulative fatigue damage models to quantify the cumulative fatigue damage of the equipment, thereby enabling condition monitoring and remaining service life estimation. However, this type of method requires remodeling for different wind turbine models, making the modeling process relatively complex and time-consuming. It is more suitable for life estimation of components such as the turbine tower and blades.

[0004] The basic principle of data-driven methods is not to delve into the internal physical mechanisms, but to use machine learning methods or big data analysis techniques to extract performance degradation-related feature patterns from long-term accumulated equipment operation status monitoring data. Data-driven methods are usually more suitable for components with complex failure mechanisms and degradation processes that can be effectively characterized by condition monitoring data. These methods are more suitable for main drive trains with complex failure mechanisms.

[0005] Constructing accurate health indicators that characterize the operating status of mechanical equipment is crucial for data-driven approaches. Traditional methods for constructing health indicators typically involve calculating the time-domain statistical characteristics of signals (such as root mean square values ​​and kurtosis) or using machine learning methods. However, existing methods have certain limitations: 1. Insufficient robustness: Traditional health indicators may experience significant fluctuations and distortions when affected by changes in equipment speed and operating conditions. Furthermore, actual vibration signals often contain substantial environmental noise, easily masking early, subtle fault characteristics with noise and harmonic components, making it difficult for the indicators to accurately reflect the true operating status of the equipment. 2. Insufficient generalization ability: Traditional health indicators exhibit poor consistency across different datasets. Performance often degrades significantly when training and test data differ in distribution, equipment type, or operating conditions, limiting cross-scenario applicability. 3. Insufficient interpretability: Health indicators constructed based on statistical features and machine learning methods often lack clear physical mechanisms and cannot intuitively demonstrate changes in specific components of the signal, making the analysis results difficult to understand and trust. Therefore, proposing a health index that is robust, applicable to different equipment and operating conditions, and supported by a sound physical mechanism is of great significance and importance for equipment condition monitoring and maintenance. Summary of the Invention

[0006] The main objective of this invention is to overcome the shortcomings and deficiencies of existing technologies and provide a method for constructing and monitoring the health indicators of wind turbine generators based on sparse dictionary distance. Combining signal sparse representation theory, a sparse dictionary distance (SDD) health indicator is proposed. The method first eliminates signal amplitude fluctuations caused by varying operating conditions through standardized preprocessing, focusing on the time-domain waveform structure of vibration signals. In the baseline dictionary construction stage, a baseline health dictionary representing the equipment's health state is learned using a combination of sparse coding and dictionary learning algorithms, capturing the operating mode under the equipment's health state. In the monitoring stage, a real-time feature dictionary representing the current operating state is calculated using the same method as in the training stage. Then, an atomic waveform difference measure based on Euclidean distance is calculated between the baseline health dictionary and the real-time feature dictionary. Based on this difference measure, a health indicator based on sparse dictionary distance is constructed and used for wind turbine generator condition monitoring. The constructed health indicator is highly robust to speed load fluctuations and noise, has more explicit physical meaning, and can sensitively capture early impact characteristics, enabling earlier detection of early faults in the equipment, achieving accurate condition monitoring of wind turbine generators, and providing earlier fault warning capabilities.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] This invention provides a method for constructing and monitoring the health indicators and status of wind turbines based on sparse dictionary distances, comprising the following steps:

[0009] S1. Collect the raw vibration acceleration signal of the wind turbine in a healthy state, and preprocess the collected raw vibration acceleration signal.

[0010] S2. Select the preprocessed vibration signal as the training set, extract signal samples from the training set to construct the training sample matrix, construct the initial dictionary through the training sample matrix, and train the benchmark health dictionary that can characterize the waveform features of the equipment health mode by combining the sparse coding algorithm and the dictionary learning algorithm.

[0011] S3. During the equipment monitoring phase, the vibration signal of the equipment at the current moment is acquired. Based on the vibration signal at the current moment, a training sample matrix and an initial dictionary for the monitoring phase are constructed. The real-time feature dictionary representing the current operating state is obtained by combining sparse coding algorithm and dictionary learning algorithm.

[0012] S4. Calculate the difference measure between the real-time feature dictionary obtained at each time moment and the benchmark health dictionary, and obtain the sparse dictionary distance for each time period based on the difference measure.

[0013] S5. The sparse dictionary distance is used to form a time series, and the time series is smoothed and used as a health indicator of the wind turbine.

[0014] S6. Calculate and determine the equipment fault threshold based on the changing trend of the sparse dictionary distance health index when the equipment is in healthy operation; compare the real-time sparse dictionary distance health index with the fault monitoring threshold to evaluate the operating status of the wind turbine; and issue an alarm when the amplitude of the sparse dictionary distance of the equipment continuously exceeds the threshold.

[0015] As a preferred technical solution, in step S1, the preprocessing includes signal filtering and normalization to obtain normalized signal samples; the normalization is specifically Z-score standardization, and the calculation formula is:

[0016] ;

[0017] in, This is the original signal segment after filtering. The mean of the original signal segment. The standard deviation of the original signal segment is used; through the Z-score standardization process, the difference in the amplitude of the vibration signal under different speed conditions is eliminated, so that the subsequent dictionary learning process focuses on the time-domain waveform structure characteristics of the signal.

[0018] As a preferred technical solution, in step S2, a training sample construction strategy based on peak search is adopted, specifically as follows:

[0019] Perform a full-segment search on the amplitude of the preprocessed vibration signal;

[0020] Set minimum peak distance and minimum peak height thresholds, sort the searched peak points according to peak size, and select the top N significant peaks;

[0021] Using each peak point as the starting reference point or center point, a data segment of fixed length L is extracted to construct a training sample matrix for dictionary learning.

[0022] As a preferred technical solution, in step S2, constructing the initial dictionary through the training sample matrix specifically involves:

[0023] The training sample matrix is ​​subjected to mean-neutralization, and its covariance matrix is ​​subjected to singular value decomposition.

[0024] The first K principal component feature vectors after singular value decomposition are selected as atoms of the initial dictionary, where K is the dictionary size parameter set before the initial training of the K-SVD dictionary learning algorithm, so as to obtain an initial dictionary containing the main vibration waveform structure of the signal.

[0025] As a preferred technical solution, in step S2, the benchmark health dictionary that can characterize the waveform features of the device health mode is trained by combining the sparse coding algorithm and the dictionary learning algorithm. Specifically, the following iterative update steps are performed using the initial dictionary constructed in step S2 as the initial value:

[0026] In the sparse coding stage, the dictionary atoms in the current iteration step are fixed, and the orthogonal matching pursuit algorithm is used to solve the sparse coefficient matrix of the training samples under the current dictionary.

[0027] During the dictionary learning and update phase, the atoms in the dictionary and their corresponding non-zero coefficients are updated column by column; for the j-th atom in the dictionary, the error matrix after removing the contribution of that atom is calculated. Only retain The columns associated with this atom constitute the shrinkage error matrix. ;right Perform singular value decomposition, select the left singular vector corresponding to the maximum singular value as the updated j-th atom, and select the product of the maximum singular value and the right singular vector as the updated coefficient;

[0028] Repeat the above sparse coding stage and dictionary learning and update stage until the set number of iterations is reached, and determine the final iteratively updated dictionary as the baseline healthy dictionary.

[0029] As a preferred technical solution, in step S4, the sparse dictionary distance is calculated using a bidirectional mapping strategy, as follows:

[0030] S41, Forward mapping calculation; based on the baseline health dictionary Each atom in In the real-time feature dictionary Search for the most similar atom in the middle and calculate the one-way mapping distance. And sum them to get the forward mapping distance;

[0031] S42, Backward mapping calculation; for real-time feature dictionary Each atom in In the benchmark health dictionary Search for the most similar atom in the middle and calculate the one-way distance. And sum them to get the backward mapping distance;

[0032] S43. Sparse dictionary distance calculation: Calculate the average of the forward mapping distance and the backward mapping distance as the final sparse dictionary distance; the calculation formula is as follows:

[0033]

[0034] in, The number of atoms in the dictionary. As a benchmark health dictionary, For real-time feature dictionary; Represents a single atom To the target dictionary One-way mapping distance;

[0035] in, The forward mapping distance represents the deviation of the signal waveform shape from the best-matching atom in the baseline health dictionary to the real-time feature dictionary. The backward mapping distance represents the deviation of the signal waveform shape from the best-matching atom in the benchmark health dictionary to the real-time feature dictionary.

[0036] As a preferred technical solution, the one-way mapping distance The Euclidean distance metric, based on atomic morphological similarity, is defined as:

[0037]

[0038] in, Dictionary atoms in It represents the cosine similarity between atoms; this formula quantifies the degree of deviation in waveform shape between atoms.

[0039] As a preferred technical solution, in step S5, the method for determining the equipment fault threshold specifically includes:

[0040] A baseline distance set is constructed using several sparse dictionary distance data points from the baseline health dictionary training phase or the initial stage of device health operation; the average value of the set is then calculated. and standard deviation Set fault monitoring thresholds for:

[0041]

[0042] When the real-time calculated health indicators exceed this threshold When this happens, it is determined that the equipment's operating status is abnormal.

[0043] As a preferred technical solution, the smoothing operation in step S5 specifically adopts the sliding window filtering method to focus on the statistical characteristics of the indicators over a period of time and eliminate random fluctuations in the indicators caused by sudden changes in equipment operating conditions.

[0044] As a preferred technical solution, the health indicator construction method is used for predictive maintenance management of industrial equipment, specifically including the detection of operating status and early fault warning of key rotating components such as the main bearing and gearbox of wind turbine generator sets and industrial bearing housings.

[0045] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0046] 1. This invention constructs a health index and condition monitoring method for wind turbines based on sparse dictionary distance. It combines the theory of sparse signal representation and has good physical mechanism support and interpretability compared with the health index constructed by traditional statistical features or machine learning methods. It can observe the atomic changes of the signal feature dictionary in real time.

[0047] 2. The sparse dictionary distance health index constructed in this invention has good generalization ability and is applicable to vibration signal operation data of different devices. Traditional health indices show significant differences in performance across different datasets, while the proposed method demonstrates better stability and overall performance across different fault types and different device datasets.

[0048] 3. The sparse dictionary distance health index constructed in this invention has good robustness and is suitable for the complex changing operating conditions of wind turbine units and the measured signal dataset with high noise levels. Traditional statistical features will show large fluctuations and oscillations in amplitude when the speed operating conditions change greatly and the noise is high. The proposed method is not sensitive to speed changes and noise interference, can adapt to more operating conditions, and has earlier fault warning capabilities, and can identify early and weak fault changes in wind turbine equipment.

[0049] 4. The unsupervised sparse dictionary-based equipment status monitoring method constructed in this invention simplifies the equipment health operation mode, does not rely on complex deep learning models or physical models, and represents the vibration signal patterns of the equipment through an intuitive and interpretable feature dictionary. The training process does not rely on a large amount of fault signal data; only a small amount of health operation data is needed to construct a stable benchmark health dictionary. This solves the problem that in reality, large equipment has limited historical operating data and a lack of fault operation data, making it difficult to establish a benchmark model for the health operation status. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 This is an overall flowchart of the health indicator construction method of the present invention.

[0052] Figure 2 This is a detailed flowchart of the method of the present invention.

[0053] Figure 3 This is a schematic diagram of the time-domain variation and anomaly threshold of the method of the present invention in a certain wind farm dataset.

[0054] Figure 4 This is a schematic diagram showing the temporal variation and anomaly threshold of wavelet packet entropy statistical health indicators in a certain wind farm dataset.

[0055] Figure 5 This is a schematic diagram showing the temporal variation and anomaly threshold of kurtosis statistical health indicators in a certain wind farm dataset.

[0056] Figure 6 This is a schematic diagram showing the time-domain variation and anomaly threshold of the root mean square value statistical health indicator in a certain wind farm dataset. Detailed Implementation

[0057] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative effort are within the scope of protection of the present application.

[0058] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.

[0059] The invention was verified using CMS vibration acceleration signal data of the high-speed shaft radial direction of the gearbox of a wind power plant generator set. The following is a detailed description of the invention in conjunction with the accompanying drawings and specific embodiments.

[0060] Experimental Data Description: This embodiment uses CMS vibration acceleration signal data from the gearbox of a wind turbine generator set for verification. The operational data covers the period from December 2019 to June 2025. During a routine inspection in August 2024, the high-speed shaft bearing of this wind turbine was found to have circumferential scratches on the rollers and raceway spalling. This embodiment will focus on verifying the accuracy of the proposed method in characterizing the condition of wind turbine equipment and its early fault warning capability.

[0061] For the dataset described above, this embodiment sets the following key step parameters:

[0062] Signal preprocessing steps: Screen out wind turbine shutdown and abnormal operation data, and set the filter frequency range to [0, 1500] Hz. The frequency band covers the main mechanical fault characteristic frequencies and their harmonic components, and filters out the interference caused by the high-frequency resonance band; screen out data during abnormal sensor periods. Considering the variable speed and load conditions of wind turbines, vibration signal data is preferred within the speed range of [1700, 1800] rpm to avoid interference caused by wind turbine shutdown and excessive speed differences.

[0063] K-SVD algorithm dictionary learning parameters: set dictionary atom length N=60 (dataset sampling rate is 5120Hz, corresponding to a time length of about 0.012 seconds, the dictionary atom length is based on the fact that it should be slightly longer than the duration of the impact signal, depending on the specific dataset), dictionary size (number of atoms) K=6, sparsity L=2, K-SVD iterations are 10;

[0064] Peak search parameters for training samples: In order to extract representative impact components, the minimum peak distance is set to 150 points, the minimum peak height coefficient is set to 0.1, and 20 peak signal sample segments are selected for each signal segment;

[0065] like Figure 1 , Figure 2As shown, the specific process of the wind turbine health index construction and condition monitoring method based on sparse dictionary distance of the present invention is as follows:

[0066] S1. Data Acquisition and Preprocessing: Vibration acceleration signals from the wind turbine generator under healthy operating conditions are acquired. Addressing the significant differences in signal amplitude caused by varying speeds and loads in industrial equipment, and the substantial interference from high-frequency resonance components in the spectrum to feature extraction, this invention first performs filtering preprocessing on the filtered signal, followed by normalization preprocessing. By converting the signal to a zero-mean, unit-variance form, the interference of signal amplitude fluctuations caused by changes in speed and operating conditions on feature extraction is eliminated, allowing subsequent processing to focus on the signal waveform characteristics.

[0067] Furthermore, this invention employs the Z-score standardization (zero-mean normalization) method, with the specific calculation formula as follows:

[0068]

[0069] in, This is the original signal segment. This is the mean of the segment. This represents the standard deviation of the segment. This step maps all signal segments to the same dimension, forcing the subsequent dictionary learning process to focus on the time-domain waveform structure characteristics of the signal (such as the waveform of an impulse) rather than the magnitude of the signal energy and amplitude, thereby significantly improving the robustness of health indicators under varying operating conditions.

[0070] For example, the acquired raw signal is segmented, with each segment having a length of 8192 sampling points, and the following preprocessing operations are performed on all signal segments:

[0071] S1.1 Low-pass filtering: Use a filter of [0, 1500] Hz to filter out high-frequency resonance band interference;

[0072] S1.2, Z-score normalization: The filtered signal is normalized to zero mean. This step is crucial because it eliminates the influence of signal energy magnitude, allowing subsequent dictionary learning to focus on waveform shape rather than amplitude.

[0073] S2. Constructing a baseline health dictionary: Vibration signals under equipment health conditions are selected as the training set; to extract the most representative local impact or harmonic features, a peak search-based sample construction strategy is adopted; the dictionary is initialized using singular value decomposition (SVD) and trained using sparse coding and dictionary learning algorithms. In a preferred embodiment, orthogonal matching pursuit (OMP) combined with the K-SVD algorithm can be used for training. After training, a baseline health dictionary that can accurately characterize the equipment health mode is obtained, and the baseline health dictionary will remain unchanged in all subsequent monitoring processes.

[0074] Furthermore, the training sample construction strategy employs a peak search-based sample extraction method. Specifically, it involves: performing a full-segment search on the preprocessed signal amplitude, setting minimum peak distance and minimum peak height thresholds to filter out significant signal peak points; and constructing a training sample matrix by truncating a fixed-length data segment (the same as the set dictionary atom size parameter) centered on each peak point. The purpose of this sample construction strategy is to ensure that the constructed training samples maintain as much similarity and consistency as possible in the time domain, aligning the phases of the signal waveforms as much as possible. This allows the dictionary learning algorithm to prioritize capturing the most concentrated energy and most fault-reflecting impact signal components in the signal, avoiding excessive phase differences in the training signal samples due to random window truncation.

[0075] Furthermore, to improve the convergence speed of dictionary learning and avoid getting trapped in local optima, the training sample matrix is ​​mean-centered during dictionary initialization, and its covariance matrix is ​​subjected to singular value decomposition (SVD). The left singular vectors corresponding to the top K largest singular values ​​are selected as initial dictionary atoms. This method can extract the principal component structure in the signal, providing a high-quality initial training dictionary for subsequent iterations of the K-SVD dictionary learning algorithm.

[0076] Furthermore, the sparse coding algorithm and dictionary learning algorithm are specifically the orthogonal matching pursuit algorithm and the K-SVD dictionary learning algorithm, using the initialized dictionary atoms as initial values. The specific steps include the following two stages:

[0077] In the sparse coding stage, the Orthogonal Matching Pursuit (OMP) algorithm is used to solve for the sparse coefficients. Compared with other sparse coding algorithms, the OMP algorithm ensures that the residuals are orthogonal to the selected atoms through orthogonal projection, resulting in faster convergence speed, greater maturity and stability, and better suitability for engineering applications.

[0078] During the dictionary update phase, the K-SVD dictionary learning algorithm is used: the atoms in the dictionary and their corresponding non-zero coefficients are updated column by column; for the j-th atom in the dictionary, the error matrix after removing the contribution of that atom is calculated. Only retain The columns associated with this atom constitute the shrinkage error matrix. ;right Perform Singular Value Decomposition (SVD), select the left singular vector corresponding to the maximum singular value as the updated j-th atom, and select the product of the maximum singular value and the right singular vector as the updated coefficient; repeat the above two stages until the set number of iterations is reached, and determine the final iteratively updated dictionary as the baseline healthy dictionary.

[0079] For example, a baseline model can be built using data from the initial stage of device health status:

[0080] S2.1 Sample Construction: After signal preprocessing, a signal segment covering various speed ranges (1710, 1720...1790, 1800) within the [1700, 1800] rpm speed range under the initial healthy state is selected to allow the benchmark health dictionary to adapt to the vibration signal pattern under different speeds. Peak search is performed, and a batch of data segments with a length of 60 is extracted as training samples with significant peak points as the starting point. A total of about 80 high-quality training sample segments are collected.

[0081] S2.2 Dictionary Initialization: To avoid local optima caused by random initialization, this embodiment uses Singular Value Decomposition (SVD) for initialization. After removing the mean from all training samples, a sample matrix is ​​constructed. SVD decomposition is then performed on this matrix, and the first K columns of the left singular matrix U (which are the same size as the initially set dictionary) are taken as the initial dictionary atoms. Each atom is then normalized using the L2 norm.

[0082] S2.3 Dictionary Training: Input the sample set and initial dictionary into the K-SVD algorithm. The algorithm iterates 10 times, alternating between sparse encoding and dictionary updates, and finally outputs a baseline health dictionary that optimally represents the health signal features in a sparse manner. The baseline health dictionary is maintained and remains unchanged during subsequent monitoring phases.

[0083] S3. Calculate the real-time feature dictionary: During the online monitoring phase of the equipment, for each new input signal segment, perform the same operations as steps S1 and S2: perform the same filtering and normalization preprocessing operations, and extract the impulse samples in the signal segment based on the same peak search method. Using the same initialization method and K-SVD parameters, train to obtain a real-time feature dictionary representing the current operating state. At this point, if an early failure occurs in the device, the atomic waveform structure in the real-time dictionary will change.

[0084] S4. Calculate the sparse dictionary distance: Calculate the difference between the real-time feature dictionary and the baseline health dictionary. To address the computational error caused by the asymmetry of the dictionary space in traditional unidirectional distance calculations, this invention employs a bidirectional mapping strategy: First, calculate the forward mapping distance, i.e., find the best matching atom of the baseline dictionary atom in the real-time dictionary and calculate its signal waveform shape deviation; then calculate the backward mapping distance, i.e., find the best matching atom of the real-time dictionary atom in the baseline dictionary and calculate its signal waveform shape deviation; finally, take the average of the bidirectional mapping distances as the sparse dictionary distance. The unidirectional mapping distance is calculated based on normalized Euclidean distance, directly quantifying the degree of deviation of the vibration signal waveform shape between the real-time feature dictionary atom and the health dictionary atom. This bidirectional averaging calculation strategy aims to eliminate the error in unidirectional sparse dictionary distance calculations. Specific steps include:

[0085] Step S41: Forward mapping calculation; based on the baseline health dictionary Each atom in In the real-time feature dictionary Search for the most similar atom in the middle and calculate the one-way mapping distance. And sum them to get the forward mapping distance;

[0086] Step S42: Backward mapping calculation; for real-time feature dictionary Each atom in In the benchmark health dictionary Search for the most similar atom in the middle and calculate the one-way distance. And sum them to get the backward mapping distance;

[0087] Step S43: Sparse dictionary distance calculation; such as As shown in the calculation formula, the average of the forward mapping distance and the backward mapping distance is calculated as the final sparse dictionary distance.

[0088] Furthermore, in the sparse dictionary distance calculation step described in step S4, the distance between individual atoms... Using Euclidean distance The Euclidean distance, a metric for distance calculation, is defined as follows:

[0089] ;

[0090] in, Dictionary In the signal samples, after normalization preprocessing, all atom vectors are processed into unit vectors, then:

[0091] ;

[0092] in, The formula can be simplified to:

[0093] ;

[0094] Among them, inner product In the unit vector context, this is cosine similarity, which quantifies the degree of deviation in waveform shape between atoms. The formula for calculating cosine similarity is:

[0095] ;

[0096] Euclidean distance measures the degree of difference between two vectors. The smaller the Euclidean distance, the more similar the two vectors are. Sparse dictionary distance actually finds the atom in the target dictionary that is most similar to the current atom and calculates its deviation. The calculation formula is defined as follows:

[0097] ;

[0098] A larger dictionary distance metric indicates a greater difference in vibration modes between the atom and the target dictionary, while a smaller dictionary distance metric indicates a smaller difference in vibration modes. Changes in the dictionary distance metric accurately reflect the difference between the device's current operating state and its healthy operating state. Compared to the original cosine similarity metric or the cosine similarity metric of angle values, dictionary distance preserves the linear characteristics of Euclidean space and is more sensitive to minor waveform distortions caused by early, weak faults.

[0099] S5. Construct a sparse dictionary distance health index: The sparse dictionary distances obtained from continuous monitoring are used to form a time series, and a sliding window filter is used for smoothing to eliminate random fluctuations. The focus is on the changes of the equipment over a period of time as the health index of the wind turbine.

[0100] Furthermore, the smoothing method described in step S5 is preferably a window moving average. The window moving average method has a clear time significance. The purpose of performing sliding window smoothing on the real-time indicator sequence is to eliminate random noise interference, so that the health indicators focus on the changing trend of the equipment over a longer period of time, and avoid signal amplitude interference caused by changes in operating conditions in a short period of time.

[0101] For example, the construction of a sparse dictionary distance health index is as follows:

[0102] S5.1 Arrange the sparse dictionary distances calculated for each time period in chronological order to form a health indicator sequence.

[0103] S5.2 To eliminate single-point noise and fluctuations caused by speed changes, a moving window (with a window length of 25, roughly corresponding to one week of data in the dataset) is used to smooth the index sequence, focusing on the changing trend of the wind turbine dictionary distance index during this period.

[0104] S6. Determine fault thresholds and abnormal alarms: Calculate and determine equipment fault thresholds based on the changing trend of sparse dictionary distance health indicators when the equipment is in healthy operating condition; compare the real-time sparse dictionary distance health indicators with the fault monitoring thresholds to evaluate the operating status of the wind turbine; when the amplitude of the sparse dictionary distance during equipment operation exceeds the threshold, an alarm is issued.

[0105] Furthermore, the fault threshold determination method described in step S6 is as follows: construct a benchmark set using a series of dictionary distance data from the initial stage of equipment healthy operation (or training phase), and calculate its average value. and standard deviation Set the fault warning threshold as follows: (or other statistical methods). When real-time health indicators continuously exceed this threshold, the equipment is deemed to be in an abnormal operating state.

[0106] Based on the obtained time-domain variation graph of health indicators, three times the standard deviation of the mean of health indicators during the healthy operation phase (using the statistical 3σ method) is selected as the fault threshold. A fault warning is issued when the real-time sparse dictionary health indicator value exceeds the fault threshold twice consecutively.

[0107] Different health indicators may perform differently on different datasets. To quantify the effectiveness and performance of health indicators on a dataset, monotonicity, trend and robustness indicators are usually used for evaluation.

[0108] The commonly used evaluation formulas for the three dimensions of monotonicity, trend, and robustness are shown below.

[0109] ;

[0110] ;

[0111] Equation 1 is the formula for calculating the monotonicity evaluation index. Let represent the health indicator value at time t, M be the total length of the health indicator sequence, and n be the order of the difference. The function is an indicator function; the function value is 1 when the condition within the parentheses is true, and 0 otherwise. As shown in Equation 2, the final monotonicity score in this embodiment uses the average of the first, second, and third order monotonicity scores. A better monotonicity score indicates stronger monotonicity of the health indicator.

[0112] ;

[0113] As shown in the Corr calculation formula above, in the trend formula, M is the total length of the health indicator series. Let t be the health indicator value at time t, where t is the time index (t=1, 2, M); M is the total length of the health indicator sequence. The trend formula 3 reflects the degree of linear correlation between the health indicator and the changes over time. The closer the trend score is to 1, the stronger the linear correlation between the health indicator and time, and the better it can reflect the continuous degradation process of the equipment.

[0114]

[0115] As shown in the Rob calculation formula above, M in the robustness formula is the total length of the health indicator sequence. This represents the residual at the corresponding time point. The value represents the health indicator after smoothing. The residual is calculated by subtracting the smoothed value from the original data value. The closer the robustness score is to 1, the stronger the indicator's ability to resist interference.

[0116] Under different fault types and operating conditions, quantitative evaluation indicators are needed to select the best-performing indicators that can accurately reflect the gear degradation state, so as to find the fault initiation point in the early stage of degradation and conduct more accurate prediction and analysis of remaining service life.

[0117] Equipment degradation typically increases with operating time. Trend indicators reflect the correlation between health indicators and equipment operating time; the better the trend of a health indicator, the more it reflects the equipment's degradation state. Different weights are assigned to the three indicators in this embodiment. Trend is given the highest weight, monotonicity the second highest, and robustness the lowest, with values ​​set to [values ​​to be filled in]. After comprehensively analyzing the monotonicity, trend, and robustness of multiple statistical features, the final score for each feature is obtained, calculated using the following formula:

[0118]

[0119] Based on the above evaluation index formula, the effects of various statistical health indicators are shown in Table 1.

[0120] Table 1. Comprehensive score of various health indicators in the wind power data set.

[0121] Indicator Name Trend Monotonicity robustness Overall score Root mean square value 0.3051 0.0028 0.8671 0.2706 cliff 0.5943 0.0081 0.6504 0.4241 Peak-to-peak value 0.1301 0.0046 0.9058 0.17 Skewness 0.6359 0.0031 0.582 0.4407 Waveform entropy 0.5972 0.0028 0.9693 0.4561 Wavelet packet entropy 0.6003 0.0116 0.9855 0.4623 Dictionary distance 0.8424 0.0108 0.999 0.6085

[0122] like Figure 3 As shown, in the early stage of monitoring, the sparse dictionary distance fluctuated at a relatively low level, indicating that the vibration mode of the real-time feature dictionary was basically consistent with that of the baseline health dictionary. As the wind turbine gearbox degraded, the sparse dictionary distance showed a significant upward trend and provided a correct early warning of the wind turbine fault after the fault occurred.

[0123] Overall performance score comparison: As shown in Table 1, compared with other traditional health indicators, the most important trend score of sparse dictionary distance is much higher, and its overall score is the highest, indicating that its overall performance is relatively better. Root mean square value and kurtosis are commonly used traditional health indicators. The overall score of root mean square value is lower, indicating that its performance in the wind turbine dataset is relatively poor, while the overall scores of commonly used kurtosis and wavelet packet entropy are better.

[0124] Comparison of time-domain trend graphs: such as Figure 3 As shown, the amplitude of the dictionary distance health index remained relatively stable before the fault occurred, and remained at a high level after the fault occurred, rarely falling below the previously set fault threshold, which is closest to the actual condition of the wind turbine; similarly, as Figure 4 As shown, the wavelet packet entropy health index does not fluctuate significantly when the wind turbine is in a healthy state, but it also fluctuates significantly after a fault occurs, deviating from the actual state of the wind turbine; for example... Figure 5 Although the kurtosis health index shown also changed significantly after the fault occurred, it still fluctuated considerably when the wind turbine was healthy. Furthermore, around February 2025, after the fault, the index value experienced a sudden drop, remaining below the previously set fault threshold for an extended period, deviating from the actual state of the wind turbine and affecting the accuracy of condition monitoring. Under varying operating conditions of the wind turbine, such as… Figure 6 As shown, the root mean square health index is completely distorted, and the overall time domain change graph does not show an obvious degradation trend, which is consistent with the trend scoring in Table 1. Moreover, after the actual fault occurred, it was far below the set threshold, which is significantly different from the actual equipment degradation state.

[0125] Comparison of Fault Warning Time: In addition to its good overall score, the dictionary distance indicator showed a significant upward trend starting in February 2024, six months earlier than the fault report date, prior to the fault report in August 2024. The first anomalies in kurtosis and wavelet packet entropy occurred in September and October 2024, respectively, later than the fault report date. Although the root mean square (RMS) value had an earlier fault report date, the overall trend of the RMS value was completely distorted, rendering the warning date meaningless. Therefore, the dictionary distance indicator provides earlier fault warnings than traditional health indicators, which is of great significance for equipment condition monitoring and maintenance.

[0126] Therefore, dictionary distance has better stability and overall performance compared to traditional health indicators, and can accurately reflect the actual operating status of the wind turbine.

[0127] Compared to the vibration signal data of degrading rotating machinery in the laboratory, the measured operating data of wind turbines exhibits more complex variations in speed and operating conditions. Experimental results from the comprehensive implementation examples show that different traditional health indicators have varying adaptability to the wind turbine operating data set. Traditional statistical features perform relatively poorly in the measured wind turbine operating data set; the most frequently used root mean square value even shows distortion, losing its ability to characterize the wind turbine's degradation state. While health indicators such as kurtosis and wavelet packet entropy are not completely distorted, they may still produce false alarms. In contrast, the sparse dictionary distance health indicator demonstrates good stability and robustness in the variable operating condition data of wind turbines, as well as earlier warning capabilities. Therefore, the sparse dictionary distance health indicator has significant engineering implications for equipment maintenance and condition monitoring.

[0128] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0129] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0130] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.

Claims

1. A method for constructing health indicators and monitoring the condition of wind turbine generators based on sparse dictionary distance, characterized in that, Includes the following steps: S1. Collect the raw vibration acceleration signal of the wind turbine in a healthy state, and preprocess the collected raw vibration acceleration signal. S2. Select the preprocessed vibration signal as the training set, extract signal samples from the training set to construct the training sample matrix, construct the initial dictionary through the training sample matrix, and train the benchmark health dictionary that can characterize the waveform features of the equipment health mode by combining the sparse coding algorithm and the dictionary learning algorithm. S3. During the equipment monitoring phase, the vibration signal of the equipment at the current moment is acquired. Based on the vibration signal at the current moment, a training sample matrix and an initial dictionary for the monitoring phase are constructed. The real-time feature dictionary representing the current operating state is obtained by combining sparse coding algorithm and dictionary learning algorithm. S4. Calculate the difference measure between the real-time feature dictionary obtained at each time moment and the benchmark health dictionary, and obtain the sparse dictionary distance for each time period based on the difference measure. S5. The sparse dictionary distance is used to form a time series, and the time series is smoothed and used as a health indicator of the wind turbine. S6. Calculate and determine the equipment fault threshold based on the changing trend of the sparse dictionary distance health index when the equipment is in healthy operation; compare the real-time sparse dictionary distance health index with the fault monitoring threshold to evaluate the operating status of the wind turbine; and issue an alarm when the amplitude of the sparse dictionary distance of the equipment continuously exceeds the threshold.

2. The method for constructing and monitoring the health indicators and status of wind turbine units based on sparse dictionary distance according to claim 1, characterized in that, In step S1, the preprocessing includes signal filtering and normalization to obtain normalized signal samples; the normalization is specifically Z-score standardization, calculated using the following formula: ; in, This is the original signal segment after filtering. The mean of the original signal segment. The standard deviation of the original signal segment is used; through the Z-score standardization process, the difference in the amplitude of the vibration signal under different speed conditions is eliminated, so that the subsequent dictionary learning process focuses on the time-domain waveform structure characteristics of the signal.

3. The method for constructing and monitoring the health indicators and status of wind turbine units based on sparse dictionary distance according to claim 1, characterized in that, In step S2, a peak search-based training sample construction strategy is adopted, specifically as follows: Perform a full-segment search on the amplitude of the preprocessed vibration signal; Set minimum peak distance and minimum peak height thresholds, sort the searched peak points according to peak size, and select the top N significant peaks; Using each peak point as the starting reference point or center point, a data segment of fixed length L is extracted to construct a training sample matrix for dictionary learning.

4. The method for constructing and monitoring the health indicators and status of wind turbine units based on sparse dictionary distance according to claim 1, characterized in that, In step S2, the construction of the initial dictionary using the training sample matrix specifically involves: The training sample matrix is ​​subjected to mean-neutralization, and its covariance matrix is ​​subjected to singular value decomposition. The first K principal component feature vectors after singular value decomposition are selected as atoms of the initial dictionary, where K is the dictionary size parameter set before the initial training of the K-SVD dictionary learning algorithm, so as to obtain an initial dictionary containing the main vibration waveform structure of the signal.

5. The method for constructing and monitoring the health indicators and status of wind turbine units based on sparse dictionary distance according to claim 1, characterized in that, In step S2, the benchmark health dictionary that can characterize the waveform features of the device health mode is trained by combining the sparse coding algorithm and the dictionary learning algorithm. Specifically, the following iterative update steps are performed using the initial dictionary constructed in step S2 as the initial value: In the sparse coding stage, the dictionary atoms in the current iteration step are fixed, and the orthogonal matching pursuit algorithm is used to solve the sparse coefficient matrix of the training samples under the current dictionary. During the dictionary learning and update phase, the atoms in the dictionary and their corresponding non-zero coefficients are updated column by column; for the j-th atom in the dictionary, the error matrix after removing the contribution of that atom is calculated. Only retain The columns associated with this atom constitute the shrinkage error matrix. ;right Perform singular value decomposition, select the left singular vector corresponding to the maximum singular value as the updated j-th atom, and select the product of the maximum singular value and the right singular vector as the updated coefficient; Repeat the above sparse coding stage and dictionary learning and update stage until the set number of iterations is reached, and determine the final iteratively updated dictionary as the baseline healthy dictionary.

6. The method for constructing and monitoring the health indicators and status of wind turbine units based on sparse dictionary distance according to claim 1, characterized in that, In step S4, the sparse dictionary distance is calculated using a bidirectional mapping strategy, as follows: S41, Forward mapping calculation; based on the baseline health dictionary Each atom in In the real-time feature dictionary Search for the most similar atom in the middle and calculate the one-way mapping distance. And sum them to get the forward mapping distance; S42, Backward mapping calculation; for real-time feature dictionary Each atom in In the benchmark health dictionary Search for the most similar atom in the middle and calculate the one-way distance. And sum them to get the backward mapping distance; S43. Sparse dictionary distance calculation: Calculate the average of the forward mapping distance and the backward mapping distance as the final sparse dictionary distance; the calculation formula is as follows: in, The number of atoms in the dictionary. As a benchmark health dictionary, For real-time feature dictionary; Represents a single atom To the target dictionary One-way mapping distance; in, The forward mapping distance represents the deviation of the signal waveform shape from the best-matching atom in the baseline health dictionary to the real-time feature dictionary. The backward mapping distance represents the deviation of the signal waveform shape from the best-matching atom in the benchmark health dictionary to the real-time feature dictionary.

7. The method for constructing and monitoring the health indicators and status of wind turbine units based on sparse dictionary distance according to claim 6, characterized in that, The one-way mapping distance The Euclidean distance metric, based on atomic morphological similarity, is defined as: in, Dictionary atoms in It represents the cosine similarity between atoms; this formula quantifies the degree of deviation in waveform shape between atoms.

8. The method for constructing and monitoring the health indicators and status of wind turbine units based on sparse dictionary distance according to claim 1, characterized in that, In step S5, the method for determining the equipment fault threshold is specifically as follows: A baseline distance set is constructed using several sparse dictionary distance data points from the baseline health dictionary training phase or the initial stage of device health operation; the average value of the set is then calculated. and standard deviation Set fault monitoring thresholds for: When the real-time calculated health indicators exceed this threshold When this happens, it is determined that the equipment's operating status is abnormal.

9. The method for constructing and monitoring the health indicators and status of wind turbine units based on sparse dictionary distance according to claim 1, characterized in that, The smoothing operation described in step S5 specifically employs a sliding window filtering method to focus on the statistical characteristics of the indicators over a period of time and eliminate random fluctuations in the indicators caused by sudden changes in equipment operating conditions.

10. The method for constructing and monitoring the health indicators and status of wind turbine units based on sparse dictionary distance according to claim 1, characterized in that, The health indicator construction method is used for predictive maintenance management of industrial equipment, specifically including the detection of operating status and early fault warning of key rotating components such as the main bearing and gearbox of wind turbine generator sets and industrial bearing housings.

Citation Information

Patent Citations

  • Rotating equipment fault diagnosis method and system based on improved sparse dictionary

    CN114136604A

  • Gearbox fault diagnosis method and system based on local feature sparse coding representation

    CN117131412A

  • Method and device for correcting and reconstructing a barrel distorted image

    WO2018119565A1