Fan bearing fault detection method based on voiceprint features

CN122333105BActive Publication Date: 2026-08-11四川华电珙县发电有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-29
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

1、现有方法在谱图增强时多采用全局统一的对比度拉升,未考虑不同频带对故障冲击的敏感程度差异,导致那些仅在少数频带上呈现右偏分布的稀疏冲击在增强中被背景成分淹没,或者导致平稳频带的噪声被不当放大

Benefits of technology

1、本发明提出基于时间偏度驱动的频带自适应对比度增强方法,通过计算每个梅尔频带能量沿时间分布的偏度系数,自适应地为各频带设置差异化的增强陡峭度因子,实现对稀疏故障冲击亮斑的精准增强,解决了传统全局增强方式容易同步放大背景噪声和稳定谐波的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122333105B_ABST
    Figure CN122333105B_ABST
Patent Text Reader

Abstract

This invention relates to the field of wind turbine bearing fault detection technology, and discloses a wind turbine bearing fault detection method based on acoustic signature features. The method includes: wind turbine bearing vibration sample acquisition, annotation, and construction of the original logarithmic Mel spectrum matrix; dual-channel spectrum enhancement preprocessing for sparse impact bright spots; training of a fault classification model based on frequency band grouping selection perception and periodic prior attention; inputting the target data of the equipment to be diagnosed into the fault diagnosis model for fault diagnosis, and outputting the fault diagnosis results. The beneficial effect of this invention lies in proposing a frequency band adaptive contrast enhancement method based on time skewness driving. By calculating the skewness coefficient of the energy distribution along the time of each Mel frequency band, a differentiated enhancement steepness factor is adaptively set for each frequency band, achieving precise enhancement of sparse fault impact bright spots, and solving the problem that traditional global enhancement methods easily amplify background noise and stable harmonics simultaneously.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wind turbine bearing fault detection technology, specifically a wind turbine bearing fault detection method based on acoustic signature characteristics. Background Technology

[0002] In the field of industrial rotating machinery condition monitoring, early fault diagnosis of fan bearings has always been a key technical challenge to ensure the safe operation of equipment. Fans operate in complex environments, and their bearing vibration signals often exhibit non-stationary characteristics and strong noise backgrounds. Fault impact components only show localized, short-term spikes in the time domain and are frequently masked by frequency harmonics, structural resonance, and environmental noise. How to effectively extract weak fault features from low signal-to-noise ratio vibration signals and achieve accurate automatic identification of fault types is a core problem that urgently needs to be solved. In recent years, artificial intelligence-based fault diagnosis methods have gradually emerged, utilizing machine learning, deep neural networks, and other technologies to model and classify vibration signals, significantly improving the automation level of diagnosis. However, existing methods generally suffer from the following problems: 1. Existing methods often use a globally uniform contrast boost when enhancing the spectrum, without considering the different sensitivities of different frequency bands to fault impacts. This results in sparse impacts that only show a right-skewed distribution in a few frequency bands being submerged by background components during enhancement, or noise in stable frequency bands being improperly amplified.

[0003] 2. Existing methods often employ moving mean filtering and other outlier-prone processing methods when removing background trends, or extract transient components only from the time dimension without simultaneously introducing cross-band consistency constraints in the frequency direction. This makes it difficult to effectively distinguish between the true broadband impulse response and random spike interference on a single frequency band.

[0004] 3. Existing methods often allow convolution kernels to share parameters indiscriminately across all Mel frequency bands when modeling in the frequency domain, without considering the differences in sensitive frequency band intervals for different fault types and the personalized frequency band distribution among different samples. This results in insufficient targeting of front-end feature extraction for discrimination, and important frequency band information may be weakened by the overall averaging effect.

[0005] 4. Existing methods typically rely on general pooling or data-driven attention to automatically learn inter-frame relationships during time aggregation, without utilizing the prior knowledge that fault shocks have quasi-periodic repetition. When faced with speed fluctuations and non-strictly constant periods, it is difficult to stably capture long-range dependencies that conform to fault rhythms. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides a method for detecting wind turbine bearing faults based on acoustic signature features.

[0007] To achieve the above objectives, the present invention employs the following technical solution: A method for detecting wind turbine bearing faults based on acoustic signature features includes the following steps: Vibration data of wind turbine bearings were collected, sliced, and labeled. The original time-domain vibration signal was converted into the original logarithmic Mel spectrum matrix jointly represented by time frames and Mel frequency bands. A dual-channel spectral enhancement model for sparse impact bright spot features of bearing faults is constructed. The first channel performs band-adaptive contrast enhancement on the original log-Mel spectrum matrix based on time skewness. The second channel extracts short-time transient residuals through band trend decomposition and cross-band consistency constraints. The two channels generate a dual-channel spectral input tensor containing complementary features. A frequency band grouping selection sensing module and a periodic prior attention mechanism are introduced to build a wind turbine bearing fault classification model; the training set and test set are divided using a dual-channel spectrogram input tensor to complete the iterative training and parameter optimization of the fault classification model, and a fault diagnosis model is obtained. The target vibration data of the wind turbine bearing to be diagnosed is collected, and after preprocessing such as slicing, log-Mel spectrum conversion and dual-channel spectrum enhancement, it is input into the trained fault diagnosis model for feature recognition and classification. Finally, the fault type and fault detection result of the wind turbine bearing are output.

[0008] Furthermore, the original time-domain vibration signal is converted into an original logarithmic Mel spectrum matrix jointly represented by time frames and Mel frequency bands, including: Continuous vibration data is sliced ​​according to a fixed duration or a fixed number of sampling points to obtain multiple original vibration sample sequences; For each original vibration sample sequence, perform framing, windowing, and spectral transformation to obtain a short-time spectral sequence; The power spectrum of each analysis frame is input into the Mel filter bank to obtain the Mel band energy sequence; The initial log-Mel spectrogram matrix is ​​resampled along the time axis to a uniform number of frames to obtain the original log-Mel spectrogram matrix.

[0009] Furthermore, time-skew-driven band-adaptive contrast enhancement includes: The original log-Mel spectrum matrix Perform in-sample linear normalization to obtain the normalized logarithmic Mel spectrum matrix. ; For each Mel band, extract a band energy sequence along the time axis; For each Mel band Determine the enhancement center threshold of this Mel band. ; The first generation is generated based on the skewness coefficient. Enhanced kurtosis factor in the Mel band ; For the normalized logarithmic Mel spectrum matrix By performing point-by-point smoothing nonlinear enhancement, a frequency band contrast enhancement spectral matrix is ​​obtained. .

[0010] Furthermore, the extraction of short-time transient residuals through frequency band trend decomposition and cross-frequency band consistency constraints includes: For each Mel band Extract the frequency band contrast enhancement spectrum matrix along the time axis The corresponding frequency band energy sequence; Perform a moving median filter on the band energy sequence of each Mel band to obtain the trend estimation result. The two-dimensional matrix composed of the trend estimation results of all Mel bands is denoted as follows. ; Enhance the frequency band contrast spectral matrix Subtract the slow-changing trend matrix Only the positive differences are retained to obtain the preliminary transient residual matrix. ; For the initial transient residual matrix Each position in Take the neighboring Mel band in the frequency direction and calculate the cross-band consistency score; The cross-band consistency score is mapped to a consistency weight, and then the consistency weight is used to adjust the initial transient residual matrix. Weighting is performed to obtain the transient residual spectrum matrix. ; For transient residual spectrum matrix Perform intra-sample linear normalization to enhance the band contrast of the spectral matrix. With transient residual spectrum matrix Stacking along the channel dimension yields a two-channel spectral input tensor. .

[0011] Furthermore, the iterative training and parameter optimization of the fault classification model include: Frequency band grouping and sensing feature extraction involves dividing the dual-channel spectrum input tensor into several frequency band groups along the frequency direction, extracting local time-frequency features from each group independently, and then adaptively generating weighting coefficients for each group based on the global frequency band statistics of the current sample. Finally, the samples are fused to obtain a frequency band selection sensing feature map. A time attention mechanism is constructed based on candidate impact time and average impact period. The frame-by-frame time feature sequence is extracted from the frequency band selection sensing feature map, and the transient residual spectrum matrix is ​​used to detect candidate impact time and estimate average impact period. The period prior bias matrix is ​​then injected into the time attention mechanism. The model is trained using a classification output and category-weighted training. The classification logic value is output through a fully connected classification layer, and the entire model is trained using category-weighted cross-entropy loss. The model parameters are updated using an adaptive moment estimation optimizer.

[0012] Furthermore, the frequency band grouping selection and perceptual feature extraction are as follows: Input the dual-channel spectrum into the tensor Classified by frequency direction One continuous frequency band group; Perform independent convolution extraction on each local input tensor to obtain the frequency band group feature map; Generate frequency band group weights based on the global frequency band statistics of the current sample; Multiply the feature map of each frequency band group by the corresponding frequency band group weight, and then stitch them together along the frequency direction in the original frequency order to obtain a weighted stitched feature map.

[0013] Furthermore, a time attention mechanism is constructed based on the candidate impact time and the average impact period, specifically as follows: Frame-by-frame compression is performed on the frequency band selection sensing feature map to obtain the temporal feature sequence; Constructing a frame-by-frame impact intensity sequence based on the transient residual spectrum matrix; Estimate the average impact period based on the candidate impact time index set; Construct a periodic prior bias matrix based on the candidate impact time index set and the average impact period; Time feature series The query matrix, key matrix, and value matrix are obtained by passing through three linear transformation layers respectively. The time-enhanced feature sequence is obtained by weighting and summing the value matrix using the time attention weight matrix.

[0014] Furthermore, frame-by-frame compression is performed on the frequency band selection sensing feature map to obtain a temporal feature sequence, specifically: Perform band selection sensing feature map along the channel direction Convolution maps the fused output channel dimension to the frame-by-frame feature dimension, and the channel number to the target dimension, resulting in a dimension of size [missing information]. The characteristic tensor; Perform average pooling along the frequency direction for each time frame to obtain the time feature sequence. The size is This compresses the two-dimensional frequency band features into a one-dimensional time feature vector; in, This represents the total number of time frames after unification. This represents the number of discrete frequency bands in the frequency direction for a single sample. This indicates the feature dimension per frame.

[0015] Furthermore, a frame-by-frame impact intensity sequence is constructed based on the transient residual spectrum matrix, specifically as follows: For the first For each time frame, the one with the largest value is selected from all Mel bands of the transient residual spectrum matrix. Each frequency band residual, then... Calculate the average of the values; frame-by-frame impact intensity sequence Perform a moving average smoothing with a length of 3 or 5, then perform peak detection to obtain a set of candidate impact moment indices. Candidate impact moment index set This represents the set of time frame locations in the current sample where the suspected impact occurred.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention proposes a band-adaptive contrast enhancement method based on time skewness driving. By calculating the skewness coefficient of the energy distribution of each Mel band along the time, it adaptively sets a differentiated enhancement steepness factor for each band, thereby achieving precise enhancement of sparse fault impact bright spots. This solves the problem that traditional global enhancement methods are prone to simultaneously amplifying background noise and stable harmonics.

[0017] 2. This invention proposes a transient residual extraction mechanism based on moving median trend decomposition and cross-band consistency constraints. It utilizes the characteristic of median filtering being insensitive to impulse outliers to robustly remove slowly varying backgrounds, and suppresses isolated narrowband noise by calculating the degree of co-occurrence of residual energy in adjacent frequency bands. It mathematically models the two major physical properties of strong impulse and broadband excitation.

[0018] 3. This invention proposes a front-end feature extraction structure for frequency band grouping selection perception. The spectrum is divided into multiple groups for independent convolution, and adaptive weights are generated based on the global frequency band statistical characteristics of each sample for weighted fusion. This allows the front end of the model to automatically strengthen the band intervals most favorable for current fault identification, changing the traditional processing method where parameter sharing of convolution kernels across all frequency bands leads to the dilution of discrimination information.

[0019] 4. This invention proposes a periodic prior attention mechanism based on candidate impact time and average impact period. The impact time is detected from the transient residual spectrum and the impact interval is estimated. Based on this, a periodic prior bias matrix is ​​constructed and a time self-attention calculation process is injected. This guides the model to aggregate cross-frame information according to the natural rhythm of fault impact, making long-range time series modeling physically reasonable. Attached Figure Description

[0020] Figure 1 This is a flowchart of the present invention; Figure 2It is the original vibration signal and the original logarithmic Mel spectrum; Figure 3 This is a schematic diagram of the original logarithmic Mel spectrum matrix; Figure 4 This is a schematic diagram of the time skewness coefficients for each Mel frequency band; Figure 5 This is a schematic diagram of the frequency band adaptive enhancement parameters; Figure 6 This is a schematic diagram of the initial transient residual matrix; Figure 7 This is a diagram illustrating cross-band consistency scores; Figure 8 This is a schematic diagram of the adaptive frequency band group weights for samples; Figure 9 A schematic diagram of the frame-by-frame impact intensity sequence and candidate impact times. Detailed Implementation

[0021] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined in this application.

[0022] like Figure 1 As shown, this invention proposes a method for detecting wind turbine bearing faults based on acoustic signature features, the main contents of which are as follows: S1. Collection, labeling, and construction of the original log-Mel spectrum matrix of wind turbine bearing vibration samples. The vibration signals of wind turbine bearings under normal, inner ring fault, outer ring fault, and rolling element fault conditions all contain frequency components, harmonic components, and environmental noise components in the time domain. Only a short-term surge corresponding to the fault impact occurs in a local time period. This invention converts the original time-domain vibration signal into an original logarithmic Mel spectrum matrix jointly represented by time frames and Mel frequency bands. The specific steps are as follows: 1) A vibration acquisition unit is arranged on the outside of the bearing housing of the fan under test. The vibration acquisition unit is used to collect radial vibration signals during the operation of the bearing.

[0023] In a convenient implementation method, one piezoelectric accelerometer can be installed radially on the bearing housing; when improved installation robustness is required, one piezoelectric accelerometer can be installed in each of two mutually perpendicular directions, and the signal with the higher signal-to-noise ratio can be selected for subsequent processing; the sampling frequency is denoted as... , This indicates the number of vibration sampling points collected per unit time, such as 12.8kHz, 25.6kHz, or higher.

[0024] 2) Slice the continuous vibration data according to a fixed duration or a fixed number of sampling points to obtain multiple original vibration sample sequences. The original vibration sample sequences are denoted as... , This represents a one-dimensional time-domain vibration sequence corresponding to a single sample, with its length denoted as . , This indicates the total number of sampling points contained in a single sample.

[0025] In one implementation, 16384, 32768, or other fixed lengths matching the equipment rotation speed can be used. To avoid the sample boundary truncating the impact response, adjacent samples can be overlapped, with an overlap ratio of 50%.

[0026] 3) Based on the test bench operating condition label, maintenance record or manual verification results, label the fault category for each original vibration sample sequence. There are 4 fault categories: normal state, inner ring fault state, outer ring fault state and rolling element fault state.

[0027] Based on this, each original vibration sample sequence corresponds to a clear category label, which facilitates subsequent supervised training.

[0028] 4) Perform framing, windowing, and spectral transformation on each original vibration sample sequence to obtain a short-time spectral sequence.

[0029] In practical implementation, the original vibration sample sequence is divided into frame lengths. and frame shift Frame segmentation is performed, where, This indicates the number of sampling points contained in each analysis frame, for example, it can be 512 or 1024; This represents the sampling point interval between the start points of two adjacent analysis frames, for example, it can be taken as... Half of it. After framing, each analysis frame is multiplied by a Hamming window, and then a Fast Fourier Transform is performed on the windowed analysis frames to obtain the corresponding power spectrum.

[0030] In one embodiment, as an example, suppose , For a length The original vibration sample sequence is extracted, starting from the starting point. Every 256 sampling points, an analysis frame of length 512 is extracted. Each frame is multiplied by a Hamming window of length 512 and then subjected to a 512-point fast Fourier transform to obtain the power spectrum corresponding to each analysis frame.

[0031] 5) Input the power spectrum of each analysis frame into the Mel filter bank to obtain the Mel band energy sequence. The total number of Mel bands is denoted as... , This represents the number of discrete frequency bands in the frequency direction for a single sample, and can be, for example, 40, 64, or 80. In a convenient implementation approach, it can be... This allows for a balance between frequency resolution and computational complexity.

[0032] Furthermore, logarithmic compression is performed on the Mel band energy to obtain the initial log-Mel spectrum matrix of a single sample. The initial log-Mel spectrum matrix is ​​denoted as... , This represents a two-dimensional matrix consisting of time frames and Mel bands, with a size of [missing information]. ,in, This indicates the number of time frames obtained after the current sample is segmented. Logarithmic compression can be achieved by "adding a minimum constant first, then taking the natural logarithm," where the minimum constant is used to avoid the inability to take the logarithm when the Mel band energy is 0.

[0033] In one embodiment, as an example, suppose The power spectrum of a certain time frame is mapped through 64 Mel filters to obtain 64 Mel band energy values; during logarithmic compression, a very small constant (e.g., ...) is first added to each energy value. Then, take the natural logarithm to obtain the log-Mel spectrum vector for that time frame; after processing all time frames sequentially, concatenate them in chronological order to obtain the initial log-Mel spectrum matrix of the sample. .

[0034] 6) Resample the initial log-Mel spectrogram matrix along the time axis to a uniform number of frames to obtain the original log-Mel spectrogram matrix. The size is ,in, This represents the total number of time frames after unification, for example, 128.

[0035] In practical implementation, linear interpolation can be used to transform the initial log-Mel spectrum matrix. The timeline from Resampled to a fixed number of frames This ensures that all samples have the same input size.

[0036] In one embodiment, as an example, suppose If a sample is divided into frames, the number of time frames obtained is... Then For the original horizontal axis coordinates, Using the target horizontal axis coordinate, linear interpolation is performed sequentially along the time direction for the energy sequence of each Mel band, resulting in a matrix size of 128×64; if another sample... Similarly, samples were resampled along the time direction for 128 time frames to ensure that the output size of all samples was the same. .

[0037] In one embodiment, such as Figure 2 As shown, the original vibration signal and the original logarithmic Mel spectrum are analyzed to display the original radial vibration time-domain waveform acquired by the piezoelectric accelerometer. The horizontal axis represents time (unit: seconds), and the vertical axis represents acceleration (unit: g). Figure 2 Multiple distinct short-term spikes are visible, reflecting the periodic impact response caused by the outer ring fault.

[0038] In this embodiment, as Figure 3 As shown, the original log-Mel spectrum matrix is ​​a two-dimensional time-frequency representation of the original vibration signal obtained after framing, windowing, fast Fourier transform, Mel filter bank mapping, and logarithmic compression. The horizontal axis represents the time frame index (dimensionless), the vertical axis represents the Mel frequency band index (dimensionless), and the color represents the logarithmic energy value (compressed and normalized, dimensionless). Figure 3 The presence of local bright spots at the time frame corresponding to the impact, while the background harmonics and noise form a relatively uniform low-energy substrate, reflects the process of converting a one-dimensional vibration signal into a time-frequency feature matrix suitable for two-dimensional convolution processing.

[0039] In one embodiment, such as Figure 4 As shown, the temporal skewness coefficients of each Mel band are analyzed, displaying the skewness coefficient of the energy distribution along the time direction for each Mel band. The horizontal axis represents the Mel band index (dimensionless), and the vertical axis represents the skewness coefficient (dimensionless). Positive values ​​(orange bars) indicate that the energy of that band is right-skewed, with a few time frames showing significantly higher energy than the background, consistent with the statistical characteristics of fault impacts. Negative values ​​or near-zero values ​​(blue bars) correspond to a stable background, indicating significant differences in the sensitivity of different bands to impacts, providing a basis for band adaptive enhancement.

[0040] In this embodiment, as Figure 5 As shown, the adaptive enhancement parameters of the frequency bands are analyzed, displaying the enhancement center threshold (blue dotted line) and enhancement kurtosis factor (dark red bar) for each Mel band. The horizontal axis represents the Mel band index (dimensionless), and the vertical axis represents the parameter values ​​(dimensionless). The enhancement center threshold is taken from the upper quartile energy of each frequency band and is used to distinguish between background and suspected impacts. The enhancement kurtosis factor is calculated from the skewness coefficient; the stronger the right skewness of the frequency band, the larger the enhancement kurtosis factor.

[0041] S2. Dual-channel spectral enhancement preprocessing for sparse impact bright spots In the original logarithmic Mel spectrum matrix, wind turbine bearing failures are usually manifested as local high-energy bright spots on a few time frames. However, on-site noise, structural resonance, and stable harmonics often occupy a large dynamic range, causing the bright spots corresponding to the fault impact to be submerged by background components.

[0042] This invention performs dual-channel spectral enhancement preprocessing for sparse impact bright spots: the first channel performs band-adaptive contrast enhancement on the original log-Mel spectrum matrix based on time skewness, highlighting local high-energy bright spots suspected of impact; the second channel, based on the band-contrast-enhanced spectrum, further extracts short-time transient residuals through band trend decomposition and cross-band consistency constraints, suppressing stationary harmonics and isolated narrowband noise; the spectra of the two channels are finally stacked into a dual-channel spectral input tensor, providing complementary time-frequency representations for subsequent fault classification models. The specific steps are as follows: S201, Time-skew-driven band-adaptive contrast enhancement Different Mel bands are sensitive to faults to varying degrees. Mel bands containing fault impacts usually exhibit a right-skewed distribution with "lower energy in most time frames and significantly higher energy in a few time frames," while stable background bands are usually closer to a symmetrical distribution or a weakly fluctuating distribution. If the same enhancement intensity is used for all Mel bands, it is easy to amplify both background noise and stable harmonics.

[0043] This step analyzes the statistical skewness in the time direction for each Mel band, and then adaptively adjusts the enhancement intensity according to the frequency band to obtain the frequency band contrast enhancement spectrum matrix. The specific steps are as follows: 1) Convert the original logarithmic Mel spectrum matrix Perform in-sample linear normalization to obtain the normalized logarithmic Mel spectrum matrix. Normalized logarithmic Mel spectrum matrix This indicates that the spectral values ​​of the current sample will be uniformly mapped to... The two-dimensional matrix after the range has a size of .

[0044] In the specific implementation, the original logarithmic Mel spectrum matrix is ​​first obtained. The minimum and maximum values ​​in the set are then mapped linearly by subtracting the minimum value from each element and dividing by the difference between the maximum and minimum values. The interval allows samples from different working conditions to share a common set of enhancement parameters during subsequent enhancement.

[0045] In one embodiment, as an example, suppose a certain sample The minimum value is -8.2, the maximum value is 2.6, and the range is 10.8; for a position with an original value of -2.0, the normalized value is... .

[0046] 2) For each Mel frequency band Extract a frequency band energy sequence along the time axis.

[0047] Specifically, fixed Mel band index Then, read the normalized logarithmic Mel spectrum matrix. The total time frame values ​​in this Mel band are obtained, with a length of The frequency band energy sequence is used to characterize the energy fluctuation of the current Mel band throughout the entire sample.

[0048] Furthermore, the mean, standard deviation, and skewness coefficient of the energy sequence in this frequency band are calculated. The mean and standard deviation serve as the basic statistical description of the energy distribution in this frequency band, which are used to help determine the degree of energy fluctuation and provide intermediate quantities required for the calculation of the skewness coefficient. The standard deviation can be used as one of the reference bases for subsequent determination of the upper limit of enhancement intensity, while the skewness coefficient is directly used to determine the enhancement steepness factor.

[0049] Specifically, the skewness coefficient is denoted as , Indicates the first The degree of skewness in the energy distribution along the time direction of each Mel band. When When the value is relatively large and positive, it indicates that there are a small number of high-energy time frames in the Mel band that are significantly higher than the background, which is more consistent with the statistical characteristics of fault impact bright spots; when When the value is close to 0 or negative, it indicates that the Mel frequency band is closer to a stable background and should not be overly enhanced. The skewness coefficient is obtained using a standard third-order center moment normalization calculation method.

[0050] In one embodiment, as an example, suppose the bandwidth energy sequence length of a certain Mel band is... The mean is 0.32 and the standard deviation is 0.15. The deviation of the energy value for each time frame from the mean is calculated, the cubed value is averaged, and then divided by the cube of the standard deviation to obtain the skewness coefficient. This indicates that the energy distribution in this frequency band is significantly skewed to the right, and there are a small number of high-energy time frames.

[0051] 3) For each Mel frequency band Determine the enhancement center threshold of this Mel band. , Indicates the first The reference value for the boundary between "normal background energy" and "high-energy bright spot" in a Mel band.

[0052] In practice, the frequency band energy sequence of the current Mel band is first sorted in ascending order, and then the value at the 75th percentile after sorting is taken as the enhancement center threshold. Therefore, only time frames that are higher than the upper quartile level of this frequency band will be significantly boosted in subsequent steps.

[0053] In one embodiment, for example, if a certain Mel band has 128 time frames, these 128 energy values ​​can be sorted in ascending order first, and then the value near the 96th position can be taken as the enhancement center threshold. If only a few frames in this frequency band experience fault impact peaks, these peaks will typically be higher than the enhancement center threshold. This allows for greater gains in subsequent enhancements.

[0054] 4) Based on the skewness coefficient Generate the first Enhanced kurtosis factor in the Mel band , This represents the slope control amount of the current Mel band in the enhancement function. The larger the value, the more likely the Mel band is to contain sparse impacts, exceeding the enhancement center threshold. The high-energy time frames will be magnified more significantly.

[0055] In one implementation, the product of the base enhancement value and the skewness coefficient can be calculated first, and then the result can be restricted to a preset upper limit. For example, if the base enhancement value is set to 2.0, its product with the skewness coefficient can be calculated. If the skewness coefficient is negative, the product is treated as 0. If the product exceeds the preset upper limit of 6.0, it is truncated to 6.0. The final result is used as the enhancement kurtosis factor for that Mel band. This allows for sufficient enhancement of the impact-containing Mel frequency band while avoiding over-enhancement caused by individual abnormal peaks.

[0056] In this implementation, the base enhancement value is recorded. upper limit For the first One Mel frequency band, take ,but .in, Indicates the first Unlimited kurtosis factor of a Mel band This represents the function that takes the maximum value. This represents the function that takes the minimum value.

[0057] 5) For the normalized logarithmic Mel spectrum matrix By performing point-by-point smoothing nonlinear enhancement, a frequency band contrast enhancement spectral matrix is ​​obtained. , represented as: ; in, Represents the frequency band contrast enhancement spectrum matrix In the The first time frame, the first Enhancement results at the Mel frequency band; This represents the global enhancement amplitude coefficient, used to control the overall enhancement intensity; for example, it can be set to 0.3. This represents the hyperbolic tangent function, used to provide a smooth, saturated nonlinear gain; This indicates that the value will be truncated to... A function on an interval.

[0058] In practical implementation, when the value of the normalized logarithmic Mel spectrum matrix at a certain location is higher than the enhancement center threshold of the Ben-Mel band... When the hyperbolic tangent function output is positive, and the more it exceeds the threshold, the greater the gain, indicating that the position is enhanced; when the position is below or close to the enhancement center threshold... When the gain is close to 0 or slightly negative, the low-brightness background will not be significantly boosted, thus achieving selective processing of "enhancing the impact of high brightness and maintaining the low-brightness background".

[0059] In one embodiment, for example, suppose the skewness coefficient of a certain Mel frequency band If the base enhancement value is 2.0, then the enhancement steepness factor is... If the enhancement center threshold of the Mel frequency band Global enhancement amplitude coefficient If the normalized energy value at a certain time frame is 0.55, then that position is significantly higher than the enhancement center threshold, and the value will be increased after enhancement; if the normalized energy value at another time frame is 0.20, then that position is lower than the enhancement center threshold, and the enhancement amplitude is very small; based on this, sparse impact bright spots can be separated from large areas of background.

[0060] It should be noted that real fault impacts usually only occur in some Mel bands, and they exhibit a clear right-skewed distribution in these Mel bands. This step does not apply uniform contrast enhancement to all Mel bands. Instead, it first analyzes the skewness coefficient of each Mel band along the time direction, and then adaptively sets the enhancement steepness factor based on the skewness coefficient. Based on this, it is possible to selectively enhance the impact time frames in suspected fault bands, while avoiding amplifying irrelevant stable bands as well.

[0061] S202. Transient residual extraction based on frequency band trend decomposition and cross-frequency band consistency constraints. After frequency band adaptive contrast enhancement, the impact bright spot is more obvious, but stable harmonics and slowly varying background may still be superimposed with the impact component. Real fault impact not only has short-term sudden increase characteristics, but also usually affects several adjacent Mel frequency bands at the same time, while random narrowband noise usually only appears as an isolated spike on a single Mel frequency band.

[0062] This step first estimates the slow variation trend of each Mel band along the time direction, then extracts the positive transient residuals, and introduces cross-band consistency constraints to suppress isolated narrowband interference, thus obtaining the transient residual spectrum matrix. The specific steps are as follows: 1) For each Mel frequency band Extract the frequency band contrast enhancement spectrum matrix along the time axis The corresponding frequency band energy sequence is used to estimate the slow-varying background trend and the rapid burst component in the current Mel band.

[0063] In the specific implementation, the Mel band index is fixed. From the frequency band contrast enhancement spectrum matrix The values ​​of all time frames in this frequency band are read and a length of [length missing] is constructed. The one-dimensional frequency band energy sequence is used for subsequent moving median filtering to estimate the slow-varying trend.

[0064] In one embodiment, as an example, suppose , Then from Extract all values ​​of the 15th Mel band at time frames 1 to 128 to obtain the sequence. .

[0065] 2) Perform a moving median filter on the band energy sequence of each Mel band to obtain the trend estimation result. The two-dimensional matrix composed of the trend estimation results of all Mel bands is denoted as... , Represents the frequency band contrast enhancement spectrum matrix The corresponding slow-changing trend matrix has a size of ,in, Indicates the first The first time frame, the first Trend estimates at the Mel frequency band.

[0066] In the specific implementation, the width of the moving median filter window is denoted as... , This indicates the range of frames used for trend estimation along the time direction, such as 7 or 9.

[0067] In one implementation, the band energy sequence for each Mel band is... With window width Slide along the time axis; for the first Take each time frame before and after that frame. Frame, along with itself The median of the frame's energy value is used as the trend estimate for that location. At the beginning and end boundaries of the sequence, the median is calculated using only the actual time frames present within the window.

[0068] In one embodiment, as an example, suppose For the 10th time frame, take 7 energy values ​​from the 7th to the 13th frames, sort them, and take the 4th value as the trend estimate; for the 1st time frame, the window only covers the 1st to 4th frames, and take the median of these 4 frames.

[0069] It should be noted that the moving median filter is used instead of the moving mean filter because the moving median filter is not sensitive to a small number of high-energy impact points and can more stably describe stationary harmonics and slowly varying backgrounds, while the moving mean filter is easily affected by impact peaks, causing the true impact to be partially absorbed into the trend term.

[0070] 3) Enhance the frequency band contrast of the spectrum matrix Subtract the slow-changing trend matrix Only the positive differences are retained to obtain the preliminary transient residual matrix. Preliminary transient residual matrix This indicates a short-term positive surge relative to a slow-changing trend, with a size of .

[0071] In the actual implementation, for each position ,like If the difference is greater than 0, it is retained; if it is less than or equal to 0, the position is set to 0. Based on this, falling edge fluctuations and local troughs will not be mistaken for fault impacts; among which, The frequency band contrast enhancement spectrum matrix represents the first... The first time frame, the first The enhanced value at the Mel frequency band. The slow-changing trend matrix represents the first... The first time frame, the first The trend estimate at each Mel frequency band; then the initial transient residual matrix. element in row i and column j .

[0072] 4) Regarding the initial transient residual matrix Each position in Take the neighboring Mel band in the frequency direction and calculate the cross-band consistency score. The cross-band consistency score is denoted as . , Indicates the first The first time frame, the first The degree of co-occurrence of the initial transient residual at a given Mel band in adjacent Mel bands.

[0073] In practical implementation, Centered on the frequency direction, select a width of The frequency band neighborhood, among which This represents the total number of frequency bands used in the statistics, for example, 5, which includes the current Mel band and two adjacent Mel bands on the left and right. At the frequency band boundaries, only the actual existing frequency bands are used. Then, the initial transient residuals of the same time frame within the neighborhood of that frequency band are squared and averaged to obtain the cross-band consistency score. Therefore, if multiple adjacent Mel bands simultaneously exhibit large residuals in the same time frame, the cross-band consistency score is high; if only a single Mel band exhibits an isolated spike, the cross-band consistency score is low.

[0074] In one embodiment, as an example, suppose , regarding position Take the Mel frequency band indices j0-2, j0-1, j0, j0+1, j0+2; if these 5 adjacent frequency bands are in the... The initial transient residuals for the time frames are 0.12, 0.08, 0.15, 0.10, and 0.05, respectively. Therefore, the cross-band consistency score is... If the residuals of adjacent frequency bands are all close to 0, the score will be significantly reduced.

[0075] 5) Map the cross-band consistency score to a consistency weight, and then use the consistency weight to adjust the initial transient residual matrix. Weighting is performed to obtain the transient residual spectrum matrix. .

[0076] In practical implementation, we can first process all Linear normalization is performed within the current sample to obtain a normalized consistency score between 0 and 1; then, the normalized consistency score is input into the logistic function to obtain the consistency weight; finally, the consistency weight is compared with the initial transient residual matrix. Multiply point by point to output the transient residual spectrum matrix. Preferably, the slope parameter of the logistic function can be 6.0, and the center point can be 0.5, so that the high consistency region is preserved and the low consistency isolated noise is suppressed.

[0077] In one embodiment, let a certain location The cross-band consistency score, after in-sample linear normalization, is 0.68. The input logistic function... The calculated consistency weight is approximately 0.746; this weight is then compared with... Multiply, if The value at that position in the weighted transient residual spectrum matrix is... .

[0078] 6) For the transient residual spectrum matrix Perform in-sample linear normalization and map it to Range; then, enhance the band contrast spectral matrix. With transient residual spectrum matrix Stacking along the channel dimension yields a two-channel spectral input tensor. The size is The first channel is a frequency band contrast enhancement spectrum matrix. The second channel is the transient residual spectrum matrix. .

[0079] In one embodiment, as an example, suppose , ;Band contrast enhancement spectrum matrix The normalized transient residual spectrum matrix is ​​128×64. Also 128×64; the input tensor of the dual-channel spectrum obtained after stacking The dimensions are 128×64×2, where channel 0 is [missing information] along the channel dimension. Channel 1 is .

[0080] It should be noted that if an outer ring fault impact actually exists at a certain time frame, then multiple consecutive adjacent Mel bands will usually show significant positive residuals simultaneously in the vicinity of that time frame. In this case, the cross-band consistency score is large, the consistency weight is close to 1, and the transient residual spectrum matrix is... This set of broadband impulse responses will be retained; if only a single Mel band has a spike in a certain time frame, while the adjacent Mel bands are basically stable, the cross-band consistency score is small, the consistency weight is low, and the spike will be significantly weakened.

[0081] It should also be noted that the impact of real wind turbine bearing failure has both short duration and cross-band excitation characteristics. This step combines the use of "trend decomposition in the time direction" and "consistency constraint in the frequency direction", which can further separate stationary harmonics and slowly varying background from the frequency band contrast enhancement spectrum matrix and significantly reduce the interference of random narrowband noise on subsequent classification.

[0082] In one embodiment, such as Figure 6 As shown, the preliminary transient residual matrix is ​​analyzed by directly subtracting the moving median trend from the frequency band contrast enhancement spectrum matrix and retaining the positive difference. The horizontal axis represents the time frame index (dimensionless), the vertical axis represents the Mel frequency band index (dimensionless), and the color represents the residual intensity (dimensionless). The residuals at the impact locations are significant, but some isolated frequency bands also show sporadic high residuals, demonstrating the separation effect of trend decomposition on the sudden increase components.

[0083] In this embodiment, as Figure 7As shown, the cross-band consistency score analysis demonstrates the degree of co-occurrence of the initial transient residuals in adjacent Mel frequency bands. The horizontal and vertical axes are the same as above, and the color represents the consistency score (square mean, dimensionless). At the actual impact moment, multiple adjacent frequency bands simultaneously exhibit high residuals, resulting in higher scores; random narrowband noise, on the other hand, scores very low. The experimental results reflect the physical characteristics of the fault impact of "cross-band co-excitation," providing a quantitative basis for subsequent suppression of isolated noise using consistency weights.

[0084] S3. Training of a fault classification model based on frequency band grouping selection perception and periodic prior attention. Dual-channel spectral input tensor At the same time, the enhanced overall time-frequency energy distribution and the detrended transient impact distribution are preserved. The sensitive frequency bands of different wind turbine bearing fault types are not exactly the same, and the fault impact usually has quasi-periodic repetitive characteristics on the time axis. If ordinary two-dimensional convolution and ordinary time pooling are used directly, it is easy to ignore the differences between different frequency bands on the one hand, and the impact interval pattern on the other hand.

[0085] This invention first performs frequency band grouping and selection of sensing features, then constructs an impact period prior by combining the transient residual spectrum matrix and guides time attention aggregation, and finally completes fault category training. The specific steps are as follows: S301, Frequency Band Grouping Selection and Sensing Feature Extraction The dual-channel spectral input tensor is divided into several frequency bands along the frequency direction. Local time-frequency features are extracted independently from each group. Then, weighting coefficients for each group are adaptively generated based on the global frequency band statistics of the current sample. Finally, these are fused to obtain a frequency band selection-aware feature map, enabling the model to automatically strengthen the frequency band ranges that are beneficial for the current sample. The specific steps are as follows: 1) Input the dual-channel spectrum into the tensor Classified by frequency direction A continuous frequency band group This indicates the total number of frequency bands, for example, 4.

[0086] In one implementation, the total number of Mel bands Divided into equal parts A series of consecutive frequency bands, each containing Adjacent Mel frequency bands; if Cannot be Divisibility allows for a slightly larger allocation of frequency band to the front group. If the total number of Mel bands... and Each frequency band group contains 16 adjacent Mel bands; The local input tensor corresponding to each frequency band group is denoted as . , Indicates the input tensor of the two-channel spectrum. In the Local spectral patches in each frequency band, with a size of [size missing]. ,in, Indicates the first The number of Mel bands contained in each frequency band group.

[0087] In one embodiment, as an example, suppose , , Each frequency band group contains One adjacent Mel frequency band; the local input tensor corresponding to the first frequency band group. Input tensors for dual-channel spectra The local spectral patch in the 1st to 16th Mel bands, with a size of Similarly, the local input tensor corresponding to the second frequency band group This is a local spectral patch on the 17th to 32nd Mel bands, with a size of [size missing]. Based on this, we obtain four local input tensors: U1, U2, U3, and U4.

[0088] 2) For each local input tensor Perform independent convolution extraction to obtain frequency band group feature maps. Frequency band group feature map Indicates the first The local time-frequency features extracted from each frequency band group have a size of [size missing]. ,in, This indicates the number of output channels for that frequency band group, for example, 8.

[0089] In specific implementation, the first Each frequency band group can use two layers of two-dimensional convolution: the kernel size of the first layer is taken as... The output channel count is 8, the stride is 1, and equal-size padding is used for boundary padding, followed by batch normalization and ReLU activation; the kernel size of the second layer remains the same. The output channel count remains at 8, the stride is 1, followed by batch normalization and ReLU activation. The convolution parameters for each frequency band are independent of each other, enabling different local patterns to be learned in different frequency bands.

[0090] In one embodiment, taking the first frequency band group as an example, its local input tensor The dimensions are 128×16×2; the first convolutional layer has a 3×3 kernel, 2 input channels, 8 output channels, a stride of 1, and equal padding. After batch normalization and ReLU, the output size is 128×16×8; the second convolutional layer has a 3×3 kernel, 8 input channels, 8 output channels, a stride of 1, and equal padding. After batch normalization and ReLU, the output is a frequency band feature map. The dimensions are 128×16×8.

[0091] 3) Generate band group weights based on the global band statistics of the current sample. The band group weights are denoted as... , Indicates the first The importance of each frequency band group in the current sample is determined by a value ranging from 0 to 1, and the sum of the weights of all frequency band groups is 1.

[0092] In the specific implementation, each local input tensor is processed separately. Calculate the mean and standard deviation across the two channels to obtain the first... The statistical vectors of each frequency band group are 4-dimensional; these statistical vectors are then concatenated sequentially to form a global statistical vector; this global statistical vector is then input into a two-layer fully connected network. The first layer's output dimension can be 16 and activated by ReLU; the second layer's output dimension is equal to... Finally, Softmax normalization is performed on the output of the second layer to obtain the weights of all frequency band groups. .

[0093] In one embodiment, as an example, suppose For the local input tensor of the first frequency band, the mean and standard deviation are calculated on the two channels respectively, resulting in a 4-dimensional statistical vector, such as... The statistical vectors of the four groups are concatenated into a 16-dimensional global statistical vector. After being processed by a fully connected layer (using a 16→16→4 dimension transformation) and the Softmax function, the weights of the four frequency band groups are obtained, for example, ρ1=0.35, ρ2=0.28, ρ3=0.22, and ρ4=0.15.

[0094] 4) Create feature maps for each frequency band group. Multiply by the corresponding frequency band weight Then, the original frequency sequences are spliced ​​along the frequency direction to obtain a weighted spliced ​​feature map.

[0095] Furthermore, the weighted splicing feature map is processed. Convolutional fusion yields a frequency band selection sensing feature map. Frequency band selection sensing feature map This represents the front-end feature result after independent convolution with frequency band grouping and adaptive frequency band weighting of samples, with a size of [size missing]. ,in, This indicates the number of output channels after fusion, for example, 16.

[0096] In one embodiment, for example, Multiply , Multiply , Multiply , Multiply Then, the features are stitched together along the frequency direction in the original frequency order to obtain a weighted stitched feature map with a size of 128×64×8. Subsequently, convolutional fusion is performed using 16 1×1 convolutional kernels to output a frequency band selection perceptual feature map. The dimensions are 128×64×16.

[0097] It should be noted that different fault types do not share the same sensitive frequency bands. This step combines "independent convolution of frequency band groups" and "sample adaptive frequency band weighting". The front end of the model can automatically strengthen the key frequency bands according to the frequency band distribution of the current samples, thereby reducing the problem of dilution of discrimination information caused by sharing the same convolution kernel across all frequency bands.

[0098] S302, Temporal attention aggregation based on candidate impact time and average impact period The frame-by-frame time feature sequence is extracted from the frequency band selection sensing feature map, and the candidate impact time is detected and the average impact period is estimated using the transient residual spectrum matrix. Based on this, a periodic prior bias matrix is ​​constructed to inject a time attention mechanism, so that the model explicitly pays attention to the time frame relationship that conforms to the fault impact rhythm when aggregating long-range time information. The specific steps are as follows:

[0099] 1) Sensing feature map of frequency band selection Perform frame-by-frame compression to obtain the temporal feature sequence. Time feature sequence This represents a one-dimensional fault representation corresponding to each time frame, with a size of [missing information]. ,in, This represents the feature dimension per frame, for example, 64.

[0100] In the specific implementation, the perceptual feature map is first selected for the frequency band. Execute along the channel direction Convolution, dimensional mapping as Dimension (e.g., 64), mapping the number of channels to the target dimension, yields a dimension of... The feature tensor; then for each time frame along the frequency direction (along Perform average pooling on the dimension to obtain the time feature sequence. The size is This compresses the two-dimensional frequency band features into a one-dimensional time feature vector.

[0101] In one embodiment, as an example, suppose The size is 128×64×16; first, the number of channels is mapped to 64 using 64 1×1 convolutional kernels, resulting in a tensor of size 128×64×64; then, average pooling is performed on each time frame along the frequency direction (64 frequency bands) to obtain the temporal feature sequence. The size is 128×64, meaning that each time frame corresponds to a 64-dimensional feature vector.

[0102] 2) Based on transient residual spectrum matrix Construct a frame-by-frame impact intensity sequence, denoted as . , Indicates length is A one-dimensional sequence, where Indicates the first The degree of impact in each time frame.

[0103] In specific implementation, for the first Each time frame, from the transient residual spectrum matrix Select the largest value from all Mel frequency bands Each frequency band residual, and then this Average the values ​​to get ;in, This indicates the number of high-response frequency bands used for statistical impact intensity in each time frame, such as 3 or 5; based on this, local impact peaks will not be diluted by averaging across all frequency bands.

[0104] Furthermore, for frame-by-frame impact intensity sequences Perform a moving average smoothing with a length of 3 or 5, then perform peak detection to obtain a set of candidate impact moment indices. Candidate impact moment index set This represents the set of time frame locations in the current sample where the suspected impact occurred.

[0105] In practical implementation, when a time frame simultaneously satisfies the conditions that "the location is a local maximum" and "the frame-by-frame impact intensity at that location is greater than the mean of the entire sequence plus 0.5 times the standard deviation", that time frame can be added to the candidate impact moment index set. Specifically, for the frame-by-frame impact intensity sequence After performing a moving average smoothing with a length of 3, iterate through each time frame index. ;like Both of the following conditions must be met: a) and (i.e., local maxima); b) (in and They are respectively (mean and standard deviation), then Add to candidate impact time index set .

[0106] In one embodiment, as an example, suppose , The threshold is If the 25th frame If the value is greater than that of frames 24 and 26, then frame 25 is included. .

[0107] 3) Based on the candidate impact time index set Estimate the average impact period, denoted as . , This represents the average time frame interval between the recurrence of adjacent shocks in the current sample.

[0108] In the specific implementation, the candidate impact moment index set The difference between adjacent time frame indices is calculated (time frame indices are retrieved in ascending order from the candidate impact moment index set, and the difference between two adjacent indices is calculated) to obtain a set of adjacent impact intervals. The median of these adjacent impact intervals is then taken as the average impact period. .

[0109] In one embodiment, as an example, the candidate impact moment index set The adjacent interval is The median is 16, which means frame.

[0110] It should be noted that using the median instead of the mean can reduce the impact of individual false positive peaks on period estimation; if the candidate impact time index set If the number of elements in the matrix is ​​less than two, the prior bias matrix of the subsequent cycles will be set to zero, causing the model to degenerate into ordinary temporal attention.

[0111] 4) Based on the candidate impact time index set and average impact period Constructing a periodic prior bias matrix Periodic prior bias matrix Represents the explicit prior of the injection-time attention computation process, with size . ,in, Indicates the first The time frame and the first The prior amount added to the attention score between time frames is calculated as follows:

[0112] ; in, This represents the periodic prior strength coefficient, used to control the degree of influence of the periodic prior on the attention distribution; for example, it can be set to 0.5. This indicates the cycle tolerance width, used to tolerate slight changes in the impact interval caused by speed fluctuations; for example, it can be 2.0 frames. Indicates the first The time frame and the first The absolute interval between time frames.

[0113] In practical implementation, if the interval between two time frames is close to the average impact period... And at least one of these time frames belongs to the candidate impact time index set. Then the periodic prior bias matrix A large positive bias is applied at this position to encourage the model to consider these two time frames as potentially belonging to the same fault rhythm chain; however, if the interval between them deviates significantly from the mean impact period... If this happens, the bias decreases rapidly.

[0114] 5) Time feature series The query matrix is ​​obtained by passing through three linear transformation layers respectively. Key matrix Sum matrix Then, the periodic prior bias matrix is... By incorporating the scaling dot product attention into the score calculation, we obtain the temporal attention weight matrix. , represented as: ; in, This represents the temporal attention weight matrix, with size . ; Represents the query matrix; Represents the key matrix; Indicates the transpose of the key matrix; This represents the dimension of the internal features of the attention mechanism; for example, it can be 64. This indicates a row-based normalization operation.

[0115] 6) Utilizing the time attention weight matrix value matrix Weighted summation is performed to obtain the time-enhanced feature sequence. Further, this time-enhanced feature sequence is passed sequentially through: a feedforward network (consisting of two fully connected layers; the first layer expands the dimension to 4 times before ReLU activation, and the second layer restores the original dimension), residual connections (element-wise addition of the feedforward network input and output), and layer normalization to obtain the updated time-enhanced feature sequence. Finally, global average pooling is performed on the time dimension to obtain the global fault feature vector. Global Fault Feature Vector This represents the overall fault characterization of the current sample, denoted by dimension . For example, 64 or 128 can be used.

[0116] It should be noted that the impact of wind turbine bearing failure usually has quasi-periodic repetitive characteristics. However, due to the influence of speed fluctuations and load changes, the impact interval is not strictly constant. This step not only uses a data-driven approach to learn the correlation between time frames, but also explicitly transforms the candidate impact time and the average impact period into a periodic prior bias matrix and directly injects it into the time attention calculation. This allows the model to more stably focus on the combination of time frames that conform to the impact rhythm, thereby improving the physical rationality and robustness of long-range time-dependent modeling.

[0117] S303, Classification Output and Category-Weighted Training 1) The global fault feature vector The input is a fully connected classification layer, and the output is logical values ​​for four categories. The logical value vector is denoted as... , This represents the unnormalized score of the current sample for four categories: normal state, inner race fault state, outer race fault state, and rolling element fault state, with a dimension of 4; then, the classification logical value vector... After performing the Softmax transformation, the predicted probabilities of the four fault categories are obtained.

[0118] In one implementation, the fully connected classification layer contains a fully connected network with an input dimension of... The output dimension is 4; this layer has no bias term and directly outputs the classification logical value vectors of the 4 categories. .

[0119] 2) Set class weights based on the number of samples in each class in the training set. The number of samples in each category is denoted as The maximum number of samples in all categories is denoted as Then the first The loss weights for each category can be calculated as follows: It is determined that, based on this, fault categories with smaller sample sizes will receive higher loss weights during the training phase, thereby mitigating the class bias problem when the number of normal samples is much greater than that of fault samples.

[0120] 3) The entire model is trained using class-weighted cross-entropy loss; class-weighted cross-entropy loss is used to compare the difference between the predicted probability and the true label, and to give a larger gradient to minority class samples during backpropagation.

[0121] In one embodiment, as an example, suppose a batch contains two samples with true labels of category 0 (normal state, weight 1.2) and category 2 (outer ring fault state, weight 2.5); if the model predicts a probability of 0.8 for the first sample in category 0 and a probability of 0.6 for the second sample in category 2, then the weighted loss for this batch is... .

[0122] 4) Update the model parameters using an adaptive moment estimator optimizer.

[0123] In one implementation, the initial learning rate can be chosen as follows: The batch size can be 16 or 32, and the total number of training rounds can be 80 to 150. After each training round, the F1 score of each of the four categories is calculated using the validation set. Then, the arithmetic mean of the four F1 scores is calculated to obtain the macro average F1 score. The model parameters corresponding to the highest macro average F1 score are saved as the final fault classification model.

[0124] In an easy-to-implement approach, all samples can be divided into a training set, a validation set, and a test set in a 7:2:1 ratio. The training set is used for parameter updates, the validation set is used for model selection, and the test set is used for final performance evaluation. After training, the final fault classification model output will have the ability to automatically identify four types of bearing conditions.

[0125] In one embodiment, such as Figure 8 As shown, the adaptive frequency band group weights of the analysis samples are displayed in a bar chart, showing the adaptive importance weights of the current sample to the four frequency band groups (each group has 16 adjacent Mel bands). The horizontal axis represents the frequency band group label (including the frequency band range), and the vertical axis represents the weight value (dimensionless). The different heights of the different groups indicate that the model automatically strengthens the frequency band intervals that are more favorable for discrimination based on the global frequency band statistical characteristics of the current sample. This reflects the role of the frequency band grouping selection perception mechanism and avoids the dilution of discriminative information caused by sharing a unified convolutional kernel across all frequency bands.

[0126] In this embodiment, as Figure 9 As shown, the frame-by-frame impact intensity sequence and candidate impact times are analyzed. The blue solid line represents the frame-by-frame impact intensity after moving average smoothing, and the gray line represents the original intensity. The horizontal axis is the time frame index (dimensionless), and the vertical axis is the impact intensity (obtained by averaging the top K largest residual values, dimensionless). The red scatter dots mark the candidate impact times detected by the local maxima and threshold criteria, and the dashed line represents the detection threshold. This sequence quantifies the impact significance at each time point from the transient residual spectrum and reliably locates the time frames where suspected impacts occurred.

[0127] S4. Online Monitoring and Fault Identification Process for Wind Turbine Bearings After model training is completed, the final fault classification model is deployed in the wind turbine condition monitoring system or edge diagnostic terminal to continuously identify the vibration signals collected in real time. The specific steps are as follows: 1) Use the same sampling frequency as during the training phase. Real-time acquisition of wind turbine bearing vibration data, using the same sample length as the training phase. The original vibration sample sequence to be identified is constructed using an overlapping method; based on this, the online identification input and the training input are consistent in data form.

[0128] 2) Repeat step S1 for the original vibration sample sequence to be identified, including framing, windowing, fast Fourier transform, Mel filter bank mapping, logarithmic compression, and time axis resampling, to obtain the original log-Mel spectrum matrix corresponding to the sample to be identified. .

[0129] 3) The original log-Mel spectrum matrix of the sample to be identified Repeat steps S201 and S202 to obtain the band contrast enhancement spectrum matrix. Transient residual spectrum matrix and dual-channel spectral input tensor .

[0130] 4) Input the dual-channel spectrum into the tensor Input the final fault classification model, and sequentially perform frequency band grouping selection sensing feature extraction, candidate impact moment detection, average impact period estimation, period prior time attention aggregation, and fully connected classification, and output the predicted probabilities of 4 fault categories.

[0131] 5) Take the category with the highest predicted probability as the fault identification result of the current identification window.

[0132] In one implementation, a fault alarm is triggered when an abnormal category is the highest probability category in three consecutive recognition windows, and the average predicted probability of these three recognition windows is greater than 0.7; if the highest probability category is normal, monitoring continues. Using consecutive windows for judgment can reduce false alarms caused by occasional noise in a single recognition window.

[0133] In one embodiment, for example, if after processing, the predicted probabilities of the outer ring fault state in the three most recent consecutive identification windows are 0.81, 0.84, and 0.79, respectively, then the average predicted probability is 0.813, which is greater than 0.7, and the maximum probability category in the three consecutive identification windows is the outer ring fault state. In this case, the system outputs an outer ring fault alarm result. If only one identification window has a high predicted probability of the outer ring fault state, while the windows before and after it return to normal, the system does not trigger an alarm to improve the stability during actual deployment.

[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting wind turbine bearing faults based on acoustic signature features, characterized in that, Includes the following steps: Vibration data of wind turbine bearings were collected, sliced, and labeled. The original time-domain vibration signal was converted into the original logarithmic Mel spectrum matrix jointly represented by time frames and Mel frequency bands. A dual-channel spectral enhancement model for sparse impact bright spot features of bearing faults is constructed. The first channel performs band-adaptive contrast enhancement on the original log-Mel spectrum matrix based on time skewness. The second channel extracts short-time transient residuals through band trend decomposition and cross-band consistency constraints. The two channels generate a dual-channel spectral input tensor containing complementary features. A frequency band grouping selection sensing module and a periodic prior attention mechanism are introduced to build a wind turbine bearing fault classification model; the training set and test set are divided using a dual-channel spectrogram input tensor to complete the iterative training and parameter optimization of the fault classification model, and a fault diagnosis model is obtained. The target vibration data of the wind turbine bearing to be diagnosed is collected, and after preprocessing such as slicing, log-Mel spectrum conversion and dual-channel spectrum enhancement, it is input into the trained fault diagnosis model for feature recognition and classification. Finally, the fault type and fault detection result of the wind turbine bearing are output. Among them, the band-adaptive contrast enhancement based on time skewness includes: The original logarithmic Mel spectrum matrix Perform in-sample linear normalization to obtain the normalized logarithmic Mel spectrum matrix. ; For each Mel band, extract a band energy sequence along the time axis; For each Mel band Determine the enhancement center threshold of this Mel band. ; The first generation is generated based on the skewness coefficient. Enhanced kurtosis factor in the Mel band ; For the normalized logarithmic Mel spectrum matrix By performing point-by-point smoothing nonlinear enhancement, a frequency band contrast enhancement spectral matrix is ​​obtained. ; Extracting short-time transient residuals through frequency band trend decomposition and cross-frequency band consistency constraints includes: For each Mel band Extract the frequency band contrast enhancement spectrum matrix along the time axis The corresponding frequency band energy sequence; Perform a moving median filter on the band energy sequence of each Mel band to obtain the trend estimation result. The two-dimensional matrix composed of the trend estimation results of all Mel bands is denoted as follows. ; Represents the frequency band contrast enhancement spectrum matrix The corresponding slow-changing trend matrix; Enhance the frequency band contrast spectral matrix Subtract the slow-changing trend matrix Only the positive differences are retained to obtain the preliminary transient residual matrix. ; For the initial transient residual matrix Each position in Take the neighboring Mel band in the frequency direction and calculate the cross-band consistency score; The cross-band consistency score is mapped to a consistency weight, and then the consistency weight is used to adjust the initial transient residual matrix. Weighting is performed to obtain the transient residual spectrum matrix. ; For transient residual spectrum matrix Perform intra-sample linear normalization to enhance the band contrast of the spectral matrix. With transient residual spectrum matrix Stacking along the channel dimension yields a two-channel spectral input tensor. .

2. The method for detecting wind turbine bearing faults based on acoustic signature features according to claim 1, characterized in that, The process of converting the original time-domain vibration signal into an original logarithmic Mel spectrum matrix jointly represented by time frames and Mel frequency bands includes: Continuous vibration data is sliced ​​according to a fixed duration or a fixed number of sampling points to obtain multiple original vibration sample sequences; For each original vibration sample sequence, perform framing, windowing, and spectral transformation to obtain a short-time spectral sequence; The power spectrum of each analysis frame is input into the Mel filter bank to obtain the Mel band energy sequence; The initial log-Mel spectrogram matrix is ​​resampled along the time axis to a uniform number of frames to obtain the original log-Mel spectrogram matrix.

3. The method for detecting wind turbine bearing faults based on acoustic signature features according to claim 1, characterized in that, Iterative training and parameter optimization of the fault classification model include: Frequency band grouping and sensing feature extraction involves dividing the dual-channel spectrum input tensor into several frequency band groups along the frequency direction, extracting local time-frequency features from each group independently, and then adaptively generating weighting coefficients for each group based on the global frequency band statistics of the current sample. Finally, the samples are fused to obtain a frequency band selection sensing feature map. A time attention mechanism is constructed based on candidate impact time and average impact period. The frame-by-frame time feature sequence is extracted from the frequency band selection sensing feature map, and the transient residual spectrum matrix is ​​used to detect candidate impact time and estimate average impact period. The period prior bias matrix is ​​then injected into the time attention mechanism. The model is trained using a classification output and category-weighted training. The classification logic value is output through a fully connected classification layer, and the entire model is trained using category-weighted cross-entropy loss. The model parameters are updated using an adaptive moment estimation optimizer.

4. The method for detecting wind turbine bearing faults based on acoustic signature features according to claim 3, characterized in that, Frequency band grouping selection and perceptual feature extraction, specifically: Input the dual-channel spectrum into the tensor Classified by frequency direction One continuous frequency band group; Perform independent convolution extraction on each local input tensor to obtain the frequency band group feature map; Generate frequency band group weights based on the global frequency band statistics of the current sample; Multiply the feature map of each frequency band group by the corresponding frequency band group weight, and then stitch them together along the frequency direction in the original frequency order to obtain a weighted stitched feature map.

5. The method for detecting wind turbine bearing faults based on acoustic signature features according to claim 3, characterized in that, A time attention mechanism is constructed based on candidate impact times and average impact periods, specifically as follows: Frame-by-frame compression is performed on the frequency band selection sensing feature map to obtain the temporal feature sequence; Constructing a frame-by-frame impact intensity sequence based on the transient residual spectrum matrix; Estimate the average impact period based on the candidate impact time index set; Construct a periodic prior bias matrix based on the candidate impact time index set and the average impact period; Time feature series The query matrix, key matrix, and value matrix are obtained by passing through three linear transformation layers respectively. The time-enhanced feature sequence is obtained by weighting and summing the value matrix using the time attention weight matrix.

6. The method for detecting wind turbine bearing faults based on acoustic signature features according to claim 5, characterized in that, Frame-by-frame compression is performed on the frequency band selection sensing feature map to obtain the temporal feature sequence, specifically: Perform band selection sensing feature map along the channel direction Convolution maps the fused output channel dimension to the frame-by-frame feature dimension, and the channel number to the target dimension, resulting in a dimension of size [missing information]. The characteristic tensor; Perform average pooling along the frequency direction for each time frame to obtain the time feature sequence. The size is This compresses the two-dimensional frequency band features into a one-dimensional time feature vector; in, This represents the total number of time frames after unification. This represents the number of discrete frequency bands in the frequency direction for a single sample. This indicates the feature dimension per frame.

7. The method for detecting wind turbine bearing faults based on acoustic signature features according to claim 5, characterized in that, The frame-by-frame impact intensity sequence is constructed based on the transient residual spectral matrix, specifically as follows: For the For each time frame, the one with the largest value is selected from all Mel bands of the transient residual spectrum matrix. Each frequency band residual, then... Calculate the average of the values; frame-by-frame impact intensity sequence Perform a moving average smoothing with a length of 3 or 5, then perform peak detection to obtain a set of candidate impact moment indices. Candidate impact moment index set This represents the set of time frame locations in the current sample where the suspected impact occurred.

Citation Information

Patent Citations

  • Rolling bearing fault diagnosis method and device and electronic equipment

    CN115791169A

  • Industrial cross-domain voiceprint fault diagnosis method based on characteristic decomposition and reconstruction

    CN119252284A