A wind turbine effective voiceprint enhancement method based on wind speed perception neural network
By using a voiceprint enhancement method based on a wind speed sensing neural network, the adaptability problem of voiceprint enhancement methods for wind turbines under wind speed variation is solved, ensuring the fidelity of fault characteristics and the reliability of diagnosis, and providing a credibility assessment of the enhancement results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 华电重庆新能源有限公司
- Filing Date
- 2026-06-23
- Publication Date
- 2026-07-21
AI Technical Summary
Existing acoustic enhancement methods for wind turbines lack the ability to perceive changes in wind speed conditions, making it difficult to adaptively adjust enhancement strategies. Furthermore, neural network methods do not fully utilize wind speed information, resulting in the weakening of fault-related weak acoustic components during noise reduction, which reduces the reliability of fault diagnosis.
A wind speed sensing neural network-based method is adopted, which uses a wind speed condition-gated dual-branch neural network for acoustic text enhancement. The wind speed information is used to dynamically modulate the enhancement mask, and the fidelity loss of the fault-sensitive frequency band and the impact feature retention loss are combined to ensure the fidelity of the fault features of key components and the protection of transient impact components.
It achieves adaptive acoustic enhancement in non-stationary environments, preserves fault-related acoustic components to the maximum extent, improves the reliability and interpretability of wind turbine condition monitoring and fault diagnosis, and provides a credibility assessment of the enhancement results.
Smart Images

Figure CN122435951A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wind turbine condition monitoring and fault diagnosis technology, and in particular to an effective acoustic enhancement method for wind turbines based on neural networks. Background Technology
[0002] Wind turbines operate long-term in high-altitude, strong-wind, and variable-load environments, making their critical components, such as gearboxes, main shaft bearings, and generators, prone to wear, cracks, and spalling. Acoustic fingerprinting, with its advantages of non-contact acquisition and flexible installation, has become an important means of condition monitoring for wind turbines.
[0003] However, the acoustic signature of wind turbine operation is highly susceptible to interference from environmental wind noise, blade aerodynamic noise, and fluctuations in operating conditions, exhibiting significant non-stationary characteristics. Especially under conditions of high wind speed or rapid wind speed changes, strong low-frequency wind noise components can severely mask weak acoustic components related to faults such as abnormal gear meshing, localized bearing impact, and generator friction, greatly reducing the reliability of subsequent feature extraction and fault identification.
[0004] Existing methods for enhancing or reducing acoustic signatures mainly include fixed-parameter filtering, wavelet denoising, empirical mode decomposition, spectral subtraction, and neural network-based enhancement methods. While these methods can suppress noise to some extent, they still have significant shortcomings: traditional methods lack the ability to perceive changes in wind speed conditions, making it difficult to adaptively adjust the enhancement strategy according to the actual wind noise intensity; some neural network methods mainly rely on the acoustic signature signal itself, failing to fully utilize wind speed information and its short-term changing trends, which are strongly correlated with wind noise; at the same time, existing methods focus more on the overall noise reduction effect, lacking specific fidelity constraints on fault-sensitive frequency bands and transient impact characteristics, which may inadvertently weaken effective fault acoustic signature components during the noise reduction process. Summary of the Invention
[0005] In view of this, the purpose of this invention is to provide an effective acoustic enhancement method for wind turbines based on a wind speed sensing neural network. This method aims to address the shortcomings of existing wind turbine acoustic enhancement methods, such as the lack of ability to sense changes in wind speed conditions, the difficulty in adaptively adjusting the enhancement strategy according to wind noise intensity, the failure of existing neural network methods to fully utilize wind speed information, and the tendency to weaken fault-related weak acoustic components during noise reduction. The goal is to achieve the goal of suppressing wind noise while preserving fault-related acoustic components to the maximum extent, thereby improving the reliability of subsequent condition monitoring and fault diagnosis.
[0006] The present invention provides an effective acoustic signature enhancement method for wind turbine generators based on a wind speed sensing neural network, comprising the following steps:
[0007] 1) Collect raw acoustic signature signals and wind speed data during the operation of the wind turbine;
[0008] 2) Divide the original voiceprint signal into several voiceprint segments, obtain the wind speed value, wind speed change rate and low frequency wind noise energy ratio of each voiceprint segment associated with the voiceprint segment, then calculate the wind noise pollution score of each voiceprint segment, and then calculate the voiceprint quality score based on the wind noise pollution score; and use the corresponding wind speed value, wind speed change rate and voiceprint quality score to construct the wind speed condition vector of each voiceprint segment.
[0009] 3) Effective voiceprint enhancement processing of voiceprint segments is performed using a wind speed condition-gated dual-branch neural network; the wind speed condition-gated dual-branch neural network includes a voiceprint spectrogram encoding branch, a wind speed condition encoding branch, a wind speed condition-gated fusion module, and an enhancement mask decoding module; the wind speed condition-gated dual-branch neural network is trained using a loss function that includes a fault-sensitive frequency band fidelity loss, the fault-sensitive frequency band fidelity loss is used to limit the degree of suppression of acoustic components in the preset fault-sensitive frequency band during the enhancement process, and the impact feature preservation loss is used to constrain the change in the impact index of the voiceprint signal before and after enhancement;
[0010] The enhancement process includes:
[0011] The speaker spectrum corresponding to the speaker segment to be enhanced is input into the speaker spectrum encoding branch to extract the speaker time-frequency features; a speaker segment is selected with a set time window, and the speaker segment to be enhanced is located at the center of the time window; the wind speed condition vectors of the selected speaker segment are arranged in time order to form wind speed condition time-series features, and input into the wind speed condition encoding branch to extract wind speed condition features; the wind speed condition gating fusion module generates modulation weights based on the wind speed condition features to dynamically modulate the speaker time-frequency features to generate joint features; then the enhancement mask decoding module generates an enhancement mask based on the joint features; the original speaker time-frequency spectrum of the enhanced speaker segment is weighted using the enhancement mask to obtain the enhanced speaker time-frequency spectrum;
[0012] 4) Reconstruct the time spectrum of the enhanced voiceprint segments into time-domain voiceprint segments, and calculate the credibility weight of each time-domain voiceprint segment based on the voiceprint quality score and the credibility of the enhancement mask. Then, perform weighted overlap and summation on each time-domain voiceprint segment to obtain the final continuously enhanced voiceprint signal.
[0013] Furthermore, step 2) of dividing the original voiceprint signal into several voiceprint segments includes:
[0014] The original voiceprint signal is segmented according to a preset frame length L and frame shift R. The k-th voiceprint segment is represented as:
[0015]
[0016] Where k = 0, 1, 2, ... K is the total number of segments;
[0017] After obtaining the k-th voiceprint segment, a window function is applied to it to obtain the windowed voiceprint segment:
[0018]
[0019] in, For window functions.
[0020] Further, in step 2), the wind speed value associated with the voiceprint segment is the wind speed value corresponding to the center time of the voiceprint segment; if the center time of the voiceprint segment is located between two adjacent wind speed sampling times, then the wind speed value is obtained by linear interpolation.
[0021] Step 2) The rate of change of wind speed associated with the voiceprint segment is calculated using the following formula:
[0022]
[0023] in, The wind speed change rate associated with the k-th voiceprint segment. The wind speed value associated with the k-th voiceprint segment. This represents the wind speed value associated with the (k-1)th voiceprint segment. Let f be the time interval between the center times of the k-th voiceprint segment and the (k-1)-th voiceprint segment; if each voiceprint segment is constructed according to a fixed frame shift R, the voiceprint signal sampling rate is f. s ,but Set the wind speed change rate associated with the first voiceprint segment to 0;
[0024] Step 2) The proportion of low-frequency wind noise energy in the acoustic signature segment is calculated using the following formula:
[0025]
[0026] in, The value represents the proportion of low-frequency wind noise energy, where t is the short-time frame index and f is the frequency index. Let f be the time spectrum of the k-th voiceprint segment. c f is the low-frequency wind noise cutoff frequency. max To analyze the upper frequency limit, ε is a minimal constant to prevent the denominator from being zero;
[0027] Step 2) The wind noise pollution score is calculated using the following formula:
[0028]
[0029] Among them, P k For the wind noise pollution score of the k-th voiceprint segment, a1, a2, and a3 are weighting coefficients, b is the bias term, and σ(·) is the Sigmoid function. The normalized wind speed value. This represents the normalized rate of change of wind speed. The normalized low-frequency wind noise energy proportion is defined as follows: wind speed value represents the strength of the wind conditions in which the current voiceprint segment is located; wind speed change rate represents the intensity of wind speed fluctuations between adjacent voiceprint segments; and low-frequency wind noise energy proportion represents the proportion of low-frequency wind noise components in the current voiceprint segment within the overall voiceprint energy. These three parameters reflect the likelihood of a voiceprint segment being contaminated by wind noise from three perspectives: external wind speed conditions, the degree of fluctuation in these conditions, and the voiceprint spectrum energy distribution. They do not imply a single causal relationship between the three. The weighting coefficients a1, a2, a3, and the bias term b can be set or trained based on historical voiceprint samples, wind speed data, manually annotated segment quality results, or validation set enhancement effects. These parameters are used to adjust the influence of each factor on the wind noise contamination score. After mapping using the Sigmoid function, the wind noise contamination score P... k The value range of P is [0,1]. k The closer the value is to 1, the higher the degree of wind noise pollution in the voiceprint segment; P k The closer it is to 0, the lower the degree of wind noise pollution in the voiceprint segment.
[0030] Step 2) The voiceprint quality score is calculated using the following formula:
[0031]
[0032] Among them, Q k This represents the quality score of the k-th voiceprint segment. Since the quality of a voiceprint segment is inversely related to the degree of wind noise pollution—that is, the higher the degree of wind noise pollution, the easier it is for the effective mechanoacoustic components in the voiceprint segment to be masked, resulting in lower segment quality; conversely, the lower the degree of wind noise pollution, the easier it is for the effective voiceprint components in the voiceprint segment to be preserved, resulting in higher segment quality—this invention adopts… As a voiceprint quality score. Therefore, Q k The value range of Q is also [0,1]. k The closer the value is to 1, the higher the validity of the fragment; Q k The closer the value is to 0, the more severe the wind noise pollution and the lower the effectiveness of the segment. The voiceprint quality score Q... k Subsequently, as a component of the wind speed condition vector, it is used to guide the wind speed condition-gated dual-branch neural network to perform differentiated enhancement processing on different quality segments, and can also serve as the basis for subsequent segment confidence weight calculation.
[0033] Step 2) The expression for the wind speed condition vector is:
[0034]
[0035] Among them, z k is the wind speed condition vector for the k-th acoustic signature segment.
[0036] Furthermore, the loss function described in step 3) is expressed as follows:
[0037]
[0038] Where λ1 and λ2 are the loss weight coefficients;
[0039] L enh To enhance the reconstruction loss, L enh The mean square error of the amplitude spectrum or the mean square error of the logarithmic power spectrum is used.
[0040] Mean square error of amplitude spectrum:
[0041]
[0042] Logarithmic power spectrum mean square error:
[0043]
[0044] Where N and F are the total number of short-time frame indices and the total number of frequency indices of the time spectrum, respectively. This is the complex spectrum estimation of the k-th voiceprint segment in the time-frequency domain after processing by the neural network. For reference only;
[0045] The fidelity loss in the fault-sensitive frequency band is calculated as follows:
[0046]
[0047] Where ε is a minimal constant to prevent the denominator from being zero; The set of fault-sensitive frequency bands is expressed as follows:
[0048]
[0049] Each element Ω in the fault-sensitive frequency band set Ω i All of these are preset continuous frequency bands that contain one or a group of potential fault characteristic frequencies of wind turbine units;
[0050] The loss is maintained due to the impact characteristics and is calculated as follows:
[0051]
[0052] in, This refers to the voiceprint fragment obtained in step 2). This represents the enhanced time-domain acoustic signature segment, the... The enhanced audioprint time spectrum was obtained by inverse short-time Fourier transform. This is a shock indicator.
[0053] Furthermore, the element Ω i The settings are based on the following fault characteristic frequencies and their surrounding neighborhoods:
[0054] The characteristic frequencies of inner ring failure, outer ring failure, rolling element failure, and cage failure of wind turbine bearings, as well as the frequency bands in which their harmonics and modulation sidebands are located;
[0055] The meshing frequencies of each gear in the wind turbine gearbox and their harmonics, as well as the frequency bands in which the modulation sidebands generated by the meshing frequencies are located;
[0056] The frequency band in which the main shaft and rotor of the wind turbine and their harmonics are located.
[0057] Furthermore, the impact index is kurtosis, peak factor, or impulse factor.
[0058] Furthermore, the acoustic signature spectrum mentioned in step 3) is obtained by the following formula:
[0059]
[0060] Among them, A k (t,f) represents the voiceprint spectrum corresponding to the kth voiceprint segment;
[0061] The expression for the time-series characteristics of wind speed conditions mentioned in step 3) is as follows:
[0062]
[0063] Where p is the time window radius; when <0 or k+p> When constructing a timing window, use boundary copying, zero padding, or simply use existing fragments.
[0064] The expression for generating the joint features mentioned in step 3) is as follows:
[0065]
[0066] Among them, H k This represents the joint feature corresponding to k voiceprint segments. This represents element-wise multiplication, where η is the gating coefficient. Indicates the time-frequency characteristics of the voiceprint; The modulation weights are generated as follows:
[0067]
[0068] Among them, W g Let b be the gated weight matrix. g Here, σ(·) is the bias term, and σ(·) is the Sigmoid activation function. Indicates wind speed operating conditions;
[0069] The complex spectrum estimation described in step 3) Calculated using the following formula:
[0070]
[0071] in, This represents the enhancement mask at the time-frequency point (t,f).
[0072] Further, step 4) reconstructs the time-domain audiogram spectrum of each enhanced audiogram segment into a time-domain audiogram segment, including:
[0073] Perform inverse short-time Fourier transform on the enhanced audioprint time spectrum:
[0074]
[0075] in, This represents the temporal voiceprint segment obtained after the k-th original voiceprint segment is enhanced by a neural network, where j = 0, 1, 2, ... ;
[0076] Time-domain audioprint fragments Apply window function :
[0077]
[0078] in, This indicates the enhanced voiceprint segment after windowing;
[0079] The confidence weight mentioned in step 4) is calculated using the following formula:
[0080]
[0081] Where, ω k ρ represents the credibility weight of the k-th enhanced voiceprint segment, γ is the weight adjustment coefficient, and its value ranges from 0 ≤ γ ≤ 1. k This indicates the mask reliability obtained by enhancing mask stability;
[0082] The continuously enhanced voiceprint signal mentioned in step 4) is calculated using the following formula:
[0083]
[0084] in, ε represents the final continuously enhanced voiceprint signal; n is the index of the discrete time series; K represents the total number of voiceprint segments; and ε is a minimal constant to prevent the denominator from being zero.
[0085] Furthermore, the effective acoustic signature enhancement method for wind turbine generators based on wind speed sensing neural networks also includes enhancing the credibility of the enhanced time-domain acoustic signature segment:
[0086]
[0087] Among them, T k Q represents the enhanced credibility of the k-th segment. k Indicates the voiceprint quality score, ρ k B indicates that the credibility of the mask is enhanced. k β1, β2, and β3 represent the fidelity of the fault-sensitive frequency band, and are weighting coefficients.
[0088] It also includes the overall enhanced credibility of the continuously enhanced voiceprint signal output:
[0089]
[0090] Where T represents the overall credibility of the final continuously enhanced voiceprint signal, ω k ε represents the credibility weight of the k-th enhanced voiceprint segment, and ε is a minimal constant to prevent the denominator from being zero;
[0091] Set an evaluation threshold. When the overall enhancement confidence T is consistently lower than the preset threshold within a preset time range, output a low confidence warning for the enhancement result.
[0092] The beneficial effects of this invention are:
[0093] 1. This invention uses wind speed and its rate of change directly as input to a neural network and dynamically influences the enhancement mask through a gating mechanism. This enables the voiceprint enhancement system to perceive and adapt to real-time wind conditions. Compared to traditional fixed algorithms or audio networks that ignore operating conditions, this invention can enhance noise reduction when wind speed increases sharply and prioritize fidelity when wind speed is stable, truly achieving adaptive processing in non-stationary environments.
[0094] 2. Addressing the specific needs of wind turbine fault diagnosis, this invention not only pursues noise reduction at the output level but also imposes targeted constraints at the training target level. Through a fault-sensitive frequency band fidelity loss function, the network is forced to retain characteristic frequency band energy related to faults in key components such as gears and bearings during the enhancement process; and through impact feature retention loss, transient impact components are protected. This effectively solves the industry problem of erroneously suppressing weak fault features during noise reduction, ensuring the "effectiveness" of the enhanced acoustic signature for subsequent fault diagnosis.
[0095] 3. This invention innovatively proposes a reliable weight reconstruction mechanism based on quality scoring and mask stability. This is not a simple signal splicing, but a data fusion and decision-making process that automatically reduces the contribution of low-quality or unstable segments to the final result, thereby significantly improving the continuity and stability of the overall output voiceprint. Simultaneously, the overall enhancement reliability T provides maintenance personnel with a quantitative assessment of the enhancement result quality, enhancing the system's interpretability and practicality.
[0096] 4. This invention forms a complete, closed-loop technical process from data synchronization, quality assessment, perception enhancement to reliable reconstruction. It deeply integrates signal processing, operating condition perception, deep learning models, and domain knowledge (fault frequency), providing a systematic solution for acoustic monitoring of wind turbine units that emphasizes both enhancement and fidelity, rather than a single algorithm module. Therefore, its final technical effect is reflected in the improvement of data quality at the front end of the entire condition monitoring chain, providing a higher quality and more reliable data foundation for subsequent fault identification and anomaly detection. Attached Figure Description
[0097] Figure 1 This is a schematic diagram of a dual-branch neural network structure for wind speed sensing.
[0098] Figure 2 A schematic diagram to enhance voiceprint reconstruction and high-reliability output. Detailed Implementation
[0099] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0100] The effective acoustic signature enhancement method for wind turbines based on a wind speed sensing neural network in this embodiment includes the following steps:
[0101] 1) Collect raw acoustic signature signals and wind speed data during wind turbine operation. During wind turbine operation, raw acoustic signature signals can be collected using acoustic sensors placed at suitable locations inside the nacelle, near the gearbox, near the generator, inside the tower, or outside the turbine, and wind speed data for the corresponding time period can be acquired simultaneously. The wind speed data can be obtained from wind speed sensors, anemometers, or the existing SCADA monitoring system of the wind turbine. Let the collected raw acoustic signature signal be represented as:
[0102]
[0103] Where x[n] represents the discrete sound pressure value corresponding to the nth sampling point, L s This represents the total number of sampling points obtained within the acquisition time T. Let the synchronously acquired wind speed sequence be:
[0104]
[0105] Where v[m] represents the wind speed value collected at the m-th time, and M is the total number of wind speed sampling points.
[0106] 2) Divide the original voiceprint signal into several voiceprint segments. Since the voiceprint sampling frequency is usually much higher than the wind speed sampling frequency, it is necessary to perform time matching between the wind speed information and the voiceprint segments. In this step, the original voiceprint signal is divided into several voiceprint segments, including:
[0107] The original voiceprint signal is segmented according to a preset frame length L and frame shift R. The k-th voiceprint segment is represented as:
[0108]
[0109] Where k = 0, 1, 2, ... K is the total number of segments;
[0110] To reduce spectral leakage caused by segmentation and truncation, after obtaining the k-th voiceprint segment, a window function is applied to it. Hamming, Hanning, or other smoothing window functions can be used to obtain the windowed voiceprint segment.
[0111]
[0112] in, For window functions.
[0113] Obtain the wind speed value, wind speed change rate, and low-frequency wind noise energy ratio associated with each soundprint segment. Then calculate the wind noise pollution score for each soundprint segment and calculate the soundprint quality score based on the wind noise pollution score. Finally, construct the wind speed condition vector for each soundprint segment using the corresponding wind speed value, wind speed change rate, and soundprint quality score.
[0114] Establishing a correspondence between acoustic signature segments and wind speed conditions enables the subsequent neural network enhancement process to perceive wind speed changes in the actual operating environment of the wind turbine. The wind speed value v associated with the acoustic signature segment mentioned in this step... k The wind speed value corresponds to the center time of the voiceprint segment; if the center time of the voiceprint segment is located between two adjacent wind speed sampling times, the wind speed value is obtained using linear interpolation. This yields the original voiceprint segment after windowing. With wind speed value Correspondence:
[0115]
[0116] The wind speed change rate is used to reflect the impact of wind speed fluctuations on the voiceprint quality. The wind speed change rate associated with the voiceprint segment mentioned in this step is calculated using the following formula:
[0117]
[0118] in, The wind speed change rate associated with the k-th voiceprint segment. The wind speed value associated with the k-th voiceprint segment. This represents the wind speed value associated with the (k-1)th voiceprint segment. Let f be the time interval between the center times of the k-th voiceprint segment and the (k-1)-th voiceprint segment; if each voiceprint segment is constructed according to a fixed frame shift R, the voiceprint signal sampling rate is f. s ,but The wind speed change rate associated with the first voiceprint segment is 0.
[0119] Wind noise in the operating environment of wind turbines is typically stronger in the low-frequency range. The proportion of low-frequency wind noise energy in the acoustic signature segment mentioned in this step is calculated using the following formula:
[0120]
[0121] in, The low-frequency wind noise energy percentage is represented by t, where t is the short-time frame index and f is the frequency index. c f is the low-frequency wind noise cutoff frequency. max To analyze the upper frequency limit, ε is a very small constant to prevent the denominator from being zero; for example, it can be ε = 10. -6 ε=10 -8 Or ε=10 -10 wait; Let be the time spectrum of the k-th voiceprint segment, which is obtained by performing a short-time Fourier transform on the voiceprint segment. The calculation expression is as follows:
[0122]
[0123] Where t represents the short-time frame index and f represents the frequency index. This time-spectrum analysis allows us to analyze the energy distribution of audioprint segments across different frequency bands.
[0124] The wind noise pollution score mentioned in this step is calculated using the following formula:
[0125]
[0126] Among them, P k For the wind noise pollution score of the k-th voiceprint segment, a1, a2, and a3 are weighting coefficients. In this embodiment, a1=0.2, a2=0.5, and a3=1 are set; b is a bias term, which is set in this embodiment to b= σ(·) is the Sigmoid function. The normalized wind speed value. This represents the normalized rate of change of wind speed. The normalized low-frequency wind noise energy proportion is defined as follows: Wind speed value represents the strength of the wind conditions in which the current voiceprint segment is located; wind speed change rate represents the intensity of wind speed fluctuations between adjacent voiceprint segments; and low-frequency wind noise energy proportion represents the percentage of low-frequency wind noise components in the overall voiceprint energy within the current voiceprint segment. These three parameters reflect the likelihood of a voiceprint segment being contaminated by wind noise from three perspectives: external wind speed conditions, the degree of fluctuation in these conditions, and the voiceprint spectrum energy distribution. They do not imply a single causal relationship between these three parameters. The weighting coefficients a1, a2, a3, and the bias term b can be set or trained based on historical voiceprint samples, wind speed data, manually annotated segment quality results, or validation set enhancement effects. These parameters are used to adjust the influence of each factor on the wind noise contamination score. After mapping using the Sigmoid function, the wind noise contamination score P... k The value range of P is [0,1]. k The closer the value is to 1, the higher the degree of wind noise pollution in the voiceprint segment; P k The closer it is to 0, the lower the degree of wind noise pollution in the voiceprint segment.
[0127] The voiceprint quality score mentioned in this step is calculated using the following formula:
[0128]
[0129] Among them, Q k This represents the quality score of the k-th voiceprint segment. Since the quality of a voiceprint segment is inversely related to the degree of wind noise pollution—that is, the higher the degree of wind noise pollution, the easier it is for the effective mechanoacoustic components in the voiceprint segment to be masked, resulting in lower segment quality; conversely, the lower the degree of wind noise pollution, the easier it is for the effective voiceprint components in the voiceprint segment to be preserved, resulting in higher segment quality—this invention adopts… As a voiceprint quality score. Therefore, Q k The value range of Q is also [0,1]. k The closer the value is to 1, the higher the validity of the fragment; Q k The closer the value is to 0, the more severe the wind noise pollution and the lower the effectiveness of the segment. The voiceprint quality score Q... k Subsequently, as a component of the wind speed condition vector, it is used to guide the wind speed condition-gated dual-branch neural network to perform differentiated enhancement processing on different quality segments, and can also serve as the basis for subsequent segment confidence weight calculation.
[0130] The expression for the wind speed condition vector mentioned in this step is:
[0131]
[0132] Among them, z k is the wind speed condition vector for the k-th acoustic signature segment.
[0133] By comprehensively considering wind speed magnitude, wind speed change rate, and the proportion of low-frequency wind noise energy in the acoustic signature, a continuous acoustic signature quality score is constructed, enabling the subsequent enhancement process to adaptively process based on the quality differences of the segments. By constructing a wind speed condition vector in the form of "wind speed value - wind speed change rate - quality score", the subsequent neural network enhancement process can not only perceive the current wind speed level, but also wind speed fluctuations and the effectiveness of acoustic signature segments, thereby improving the adaptive capability of acoustic signature enhancement in complex wind noise environments of wind turbine units.
[0134] 3) Effective voiceprint enhancement processing is performed on voiceprint segments using a wind speed condition-gated dual-branch neural network. The wind speed condition-gated dual-branch neural network includes a voiceprint spectrogram encoding branch, a wind speed condition encoding branch, a wind speed condition-gated fusion module, and an enhancement mask decoding module. In this embodiment, the voiceprint spectrogram encoding branch adopts a U-Net-type convolutional encoder-decoder network to extract multi-scale time-frequency features from the voiceprint spectrogram and retain shallow spectral details. The wind speed condition encoding branch adopts a temporal convolutional network (TCN) to perform temporal modeling of wind speed values, wind speed change rates, and voiceprint quality scores corresponding to multiple consecutive voiceprint segments, and to extract short-term variation trends and fluctuation characteristics of wind speed conditions.
[0135] The wind speed-gated dual-branch neural network is trained using a loss function that includes a fault-sensitive band fidelity loss and an impact feature preservation loss. The fault-sensitive band fidelity loss limits the degree to which the network suppresses acoustic components within a preset fault-sensitive band during enhancement, while the impact feature preservation loss constrains the changes in the impact index of the acoustic signature signal before and after enhancement. Specifically, the loss function is expressed as follows:
[0136]
[0137] Where λ1 and λ2 are the loss weight coefficients.
[0138] The fidelity loss in the fault-sensitive frequency band is calculated as follows:
[0139]
[0140] Where ε is a minimal constant to prevent the denominator from being zero; The set of fault-sensitive frequency bands is expressed as follows:
[0141]
[0142] Each element Ω in the fault-sensitive frequency band set Ω i All are preset continuous frequency bands containing one or more potential fault characteristic frequencies of wind turbine units. In this embodiment, the element Ω iThe settings are based on the following fault characteristic frequencies and their surrounding neighborhoods:
[0143] The characteristic frequencies of inner ring failure, outer ring failure, rolling element failure, and cage failure of wind turbine bearings, as well as the frequency bands in which their harmonics and modulation sidebands are located;
[0144] The meshing frequencies of each gear in the wind turbine gearbox and their harmonics, as well as the frequency bands in which the modulation sidebands generated by the meshing frequencies are located;
[0145] The frequency band in which the main shaft and rotor of the wind turbine and their harmonics are located.
[0146] This is the complex spectrum estimation of the k-th voiceprint segment in the time-frequency domain after processing by the neural network.
[0147] The loss is maintained due to the impact characteristics and is calculated as follows:
[0148]
[0149] in, The voiceprint segment obtained from step 2) is the original voiceprint segment. For the enhanced time-domain audioprint segment, The impact factor is a kurtosis, peak factor, or impulse factor. In this embodiment, the impact factor is kurtosis, peak factor, or impulse factor.
[0150] L enh To enhance the reconstruction loss, it is used to constrain the enhanced speaker signal of the neural network. Compared with reference samples less affected by wind noise The purpose of identifying the differences between them is to guide the neural network to learn how to recover effective voiceprint components from noisy signals. When there is a lack of reference samples that are less affected by wind noise, high-quality voiceprint segments with low wind speeds from the same generator unit can be selected as reference samples. Alternatively, pseudo-reference samples can be constructed using segments with high voiceprint quality scores. This is to achieve weakly supervised training. In specific implementation, L... enh The mean square error of the amplitude spectrum or the mean square error of the logarithmic power spectrum can be used.
[0151] Mean square error of amplitude spectrum:
[0152]
[0153] Logarithmic power spectrum mean square error:
[0154]
[0155] Where N and F are the total number of short-time frame indices and the total number of frequency indices of the time spectrum, respectively. This is the complex spectrum estimation of the k-th voiceprint segment in the time-frequency domain after processing by the neural network. For reference.
[0156] The enhancement process includes:
[0157] Input the speaker spectrogram corresponding to the speaker fragment to be enhanced into the speaker spectrogram encoding branch to obtain the speaker time-frequency features:
[0158]
[0159] in, E represents the time-frequency characteristics of a voiceprint. x (·) indicates a U-Net-type convolutional encoder-decoder network. The expression for the voiceprint spectrum corresponding to the k-th voiceprint segment is as follows:
[0160]
[0161] It is used to characterize the energy distribution of the acoustic signature of a wind turbine in both time and frequency dimensions.
[0162] A voiceprint segment is selected within a set time window, with the segment to be enhanced centered within the time window. The wind speed condition vectors of the selected voiceprint segments are arranged in chronological order to form a wind speed condition time-series feature, which is then input into the wind speed condition coding branch to obtain the wind speed condition feature. The expression for the wind speed condition time-series feature is as follows:
[0163]
[0164] Where p is the time window radius; when <0 or k+p> When constructing a timing window, use boundary copying, zero padding, or simply use existing fragments.
[0165] The wind speed operating condition characteristics are generated as follows:
[0166]
[0167] Among them, E v (·) denotes a temporal convolutional network.
[0168] The wind speed condition gating fusion module generates modulation weights based on the wind speed condition characteristics and dynamically modulates the acoustic signature time-frequency characteristics to generate joint features. The expression for generating the joint features is as follows:
[0169]
[0170] Among them, H k This represents the joint feature corresponding to k voiceprint segments. This represents element-wise multiplication, where η is the gating coefficient. Indicates the time-frequency characteristics of the voiceprint; The modulation weights are generated as follows:
[0171]
[0172] Among them, W g Let b be the gated weight matrix. g Here, σ(·) is the bias term, and σ(·) is the Sigmoid activation function. This indicates the characteristics of wind speed operating conditions.
[0173] The enhanced mask decoding module then generates an enhanced mask based on the joint features, as follows:
[0174] Where D(·) represents the enhanced mask decoding network, M k (t,f) represents the enhancement mask of the k-th voiceprint segment at the time-frequency point (t,f). The value of the enhancement mask ranges from 0 to 1. The closer the value is to 1, the more the time-frequency component should be preserved; the closer the value is to 0, the more the time-frequency component should be suppressed.
[0175] Using enhanced masks The original time spectrum of the enhanced voiceprint segment By weighting, we obtain the complex spectrum estimate. :
[0176]
[0177] Due to enhanced mask It is generated by combining acoustic signature spectrogram features and wind speed condition gating features. Therefore, it can dynamically change with changes in wind speed value, wind speed change rate and segment quality score, thereby achieving differentiated enhancement of different wind speed conditions, different wind noise pollution levels and different acoustic signature effective segments.
[0178] 4) Reconstruct the time spectrum of the enhanced voiceprint segments into time-domain voiceprint segments, and calculate the credibility weight of each time-domain voiceprint segment based on the voiceprint quality score and the credibility of the enhancement mask. Then, perform weighted overlap and summation on each time-domain voiceprint segment to obtain the final continuously enhanced voiceprint signal.
[0179] This step involves reconstructing the time-domain audioprint segment from the enhanced audioprint spectrum of each audioprint segment, including:
[0180] Perform inverse short-time Fourier transform on the enhanced audioprint time spectrum:
[0181]
[0182] in, This represents the temporal voiceprint segment obtained after the k-th original voiceprint segment is enhanced by a neural network, where j = 0, 1, 2, ... .
[0183] To ensure good smoothness at the segment connections in the reconstructed continuous acoustic signature signal, the time-domain acoustic signature segments are... Apply window function :
[0184]
[0185] in, This indicates the enhanced voiceprint segment after windowing.
[0186] The confidence weight mentioned in this step is calculated using the following formula:
[0187]
[0188] Where, ω k ρ represents the credibility weight of the k-th enhanced voiceprint segment, γ is the weight adjustment coefficient, and its value ranges from 0 ≤ γ ≤ 1. k This indicates the mask reliability obtained by enhancing mask stability. The ρ k Based on the enhanced mask matrix M k The degree of variation between adjacent time frames or adjacent frequency points is determined and can be normalized to the range [0,1]. For example, the mask confidence level can be calculated using the following method:
[0189] For an enhancement mask matrix M of size T×F k Calculate the variance of the local neighborhood (e.g., a 3×3 window) at each time frequency point (t,f). ;
[0190] The average local variance of the mask is obtained by averaging the local variances across all time-frequency points.
[0191]
[0192] Normalization using an exponential decay function:
[0193]
[0194] This is a scaling factor greater than 0, used to adjust the sensitivity to the degree of fluctuation. In this embodiment, it is taken as... =1. The value of can be adjusted according to the actual data distribution and the required stringency of the judgment. For example, if it is desired that the model be more sensitive to mask fluctuations (i.e., small fluctuations lead to a significant decrease in confidence), the value can be appropriately increased. The value (e.g., taking) = 2.0 or 5.0); conversely, if you wish to tolerate larger fluctuations, you can reduce the value. The value (e.g., taking) = 0.5).
[0195] The higher the voiceprint quality score and the more stable the enhancement mask, the greater the credibility weight of the segment; conversely, the contribution of the segment to the final reconstruction result is reduced.
[0196] The continuously enhanced voiceprint signal mentioned in step 4) is calculated using the following formula:
[0197]
[0198] in, ε represents the final continuously enhanced voiceprint signal; n is the index of the discrete time series; K represents the total number of voiceprint segments; and ε is a minimal constant to prevent the denominator from being zero.
[0199] As a further improvement to the above embodiments, the effective acoustic signature enhancement method for wind turbine generators based on wind speed sensing neural networks also includes enhancing the credibility of the enhanced time-domain acoustic signature segment:
[0200]
[0201] Among them, T k Q represents the enhanced credibility of the k-th segment, where β1, β2, and β3 are weighting coefficients. In this embodiment, β1 = 0.2, β2 = 0.3, and β3 = 0.5. k Indicates the voiceprint quality score, ρ k B indicates that the credibility of the mask is enhanced. k In this embodiment, B represents the fault-sensitive frequency band fidelity. k The specific expression is:
[0202]
[0203] in, ε represents the fault-sensitive frequency band fidelity loss corresponding to the k-th voiceprint segment, C is the normalization coefficient, which can be 1; ε is a minimal constant to prevent the denominator from being zero. The closer the value is to 1, the higher the fidelity.
[0204] It also includes the overall enhanced credibility of the continuously enhanced voiceprint signal output:
[0205]
[0206] Where T represents the overall credibility of the final continuously enhanced voiceprint signal, ω k ε represents the credibility weight of the k-th enhanced voiceprint segment, and ε is a minimal constant to prevent the denominator from being zero.
[0207] In this embodiment, T represents the self-evaluation of the current enhancement process. An evaluation threshold is set, and when the overall enhancement confidence T remains below the preset threshold for a continuous period of time, a low-confidence enhancement result prompt is output, indicating that the acoustic sensor status may need to be checked or specific operating conditions may need to be addressed. The long-term collected T-value data can be used to analyze the model's performance under different wind speeds and different turbine units, providing data guidance for iterative optimization of the model (such as adding training data for weak operating conditions).
[0208] The overall enhanced confidence level T also serves as a confidence input for downstream condition monitoring modules. Downstream diagnostic models or analysts can use the T value to weight the confidence level of their diagnostic conclusions. For example, when the T value is high, fault warnings based on the enhanced acoustic signature can be fully trusted; when the T value is low, it may be necessary to combine data from other sensors (such as vibration sensors) for comprehensive judgment, or trigger manual review. This can improve the decision-making intelligence and reliability of the entire condition monitoring system.
[0209] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for effective acoustic signature enhancement of wind turbine generators based on a wind speed sensing neural network, characterized in that: Includes the following steps: 1) Collect raw acoustic signature signals and wind speed data during the operation of the wind turbine; 2) Divide the original voiceprint signal into several voiceprint segments, obtain the wind speed value, wind speed change rate and low frequency wind noise energy ratio of each voiceprint segment associated with the voiceprint segment, then calculate the wind noise pollution score of each voiceprint segment, and then calculate the voiceprint quality score based on the wind noise pollution score; and use the corresponding wind speed value, wind speed change rate and voiceprint quality score to construct the wind speed condition vector of each voiceprint segment. 3) Effective voiceprint enhancement processing of voiceprint segments is performed using a wind speed condition-gated dual-branch neural network; the wind speed condition-gated dual-branch neural network includes a voiceprint spectrogram encoding branch, a wind speed condition encoding branch, a wind speed condition-gated fusion module, and an enhancement mask decoding module. The wind speed condition-gated dual-branch neural network is trained using a loss function that includes fault-sensitive band fidelity loss and impact feature retention loss. The fault-sensitive band fidelity loss is used to limit the degree of suppression of acoustic components in the preset fault-sensitive band during the enhancement process, and the impact feature retention loss is used to constrain the change in the impact index of the acoustic signal before and after enhancement. The enhancement process includes: The speaker spectrum corresponding to the speaker segment to be enhanced is input into the speaker spectrum encoding branch to extract the speaker time-frequency features; a speaker segment is selected with a set time window, and the speaker segment to be enhanced is located at the center of the time window; the wind speed condition vectors of the selected speaker segment are arranged in time order to form wind speed condition time-series features, and input into the wind speed condition encoding branch to extract wind speed condition features; the wind speed condition gating fusion module generates modulation weights based on the wind speed condition features to dynamically modulate the speaker time-frequency features to generate joint features; then the enhancement mask decoding module generates an enhancement mask based on the joint features; the original speaker time-frequency spectrum of the enhanced speaker segment is weighted using the enhancement mask to obtain the enhanced speaker time-frequency spectrum; 4) Reconstruct the time spectrum of the enhanced voiceprint segments into time-domain voiceprint segments, and calculate the credibility weight of each time-domain voiceprint segment based on the voiceprint quality score and the credibility of the enhancement mask. Then, perform weighted overlap and summation on each time-domain voiceprint segment to obtain the final continuously enhanced voiceprint signal.
2. The method for effective acoustic signature enhancement of wind turbine generators based on wind speed sensing neural networks according to claim 1, characterized in that: Step 2) Dividing the original voiceprint signal into several voiceprint segments includes: The original voiceprint signal is segmented according to a preset frame length L and frame shift R. The k-th voiceprint segment is represented as: Where k = 0, 1, 2, ... K is the total number of segments; After obtaining the k-th voiceprint segment, a window function is applied to it to obtain the windowed voiceprint segment: in, For window functions.
3. The method for effective acoustic signature enhancement of wind turbine generators based on wind speed sensing neural networks according to claim 1, characterized in that: Step 2) The wind speed value associated with the voiceprint segment is the wind speed value corresponding to the center time of the voiceprint segment; if the center time of the voiceprint segment is located between two adjacent wind speed sampling times, the wind speed value is obtained by linear interpolation. Step 2) The rate of change of wind speed associated with the voiceprint segment is calculated using the following formula: in, The wind speed change rate associated with the k-th voiceprint segment. The wind speed value associated with the k-th voiceprint segment. This represents the wind speed value associated with the (k-1)th voiceprint segment. The time interval between the center moments of the k-th voiceprint segment and the (k-1)-th voiceprint segment; Step 2) The proportion of low-frequency wind noise energy in the acoustic signature segment is calculated using the following formula: in, The value represents the proportion of low-frequency wind noise energy, where t is the short-time frame index and f is the frequency index. Let f be the time spectrum of the k-th voiceprint segment. c f is the low-frequency wind noise cutoff frequency. max To analyze the upper frequency limit, ε is a minimal constant to prevent the denominator from being zero; Step 2) The wind noise pollution score is calculated using the following formula: Among them, P k For the wind noise pollution score of the k-th voiceprint segment, a1, a2, and a3 are weighting coefficients, b is the bias term, and σ(·) is the Sigmoid function. The normalized wind speed value. This represents the normalized rate of change of wind speed. The normalized proportion of low-frequency wind noise energy; Step 2) The voiceprint quality score is calculated using the following formula: Among them, Q k This represents the quality score of the k-th voiceprint segment; Step 2) The expression for the wind speed condition vector is: Among them, z k is the wind speed condition vector for the k-th acoustic signature segment.
4. The method for effective acoustic signature enhancement of wind turbine generators based on wind speed sensing neural networks according to claim 3, characterized in that: The loss function described in step 3) is expressed as follows: Where λ1 and λ2 are the loss weight coefficients; L enh To enhance the reconstruction loss, L enh The mean square error of the amplitude spectrum or the mean square error of the logarithmic power spectrum is used, as follows: Mean square error of amplitude spectrum: Logarithmic power spectrum mean square error: Where N and F are the total number of short-time frame indices and the total number of frequency indices of the time spectrum, respectively. This is the complex spectrum estimation of the k-th voiceprint segment in the time-frequency domain after processing by the neural network. For reference only; The fidelity loss in the fault-sensitive frequency band is calculated as follows: Where ε is a minimal constant to prevent the denominator from being zero; The set of fault-sensitive frequency bands is expressed as follows: Each element Ω in the fault-sensitive frequency band set Ω i All of these are preset continuous frequency bands that contain one or a group of potential fault characteristic frequencies of wind turbine units; The loss is maintained due to the impact characteristics and is calculated as follows: in, The voiceprint fragment obtained in step 2); This represents the enhanced time-domain acoustic signature segment, the... The enhanced audioprint time spectrum was obtained by inverse short-time Fourier transform. This is a shock indicator.
5. The method for effective acoustic signature enhancement of wind turbine generators based on wind speed sensing neural networks according to claim 4, characterized in that: The element Ω i The settings are based on the following fault characteristic frequencies and their surrounding neighborhoods: The characteristic frequencies of inner ring failure, outer ring failure, rolling element failure, and cage failure of wind turbine bearings, as well as the frequency bands in which their harmonics and modulation sidebands are located; The meshing frequencies of each gear in the wind turbine gearbox and their harmonics, as well as the frequency bands in which the modulation sidebands generated by the meshing frequencies are located; The frequency band in which the main shaft and rotor of the wind turbine and their harmonics are located.
6. The method for effective acoustic signature enhancement of wind turbine generators based on wind speed sensing neural networks according to claim 4, characterized in that: The impact index is kurtosis, peak factor, or impulse factor.
7. The method for effective acoustic signature enhancement of wind turbine generators based on wind speed sensing neural networks according to claim 4, characterized in that: The acoustic signature spectrum mentioned in step 3) is obtained by the following formula: Among them, A k (t,f) represents the voiceprint spectrum corresponding to the kth voiceprint segment; The expression for the time-series characteristics of wind speed conditions mentioned in step 3) is as follows: Where p is the time window radius; when <0 or k+p> When constructing a timing window, use boundary copying, zero padding, or simply use existing fragments. The expression for generating the joint features mentioned in step 3) is as follows: Among them, H k This represents the joint feature corresponding to k voiceprint segments. This represents element-wise multiplication, where η is the gating coefficient. Indicates the time-frequency characteristics of the voiceprint; The modulation weights are generated as follows: Among them, W g Let b be the gated weight matrix. g Here, σ(·) is the bias term, and σ(·) is the Sigmoid activation function. Indicates wind speed operating conditions; The complex spectrum estimation described in step 3) Calculated using the following formula: in, This represents the enhancement mask at the time-frequency point (t,f).
8. The method for effective acoustic signature enhancement of wind turbine generators based on wind speed sensing neural networks according to claim 7, characterized in that: Step 4) Reconstructing the time spectrum of the enhanced voiceprint segments into time-domain voiceprint segments includes: Perform inverse short-time Fourier transform on the enhanced audioprint time spectrum: in, This represents the temporal voiceprint segment obtained after the k-th original voiceprint segment is enhanced by a neural network, where j = 0, 1, 2, ... ; Time-domain audioprint fragments Apply window function : in, This indicates the enhanced voiceprint segment after windowing; The confidence weight mentioned in step 4) is calculated using the following formula: Where, ω k ρ represents the credibility weight of the k-th enhanced voiceprint segment, γ is the weight adjustment coefficient, and its value ranges from 0 ≤ γ ≤ 1. k This indicates the mask reliability obtained by enhancing mask stability; The continuously enhanced voiceprint signal mentioned in step 4) is calculated using the following formula: in, ε represents the final continuously enhanced voiceprint signal; n is the index of the discrete time series; K represents the total number of voiceprint segments; and ε is a minimal constant to prevent the denominator from being zero.
9. The method for effective acoustic signature enhancement of wind turbine generators based on wind speed sensing neural networks according to any one of claims 1-8, characterized in that: This also includes enhancing the credibility of the output enhanced time-domain acoustic signature segment: Among them, T k Q represents the enhanced credibility of the k-th segment. k Indicates the voiceprint quality score, ρ k B indicates that the credibility of the mask is enhanced. k β1, β2, and β3 represent the fidelity of the fault-sensitive frequency band, and are weighting coefficients. It also includes the overall enhanced credibility of the continuously enhanced voiceprint signal output: Where T represents the overall credibility of the final continuously enhanced voiceprint signal, ω k ε represents the credibility weight of the k-th enhanced voiceprint segment, and ε is a minimal constant to prevent the denominator from being zero; Set an evaluation threshold. When the overall enhancement confidence T is consistently lower than the preset threshold within a preset time range, output a low confidence warning for the enhancement result.