Wind turbine health diagnosis management system based on multi-mode fusion analysis of voiceprint recognition

CN122812809APending Publication Date: 2026-09-25GD POWER DEVELOPMENT CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610962497.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

其中,振动监测虽然能够反映整体结构状态,但对局部早期故障的敏感性不足,且易受到安装位置和传感路径的影响;而声学方法多采用全时段特征提取或固定阈值判别,缺乏对控制动作触发过程的针对性分析,难以捕捉异常声由掩蔽状态向可辨状态转变的关键瞬间

Benefits of technology

本发明,不再依赖传统的整体能量或单一频谱分析方法,而是以控制动作触发为时间锚点,结合瞬态能量响应函数与谱熵变化率的双判据机制,精准识别异常声由稳定掩蔽状态向短时可辨状态的跃迁过程,该方式能够有效捕捉在稳态噪声背景下被淹没的微弱异常信号,避免因长期平均或全时段分析导致的异常稀释问题,从根本上提升早期故障征兆的检出能力,特别适用于风电机组复杂噪声环境下的隐蔽性缺陷识别。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122812809A_ABST
    Figure CN122812809A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of fan health monitoring, in particular to a wind turbine health diagnosis management system based on acoustic fingerprint recognition multi-mode fusion analysis, which executes the following when running: input the acoustic fingerprint of the nacelle before and after the switching of the wind turbine control action, extract the leakage segment of the abnormal sound from the stable masking state to the short-time distinguishable state, and output the transient leakage unit; input the transient leakage unit, combine the structural response and operation feedback corresponding to the leakage segment, eliminate the pseudo-abnormal caused by echo transmission and action rebound, and output the real abnormal unit; input the real abnormal unit, generate a health management unit according to its repeated occurrence rule and deterioration direction, and perform observation review, operation intervention or repair treatment according to the health management unit. Compared with the existing discrimination method based only on acoustic signals, the present application reduces the false positive rate and improves the reliability of the abnormal identification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wind turbine health monitoring technology, and in particular to a wind turbine health diagnosis and management system based on voiceprint recognition multi-mode fusion analysis. Background Technology

[0002] During long-term operation, critical components of wind turbines, such as main shaft bearings, gearboxes, yaw systems, and pitch mechanisms, are prone to wear, loosening, or fatigue damage under alternating loads and complex operating conditions. These early faults are often accompanied by weak abnormal sound signals, but due to the strong aerodynamic noise, mechanical noise, and electromagnetic interference inside the nacelle, these abnormal sounds are usually masked by steady-state background noise, making them difficult to identify directly using conventional methods. Therefore, how to extract and identify concealed abnormal sounds in complex noise environments has become a key technical problem in the field of wind turbine health monitoring.

[0003] In existing technologies, condition monitoring of wind turbines mainly relies on vibration analysis or simple energy and spectrum analysis methods based on acoustic signals. While vibration monitoring can reflect the overall structural condition, it lacks sensitivity to early local faults and is easily affected by installation location and sensor path. Acoustic methods often employ full-time feature extraction or fixed threshold discrimination, lacking targeted analysis of the control action triggering process and failing to capture the critical moment when abnormal sounds transition from a masked state to a discernible state. Furthermore, existing methods typically determine anomalies based on a single signal source, lacking joint analysis of the physical causal relationship between acoustic signals, structural response, and control behavior. This leads to a high false alarm rate in practical applications, where echo or rebound noise caused by control actions is easily misjudged as fault anomalies. Summary of the Invention

[0004] This invention provides a wind turbine health diagnosis and management system based on voiceprint recognition multi-modal fusion analysis. It can accurately capture abnormal sounds in the context of control action triggering, combine sound-vibration-control multi-source information to make causal judgments, and realize hierarchical health management based on historical patterns.

[0005] A wind turbine health diagnosis and management system based on voiceprint recognition and multi-modal fusion analysis performs the following during system operation: Input the nacelle acoustic signature before and after the wind turbine control action switching, extract the leakage segment of the abnormal sound from the stable masked state to the short-term discernible state, and output the transient leakage unit. Input the transient leakage unit, combine the structural response and operational feedback corresponding to the leakage segment, eliminate the false anomalies caused by echo propagation and action rebound, and output the real anomaly unit; Input the real abnormal unit, generate a health management unit according to its recurrence pattern and deterioration direction, and perform observation review, operation intervention or maintenance based on the health management unit.

[0006] Optionally, the nacelle acoustic signature before and after the switching of the wind turbine control action includes collecting nacelle audio streams for preset durations before and after the triggering time of the pitch, yaw, or braking control commands, performing frame-by-frame windowing on the audio streams to obtain audio signals, extracting Mel-frequency cepstral coefficients from the audio signals, and constructing a steady-state masking background noise model.

[0007] Optionally, the transient energy response function and spectral entropy change rate of each frame of the audio signal relative to the steady-state masking background noise model are calculated. When the transient energy response function exceeds the adaptive threshold and the spectral entropy change rate jumps from the low-order steady interval to the high-order discrete interval, it is determined as the starting point of the leakage segment where the abnormal sound changes from a stable masking state to a short-term discernible state.

[0008] Optionally, the starting point of the leaked segment is used as a reference to extract a time-domain segment including the complete transient decay process, and the time-frequency domain acoustic feature vector of the time-domain segment is associated to output the transient leak unit. The transient leak unit includes the leaked time-domain segment, the corresponding time-frequency domain acoustic feature vector, the start time marker, and the duration marker.

[0009] Optionally, the nacelle vibration acceleration signal and main control operation feedback data within the time window corresponding to the transient leakage unit are acquired. The main control operation feedback data includes at least the pitch angle change rate, generator torque command response and yaw brake pressure status. The time window is obtained based on the start time marker and duration marker.

[0010] Optionally, the envelope waveform of the cabin vibration acceleration signal is extracted. If the sound pressure level attenuation trend of the transient leakage unit is consistent with the mechanical attenuation trend of the envelope waveform, and the time difference between the start time of the transient leakage unit and the end time of the control action in the main control operation feedback data is less than a preset rebound delay threshold, then the transient leakage unit is determined to be a pseudo-abnormality caused by echo transmission or action rebound and is removed.

[0011] Optionally, if the acoustic signature of the transient leakage unit is not causally related to the cabin vibration acceleration signal, and the duration of the transient leakage unit exceeds the normal transient response envelope range corresponding to the control action, it is determined to be a real abnormal unit and output.

[0012] Optionally, it also includes a pre-built historical abnormal voiceprint library. The voiceprint feature vector corresponding to the real abnormal unit is stored in the historical abnormal voiceprint library, and similarity comparison and temporal correlation analysis are performed with the real abnormal units captured under the same type of control action in the historical period. The frequency of recurrence of the real abnormal unit within a unit time or unit number of actions is counted to generate recurrence pattern parameters.

[0013] Optionally, the time-frequency domain feature trend sequence of the abnormal sound in the real abnormal unit is extracted, and the energy accumulation slope, center frequency offset and duration broadening ratio of the time-frequency domain feature trend sequence in the process of continuous occurrence are calculated. Based on the energy accumulation slope, center frequency offset and duration broadening ratio, a degradation pointing index is constructed.

[0014] Optionally, when the recurrence pattern parameter is lower than the preset occasional threshold and the degradation index is in a stable range, a health management unit of observation and review type is generated, and a prompt message is output to trigger manual auscultation or vibration re-examination. When the recurrence pattern parameter is higher than the preset occasional threshold and the degradation index is in the slow change range, a health management unit of the operation intervention type is generated, and the command is output to limit the unit power ramp rate or adjust the pitch rate. When the degradation index exceeds a preset severe threshold and shows an accelerating upward trend, a health management unit of the maintenance type is generated, and an alarm is output to trigger a shutdown inspection or component replacement work order.

[0015] The beneficial effects of this invention are: This invention no longer relies on traditional overall energy or single spectrum analysis methods. Instead, it uses the control action trigger as the time anchor point and combines a dual criterion mechanism of transient energy response function and spectral entropy change rate to accurately identify the transition process of abnormal sound from a stable masked state to a short-term discernible state. This method can effectively capture weak abnormal signals that are submerged in the background of steady-state noise, avoid the abnormal dilution problem caused by long-term averaging or full-time analysis, and fundamentally improve the detection capability of early fault symptoms. It is particularly suitable for the identification of hidden defects in the complex noise environment of wind turbine units.

[0016] This invention introduces cabin vibration acceleration signals and main control operation feedback data. Through consistency analysis of sound pressure attenuation trends and mechanical vibration envelopes, and joint determination of their relationship with control action timing, a false anomaly elimination mechanism based on physical mechanisms is established. Simultaneously, by identifying the causal relationship between anomalous sounds and structural responses through acoustic-vibration correlation intensity indices, it effectively eliminates accompanying noises such as reverberation propagation and motion rebound. Compared to existing discrimination methods based solely on acoustic signals, this invention reduces the false alarm rate and improves the reliability of anomaly identification results.

[0017] This invention constructs a historical abnormal acoustic signature database, combines similarity comparison and temporal correlation analysis to quantify the recurring patterns of anomalies, and further introduces energy accumulation slope, center frequency offset, and duration broadening ratio to construct a degradation direction index, characterizing the evolution trend of anomalies from three dimensions: intensity, frequency, and time. Based on this, a hierarchical decision-making mechanism for observation and verification, operational intervention, and maintenance is established, realizing a shift from passive alarm to trend-driven proactive operation and maintenance. This allows for early intervention before anomalies develop into serious faults, significantly improving the operational safety of wind turbines and reducing maintenance costs. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only for this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a system execution logic block diagram according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the system execution steps according to an embodiment of the present invention. Detailed Implementation

[0020] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. For some well-known technologies, those skilled in the art may also use other alternative methods to implement the invention. Moreover, the accompanying drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.

[0021] like Figures 1-2 As shown, the wind turbine health diagnosis and management system based on voiceprint recognition multi-modal fusion analysis executes the following during system operation: Input the nacelle acoustic signature before and after the wind turbine control action switching, extract the leakage segment where the abnormal sound changes from a stable masked state to a short-term identifiable state, and output the transient leakage unit. Details are as follows: 1. Acquisition of cabin audio stream and construction of steady-state masked background noise model.

[0022] 1.1 Most abnormal sounds in wind turbine generators are not continuous, but are triggered or released at the moment of control action switching. During steady-state operation, a large amount of structural vibration, aerodynamic noise, and mechanical noise are in a relatively stable superposition state, and weak anomalies are often masked. However, when control commands are triggered, load path reconstruction, gap state changes, or mechanical constraint release occur, causing the previously masked abnormal sounds to become visible for a short time. Therefore, audio interception is performed centered on the triggering time of pitch control commands, yaw control commands, or braking control commands as a reference. This aligns the occurrence window of anomalies from hidden to visible in terms of physical mechanisms, improving the targeting and accuracy of anomaly capture and avoiding noise interference caused by blindly searching throughout the entire time period.

[0023] Specifically, during the operation of the wind turbine, the triggering time of the pitch control command, yaw control command, or brake control command is... Centered on the target, extract preset durations forward and backward respectively. and Obtain the cabin audio stream within the corresponding time interval. Construct continuous audio data segments before and after the control action switching.

[0024] 1.2. Cabin acoustic signature signals are inherently non-stationary signals, and their statistical characteristics change over time. Directly analyzing the entire audio stream can mask local transient features. Frame segmentation, however, can divide long-term non-stationary signals into multiple short-term approximately stationary segments, giving physical meaning to the spectral analysis within each frame. Simultaneously, windowing can reduce truncation effects at frame boundaries, decrease spectral leakage, and make the frequency domain features more concentrated and stable. Therefore, further analysis of the audio stream... Perform frame segmentation and windowing processing to obtain the first... Frame audio signal , is represented as: ;in, This represents the original cabin audio signal. Indicates the first Frame discrete audio sequence, Represents the window function. Indicates the frame shift length. Indicates the sequence number of the sampling point within the frame.

[0025] 1.3 Further, extract the Mel-frequency cepstral coefficient (MFCC) feature vector from each frame of the audio signal. Construct a steady-state masking background noise model Its statistical expression is as follows: ;in, The mean vector representing the features of MFCC. The covariance matrix represents the MFCC features; the steady-state masked background noise model is used to characterize the distribution of background acoustic signature features during the stable operation phase before control action switching. The steady-state masked background noise model uses the mean vector and covariance matrix of the MFCC features, and needs to simultaneously characterize the central distribution characteristics and fluctuation range characteristics of the background acoustic signature. The mean vector reflects the average state of each dimension of the MFCC features under stable operation conditions and can be regarded as the baseline shape of a typical background acoustic signature; the covariance matrix describes the dispersion and correlation between each feature dimension, reflecting the normal fluctuation range and multi-dimensional coupling relationship. Through these two statistics, a noise model with distribution constraints is constructed, enabling any subsequent frame of acoustic signature to not only determine whether it deviates from the mean, but also whether its deviation exceeds the normal fluctuation structure.

[0026] 2. Determination of the starting point of transient leakage segments.

[0027] 2.1 For each frame of audio signal, calculate its transient energy response function relative to the steady-state masked background noise model. Its definition is: ; in, Let be the transient energy response function, reflecting the degree to which the energy of the current frame deviates from the background, representing the th . The normalized ratio of the actual energy of a frame's audio signal to the expected energy of the steady-state masked background noise model is used to measure whether an anomalous energy transition occurs in that frame. Indicates the number of sampling points per frame. This represents the expected value based on the steady-state masking background noise model. Indicates the first The first frame The amplitude of each sampling point Indicates the first The actual temporal energy of a frame This represents the average energy level that the frame should have under steady-state background conditions.

[0028] Calculating the transient energy response function for each frame of audio signal essentially involves comparing the actual energy of the current frame with the energy level that the frame should have under a steady-state masking background, thereby measuring whether an abnormal energy surge has occurred in that frame. First, for the... Frame audio signal The actual temporal energy of the frame is obtained by squaring and summing the points, reflecting the overall level of acoustic vibration intensity in the current frame. Using the established steady-state masking background noise model, the energy of similar frames under normal and stable operating conditions is statistically modeled to obtain the corresponding expected energy value. The normalized ratio of the actual energy of the current frame to the expected energy is calculated to obtain a dimensionless response index, which is used to characterize the degree of deviation of the current frame from the steady-state background.

[0029] Specifically, the steady-state masked background noise model does not directly store energy values. Instead, it indirectly characterizes the background acoustic signature through the statistical distribution of MFCC features. Therefore, when calculating the expected energy, it is usually based on a set of historical stable frames—samples that have participated in constructing the mean vector and covariance matrix—to statistically obtain the corresponding average frame energy. Alternatively, a mapping relationship from MFCC features to energy is established to obtain a reasonable energy level under this type of acoustic signature distribution. This ensures that the expected energy not only reflects the overall intensity but also implies the consistency of the acoustic signature structure, making anomaly detection more targeted. By comparing the actual energy with the expected background energy, it effectively amplifies transient energy surges caused by structural release or impact when control actions are triggered. If the energy of a frame is within the normal fluctuation range, the ratio is close to 1; however, when an abnormal sound is released from the masked state, its energy will be significantly higher than the expected background energy, causing the ratio to rise rapidly, thus achieving sensitive capture of the anomaly's initiation point.

[0030] 2.2 Simultaneously calculate the rate of change of spectral entropy for the corresponding frame. , which is represented as: ; ; in, Indicates the first Spectral entropy of a frame Indicates the first Frame in Normalized energy distribution of each frequency band Indicates the number of frequency band divisions. The spectral entropy change rate represents the rate at which the spectral structural complexity of an audio signal changes between adjacent frames. Spectral entropy measures the uniformity of energy distribution across frequency bands in an audio frame. Lower spectral entropy indicates a more stable and ordered signal structure, corresponding to a steady-state masking state, when energy is concentrated in a few bands. Higher spectral entropy indicates a more discrete and complex signal, corresponding to abnormal sound release or disturbance. The spectral entropy change rate further describes whether this change from order to disorder or from steady-state to discrete involves abrupt changes. Relying solely on energy changes is susceptible to overall noise enhancement or environmental disturbances, making it difficult to distinguish between normal amplification and structural abrupt changes. Therefore, introducing the spectral entropy change rate allows for the identification of abnormal sounds at the spectral structure level. When abnormal sounds are released from a masking state, they often disrupt the original stable spectral distribution, causing a rapid jump in spectral entropy. This allows for more accurate detection of the anomaly's initiation point and effectively reduces the false positive rate.

[0031] 2.3 Constructing an adaptive threshold Its definition is: ;in, This represents the mean of the energy response function during the steady-state phase. The standard deviation of the energy response function in the steady-state phase is represented by the standard deviation of the energy response function in the steady-state phase. This represents the adjustment coefficient, which ranges from 2.5 to 3. If the system noise fluctuates significantly, it can be gradually increased. .

[0032] 2.4 When both conditions are met and The current frame is determined to be the starting point of a leaked segment where the abnormal sound transitions from a stable masked state to a short-term discernible state.

[0033] The energy dimension shows a significant increase in vibration intensity relative to the steady-state background, indicating the presence of a new sound source or the release of an existing masked sound source. The frame's transformation from a stable, ordered frequency band distribution to a discrete, complex one, based on its spectral structure, indicates a sudden change in the acoustic signature structure. When both a sudden increase in energy and a structural abrupt change occur simultaneously, it aligns with the physical process of an anomalous sound being excited from a masked state and briefly manifesting. If only one condition is met, it might correspond to increased environmental noise (energy increase only) or disturbances in normal operating conditions (spectral structure change only), failing to constitute a valid anomaly. This represents the threshold value for the rate of change of spectral entropy, used to distinguish between low-order stationary intervals and high-order discrete intervals. The mean value of the steady-state rate of change of spectral entropy is taken. Add 1.5 times the standard deviation .

[0034] 3. Construct a transient leakage unit.

[0035] The frame index corresponding to the start point of the leaked fragment Based on this, a continuous time-domain segment containing the complete transient decay process is extracted to form the leaked time-domain segment. Its time range is defined as: ;in, Indicates the start time of the leaked segment. The transient decay duration is represented by the energy fallback determination method, which involves continuously tracking the transient energy response function of each subsequent frame starting from the beginning of the leak segment. When it gradually decays from its peak and first falls back to the energy threshold range corresponding to the steady-state masking background, it enters the [phase of a state where the energy level decreases]. When the transient process remains stable within a certain interval and over several consecutive frames, it is considered to have ended. The corresponding time length is the transient decay duration. .

[0036] Further, the corresponding time-frequency domain acoustic signature feature vector is extracted from the leaked time-domain segment. , which is represented as: ;in, Represents the MFCC eigenvector. To represent short-time spectral characteristics, a Fast Fourier Transform is performed directly on each frame of the audio signal to obtain the corresponding spectral amplitude or power distribution. This extracts the frequency band energy distribution and spectral centroid from the spectrum to form short-time spectral features that reflect the frequency domain structure information of that frame. The zero-crossing rate is represented by the number of times the signal waveform crosses the zero amplitude point in each frame of audio signal. This number is then normalized according to the frame length to obtain the symbol change frequency per unit time.

[0037] Ultimately, this will leak a fragment of the time domain. With the corresponding time-frequency domain acoustic feature vector The time stamp is encapsulated and bound to form a structured association, thus creating a transient leakage unit. : ;in, Indicates a transient leakage unit. As a start time marker As a duration marker.

[0038] The transient leakage unit serves as the input basis for subsequent identification of real anomaly units and multi-mode fusion analysis, enabling accurate extraction of anomalous sounds from stable masked states to short-term discernible states.

[0039] The transient leakage unit is input, and combined with the structural response and operational feedback corresponding to the leakage segment, false anomalies caused by echo propagation and motion rebound are eliminated, and the true anomaly unit is output. Specifically: 1. Obtain main control operation feedback data.

[0040] For the obtained transient leakage unit In its corresponding time window Inside, synchronous acquisition of cabin vibration acceleration signals and the main control operation feedback data set .

[0041] Among them, the main control operation feedback data includes at least the pitch angle change rate. Generator torque command response and yaw brake pressure status This constitutes a multidimensional running state vector: .

[0042] The pitch angle change rate is obtained by differential calculation after real-time acquisition of pitch angle data by angle sensors in the pitch control system. The generator torque command response is the torque control command issued by the main control system to the generator and its execution feedback, recorded by the converter control system, including the target torque and actual torque response, used to characterize power regulation behavior. The yaw brake pressure state is the hydraulic or pneumatic pressure state of the brake in the yaw system, acquired by pressure sensors, used to reflect the yaw lock-up or release process. The vibration acceleration signal and the main control operation feedback data are aligned with the transient leakage unit on the time axis to ensure consistency in subsequent correlation analysis.

[0043] The core idea of ​​this section is to move beyond relying solely on acoustic signals to identify anomalies. Instead, it introduces multi-source information from structural responses and control behaviors to perform reverse verification and screening of the anomaly sources. Specifically, using the identified transient leakage unit as a time anchor, engine room vibration signals and main control system operating data are extracted synchronously within the same time window. The equipment actions and structural responses at the time of the sound are compared and analyzed on the same timeline, providing a basis for subsequent judgment on whether the sound represents a real fault.

[0044] Further analysis involves extracting the envelope from the vibration acceleration signal to obtain the energy decay process of the mechanical system within that time period. Simultaneously, by combining the sound pressure level change trend, it is determined whether the sound decay pattern and the mechanical vibration decay pattern are consistent. If they are highly consistent, and the sound appears immediately after the control action ends, then the sound is likely not a new anomaly, but rather accompanying noise caused by structural conduction or mechanism release, and is thus identified as a false anomaly and eliminated. Conversely, if a transient leakage unit does not show a clear correlation between its acoustic signature and the vibration signal—for example, the sound changes but the structural vibration does not respond accordingly—it indicates that the sound is not caused by the overall mechanical structure response, but may originate from a local component anomaly. Furthermore, if the duration of the sound is significantly longer than the transient response range caused by normal control actions, it further indicates that it does not belong to normal operating condition disturbances. Therefore, such signals are identified as genuine anomaly units.

[0045] 2. Detection and removal of pseudo-abnormal units.

[0046] 2.1 Regarding the cabin vibration acceleration signal Envelope extraction processing is performed on the nacelle vibration acceleration signal. Bandpass filtering is applied to remove low-frequency trend terms and high-frequency noise interference, concentrating the signal in the effective mechanical response frequency band. The filtered signal is then rectified, and the absolute value is taken to unify positive and negative vibrations into an energy intensity expression. Low-pass filtering or moving average processing is then applied to the rectified signal to eliminate high-frequency fluctuations, resulting in a smoothly varying envelope curve. This envelope curve is used as a representation of the decay of mechanical vibration energy over time, thus obtaining its envelope waveform. It is used to characterize the decay process of mechanical vibration energy.

[0047] 2.2 The key to this invention lies in determining whether the change in sound originates from the mechanical structure response. Sound pressure level (SPL) is the logarithmic expression of sound intensity, stably reflecting the change in sound energy over time. By converting the instantaneous sound pressure of the transient leakage segment into a SPL curve, an acoustic attenuation trend consistent with human perception and more sensitive to dynamic range changes can be obtained. Comparing this trend with the mechanical attenuation trend of the vibration envelope, if the two are consistent, it indicates that the sound is an accompanying sound caused by structural vibration transmission or rebound; otherwise, it may be an independent abnormal sound source. Therefore, the SPL change trend is introduced to further calculate the change trend of the SPL corresponding to the transient leakage unit over time. And construct the acoustic attenuation function. ; in, Indicates sound pressure level. This represents the instantaneous sound pressure level of the transient leakage segment. Represents the reference sound pressure level, taken as... Sound pressure level (SPL) is the minimum sound pressure level that the human ear can perceive in the air, and it is used as a benchmark for calculating sound pressure levels.

[0048] 2.3 Calculate the consistency index between the sound pressure level attenuation trend and the vibration envelope attenuation trend. : ; in, Represents the correlation coefficient. This represents the average sound pressure level. This represents the average value of the vibration envelope.

[0049] 2.4 Simultaneously, calculate the start-up time of the transient leakage element. With the end of the control action Time difference: ;in, Indicates the start time of the transient leakage element. This indicates the end time of the corresponding control action in the master control operation feedback.

[0050] 2.5. When the following conditions are met: (The decay trend is consistent); (Time difference is less than the rebound delay threshold); The transient leakage unit is then determined to be a pseudo-abnormal unit caused by echo propagation or action rebound, and is thus eliminated; among which, The correlation coefficient represents the attenuation consistency threshold, ranging from 0.7 to 0.9, with 0.8 being preferred. A correlation coefficient exceeding 0.8 indicates strong consistency in the changing trends of the two sets of signals, effectively characterizing the attenuation process originating from the same source. This represents the rebound delay threshold, with a value ranging from 100ms to 150ms, preferably 130ms. The elastic release and damping decay of the mechanical structure after the control action ends usually occur within a time scale of hundreds of milliseconds. This range can cover the rebound and sound wave propagation delay of most mechanisms, while avoiding misjudging independent anomalies with long delays as rebound noise.

[0051] When a wind turbine performs control actions such as pitch, yaw, or braking, the structural system undergoes a release, rebound, and decay process. During this process, the mechanical vibration energy gradually decays, and a corresponding acoustic response is generated through structural transmission. Therefore, the decay trend of the sound pressure level is highly consistent with the decay trend of the vibration envelope in terms of time. In addition, the sound caused by structural reverberation or mechanism rebound usually appears shortly after the control action ends, and its time delay has a clear physical constraint. If both the trend consistency and time proximity conditions are met, it indicates that the sound has a direct causal relationship with the structural response and is a concomitant effect of the control action, rather than an independent anomaly source. Therefore, it should be eliminated as a pseudo-anomaly. Conversely, if neither of these conditions is met, a clear physical correlation cannot be established, and it should be retained as a potential real anomaly.

[0052] 3. Identification and output of real abnormal units.

[0053] 3.1 For transient leakage units that were not eliminated, further determine the causal relationship between their acoustic signature characteristics and vibration response. Construct an acoustic-vibration correlation criterion. : ;in, Indicates the intensity index of acoustic-vibration correlation. This represents the acoustic signature feature vector of the transient leakage unit. Indicates vibration acceleration signal, The correlation calculation function first transforms the acoustic signature changes in the transient leakage unit into a time-varying feature sequence. Then, it aligns and compares this feature sequence with the vibration acceleration signal within the same time window, and uses normalized correlation analysis to obtain the acoustic-vibration correlation strength index between the two. The specific steps are as follows: 3.1.1 Extract the voiceprint feature sequence within the time window corresponding to the transient leakage unit. The aforementioned transient leakage unit contains MFCC feature vector, short-time spectral features, and zero-crossing rate. First, arrange the voiceprint features corresponding to each frame within the time window in the order of the frame sequence to form a multi-dimensional voiceprint feature sequence that changes over time. That is, instead of taking a single feature value of a certain frame, the features of multiple consecutive frames within the entire leakage segment are spliced ​​together to preserve its dynamic evolution process.

[0054] 3.1.2. Dimensionality reduction processing is performed on the multidimensional acoustic signature feature sequence to generate a single-channel acoustic signature representation sequence that can be used for correlation analysis. The vibration acceleration signal itself is a one-dimensional time series, while acoustic signature features are usually multidimensional vectors. Therefore, the multidimensional features are first compressed into a main acoustic signature sequence that can represent the main changing trend of the transient leakage unit. That is, the first cepstral coefficient of MFCC, the main frequency band energy or the spectral centroid, which are most sensitive to anomalies, are directly selected. Principal components are extracted from the multidimensional acoustic signature features, and the first principal component is taken as the main acoustic signature sequence. The MFCC features, short-time spectral features and zero-crossing rate are weighted and fused according to preset weights to form a comprehensive acoustic signature change sequence, resulting in an acoustic signature sequence F(t) corresponding to time.

[0055] 3.1.3. The vibration acceleration signal undergoes time alignment and preprocessing consistent with the acoustic print sequence. Since vibration signals typically have a high sampling rate, while acoustic print feature sequences are extracted frame-by-frame from low-sampling-rate features, the vibration acceleration signal needs to be segmented within the same time window to correspond to the acoustic print frame sequence on a time scale. Specifically, the vibration acceleration signal is synchronously framed using the same frame length and frame shift as the acoustic print frame, and vibration characterization values, including root mean square value, envelope mean, and peak value, are extracted within each frame to form a vibration response sequence of the same length as the acoustic print sequence. This ensures a one-to-one correspondence between the acoustic print sequence and the vibration sequence on a time index.

[0056] 3.1.4 Normalize the voiceprint sequence and vibration sequence. Since the two have different dimensions and numerical ranges, direct comparison will lead to distorted results. Therefore, subtract their respective means and normalize them according to their respective standard deviations so that both sequences are transformed into dimensionless sequences with zero mean and unit scale. After normalization, the correlation calculation reflects the consistency of the two trends, without being affected by the original amplitude.

[0057] 3.1.5. Under the condition of allowing a certain time delay, perform sliding alignment correlation analysis. Because there may be a propagation delay or response lag between the occurrence of abnormal sound and structural vibration, it is not mandatory for the two to change synchronously at exactly the same moment. In actual execution, the normalized acoustic signature sequence is used as a reference, and the vibration sequence is allowed to slide back and forth within a preset small delay window. The correlation between the two is calculated at each delay position, and the maximum value is taken as the final acoustic-vibration correlation strength index. .

[0058] 3.1.6, According to The size determines the strength of the association. If A relatively large value indicates a high degree of consistency between the acoustic signature sequence and the vibration response sequence in terms of their changing trends, suggesting that this transient leakage unit is likely caused by structural vibration, reverberation transmission, or mechanical action. The smaller value indicates a lack of obvious synchronicity between the two, suggesting that the sound is not directly driven by the vibration of the overall structure, but is more likely an independent abnormal sound source. Therefore, it can be used as one of the important criteria for judging a real anomaly.

[0059] 3.2 Simultaneously define the normal transient response envelope time range corresponding to the control action. And calculate the duration of the transient leakage element. .

[0060] 3.3. When the following conditions are met: (No significant causal relationship between acoustic and vibrational frequencies); (Duration exceeds the normal transient range); Then the transient leakage unit is determined to be a real abnormal unit. And outputs data for subsequent health management analysis, including... The correlation threshold represents the acoustic-vibration correlation, ranging from 0.3 to 0.5, with 0.4 being preferred. When the correlation is below 0.4, it is considered that there is no significant linear correlation between the two sequences, effectively distinguishing between homologous responses and independent anomalies. This indicates the duration of the normal transient response of the control action, with a value range of 200 to 300 ms. After the wind turbine performs pitch, yaw, or braking actions, its structural vibration and associated acoustic response usually complete the energy release and attenuation within hundreds of milliseconds. Signals exceeding this time scale are more likely to originate from abnormal conditions rather than normal operating disturbances.

[0061] The above is a joint criterion based on the two dimensions of causal break and time boundary crossing. First, The lower sound signature indicates that there is no clear synchronization between the sound signature change process and the overall vibration response of the cabin. This means the sound is not a concomitant response caused by structural vibration transmission or control actions, but more likely originates from local component anomalies, such as local loosening, wear, or cracks—physically, they are non-homogeneous signals. Secondly, the duration exceeds the normal transient response range corresponding to the control action, indicating that the sound did not decay rapidly after the control action ended, but rather exhibited abnormally prolonged characteristics. This contradicts the short-term decay pattern of normal mechanical rebound or reverberation. Therefore, only when both the lack of a causal relationship between sound and vibration and the abnormally prolonged duration are simultaneously established can the interference from control actions and the influence of structural transmission be ruled out mechanistically, thus confirming it as a genuine anomalous unit.

[0062] Through the above steps, pseudo-anomalies caused by structural reverberation or mechanical springback are eliminated, and real anomaly units with independent anomaly sources are accurately extracted.

[0063] Input the actual abnormal unit, generate a health management unit according to its recurrence pattern and deterioration direction, and perform observation review, operational intervention, or maintenance based on the health management unit. Specifically: 1. Construct parameters for recurring patterns.

[0064] 1.1 The actual abnormal unit Voiceprint feature vector Store in historical abnormal voiceprint database The database is categorized and indexed according to the type of control action, including acceleration, yaw, and braking. The historical abnormal voiceprint database is constructed as follows: The first step is data entry and structured encapsulation.

[0065] Once a genuine abnormal unit is identified, its core information is uniformly encapsulated, including the voiceprint feature vector, corresponding time stamp, control action type, and necessary operating context information, including speed and load. The above information is written into the database according to a unified field format to form a standardized abnormal voiceprint recording unit, ensuring that data from different sources have a consistent data structure.

[0066] The second step is to construct a primary category index for control action types.

[0067] Using the control action type as the primary index key, all abnormal voiceprint records are classified into primary categories. Specifically, an action tag field is added to each record, with values ​​limited to one of three categories: pitch, yaw, or braking. A partition or index structure is created in the database based on this field, so that in subsequent queries, the corresponding historical abnormal unit set can be quickly filtered out directly by control action type, enabling comparison with the same operating conditions.

[0068] The third step is to construct an auxiliary index based on the similarity of voiceprint features.

[0069] Based on the primary classification, an auxiliary index is further constructed for the voiceprint feature vector to accelerate similarity retrieval. Specifically, after normalizing the voiceprint feature vector, its principal components or low-dimensional embedding vectors are extracted and stored using a vector index method. This allows for quick location of similar candidate anomalies in the feature space without traversing all historical data when performing similarity comparison.

[0070] The fourth step is to construct a time-series association index.

[0071] To support the statistical analysis of recurring patterns and the analysis of degradation trends, a time-dimensional index is also established. A timestamp is retained in each abnormal record and they are organized in chronological order. At the same time, a time series index structure can be built by device number + control action type, so that abnormal records under the same device and the same type of action can form a continuous sequence.

[0072] Through the above construction, the historical abnormal voiceprint database not only realizes the classification index according to the control action type, but also has the ability to retrieve by voiceprint feature similarity and perform time series association analysis, thus providing a data foundation for subsequent similarity comparison, repetition pattern extraction and degradation direction analysis.

[0073] 1.2, within the preset historical period Within, extract the set of historical real abnormal units under the same type of control action. And calculate the similarity between the current anomalous unit and historical samples: ;in, Indicates the relationship with the first The similarity of historical anomalous units, The voiceprint feature vector representing the historical anomaly unit is calculated using the dot product operation for each term. Represents the vector norm.

[0074] Specifically, the process involves first filtering the set and then comparing each item one by one, with the following steps: The first step is to use the current actual time t when the abnormal unit occurs as a benchmark to trace back a preset historical period. Determine the time window In the historical abnormal voiceprint database, the timestamp field of each record is used for filtering to extract all abnormal records within the time window, forming a candidate set.

[0075] The second step is to extract sets of similar anomalies based on the type of control action.

[0076] Based on the aforementioned candidate set, a secondary screening is performed on the candidate set according to the control action type (e.g., pitch, yaw, or braking) of the current anomalous unit, retaining only historical anomalous units with the same action type. Through this process, a set of historical true anomalous units under the same type of control action is obtained, i.e. This ensures that subsequent comparisons are conducted under the same operating conditions, avoiding interference from differences in voiceprints caused by different control behaviors.

[0077] The third step is to extract voiceprint features.

[0078] The corresponding voiceprint feature vectors are extracted from the current abnormal unit and the historical abnormal unit set, respectively. These vectors are formed by the combination of the aforementioned MFCC, short-time spectral features and zero-crossing rate. Before calculating the similarity, all feature vectors are uniformly processed, including removing the mean and standardizing, to make different samples comparable and avoid the influence of differences in dimensions or amplitudes on the results.

[0079] The fourth step is to calculate the similarity between the current anomaly and historical samples one by one.

[0080] Using the voiceprint feature vector of the current anomalous unit as a benchmark, similarity is calculated one by one with the voiceprint feature vector of each historical anomalous unit in the set. First, the inner product between the two vectors is calculated to reflect their consistency in each feature dimension; then, the magnitude of each vector is calculated separately, and the inner product result is divided by the product of the two magnitudes to obtain the normalized similarity value. This value essentially reflects the angle relationship between the two voiceprints in the feature space. The closer the value is to 1, the more similar the two are.

[0081] 1.3, when When an event is identified as a similar abnormal event, its frequency of occurrence within a unit of time or unit of control actions is counted, and a recurrence pattern parameter is defined. : ; in, This indicates the number of times the same type of anomaly has occurred within a historical period. To indicate the statistical time duration, this invention selects 72 hours. This indicates the total number of corresponding control actions. The threshold for voiceprint similarity is set to 0.8. The recurrence pattern parameter is used to characterize the frequency and periodicity of abnormal sounds.

[0082] 2. Construct a degradation index.

[0083] By observing the evolution trend of the same type of anomaly during repeated occurrences, we can determine whether the equipment is gradually deteriorating. We can characterize the trajectory of the abnormal sound from three dimensions: From an energy perspective, by calculating the energy accumulation slope, we can reflect whether the abnormal sound is getting stronger. If the energy gradually increases each time it occurs, it indicates that the abnormal source is continuously intensifying, and the corresponding fault is developing, rather than being an occasional disturbance.

[0084] From the perspective of frequency structure, the center frequency offset can be used to determine whether the spectral position of abnormal sound is changing. For example, a shift from low frequency to high frequency may indicate a change in structural stiffness or a change in contact state. Continuous frequency drift is usually a typical feature of component state evolution, rather than random fluctuations under stable operating conditions.

[0085] From a time perspective, the duration broadening ratio reflects whether the duration of abnormal sounds has increased. The transient response caused by normal control actions is usually short-lived, but if the duration of abnormal sounds gradually increases, it indicates that the system's damping characteristics or energy release mechanism have changed, and the scope of the abnormal impact is expanding.

[0086] Specifically, for real anomalous units, the time-frequency domain characteristic trend sequence of their occurrence in a series of consecutive occurrences is extracted, and a degradation index with three dimensions of energy change, frequency shift and duration change is constructed.

[0087] 2.1 Energy Accumulation Slope : ;in, This represents the abnormal sound energy at the nth occurrence. This indicates the abnormal sound energy at the time of its first occurrence. Indicates the number of consecutive occurrences.

[0088] 2.2 Center frequency offset : ;in, Indicates the first The spectral centroid at the time of its first occurrence Indicates the spectral centroid when it first appears.

[0089] 2.3 Duration Stretch Ratio : ;in, Indicates the first The spectral centroid at the time of its first occurrence Indicates the spectral centroid when it first appears.

[0090] 2.4 Based on the above three indicators, a degradation index is constructed. : ;in, This indicates a degradation index. These represent weighting coefficients, with values ​​of 0.4, 0.3, and 0.3 respectively. The degradation direction index is used to comprehensively characterize the degradation trend of abnormal sounds in terms of intensity, frequency structure, and duration.

[0091] 3. Generation and hierarchical handling of health management units.

[0092] Based on recurrence pattern parameters With degradation index Construct hierarchical decision-making logic to generate corresponding types of health management units. : 3.1 When the following conditions are met: Then, a health management unit of the observation and review type will be generated: ;in, This represents the incident threshold, ranging from 0.05 to 0.15, with 0.1 being preferred. When the probability of a certain type of anomaly occurring in the same type of control action is less than 10%, it is considered to be a random disturbance or occasional noise. This represents the degradation stability interval, with values ​​ranging from [value 1] to [value 2]. Within this range, the changes in the three dimensions of energy, frequency, and duration are all small, and no obvious trend of growth is observed, which is consistent with the characteristics of stable fluctuations or occasional anomalies. 3.2, When the following conditions are met: Then, a health management unit of the operational intervention type is generated: ;in, This indicates a slow degradation stage, meaning the anomaly has begun to evolve but has not yet reached a severe level. The value is... Within this range, at least one dimension of energy, frequency, or duration shows a continuous trend of change, but the overall situation is still within a controllable range and is suitable for suppression through operational intervention. 3.3, When the following conditions are met: Then, a health management unit of the maintenance / disposal type will be generated: ;in, The threshold value is 0.75. When the degradation index exceeds this value, it means that multiple dimensions have simultaneously deviated significantly from the initial state, indicating that the anomaly has obvious fault characteristics and should enter the maintenance and repair stage. This indicates the growth rate of the degradation index. The acceleration threshold is used to determine whether the degradation has entered a rapid deterioration stage. It is defined as the rate of change of the degradation index per unit time, ranging from 0.01 to 0.03 / hour. In the normal slow degradation process, the change of D is gradual. When the growth rate exceeds this threshold, it indicates that the abnormal development speed has accelerated significantly, and there is a risk of sudden damage expansion, requiring immediate triggering of maintenance strategies. Control whether it occurs frequently, At what stage should the degree of degradation be controlled? Is the situation deteriorating rapidly?

[0093] The corresponding output strategies are as follows: Output a prompt message to trigger manual auscultation or vibration re-examination; Output control commands to limit the unit's power ramp rate or adjust the pitch rate; Output alarm information to trigger a shutdown inspection or component replacement work order.

[0094] Through the above, a tiered health management decision-making system based on abnormal recurrence patterns and deterioration trends can be achieved.

[0095] This invention encompasses any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of this invention. To provide the public with a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments; however, those skilled in the art will fully understand the invention even without these details. Furthermore, to avoid unnecessary misunderstanding of the essence of this invention, well-known methods, processes, procedures, components, and circuits are not described in detail.

[0096] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A wind turbine health diagnosis and management system based on voiceprint recognition and multi-modal fusion analysis, characterized in that, The system executes the following during runtime: Input the nacelle acoustic signature before and after the wind turbine control action switching, extract the leakage segment of the abnormal sound from the stable masked state to the short-term discernible state, and output the transient leakage unit. Input the transient leakage unit, combine the structural response and operational feedback corresponding to the leakage segment, eliminate the false anomalies caused by echo propagation and action rebound, and output the real anomaly unit; Input the real abnormal unit, generate a health management unit according to its recurrence pattern and deterioration direction, and perform observation review, operation intervention or maintenance based on the health management unit.

2. The wind turbine health diagnosis and management system based on voiceprint recognition multi-modal fusion analysis according to claim 1, characterized in that, The nacelle acoustic signatures before and after the switching of wind turbine control actions include collecting nacelle audio streams for preset durations before and after the triggering time of pitch, yaw, or braking control commands, performing frame-by-frame windowing on the audio streams to obtain audio signals, extracting Mel-frequency cepstral coefficients from the audio signals, and constructing a steady-state masking background noise model.

3. The wind turbine health diagnosis and management system based on voiceprint recognition multi-modal fusion analysis according to claim 2, characterized in that, Calculate the transient energy response function and spectral entropy change rate of the audio signal in each frame relative to the steady-state masking background noise model. When the transient energy response function exceeds the adaptive threshold and the spectral entropy change rate jumps from the low-order steady interval to the high-order discrete interval, it is determined as the starting point of the leakage segment where the abnormal sound changes from a stable masking state to a short-term discernible state.

4. The wind turbine health diagnosis and management system based on voiceprint recognition multi-modal fusion analysis according to claim 3, characterized in that, The starting point of the leaked segment is used as a reference to extract a time-domain segment that includes the complete transient decay process, and the time-frequency domain acoustic feature vector of the time-domain segment is associated with it to output the transient leak unit. The transient leak unit includes the leaked time-domain segment, the corresponding time-frequency domain acoustic feature vector, the start time marker, and the duration marker.

5. The wind turbine health diagnosis and management system based on voiceprint recognition multi-modal fusion analysis according to claim 1, characterized in that, The engine room vibration acceleration signal and main control operation feedback data within the time window corresponding to the transient leakage unit are acquired. The main control operation feedback data includes at least the pitch angle change rate, generator torque command response and yaw brake pressure status. The time window is obtained based on the start time mark and duration mark.

6. The wind turbine health diagnosis and management system based on voiceprint recognition multi-modal fusion analysis according to claim 5, characterized in that, Extract the envelope waveform of the cabin vibration acceleration signal. If the sound pressure level attenuation trend of the transient leakage unit is consistent with the mechanical attenuation trend of the envelope waveform, and the time difference between the start time of the transient leakage unit and the end time of the control action in the main control operation feedback data is less than the preset rebound delay threshold, then the transient leakage unit is determined to be a pseudo-abnormality caused by echo transmission or action rebound and is removed.

7. The wind turbine health diagnosis and management system based on voiceprint recognition multi-modal fusion analysis according to claim 6, characterized in that, When the acoustic signature of the transient leakage unit is not causally related to the cabin vibration acceleration signal, and the duration of the transient leakage unit exceeds the normal transient response envelope range corresponding to the control action, it is determined to be a real abnormal unit and outputs an error.

8. The wind turbine health diagnosis and management system based on voiceprint recognition multi-modal fusion analysis according to claim 1, characterized in that, It also includes a pre-built historical abnormal voiceprint library. The voiceprint feature vector corresponding to the real abnormal unit is stored in the historical abnormal voiceprint library, and similarity comparison and temporal correlation analysis are performed with the real abnormal units captured under the same type of control action in the historical period. The frequency of recurrence of the real abnormal unit within a unit time or unit number of actions is counted to generate recurrence pattern parameters.

9. The wind turbine health diagnosis and management system based on voiceprint recognition multi-modal fusion analysis according to claim 8, characterized in that, Extract the time-frequency domain characteristic trend sequence of the abnormal sound in the real abnormal unit, calculate the energy accumulation slope, center frequency offset and duration broadening ratio of the time-frequency domain characteristic trend sequence in the process of continuous occurrence, and construct the degradation pointing index based on the energy accumulation slope, center frequency offset and duration broadening ratio.

10. The wind turbine health diagnosis and management system based on voiceprint recognition multi-modal fusion analysis according to claim 9, characterized in that, When the recurrence pattern parameter is lower than the preset occasional threshold and the degradation index is in a stable range, a health management unit of observation and review type is generated, and a prompt message is output to trigger manual auscultation or vibration re-examination. When the recurrence pattern parameter is higher than the preset occasional threshold and the degradation index is in the slow change range, a health management unit of the operation intervention type is generated, and the command is output to limit the unit power ramp rate or adjust the pitch rate. When the degradation index exceeds a preset severe threshold and shows an accelerating upward trend, a health management unit of the maintenance type is generated, and an alarm is output to trigger a shutdown inspection or component replacement work order.