An abnormal detection method and system for wind turbine blades based on voiceprint acquisition

By characterizing and extracting the voiceprint data of wind turbine blades, combined with fault type weight analysis, the problems of low efficiency and insufficient accuracy of wind turbine blade fault detection in the existing technology are solved, efficient and accurate blade abnormality detection is achieved, and the stable operation of wind turbines is ensured.

CN120012032BActive Publication Date: 2025-07-18BEIJING ZHONGKE DONGREN TECH CO LTD

Patent Information

Application Number
CN202510505980.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-18
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The existing wind turbine blade fault detection methods are inefficient and insufficiently accurate, making it difficult to detect early potential faults, especially relying on manual inspection and sensor detection to be susceptible to environmental interference.

Method used

The voiceprint acquisition technology is used to characterize and extract the voiceprint data of wind turbine blades. Combined with fault type weight analysis, the abnormal status of the blade is identified through the acoustic feature extraction module and the fault diagnosis subnet to achieve accurate blade fault detection.

Benefits of technology

It realizes efficient and accurate detection of abnormal wind turbine blades, improves power generation efficiency and safety, and reduces economic losses caused by failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012032B_ABST
    Figure CN120012032B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for abnormal detection of wind turbine blades based on voiceprint acquisition, including: first, performing voiceprint feature encoding on blade voiceprint data to obtain voiceprint embedding features, and extracting abnormal feature vectors therefrom; determining at least two fault types, and identifying abnormal state scores for the abnormal feature vectors according to each fault type; analyzing the voiceprint embedding features to obtain fault type weights; combining the weights to perform weight assignment on the abnormal state scores to obtain a target abnormal state detection result; and finally determining the blade abnormal detection result accordingly. With such a design, effective detection of abnormalities in wind turbine blades is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voiceprint data processing, and more particularly, to a method and system for detecting abnormalities in wind turbine blades based on voiceprint acquisition. Background Art

[0002] Wind turbine blades are in a complex and harsh environment for a long time and are prone to failures such as blade cracks and wear, which affect power generation efficiency and safety. Existing blade fault detection methods have certain limitations. Some rely on manual regular inspections, which are inefficient and difficult to detect early potential faults; some are based on sensors such as vibration and temperature, which are easily affected by the environment and have poor accuracy. Summary of the Invention

[0003] The purpose of the present invention is to provide a method and system for detecting abnormalities in wind turbine blades based on voiceprint acquisition.

[0004] In a first aspect, an embodiment of the present invention provides a method for detecting abnormalities in wind turbine blades based on voiceprint acquisition, including:

[0005] Performing voiceprint feature encoding on the voiceprint data of the wind turbine blade to obtain a voiceprint embedding feature corresponding to the voiceprint data of the blade;

[0006] Performing abnormal feature extraction on the voiceprint embedding feature to obtain an abnormal feature vector corresponding to the voiceprint embedding feature;

[0007] Determining at least two fault types related to the voiceprint data of the blade, and respectively performing abnormal state recognition on the abnormal feature vector according to each fault type to obtain an abnormal state score corresponding to each fault type;

[0008] Performing fault weight analysis on the voiceprint embedding feature to obtain a fault type weight corresponding to each fault type;

[0009] Based on the fault type weight corresponding to each fault type, performing weight value assignment on the abnormal state score corresponding to each fault type to obtain a target abnormal state detection result corresponding to each fault type; a weight frequency domain response value corresponding to a fault type is used for weight value assignment of the abnormal state score corresponding to the corresponding fault type;

[0010] Determining the abnormal detection result of the wind turbine blade according to the target abnormal state detection result corresponding to each fault type.

[0011] In a second aspect, an embodiment of the present invention provides a server system, including a server, and the server is used to execute the method described in the first aspect.

[0012] Compared with the prior art, the beneficial effects provided by the present invention include: adopting an abnormal detection method and system for wind turbine blades based on voiceprint acquisition disclosed by the present invention, obtaining voiceprint embedding features by performing voiceprint feature encoding on the voiceprint data of the blades, and extracting abnormal feature vectors therefrom; determining at least two fault types, and identifying abnormal state scores for the abnormal feature vectors according to each fault type; analyzing the voiceprint embedding features to obtain fault type weights; combining the weights to perform weight assignment on the abnormal state scores to obtain a target abnormal state detection result; and finally determining the abnormal detection result of the blades accordingly. This method realizes the effective detection of abnormalities in wind turbine blades. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0014] Figure 1 It is a schematic flow chart of the steps of the abnormal detection method for wind turbine blades based on voiceprint acquisition provided by the embodiments of the present invention;

[0015] Figure 2 It is a schematic block diagram of the structure of the computer device provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and illustrated in the drawings here can be arranged and designed in various different configurations.

[0017] The following will describe in detail the specific embodiments of the present invention with reference to the accompanying drawings.

[0018] To solve the technical problems in the foregoing background art, Figure 1 It is a schematic flow chart of the abnormal detection method for wind turbine blades based on voiceprint acquisition provided by the embodiments of the present disclosure. The following will introduce this abnormal detection method for wind turbine blades based on voiceprint acquisition in detail.

[0019] Step S201: Perform voiceprint feature encoding on the voiceprint data of the wind turbine blades to obtain the voiceprint embedding features corresponding to the voiceprint data of the blades;

[0020] Step S202: Extract abnormal features from the voiceprint embedding features to obtain the abnormal feature vector corresponding to the voiceprint embedding features;

[0021] Step S203: Determine at least two fault types associated with the blade voiceprint data, and respectively perform abnormal state recognition on the abnormal feature vector according to each fault type to obtain the abnormal state score corresponding to each fault type;

[0022] Step S204: Conduct fault weight analysis on the voiceprint embedding features to obtain the fault type weight corresponding to each fault type;

[0023] Step S205: Based on the fault type weight corresponding to each fault type, perform weight value assignment on the abnormal state score corresponding to each fault type to obtain the target abnormal state detection result corresponding to each fault type; the weight frequency domain response value corresponding to a fault type is used for weight value assignment of the abnormal state score corresponding to the corresponding fault type;

[0024] Step S206: Determine the blade abnormal detection result of the wind turbine according to the target abnormal state detection result corresponding to each fault type.

[0025] In an embodiment of the present invention, by way of example, in a large-scale wind farm, there are numerous wind turbines distributed. During the rotation of the blades of each generator, unique sound signals are generated due to their own states and surrounding environmental factors. These signals are collected by high-precision microphones arranged at appropriate positions in the wind turbine generator tower barrel or blade cavity, forming blade voiceprint data, and transmitted to the server. After the server obtains the blade voiceprint data, it will call a blade fault detection model for blade anomaly detection, and this model includes a first acoustic feature extraction module. Taking a certain wind turbine as an example, the voiceprint preprocessing layer in the first acoustic feature extraction module starts to work. It will determine the acoustic signal characteristics associated with the blade voiceprint data based on its own algorithm, such as the frequency range of the sound, amplitude changes, etc. Then, according to these characteristics, the blade voiceprint data is frame-processed to obtain multiple frame voiceprint features, among which there are target frame voiceprint features. The server further determines the corresponding time-frequency analysis domain according to the acoustic signal characteristics to which the target frame voiceprint features belong. For example, if the target frame voiceprint features are mainly concentrated in the medium and high frequency bands, then the corresponding time-frequency analysis domain is the range related to the medium and high frequencies. This time-frequency analysis domain contains multiple frequency domain response values, and each frequency domain response value has a corresponding frame feature range. Next, the voiceprint preprocessing layer determines the frame feature range to which the target frame voiceprint features belong, and determines the frequency domain response values corresponding to this range as the time-frequency feature map corresponding to the target frame voiceprint features. Subsequently, the first acoustic feature extraction module performs time-frequency feature reconstruction on the time-frequency feature maps corresponding to all frame voiceprint features. These time-frequency feature maps are combined according to specific rules to obtain the time-frequency feature data of the blade voiceprint data. Finally, the time-frequency feature data is encoded with voiceprint features to obtain the voiceprint embedding features of the blade voiceprint data. After the server completes the extraction of the voiceprint embedding features, it will use the acoustic feature encoder in the first acoustic feature extraction module to extract anomaly features from the voiceprint embedding features. When training this blade fault detection model, the server will obtain a first voiceprint sample set, and the first voiceprint training data in it is not configured with anomaly type labels. At the same time, the server will also obtain multiple self-supervised learning tasks for the first original acoustic feature extraction module, such as time-frequency map mask reconstruction tasks, time-frequency anomaly replacement detection tasks, voiceprint denoising contrast tasks, and expert knowledge transfer tasks, etc. Taking the time-frequency map mask reconstruction task as an example, the server performs time-frequency transformation processing on the first voiceprint training data based on the first original acoustic feature extraction module to obtain the first sample time-frequency feature data, which includes sample time-frequency feature maps at multiple time frame positions. Then, the server selects the first time frame positions to be masked in the time-frequency region from these time frame positions according to a preset rule. Assuming the preset rule is to randomly select a certain proportion of time frame positions, the server may randomly select several time frame positions as the first time frame positions.Next, the server performs time-frequency region masking processing on the sample time-frequency feature map at the first time frame position in the first sample time-frequency feature data according to the masking region identifier to obtain masked time-frequency sample data. After that, the server performs voiceprint feature encoding on the masked time-frequency sample data based on the first original acoustic feature extraction module to obtain the first training voiceprint embedding feature, and performs abnormal feature extraction on it to obtain the first acoustic feature representation. Then, based on the feature reconstruction module related to this task, a reconstructed time-frequency map masking process is performed on the first acoustic feature representation to obtain a predicted time-frequency feature map. Finally, according to the sample time-frequency feature map and the predicted time-frequency feature map at the first time frame position in the first sample time-frequency feature data, the error parameter corresponding to this task is determined. By continuously adjusting these error parameters, the model parameters of the first original acoustic feature extraction module are optimized, and finally a first acoustic feature extraction module with good performance is obtained, which is used to accurately extract abnormal feature vectors. In the actual operation of a wind farm, there may be various types of blade failures, such as blade cracks, surface wear, loose connection components, etc. The server will determine at least two failure types related to the current blade voiceprint data. The second acoustic feature extraction module in the blade failure detection model includes a fault diagnosis subnet for each failure type. Taking the voiceprint data of a certain wind turbine blade as an example, the server determines the second acoustic feature extraction module from the blade failure detection model, and then based on the fault diagnosis subnet corresponding to each failure type, performs abnormal state recognition on the previously extracted abnormal feature vectors. When training these fault diagnosis subnets, the server obtains a second voiceprint sample set, in which the second voiceprint training data is associated with fault type labels, and these labels correspond to at least two fault type encodings, and each fault type corresponds to one encoding. The server uses the pre-trained first acoustic feature extraction module to perform voiceprint feature encoding on the second voiceprint training data to obtain the fifth training voiceprint embedding feature, and performs abnormal feature extraction to obtain the fifth acoustic feature representation. Then, the server obtains at least two original fault diagnosis subnets, respectively performs fault type discrimination on the fifth acoustic feature representation, and obtains at least two fault type probabilities corresponding to the at least two original fault diagnosis subnets. According to these probabilities and the fault type encodings, the second target error parameter is determined, and the model parameters of the original fault diagnosis subnet are optimized to obtain an optimized fault diagnosis subnet, so as to accurately identify the abnormal feature vectors for each fault type and obtain the abnormal state score corresponding to each fault type. For example, for the fault type of blade crack, the fault diagnosis subnet analyzes and calculates to give an abnormal state score reflecting the possibility of the current blade voiceprint data corresponding to the blade crack fault. The server determines the third acoustic feature extraction module from the blade failure detection model, and this module includes a multi-branch feature extractor and a dynamic weight allocator. The server performs multi-modal acoustic feature extraction on the voiceprint embedding feature based on the multi-branch feature extractor.For example, analyze from multiple dimensions such as the frequency, phase, and amplitude of the voiceprint to obtain the multimodal feature mapping results corresponding to the voiceprint embedding features. Then, the dynamic weight allocator calculates the fault weights for these multimodal feature mapping results to obtain the fault type weights corresponding to each fault type. When training the third acoustic feature extraction module, the server obtains the third voiceprint sample set, and the third voiceprint training data therein is associated with abnormal type labels. The server uses the pre-trained first acoustic feature extraction module to perform voiceprint feature encoding on the third voiceprint training data to obtain the sixth training voiceprint embedding feature, and performs abnormal feature extraction to obtain the sixth acoustic feature representation. Then, use the fault diagnosis subnet in the pre-trained second acoustic feature extraction module to identify the abnormal state of the sixth acoustic feature representation to obtain the training abnormal state scores corresponding to each fault type. At the same time, based on the third original acoustic feature extraction module, perform fault weight analysis on the sixth training voiceprint embedding feature to obtain the training fault type weights corresponding to each fault type. According to the training abnormal state scores and the training fault type weights, determine the training blade abnormal detection results of the reference wind turbine associated with the third voiceprint training data. Then, determine the third target error parameter according to the abnormal type label and the training blade abnormal detection results, and optimize the model parameters of the third original acoustic feature extraction module to obtain a third acoustic feature extraction module with good performance, and accurately calculate the fault type weights corresponding to each fault type. For example, after analysis and calculation, it is determined that the weight of the blade crack fault type is 0.6, and the weight of the surface wear fault type is 0.4, etc. After the server obtains the abnormal state scores and fault type weights corresponding to each fault type, it performs weight allocation. For example, for the blade crack fault type, assume its abnormal state score is 80 points and the fault type weight is 0.6; for the surface wear fault type, the abnormal state score is 70 points and the fault type weight is 0.4. Then the target abnormal state detection result corresponding to the blade crack fault type is 80×0.6 = 48 points, and the target abnormal state detection result corresponding to the surface wear fault type is 70×0.4 = 28 points. The weight frequency domain response value here is used to perform weighted calculation on the abnormal state scores of the corresponding fault types, so as to obtain a target abnormal state detection result that can more accurately reflect the actual abnormal degree of the fault type. The server synthesizes the target abnormal state detection results corresponding to all fault types to determine the abnormal detection result of the blade. For example, if the target abnormal state detection result of the blade crack fault type exceeds the preset severe abnormal threshold, and the target abnormal state detection results of other fault types are relatively low, the server can judge that there is a greater possibility of blade crack fault in the wind turbine blade, and it is necessary to arrange maintenance personnel to check and repair the blade in time. If the target abnormal state detection results of all fault types are within the normal range, it indicates that the current blade is in good operating condition.In this way, the server can provide accurate blade anomaly detection information for the operation and maintenance personnel of the wind farm, ensuring the stable operation of the wind turbine, improving power generation efficiency, and reducing safety risks and economic losses caused by blade failures.

[0026] In an embodiment of the present invention, the acoustic feature encoding of the blade acoustic data of the wind turbine to obtain the acoustic embedding feature corresponding to the blade acoustic data can be implemented through the following examples.

[0027] Perform time-frequency transformation processing on the blade acoustic data to obtain the time-frequency feature data of the blade acoustic data;

[0028] Perform acoustic feature encoding on the time-frequency feature data to obtain the acoustic embedding feature of the time-frequency feature data.

[0029] In an embodiment of the present invention, for example, after the server receives the blade acoustic data of a certain wind turbine, it starts the first step of acoustic feature encoding - performing time-frequency transformation processing on the blade acoustic data to obtain time-frequency feature data. For example, there have been some abnormal vibrations in the blades of this wind turbine recently, resulting in fluctuations in the collected acoustic data that are different from the normal state in the frequency and time dimensions. The server uses a specific time-frequency transformation algorithm, such as empirical mode decomposition, to convert these acoustic data from a simple time-domain signal into time-frequency feature data that contains both time and frequency information. For example, through the transformation, it can be seen that there is an energy concentration phenomenon in the high-frequency band of the acoustic data during a certain period, which may imply that the blade has encountered a specific abnormal condition at that moment. The server performs acoustic feature encoding on the obtained time-frequency feature data to obtain the acoustic embedding feature of the time-frequency feature data. This process is like creating a unique "digital identity" for the time-frequency feature data. The server uses an acoustic feature encoding algorithm to extract and encode the key features in the time-frequency feature data into a compact vector representation form, that is, the acoustic embedding feature. Taking the above-mentioned wind turbine as an example, the server extracts key information such as the frequency range and intensity change reflecting the abnormal vibration in the time-frequency feature data and converts it into a specific acoustic embedding feature vector. This vector contains the core features of the blade acoustic, and based on this vector, the server can subsequently carry out a series of important operations such as abnormal feature extraction and fault type identification, accurately judge the operating state of the blade, and ensure the stable and efficient operation of the wind turbine.

[0030] In an embodiment of the present invention, the performing time-frequency transformation processing on the blade acoustic data to obtain the time-frequency feature data of the blade acoustic data can be implemented through the following examples.

[0031] When obtaining the blade voiceprint data of a wind turbine, obtain a blade fault detection model for blade anomaly detection of the wind turbine; the blade fault detection model includes a first acoustic feature extraction module;

[0032] Based on the first acoustic feature extraction module, determine the acoustic signal characteristics related to the blade voiceprint data, and extract the framed voiceprint features of the blade voiceprint data according to the acoustic signal characteristics;

[0033] Based on the first acoustic feature extraction module, perform time-frequency feature conversion on the framed voiceprint features to obtain a time-frequency feature map corresponding to the framed voiceprint features;

[0034] Based on the first acoustic feature extraction module, perform time-frequency feature reconstruction on the time-frequency feature map corresponding to the framed voiceprint features to obtain the time-frequency feature data of the blade voiceprint data.

[0035] In an embodiment of the present invention, by way of example, in a large wind farm, the server continuously receives the blade voiceprint data from each wind turbine. When the server obtains the blade voiceprint data of a certain wind turbine, it immediately calls from the storage a blade fault detection model for blade anomaly detection of this wind turbine. This model includes a first acoustic feature extraction module. The first acoustic feature extraction module starts to work. It analyzes the incoming blade voiceprint data and determines the associated acoustic signal characteristics. For example, it identifies characteristics such as the signal intensity change pattern in a specific frequency range and the periodicity of the signal in the blade voiceprint data. Based on these characteristics, the module performs framing processing on the blade voiceprint data and extracts the framed voiceprint features. Suppose the duration of the blade voiceprint data is 10 seconds. The module divides it into frames every 0.1 second, obtaining 100 framed voiceprint features. Each frame carries the voiceprint characteristics within this 0.1 second. Then, the first acoustic feature extraction module performs time-frequency feature conversion on these framed voiceprint features. It converts them into a time-frequency feature map based on the frequency, amplitude, and other information of each framed voiceprint feature. For example, for a certain framed voiceprint feature, the module uses a specific algorithm to display the frequency information at different times within this frame in the form of a two-dimensional graph, forming a time-frequency feature map. The colors or grayscales at different positions in the map represent different frequency intensities. Finally, the first acoustic feature extraction module performs time-frequency feature reconstruction on the time-frequency feature maps corresponding to all the framed voiceprint features. It arranges and combines the time-frequency feature maps of each frame in chronological order to construct a multi-dimensional matrix, which is the time-frequency feature data of the blade voiceprint data.

[0036] In an embodiment of the present invention, the first acoustic feature extraction module includes a voiceprint preprocessing layer; the framed voiceprint features include a plurality of framed voiceprint features, and the plurality of framed voiceprint features include a target framed voiceprint feature;

[0037] The time-frequency feature conversion is performed on the framed voiceprint features based on the first acoustic feature extraction module to obtain a time-frequency feature map corresponding to the framed voiceprint features, which can be implemented through the following examples.

[0038] According to the acoustic signal characteristics to which the target framed voiceprint features belong, determine the time-frequency analysis domain corresponding to the target framed voiceprint features; the time-frequency analysis domain includes a plurality of frequency domain response values, and each frequency domain response value has a corresponding framed feature range;

[0039] Based on the voiceprint preprocessing layer, determine the framed feature range to which the target framed voiceprint features belong, and determine the frequency domain response value corresponding to the framed feature range to which the target framed voiceprint features belong as the time-frequency feature map corresponding to the target framed voiceprint features.

[0040] In an embodiment of the present invention, by way of example, in a certain wind farm, after the server obtains the blade voiceprint data of a wind turbine, it calls the first acoustic feature extraction module including a voiceprint preprocessing layer. This module performs framing processing on the blade voiceprint data to obtain a plurality of framed voiceprint features, and one of them is selected as the target framed voiceprint feature. The server starts to process this target framed voiceprint feature based on the first acoustic feature extraction module. First, according to the acoustic signal characteristics possessed by the target framed voiceprint feature itself, such as its frequency distribution range, energy concentration frequency band, etc., to determine its corresponding time-frequency analysis domain. For example, if the energy of the target framed voiceprint feature is mainly concentrated in the frequency band of 500 Hz - 1000 Hz, the server determines its corresponding time-frequency analysis domain accordingly. This time-frequency analysis domain contains a plurality of frequency domain response values, and each frequency domain response value corresponds to a specific framed feature range. Then, the voiceprint preprocessing layer comes into play. It further analyzes the target framed voiceprint feature to accurately determine the framed feature range to which the target framed voiceprint feature belongs. Suppose after analysis, it is determined that the target framed voiceprint feature is in the frequency range of 600 Hz - 800 Hz and the amplitude is in a certain specific interval, which is the framed feature range to which it belongs. Then, the server extracts the frequency domain response value corresponding to this framed feature range and further determines it as the time-frequency feature map corresponding to the target framed voiceprint feature. This time-frequency feature map intuitively presents the characteristics of the target framed voiceprint feature within a specific frequency and time range, providing key information for subsequent comprehensive analysis and anomaly detection of the blade voiceprint data. Through such fine processing, the server can extract key features from the complex blade voiceprint data, laying a foundation for accurately judging whether there is an anomaly in the wind turbine blade.

[0041] In an embodiment of the present invention, the abnormal feature vector is determined by an acoustic feature encoder in a first acoustic feature extraction module for extracting abnormal features from the voiceprint embedding feature; the first acoustic feature extraction module belongs to a blade fault detection model for performing blade abnormal detection on the wind turbine; the embodiment of the present invention also provides the following implementation manners:

[0042] Obtain a first voiceprint sample set; the first voiceprint training data included in the first voiceprint sample set is voiceprint training data without an abnormal type label configured;

[0043] Obtain multiple self-supervised learning tasks for a first original acoustic feature extraction module;

[0044] Based on the first original acoustic feature extraction module, perform data processing on the first voiceprint training data to obtain an acoustic feature representation associated with each self-supervised learning task, and determine multiple error parameters corresponding to the multiple self-supervised learning tasks according to the acoustic feature representation; one self-supervised learning task corresponds to one error parameter;

[0045] Determine a first target error parameter according to the multiple error parameters, optimize the model parameters of the first original acoustic feature extraction module according to the first target error parameter, and determine the optimized first original acoustic feature extraction module as the first acoustic feature extraction module.

[0046] In an embodiment of the present invention, exemplarily, the server first needs to obtain a first voiceprint sample set, and no abnormal type labels are configured for the first voiceprint training data in this sample set. For example, this data may come from the voiceprints collected when each wind turbine blade is operating normally at different time periods and under different wind conditions in a power generation field, and also includes a small amount of voiceprint data that may be abnormal but the abnormal type has not been determined. Next, the server obtains multiple self-supervised learning tasks for the first original acoustic feature extraction module. These tasks are designed to enable the module to automatically learn useful features from unlabeled data. Taking the time-frequency map mask reconstruction task as an example, the server performs time-frequency transformation processing on the first voiceprint training data based on the first original acoustic feature extraction module, and converts the voiceprint data into first sample time-frequency feature data including sample time-frequency feature maps at multiple time frame positions. The server selects a first time frame position to be subjected to time-frequency region masking from numerous time frame positions according to a predetermined preset rule. For example, if the preset rule is to select one out of every 10 time frame positions, the server will select the corresponding positions accordingly. Then, the server performs time-frequency region masking processing on the sample time-frequency feature map at the first time frame position in the first sample time-frequency feature data according to the masking region identifier, obtaining masked time-frequency sample data. After that, the server performs voiceprint feature encoding on the masked time-frequency sample data again based on the first original acoustic feature extraction module, obtains a first training voiceprint embedding feature, and then performs abnormal feature extraction on it to obtain a first acoustic feature representation. The server uses a feature reconstruction module associated with this task to perform reconstructed time-frequency map masking processing on the first acoustic feature representation, obtaining a predicted time-frequency feature map. Finally, by comparing the sample time-frequency feature map at the first time frame position in the first sample time-frequency feature data with the predicted time-frequency feature map, the error parameter corresponding to this task is determined. For other self-supervised learning tasks, such as time-frequency anomaly replacement detection tasks, voiceprint denoising contrast tasks, etc., the server also processes them according to a similar process, respectively obtaining their corresponding error parameters. After the server collects multiple error parameters corresponding to all self-supervised learning tasks, it determines a first target error parameter through a specific algorithm. This algorithm may perform operations such as weighted averaging on all error parameters. Based on the first target error parameter, the server optimizes the model parameters of the first original acoustic feature extraction module, adjusting parameters such as weights and biases inside the module. After a series of optimizations, the first original acoustic feature extraction module with improved performance is determined as the first acoustic feature extraction module, which is used to accurately extract abnormal feature vectors from voiceprint embedding features subsequently to more precisely detect abnormal conditions of wind turbine blades.

[0047] In an embodiment of the present invention, the multiple self-supervised learning tasks include a time-frequency map mask reconstruction task;

[0048] The above-mentioned first original acoustic feature extraction module processes the first voiceprint training data to obtain an acoustic feature representation associated with each self-supervised learning task, and determines a plurality of error parameters corresponding to the plurality of self-supervised learning tasks according to the acoustic feature representation. The following example can be used for implementation.

[0049] Perform time-frequency transformation processing on the first voiceprint training data based on the first original acoustic feature extraction module to obtain first sample time-frequency feature data of the first voiceprint training data; the first sample time-frequency feature data includes sample time-frequency feature maps at a plurality of time frame positions.

[0050] Select a first time frame position to be masked in the time-frequency region from the plurality of time frame positions corresponding to the first sample time-frequency feature data according to a preset rule.

[0051] Perform time-frequency region masking processing on the sample time-frequency feature map at the first time frame position in the first sample time-frequency feature data according to the masking region identifier to obtain masked time-frequency sample data.

[0052] Perform voiceprint feature encoding on the masked time-frequency sample data based on the first original acoustic feature extraction module to obtain a first training voiceprint embedding feature of the masked time-frequency sample data, and perform abnormal feature extraction on the first training voiceprint embedding feature to obtain a first acoustic feature representation corresponding to the first training voiceprint embedding feature; the first acoustic feature representation belongs to the acoustic feature representation.

[0053] Based on a feature reconstruction module associated with the time-frequency map mask reconstruction task, perform reconstructed time-frequency map masking processing on the first acoustic feature representation to obtain a predicted time-frequency feature map corresponding to the masking region identifier in the masked time-frequency sample data.

[0054] Determine an error parameter corresponding to the time-frequency map mask reconstruction task according to the sample time-frequency feature map at the first time frame position in the first sample time-frequency feature data and the predicted time-frequency feature map.

[0055] In an embodiment of the present invention, exemplarily, in a monitoring system of a wind farm, the server is responsible for processing a large amount of data related to wind turbine blades. At this time, the server is executing a time-frequency map mask reconstruction task based on a first original acoustic feature extraction module, which is one of multiple self-supervised learning tasks. The server first performs time-frequency transformation processing on the first voiceprint training data based on the first original acoustic feature extraction module. Assume that the first voiceprint training data is the voiceprint record of a certain wind turbine blade within a period of time. After processing, first sample time-frequency feature data is obtained, and each time frame position records the sample time-frequency feature map corresponding to the corresponding moment, presenting the characteristics of the voiceprint at that moment in the time and frequency dimensions. Then, the server selects, according to a preset rule, the first time frame positions to be subjected to time-frequency region masking from these multiple time frame positions. The preset rule may be to randomly select 10% of the time frame positions. The server determines specific first time frame positions from among numerous time frame positions through a random algorithm. After that, the server performs time-frequency region masking processing on the sample time-frequency feature map at the first time frame position in the first sample time-frequency feature data according to the masking region identifier. For example, if the masking region identifier specifies a region in a certain specific frequency range, the server masks the corresponding region in this frequency range in the sample time-frequency feature map to obtain masked time-frequency sample data. Then, the server performs voiceprint feature encoding on the masked time-frequency sample data again based on the first original acoustic feature extraction module to generate a first training voiceprint embedding feature, and then performs abnormal feature extraction on it to obtain a first acoustic feature representation corresponding to the first training voiceprint embedding feature, and this representation belongs to a part of the acoustic feature representation. Next, the server calls a feature reconstruction module associated with this task to perform reconstruction time-frequency map masking processing on the first acoustic feature representation. The feature reconstruction module attempts to restore the time-frequency features of the masked region based on the information in the first acoustic feature representation to obtain a predicted time-frequency feature map corresponding to the masking region identifier in the masked time-frequency sample data. Finally, the server calculates the difference between the original sample time-frequency feature map at the first time frame position in the first sample time-frequency feature data and the reconstructed predicted time-frequency feature map, and determines the error parameter corresponding to the time-frequency map mask reconstruction task based on this. This error parameter reflects the accuracy of the first original acoustic feature extraction module in processing this task and provides an important basis for subsequent optimization modules.

[0056] In an embodiment of the present invention, the multiple self-supervised learning tasks include a time-frequency anomaly replacement detection task;

[0057] Performing data processing on the first voiceprint training data based on the first original acoustic feature extraction module to obtain an acoustic feature representation associated with each self-supervised learning task, and determining multiple error parameters corresponding to the multiple self-supervised learning tasks according to the acoustic feature representation can be implemented through the following example.

[0058] Based on the first original acoustic feature extraction module, perform time-frequency transformation processing on the first voiceprint training data to obtain the first sample time-frequency feature data of the first voiceprint training data; the first sample time-frequency feature data includes sample time-frequency feature maps at multiple time frame positions;

[0059] Select a second time frame position to be subjected to time-frequency feature replacement from the multiple time frame positions corresponding to the first sample time-frequency feature data according to a preset rule;

[0060] According to preset time-frequency information, perform time-frequency feature replacement processing on the sample time-frequency feature map at the second time frame position in the first sample time-frequency feature data to obtain replacement time-frequency sample data; the preset time-frequency information is different from the sample time-frequency feature map at the second time frame position; the preset time-frequency information is an abnormal time-frequency feature segment extracted from historical fault data or an unnatural time-frequency feature generated by Gaussian noise;

[0061] Based on the first original acoustic feature extraction module, perform voiceprint feature encoding on the replacement time-frequency sample data to obtain the second training voiceprint embedding feature of the replacement time-frequency sample data, and perform abnormal feature extraction on the second training voiceprint embedding feature to obtain the second acoustic feature representation corresponding to the second training voiceprint embedding feature; the second acoustic feature representation belongs to the acoustic feature representation;

[0062] Based on a feature reconstruction module associated with the time-frequency anomaly replacement detection task, determine the target detection time frame position for the second acoustic feature representation to obtain the time-frequency feature replacement prediction results corresponding to the multiple time frame positions;

[0063] According to the second time frame position and the time-frequency feature replacement prediction results corresponding to the multiple time frame positions, determine the error parameter corresponding to the time-frequency anomaly replacement detection task.

[0064] In an embodiment of the present invention, exemplarily, at the data processing server side of a large-scale wind farm, a time-frequency anomaly replacement detection task for the acoustic fingerprint data of wind turbine blades is being carried out, which is one of multiple self-supervised learning tasks. The server first performs time-frequency transformation processing on the first acoustic fingerprint training data using the first original acoustic feature extraction module. For example, the training data comes from the acoustic fingerprints of blades of multiple wind turbines at different times in a certain wind farm. After processing, the first sample time-frequency feature data is obtained, which is like an ordered set, and each element is a sample time-frequency feature map at a time frame position, showing the frequency characteristics of the acoustic fingerprint at different moments. Next, the server selects, according to a preset rule, a second time frame position to be subjected to time-frequency feature replacement from among numerous time frame positions. Suppose the preset rule is to select positions where the serial number of the time frame position is odd and the frequency fluctuation exceeds a certain threshold. The server determines the second time frame positions that meet the conditions through analysis and screening of the sample time-frequency feature data. After that, the server performs time-frequency feature replacement processing on the sample time-frequency feature map at the second time frame position in the first sample time-frequency feature data according to the preset time-frequency information. For example, an abnormal time-frequency feature segment is extracted from historical fault data as the preset time-frequency information, and this information is different from the sample time-frequency feature map at the current second time frame position. The server replaces the corresponding part of the original sample time-frequency feature map with this abnormal time-frequency feature segment to obtain the replaced time-frequency sample data. Subsequently, the server again performs acoustic fingerprint feature encoding on the replaced time-frequency sample data based on the first original acoustic feature extraction module to obtain the second training acoustic fingerprint embedding feature, and performs abnormal feature extraction on it, thereby obtaining the second acoustic feature representation corresponding to the second training acoustic fingerprint embedding feature, and this representation is part of the overall acoustic feature representation. Next, the server uses a feature reconstruction module associated with this task to determine the target detection time frame position for the second acoustic feature representation. The feature reconstruction module analyzes and judges the time-frequency feature replacement situation corresponding to multiple time frame positions based on the information in the second acoustic feature representation to obtain the time-frequency feature replacement prediction result. Finally, the server calculates the difference between the initially selected second time frame position and the time-frequency feature replacement prediction results corresponding to the obtained multiple time frame positions, and thereby determines the error parameter corresponding to the time-frequency anomaly replacement detection task. This error parameter reflects the processing accuracy of the first original acoustic feature extraction module in this task and provides a key basis for optimizing the performance of the subsequent module.

[0065] In an embodiment of the present invention, the multiple self-supervised learning tasks include an acoustic fingerprint denoising contrast task;

[0066] Performing data processing on the first acoustic fingerprint training data based on the first original acoustic feature extraction module to obtain an acoustic feature representation associated with each self-supervised learning task, and determining multiple error parameters corresponding to the multiple self-supervised learning tasks according to the acoustic feature representation can be implemented through the following example.

[0067] Inject Gaussian white noise features into the framed voiceprint features of the first voiceprint training data according to the signal-to-noise ratio to obtain noisy voiceprint training data, and perform time-frequency transformation processing on the noisy voiceprint training data to obtain the noise time-frequency feature data of the noisy voiceprint training data;

[0068] Perform time-frequency transformation processing on the first voiceprint training data to obtain the first sample time-frequency feature data of the first voiceprint training data;

[0069] Based on the first original acoustic feature extraction module, perform voiceprint feature encoding on the noise time-frequency feature data to obtain the third training voiceprint embedding feature of the noise time-frequency feature data, and perform abnormal feature extraction on the third training voiceprint embedding feature to obtain the third acoustic feature representation corresponding to the third training voiceprint embedding feature; the third acoustic feature representation belongs to the acoustic feature representation;

[0070] Based on the first original acoustic feature extraction module, perform voiceprint feature encoding on the first sample time-frequency feature data to obtain the fourth training voiceprint embedding feature of the first sample time-frequency feature data, and perform abnormal feature extraction on the fourth training voiceprint embedding feature to obtain the fourth acoustic feature representation corresponding to the fourth training voiceprint embedding feature; the fourth acoustic feature representation belongs to the acoustic feature representation;

[0071] Determine the error parameter corresponding to the voiceprint denoising comparison task according to the third acoustic feature representation and the fourth acoustic feature representation.

[0072] In an embodiment of the present invention, exemplarily, in the operation and maintenance data processing center of a large-scale wind farm, the server is performing a voiceprint denoising comparison task on the voiceprint data of the wind turbine blades, which is an important part of multiple self-supervised learning tasks. The server first processes the first voiceprint training data. These training data are the voiceprints of the blades collected from different wind turbines in the power generation field under different working conditions. The server injects Gaussian white noise features into the framed voiceprint features of the first voiceprint training data according to a specific signal-to-noise ratio. For example, setting the signal-to-noise ratio to 20 dB, adding white noise that conforms to the Gaussian distribution to each framed voiceprint feature, so as to obtain the noisy voiceprint training data. Subsequently, perform time-frequency transformation processing on the noisy voiceprint training data to obtain the noise time-frequency feature data of the noisy voiceprint training data, and this data presents the characteristics of the voiceprint with noise interference in the time and frequency dimensions. At the same time, the server also performs time-frequency transformation processing on the first voiceprint training data without adding noise to obtain the first sample time-frequency feature data of the first voiceprint training data, and this data represents the time-frequency characteristics of the original pure voiceprint. Next, the server further processes the noise time-frequency feature data and the first sample time-frequency feature data respectively based on the first original acoustic feature extraction module. For the noise time-frequency feature data, after voiceprint feature encoding, the third training voiceprint embedding feature is obtained, and then abnormal feature extraction is performed to obtain the third acoustic feature representation corresponding to the third training voiceprint embedding feature, and this representation belongs to a part of the overall acoustic feature representation. Similarly, perform voiceprint feature encoding on the first sample time-frequency feature data to obtain the fourth training voiceprint embedding feature, and then perform abnormal feature extraction to obtain the fourth acoustic feature representation corresponding to the fourth training voiceprint embedding feature. Finally, the server determines the error parameter corresponding to the voiceprint denoising comparison task according to the third acoustic feature representation and the fourth acoustic feature representation. The server calculates the difference between these two acoustic feature representations, such as calculating metrics such as the distance between the two in the feature vector space and the deviation of key feature parameters, so as to quantify the performance difference of the first original acoustic feature extraction module when processing noisy and pure voiceprint data, and this difference value is the error parameter corresponding to the voiceprint denoising comparison task. This error parameter will provide an important basis for the subsequent optimization of the first original acoustic feature extraction module to improve its processing ability of the blade voiceprint data in a noisy environment and ensure the accuracy of the abnormal detection of the wind turbine blades.

[0073] In an embodiment of the present invention, the multiple self-supervised learning tasks include an expert knowledge transfer task;

[0074] Based on the first original acoustic feature extraction module to process the first voiceprint training data, obtain the acoustic feature representation associated with each self-supervised learning task, and determine the multiple error parameters corresponding to the multiple self-supervised learning tasks according to the acoustic feature representation. The following examples can be used for implementation.

[0075] Based on a feature reconstruction module related to the expert knowledge transfer task, perform an abnormal risk rating on the fourth acoustic feature representation to obtain the abnormal risk score corresponding to the first voiceprint training data;

[0076] Obtain a pre-trained benchmark model related to the first original acoustic feature extraction module, and perform an abnormal risk rating on the first voiceprint training data according to the pre-trained benchmark model to obtain the benchmark abnormal state score corresponding to the first voiceprint training data;

[0077] Determine the error parameter corresponding to the expert knowledge transfer task according to the abnormal risk score and the benchmark abnormal state score.

[0078] In an embodiment of the present invention, exemplarily, on a data processing server in a wind farm, an expert knowledge transfer task is being carried out, which is part of multiple self-supervised learning tasks. The server first performs time-frequency transformation processing on the first voiceprint training data. This first voiceprint training data is collected from the blades of multiple generators in the wind farm and covers voiceprint information in different operating states. After processing, the first sample time-frequency feature data of the first voiceprint training data is obtained, which clearly shows the distribution characteristics of the voiceprint in the time and frequency dimensions. Then, the server uses the first original acoustic feature extraction module to encode the voiceprint features of the first sample time-frequency feature data to obtain the fourth training voiceprint embedding feature, and then extracts abnormal features from it to obtain the fourth acoustic feature representation corresponding to the fourth training voiceprint embedding feature. This representation is an important part of the overall acoustic feature representation. After that, the server calls the feature reconstruction module associated with the expert knowledge transfer task. Based on the knowledge and algorithms it contains, this module performs an abnormal risk rating on the fourth acoustic feature representation. For example, the module will analyze various feature indicators in the fourth acoustic feature representation, such as the energy change in a specific frequency band and the periodicity of the voiceprint. According to these analysis results, an abnormal risk score corresponding to the first voiceprint training data is given. Assuming the score range is 0-100 points, the higher the score, the greater the abnormal risk. After evaluation, the first voiceprint training data obtains a score of 40 points. At the same time, the server obtains a pre-trained benchmark model related to the first original acoustic feature extraction module. This pre-trained benchmark model has been trained on a large amount of data and has accumulated rich judgment experience. The server uses this pre-trained benchmark model to perform an abnormal risk rating on the same first voiceprint training data to obtain the benchmark abnormal state score corresponding to the first voiceprint training data. For example, the score given by the benchmark model is 35 points. Finally, the server determines the error parameter corresponding to the expert knowledge transfer task according to the abnormal risk score and the benchmark abnormal state score. The server quantifies this error parameter by calculating the difference, ratio, etc. between the two. For example, by calculating the absolute value of the difference between the two (40 - 35 = 5), this value is used as part of the error parameter, and then combined with other relevant calculation rules, an error parameter that can reflect the gap between the first original acoustic feature extraction module and the expert knowledge (pre-trained benchmark model) in this task is finally determined. This error parameter will guide the subsequent optimization of the first original acoustic feature extraction module, enabling it to better utilize expert knowledge and improve the accuracy of abnormal detection of voiceprint data of wind turbine blades.

[0079] In an embodiment of the present invention, the abnormal feature vector is determined by an acoustic feature encoder in a first acoustic feature extraction module for extracting abnormal features from the voiceprint embedding features; the first acoustic feature extraction module belongs to a blade fault detection model for performing blade abnormality detection on the wind turbine; the blade fault detection model further includes a second acoustic feature extraction module with multi-level feature fusion after the acoustic feature encoder;

[0080] The abnormal state recognition is respectively performed on the abnormal feature vector according to each fault type to obtain the abnormal state score corresponding to each fault type, which can be implemented through the following examples.

[0081] Determine the second acoustic feature extraction module from the blade fault detection model; the second acoustic feature extraction module includes a fault diagnosis subnet corresponding to each fault type;

[0082] Based on the fault diagnosis subnet corresponding to each fault type, perform abnormal state recognition on the abnormal feature vector to obtain the abnormal state score corresponding to each fault type; one fault diagnosis subnet is used to obtain the abnormal state score corresponding to one fault type.

[0083] In an embodiment of the present invention, exemplarily, the server obtains the voiceprint embedding features obtained after being processed by the first acoustic feature extraction module, and the acoustic feature encoder in this module extracts abnormal features therefrom to determine the abnormal feature vector. And this blade fault detection model further includes a second acoustic feature extraction module located after the acoustic feature encoder, which adopts a multi-level feature fusion technology and can analyze blade faults more accurately. When it is necessary to perform abnormal state recognition on the abnormal feature vector according to different fault types to obtain the abnormal state scores corresponding to each fault type, the server first determines the second acoustic feature extraction module from the blade fault detection model. This module contains a fault diagnosis subnet for each fault type. For example, there is a dedicated fault diagnosis subnet for the blade crack fault type, and there is also a corresponding subnet for the blade surface wear fault type, etc. Taking a certain wind turbine as an example, the server obtains the abnormal feature vector after processing the voiceprint data of the fan blade. For the blade crack fault type, the server calls the corresponding fault diagnosis subnet in the second acoustic feature extraction module. This subnet deeply analyzes the abnormal feature vector based on the pre-trained algorithms and parameters. It may examine specific feature indicators related to blade cracks in the abnormal feature vector, such as the energy change within a specific frequency range, the mutation of the voiceprint signal, etc. Through the comprehensive evaluation of these indicators, the fault diagnosis subnet outputs a value as the abnormal state score corresponding to the blade crack fault type. Assuming a full score of 100 points, after analysis and calculation, the fault diagnosis subnet gives an abnormal state score of 60 points for the blade crack fault type, indicating that the possibility of the blade having a crack fault is at a medium level. Similarly, for the blade surface wear fault type, the server calls the corresponding fault diagnosis subnet, and processes the same abnormal feature vector according to the unique analysis logic and algorithm of this subnet, and obtains the abnormal state score corresponding to the blade surface wear fault type, such as 40 points, indicating that the possibility of the blade having a surface wear fault is relatively low. In this way, the server uses different fault diagnosis subnets to perform abnormal state recognition on the abnormal feature vector respectively for each fault type, and obtains the abnormal state scores corresponding to each fault type, providing a key basis for accurately judging the fault condition of the wind turbine blade.

[0084] In an embodiment of the present invention, the following implementation manners are further provided.

[0085] Obtain a second voiceprint sample set; the second voiceprint training data included in the second voiceprint sample set is associated with a fault type label; the fault type label includes at least two fault type codes corresponding to the at least two fault types, one fault type corresponds to one fault type code, and the at least two fault type codes are determined according to the fault type to which the second voiceprint training data belongs;

[0086] Obtain a pre-trained first acoustic feature extraction module, perform voiceprint feature encoding on the second voiceprint training data based on the first acoustic feature extraction module to obtain a fifth training voiceprint embedding feature of the second voiceprint training data, and perform abnormal feature extraction on the fifth training voiceprint embedding feature to obtain a fifth acoustic feature representation corresponding to the fifth training voiceprint embedding feature;

[0087] Obtain at least two original fault diagnosis subnets corresponding to the at least two fault types, and perform fault type discrimination on the fifth acoustic feature representation according to the at least two to obtain at least two fault type probabilities corresponding to the at least two original fault diagnosis subnets; one fault type corresponds to one original fault diagnosis subnet, and one original fault diagnosis subnet is used to obtain one fault type probability;

[0088] Determine a second target error parameter according to the at least two fault type probabilities and the at least two fault type encodings, optimize the model parameters of the at least two original fault diagnosis subnets according to the second target error parameter, determine the optimized at least two original fault diagnosis subnets as the at least two fault diagnosis subnets, and determine the second acoustic feature extraction module according to the at least two fault diagnosis subnets.

[0089] In an embodiment of the present invention, exemplarily, in a data processing server of a wind farm, in order to more accurately construct a second acoustic feature extraction module for blade anomaly detection, the server performs the following series of operations. The server first obtains a second voiceprint sample set, and the second voiceprint training data in this sample set are all associated with fault type labels. For example, during the long-term monitoring of a large wind farm, a large amount of voiceprint data of different wind turbine blades under various fault conditions are collected to form the second voiceprint sample set. These fault types include blade cracks, blade surface wear, blade imbalance, etc., and each fault type corresponds to a specific fault type code. For example, blade cracks correspond to the code "001", and blade surface wear corresponds to the code "002", and these codes are determined according to the actual fault types to which the second voiceprint training data belong. Then, the server calls the pre-trained first acoustic feature extraction module to process the second voiceprint training data. Taking one piece of second voiceprint training data about blade surface wear as an example, the first acoustic feature extraction module first performs voiceprint feature encoding on it to obtain the fifth training voiceprint embedding feature, and then further extracts anomaly features from this embedding feature to obtain the fifth acoustic feature representation. After that, the server obtains at least two original fault diagnosis subnets corresponding to at least two fault types. For example, for the two fault types of blade cracks and blade surface wear, there are corresponding original fault diagnosis subnets respectively. The server uses these two original fault diagnosis subnets to discriminate the fault types of the fifth acoustic feature representation obtained above. Each original fault diagnosis subnet analyzes the feature information related to its own fault type in the fifth acoustic feature representation according to its own algorithm and parameters, so as to obtain the corresponding fault type probability. For example, the original fault diagnosis subnet corresponding to blade cracks analyzes that the probability that this voiceprint data belongs to the blade crack fault type is 0.3, and the original fault diagnosis subnet corresponding to blade surface wear obtains that the probability of belonging to the blade surface wear fault type is 0.6. Finally, the server determines the second target error parameter according to these at least two fault type probabilities and the corresponding at least two fault type codes. It obtains a value that can reflect the performance deviation of the original fault diagnosis subnet, that is, the second target error parameter, through a specific calculation method, such as comparing the difference between the fault type probability and the actual situation represented by the fault type code. Based on this error parameter, the server adjusts and optimizes the model parameters of at least two original fault diagnosis subnets. For example, modify the weight coefficients in the subnet, adjust the activation function of the neurons, etc. After optimization, these two original fault diagnosis subnets are determined as the official fault diagnosis subnets, and then the second acoustic feature extraction module is determined based on these optimized fault diagnosis subnets, so that it can more accurately identify the abnormal states of different fault types in the subsequent anomaly detection of wind turbine blades.

[0090] In an embodiment of the present invention, the voiceprint embedding feature is determined by performing voiceprint feature encoding on time-frequency feature data by a voiceprint feature encoding layer in a first acoustic feature extraction module; the time-frequency feature data is determined by performing time-frequency transformation processing on the blade voiceprint data by a voiceprint preprocessing layer in the first acoustic feature extraction module; the first acoustic feature extraction module belongs to a blade fault detection model for performing blade anomaly detection on the wind turbine; the blade fault detection model further includes a third acoustic feature extraction module with multi-level feature fusion after the acoustic feature encoding layer;

[0091] Performing fault weight analysis on the voiceprint embedding feature to obtain the fault type weight corresponding to each fault type may be implemented through the following examples.

[0092] Determine the third acoustic feature extraction module from the blade fault detection model;

[0093] Based on the third acoustic feature extraction module, perform fault weight analysis on the voiceprint embedding feature to obtain the fault type weight corresponding to each fault type.

[0094] In an embodiment of the present invention, exemplarily, the server first receives voiceprint data from the blades of a wind turbine. Taking a wind turbine under a specific working condition as an example, the voiceprint data of its blades is transmitted to the server. The first acoustic feature extraction module starts to work, and the voiceprint preprocessing layer in the module performs time-frequency transformation processing on the blade voiceprint data. This is like converting the voiceprint data from one "language" to another "language" that is more convenient for analysis, converting the simple voiceprint signal that changes with time into time-frequency feature data, and making the characteristics of the voiceprint clearly presented in both the time and frequency dimensions. For example, the voiceprint preprocessing layer organizes the frequency information at different times in the voiceprint data through a specific algorithm to obtain time-frequency feature data containing the corresponding frequency distribution at each time point. Then, the acoustic feature encoding layer in the first acoustic feature extraction module performs voiceprint feature encoding on the time-frequency feature data to determine the voiceprint embedding feature. The acoustic feature encoding layer extracts the most representative key information from the time-frequency feature data and encodes it into a compact vector form, that is, the voiceprint embedding feature. This voiceprint embedding feature contains the core features of the blade voiceprint data and provides an important basis for subsequent fault analysis. In the blade fault detection model, a third acoustic feature extraction module with multi-level feature fusion is also provided after the acoustic feature encoding layer. When it is necessary to perform fault weight analysis on the voiceprint embedding feature to obtain the fault type weight corresponding to each fault type, the server determines the third acoustic feature extraction module from the blade fault detection model. Then, the server performs fault weight analysis on the voiceprint embedding feature based on the third acoustic feature extraction module. For example, this module may use a multi-branch feature extractor to extract multi-modal acoustic features from different dimensions of the voiceprint embedding feature, such as frequency, phase, amplitude, etc., to obtain the multi-modal feature mapping result corresponding to the voiceprint embedding feature. Then, the dynamic weight allocator in the module performs comprehensive analysis and calculation on these multi-modal feature mapping results. It assigns corresponding weights to each fault type according to the preset algorithm and rules, combined with the correlation degree between different fault types and these features. Assuming that the blades of the wind turbine may have three fault types: blade crack, blade wear, and blade icing, after the analysis and calculation of the third acoustic feature extraction module, it is determined that the weight of the blade crack fault type is 0.5, the weight of the blade wear fault type is 0.3, and the weight of the blade icing fault type is 0.2. These weights reflect the relative likelihood of each fault type occurring under the current voiceprint embedding feature and provide a key quantitative basis for accurately judging blade faults in the future.

[0095] In an embodiment of the present invention, the blade fault detection model further includes a second acoustic feature extraction module for obtaining the abnormal state score corresponding to each fault type;

[0096] The embodiment of the present invention also provides the following implementation manners:

[0097] Obtain a third voiceprint sample set; the third voiceprint training data included in the third voiceprint sample set is associated with an abnormal type label;

[0098] Obtain a pre-trained first acoustic feature extraction module, based on the first acoustic feature extraction module, perform voiceprint feature encoding on the third voiceprint training data to obtain the sixth training voiceprint embedding feature of the third voiceprint training data, and perform abnormal feature extraction on the sixth training voiceprint embedding feature to obtain the sixth acoustic feature representation corresponding to the sixth training voiceprint embedding feature;

[0099] Obtain a pre-trained second acoustic feature extraction module, and based on at least two fault diagnosis subnets in the second acoustic feature extraction module, perform abnormal state recognition on the sixth acoustic feature representation to obtain the training abnormal state score corresponding to each fault type; one fault diagnosis subnet is used to obtain the training abnormal state score corresponding to one fault type;

[0100] Based on the third original acoustic feature extraction module, perform fault weight analysis on the sixth training voiceprint embedding feature to obtain the training fault type weight corresponding to each fault type;

[0101] According to the training abnormal state score corresponding to each fault type and the training fault type weight corresponding to each fault type, determine the training blade abnormality detection result of the reference wind turbine associated with the third voiceprint training data;

[0102] Determine a third target error parameter according to the abnormal type label and the training blade abnormality detection result, optimize the model parameters of the third original acoustic feature extraction module according to the third target error parameter, determine the optimized third original acoustic feature extraction module as the third acoustic feature extraction module, and determine the blade fault detection model according to the first acoustic feature extraction module, the second acoustic feature extraction module, and the third acoustic feature extraction module.

[0103] In an embodiment of the present invention, exemplarily, the server first obtains a third voiceprint sample set, and the third voiceprint training data in the sample set are all associated with abnormal type labels. For example, this data is sourced from the long-term accumulated voiceprint records of different wind turbine blades in a power generation field, covering various normal and abnormal operating states. The abnormal type labels clearly mark the abnormal conditions corresponding to each piece of data, such as blade cracks, wear, deformation, etc. Then, the server invokes a pre-trained first acoustic feature extraction module to process the third voiceprint training data. Taking one piece of third voiceprint training data associated with the abnormal type label of blade wear as an example, the first acoustic feature extraction module first performs voiceprint feature encoding on it to obtain a sixth training voiceprint embedding feature, and then performs abnormal feature extraction on this embedding feature to obtain a sixth acoustic feature representation. After that, the server obtains a pre-trained second acoustic feature extraction module. This module contains at least two fault diagnosis subnets, each subnet targeting a specific fault type. The server based on the fault diagnosis subnets in this module, identifies the abnormal state of the sixth acoustic feature representation. For example, for the blade crack fault diagnosis subnet and the blade wear fault diagnosis subnet, they respectively analyze the feature information related to their respective fault types in the sixth acoustic feature representation according to their own algorithms and parameters, and obtain the training abnormal state scores corresponding to each fault type. Suppose the training abnormal state score given by the blade crack fault diagnosis subnet is 30 points, and the score given by the blade wear fault diagnosis subnet is 70 points, which indicates that under the current voiceprint data, the possibility of blade wear abnormality is relatively high. At the same time, the server performs fault weight analysis on the sixth training voiceprint embedding feature based on the third original acoustic feature extraction module to obtain the training fault type weights corresponding to each fault type. For example, after analysis, it is determined that the training fault type weight of the blade crack fault type is 0.4, and the weight of the blade wear fault type is 0.6. Next, the server determines the training blade abnormal detection result of the reference wind turbine associated with the third voiceprint training data according to the training abnormal state score and the training fault type weight corresponding to each fault type. For example, through a specific calculation method, multiply the score of 30 points for blade crack by the weight of 0.4, multiply the score of 70 points for blade wear by the weight of 0.6, and then combine the calculation results of other possible fault types to obtain the training blade abnormal detection result of the blades of this reference wind turbine. Finally, the server determines the third target error parameter according to the abnormal type label and the training blade abnormal detection result. For example, compare the difference between the blade wear abnormality marked by the abnormal type label and the calculated training blade abnormal detection result, and obtain the third target error parameter through a specific algorithm. Based on this error parameter, the server optimizes the model parameters of the third original acoustic feature extraction module, such as adjusting the weights, thresholds, etc. of the internal algorithm. After the optimization is completed, it is determined as the third acoustic feature extraction module.Combining the first acoustic feature extraction module and the second acoustic feature extraction module, a blade fault detection model with better performance is finally determined to improve the accuracy of abnormal detection of wind turbine blades.

[0104] In the embodiment of the present invention, the third acoustic feature extraction module includes a multi-branch feature extractor and a dynamic weight allocator;

[0105] Performing fault weight analysis on the voiceprint embedding features based on the third acoustic feature extraction module to obtain the fault type weights corresponding to each fault type can be implemented through the following examples.

[0106] Performing multi-modal acoustic feature extraction on the voiceprint embedding features based on the multi-branch feature extractor to obtain the multi-modal feature mapping results corresponding to the voiceprint embedding features;

[0107] Performing fault weight calculation on the multi-modal feature mapping results based on the dynamic weight allocator to obtain the fault type weights corresponding to each fault type.

[0108] In the embodiment of the present invention, by way of example, in the central monitoring server of a large wind farm, a key link in the fault detection of wind turbine blades is to perform fault weight analysis on the voiceprint embedding features through the third acoustic feature extraction module to determine the weights of each fault type.

[0109] The server calls the third acoustic feature extraction module from the blade fault detection model, and this module is composed of a multi-branch feature extractor and a dynamic weight allocator. Taking the voiceprint embedding feature data of a certain wind turbine blade as an example, this data has been obtained through pre-processing and transmitted to the third acoustic feature extraction module.

[0110] First, the multi-branch feature extractor starts to work. It performs multi-modal acoustic feature extraction on the voiceprint embedding features from different dimensions. One branch focuses on frequency features, analyzing the distribution and changes of different frequency components in the voiceprint embedding features, such as identifying the energy concentration regions in specific frequency bands, and these regions may be related to specific faults. Another branch focuses on phase features, capturing the phase changes of the voiceprint signal at different time points, and abnormal changes in the phase may also indicate faults in the blade. There is also a branch for amplitude features, studying the magnitude and fluctuation law of the amplitude of the voiceprint signal. Through these multi-dimensional extractions, the multi-modal feature mapping results corresponding to the voiceprint embedding features are obtained. This result integrates the detailed feature information of the voiceprint in multiple modalities such as frequency, phase, and amplitude, providing a rich data basis for subsequent fault weight calculation.

[0111] Subsequently, the dynamic weight allocator takes over the multi-modal feature mapping results. It comprehensively considers and calculates these multi-modal features based on pre-set complex algorithms and rules. For example, for the three common fault types of blade crack, blade wear, and blade imbalance, the algorithm evaluates the correlation degree between each feature in the multi-modal feature mapping results and different fault types. If the energy concentration phenomenon in a certain high-frequency band in the frequency feature is closely related to the blade crack fault, then this frequency feature will be assigned a higher weight when calculating the weight of the blade crack fault type. The dynamic weight allocator comprehensively considers the influences of all relevant features, calculates the weights for each fault type separately, and finally obtains the fault type weights corresponding to each fault type. Suppose after calculation, the weight of the blade crack fault type is 0.5, indicating that under the current situation reflected by the voiceprint embedding features, the possibility of the blade having a crack fault is relatively high; the weight of the blade wear fault type is 0.3; and the weight of the blade imbalance fault type is 0.2. These weights provide a quantitative basis for accurately judging the possibility of blade fault types subsequently, and help the operation and maintenance personnel to inspect and maintain the wind turbine blades more pertinently.

[0112] In the embodiment of the present invention, the at least two fault types include a target fault type;

[0113] The embodiment of the present invention also provides the following implementation manners.

[0114] Obtain a blade fault detection model for blade anomaly detection of a wind turbine, and obtain an edge computing optimization architecture for expert knowledge transfer from the blade fault detection model;

[0115] Obtain a fourth voiceprint sample set related to the target fault type; the fourth voiceprint sample set includes fourth voiceprint training data;

[0116] Based on the blade fault detection model, perform data processing on the fourth voiceprint training data to obtain a blade anomaly detection result corresponding to the fourth voiceprint training data, and determine the blade anomaly detection result corresponding to the fourth voiceprint training data as a simulation diagnosis label for training the edge computing optimization architecture;

[0117] Based on the edge computing optimization architecture, perform data processing on the fourth voiceprint training data to obtain an edge blade anomaly detection result corresponding to the fourth voiceprint training data;

[0118] Optimize the architecture parameters of the edge computing optimization architecture according to the edge blade anomaly detection result and the simulation diagnosis label, and determine the optimized edge computing optimization architecture as the lightweight diagnosis network of the blade fault detection model; the lightweight diagnosis network is used for blade anomaly detection of blade voiceprint data under the target fault type.

[0119] In an embodiment of the present invention, it is assumed that among at least two fault types of a wind turbine blade, blade crack is set as a target fault type. The server first obtains a blade fault detection model for blade anomaly detection, and simultaneously obtains an edge computing optimization architecture for expert knowledge migration from the model. The blade fault detection model has been trained on a large amount of data and has accumulated rich experience in fault detection, while the edge computing optimization architecture aims to lightweight the model so that it can run efficiently on edge devices. Next, the server obtains a fourth voiceprint sample set associated with the target fault type (blade crack), and the sample set contains fourth voiceprint training data. These data are from voiceprints collected when crack faults are suspected to occur in wind turbine blades under different working conditions, covering voiceprint information corresponding to cracks of different degrees and different positions. Then, the server processes the fourth voiceprint training data based on the blade fault detection model. The various modules in the model work together, from voiceprint feature encoding to abnormal feature extraction, to fault type discrimination and other steps, and finally obtain the blade anomaly detection result corresponding to the fourth voiceprint training data. For example, after analysis and calculation, the model determines that the blades corresponding to some data have a high probability of crack faults, and the blades corresponding to some data have a low probability of crack faults. These results are determined as simulation diagnostic labels for training the edge computing optimization architecture, and serve as reference standards for subsequent evaluation of the performance of the edge computing optimization architecture. Afterwards, the server uses the edge computing optimization architecture to process the fourth voiceprint training data. The architecture analyzes the voiceprint data based on its own algorithm, and attempts to detect whether the blade has the target fault type (blade crack), thereby obtaining the edge blade anomaly detection results corresponding to the fourth voiceprint training data. For example, the edge computing optimization architecture may determine that the blades corresponding to some data have potential crack risks, but there is a certain difference from the judgment of the blade fault detection model. Finally, the server optimizes the architecture parameters of the edge computing optimization architecture based on the edge blade anomaly detection results and simulation diagnostic labels. By comparing the differences between the two, the error parameters are calculated, and the parameters such as the algorithm and weight within the architecture are adjusted based on these parameters. For example, if the edge computing optimization architecture misjudges some blade cracks, the server will adjust the relevant parameters to make the architecture more accurate in subsequent detections. After optimization, the edge computing optimization architecture is determined as a lightweight diagnostic network for the blade fault detection model. This lightweight diagnostic network is specifically used to quickly and accurately detect blade anomalies in blade soundprint data under the target fault type (blade crack), and can be deployed on edge devices close to wind turbines to improve the real-time and efficiency of fault detection. Based on the above, the corresponding causes of blade fault types in the embodiment of the present invention can be found in Table 1.

[0120] Table 1

[0121]

[0122] In an embodiment of the present invention, the fourth voiceprint training data includes voiceprint training data of a fault condition generated under the target fault type;

[0123] The obtaining of the fourth voiceprint sample set related to the target fault type can be implemented through the following examples.

[0124] Obtain the voiceprint training data of the fault condition under the target fault type;

[0125] Determine the pending voiceprint training data from the voiceprint training data set for voiceprint feature screening;

[0126] Based on the voiceprint sample expansion module, perform an acoustic fingerprint comparison between the pending voiceprint training data and the voiceprint training data of the fault condition to obtain the matching degree between the voiceprint training data of the fault condition and the pending voiceprint training data;

[0127] If the matching degree reaches the data enhancement matching degree threshold, determine the pending voiceprint training data as the extended voiceprint training data related to the target fault type;

[0128] Determine the extended voiceprint training data and the voiceprint training data of the fault condition as the fourth voiceprint training data related to the target fault type.

[0129] In an embodiment of the present invention, by way of example, it is assumed that the target fault type is blade surface corrosion, and the server conducts the acquisition work of the fourth voiceprint sample set around this fault type. First, the server acquires the voiceprint training data under the target fault type (blade surface corrosion). These data are from the actual operating conditions when the wind turbine blade has a surface corrosion fault. For example, for a wind turbine in a coastal area, due to the erosion of sea breeze, the blade surface is corroded. The server collects the voiceprint data generated by the blade under different corrosion degrees and different operating environments as the voiceprint training data under fault conditions. Next, the server determines the pending voiceprint training data from the voiceprint training data set used for voiceprint feature screening. This set contains various voiceprint data accumulated by the power plant over a long time, covering different wind turbines and different operating states. The server selects the voiceprint data that may be related to the blade surface corrosion fault from the massive data as the pending voiceprint training data through specific screening rules. For example, the data operating in a similar environment and having a similar trend in certain frequency ranges or energy distributions of voiceprint features are selected. Then, the server conducts an acoustic fingerprint comparison between the pending voiceprint training data and the voiceprint training data under fault conditions based on the voiceprint sample expansion module. The voiceprint sample expansion module conducts a detailed comparison of the two sets of data from the features in each dimension of the voiceprint data, such as frequency, phase, amplitude, etc., and calculates the matching degree between the voiceprint training data under fault conditions and the pending voiceprint training data. For example, through a complex algorithm, it is obtained that the matching degree of a certain pending voiceprint training data and the voiceprint training data under fault conditions in key frequency features is 80%. If the matching degree reaches the data enhancement matching degree threshold, assuming the threshold is set to 70%, then the server determines this pending voiceprint training data as the extended voiceprint training data related to the target fault type (blade surface corrosion). This means that although this data is not explicitly marked as blade surface corrosion fault data, judged from the voiceprint features, it is very likely to be related to this fault. Finally, the server jointly determines the extended voiceprint training data and the voiceprint training data under fault conditions as the fourth voiceprint training data related to the target fault type (blade surface corrosion). The fourth voiceprint training data composed of these data more comprehensively covers the voiceprint features related to the blade surface corrosion fault, provides a rich and representative data basis for the subsequent analysis based on the blade fault detection model and the training of the edge computing optimization architecture, and helps to improve the accuracy and reliability of the blade surface corrosion fault detection.

[0130] An embodiment of the present invention provides a computer device 100. The computer device 100 includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned method for detecting abnormalities in wind turbine blades based on voiceprint acquisition. As Figure 2 shown, Figure 2The block diagram of the computer device 100 provided by the embodiment of the present invention. The computer device 100 includes a memory 111, a processor 112, and a communication unit 113. To achieve data transmission or interaction, the elements of the memory 111, the processor 112, and the communication unit 113 are electrically connected to each other directly or indirectly.

[0131] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Numerous modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best illustrate the principles of the disclosure and its practical applications, to thereby enable those skilled in the art to best utilize the disclosure and to utilize various embodiments with various modifications as are suited to the particular application contemplated.

Claims

1. A method for detecting abnormalities in wind turbine blades based on voiceprint acquisition, characterized in that, Including: Performing voiceprint feature encoding on the voiceprint data of the wind turbine blade to obtain the voiceprint embedding feature corresponding to the blade voiceprint data; Performing abnormal feature extraction on the voiceprint embedding feature to obtain the abnormal feature vector corresponding to the voiceprint embedding feature; Determining at least two fault types associated with the blade voiceprint data, and respectively performing abnormal state recognition on the abnormal feature vector according to each fault type to obtain the abnormal state score corresponding to each fault type; Performing fault weight analysis on the voiceprint embedding feature to obtain the fault type weight corresponding to each fault type; Performing weight value assignment on the abnormal state score corresponding to each fault type based on the fault type weight corresponding to each fault type to obtain the target abnormal state detection result corresponding to each fault type; the weight frequency domain response value corresponding to one fault type is used for weight value assignment of the abnormal state score corresponding to the corresponding fault type; Determining the blade abnormal detection result of the wind turbine according to the target abnormal state detection result corresponding to each fault type; The abnormal feature vector is determined by performing abnormal feature extraction on the voiceprint embedding feature by an acoustic feature encoder in the first acoustic feature extraction module; The first acoustic feature extraction module belongs to a blade fault detection model for performing blade abnormal detection on the wind turbine; The method further includes: Obtaining a first voiceprint sample set; the first voiceprint training data included in the first voiceprint sample set is voiceprint training data without abnormal type labels; Obtaining multiple self-supervised learning tasks for the first original acoustic feature extraction module; Based on the first original acoustic feature extraction module, performing data processing on the first voiceprint training data to obtain the acoustic feature representation associated with each self-supervised learning task, and determining multiple error parameters corresponding to the multiple self-supervised learning tasks according to the acoustic feature representation; one self-supervised learning task corresponds to one error parameter; Determining a first target error parameter according to the multiple error parameters, optimizing the model parameters of the first original acoustic feature extraction module according to the first target error parameter, and determining the optimized first original acoustic feature extraction module as the first acoustic feature extraction module.

2. The method according to claim 1, wherein The performing voiceprint feature encoding on the voiceprint data of the wind turbine blade to obtain the voiceprint embedding feature corresponding to the blade voiceprint data includes: When obtaining the voiceprint data of the wind turbine blade, obtaining a blade fault detection model for performing blade abnormal detection on the wind turbine; the blade fault detection model includes a first acoustic feature extraction module; Based on the first acoustic feature extraction module, determining the acoustic signal characteristics associated with the blade voiceprint data, and extracting the framed voiceprint feature of the blade voiceprint data according to the acoustic signal characteristics; the first acoustic feature extraction module includes a voiceprint preprocessing layer; the framed voiceprint feature includes multiple framed voiceprint features, and the multiple framed voiceprint features include a target framed voiceprint feature; Determine the time-frequency analysis domain corresponding to the target framed voiceprint feature according to the acoustic signal characteristics to which the target framed voiceprint feature belongs; the time-frequency analysis domain includes a plurality of frequency-domain response values, and each frequency-domain response value has a corresponding framed feature range; Based on the voiceprint preprocessing layer, determine the framed feature range to which the target framed voiceprint feature belongs, and determine the frequency-domain response values corresponding to the framed feature range to which the target framed voiceprint feature belongs as the time-frequency feature map corresponding to the target framed voiceprint feature; Based on the first acoustic feature extraction module, perform time-frequency feature reconstruction on the time-frequency feature map corresponding to the framed voiceprint feature to obtain the time-frequency feature data of the blade voiceprint data; Perform voiceprint feature encoding on the time-frequency feature data to obtain the voiceprint embedding feature of the time-frequency feature data.

3. The method according to claim 1, characterized in that, The plurality of self-supervised learning tasks include a time-frequency map mask reconstruction task; The first original acoustic feature extraction module processes the first voiceprint training data to obtain an acoustic feature representation associated with each self-supervised learning task, and determines a plurality of error parameters corresponding to the plurality of self-supervised learning tasks according to the acoustic feature representation, including: Perform time-frequency transformation processing on the first voiceprint training data based on the first original acoustic feature extraction module to obtain the first sample time-frequency feature data of the first voiceprint training data; the first sample time-frequency feature data includes sample time-frequency feature maps at a plurality of time frame positions; Select a first time frame position to be subjected to time-frequency region masking from the plurality of time frame positions corresponding to the first sample time-frequency feature data according to a preset rule; Perform time-frequency region masking processing on the sample time-frequency feature map at the first time frame position in the first sample time-frequency feature data according to the masking region identifier to obtain masked time-frequency sample data; Perform voiceprint feature encoding on the masked time-frequency sample data based on the first original acoustic feature extraction module to obtain the first training voiceprint embedding feature of the masked time-frequency sample data, and perform abnormal feature extraction on the first training voiceprint embedding feature to obtain the first acoustic feature representation corresponding to the first training voiceprint embedding feature; the first acoustic feature representation belongs to the acoustic feature representation; Based on a feature reconstruction module associated with the time-frequency map mask reconstruction task, perform reconstructed time-frequency map masking processing on the first acoustic feature representation to obtain a predicted time-frequency feature map corresponding to the masking region identifier in the masked time-frequency sample data; Determine the error parameter corresponding to the time-frequency map mask reconstruction task according to the sample time-frequency feature map at the first time frame position in the first sample time-frequency feature data and the predicted time-frequency feature map.

4. The method according to claim 1, characterized in that, The plurality of self-supervised learning tasks include a time-frequency anomaly replacement detection task; The first original acoustic feature extraction module processes the first voiceprint training data to obtain an acoustic feature representation associated with each self-supervised learning task, and determines a plurality of error parameters corresponding to the plurality of self-supervised learning tasks according to the acoustic feature representation, including: Based on the first original acoustic feature extraction module, perform time-frequency transformation processing on the first voiceprint training data to obtain the first sample time-frequency feature data of the first voiceprint training data; the first sample time-frequency feature data includes sample time-frequency feature maps at multiple time frame positions; From the multiple time frame positions corresponding to the first sample time-frequency feature data, select a second time frame position to be subjected to time-frequency feature replacement according to a preset rule; Perform time-frequency feature replacement processing on the sample time-frequency feature map at the second time frame position in the first sample time-frequency feature data according to preset time-frequency information to obtain replacement time-frequency sample data; the preset time-frequency information is different from the sample time-frequency feature map at the second time frame position; the preset time-frequency information is an abnormal time-frequency feature segment extracted from historical fault data or an unnatural time-frequency feature generated by Gaussian noise; Based on the first original acoustic feature extraction module, perform voiceprint feature encoding on the replacement time-frequency sample data to obtain a second training voiceprint embedding feature of the replacement time-frequency sample data, and perform abnormal feature extraction on the second training voiceprint embedding feature to obtain a second acoustic feature representation corresponding to the second training voiceprint embedding feature; the second acoustic feature representation belongs to the acoustic feature representation; Based on a feature reconstruction module associated with the time-frequency anomaly replacement detection task, determine the target detection time frame position for the second acoustic feature representation to obtain a time-frequency feature replacement prediction result corresponding to the multiple time frame positions; According to the second time frame position and the time-frequency feature replacement prediction result corresponding to the multiple time frame positions, determine an error parameter corresponding to the time-frequency anomaly replacement detection task.

5. The method according to claim 1, characterized in that Among the multiple self-supervised learning tasks, there is a voiceprint denoising contrast task; The data processing of the first voiceprint training data based on the first original acoustic feature extraction module to obtain an acoustic feature representation associated with each self-supervised learning task, and determining multiple error parameters corresponding to the multiple self-supervised learning tasks according to the acoustic feature representation includes: Inject Gaussian white noise features into the framed voiceprint features of the first voiceprint training data according to the signal-to-noise ratio to obtain noise-containing voiceprint training data, and perform time-frequency transformation processing on the noise-containing voiceprint training data to obtain the noise time-frequency feature data of the noise-containing voiceprint training data; Perform time-frequency transformation processing on the first voiceprint training data to obtain the first sample time-frequency feature data of the first voiceprint training data; Based on the first original acoustic feature extraction module, perform voiceprint feature encoding on the noise time-frequency feature data to obtain a third training voiceprint embedding feature of the noise time-frequency feature data, and perform abnormal feature extraction on the third training voiceprint embedding feature to obtain a third acoustic feature representation corresponding to the third training voiceprint embedding feature; the third acoustic feature representation belongs to the acoustic feature representation; Based on the first original acoustic feature extraction module, perform voiceprint feature encoding on the first sample time-frequency feature data to obtain the fourth training voiceprint embedding feature of the first sample time-frequency feature data, and perform abnormal feature extraction on the fourth training voiceprint embedding feature to obtain the fourth acoustic feature representation corresponding to the fourth training voiceprint embedding feature; the fourth acoustic feature representation belongs to the acoustic feature representation; Determine the error parameter corresponding to the voiceprint denoising comparison task according to the third acoustic feature representation and the fourth acoustic feature representation; The multiple self-supervised learning tasks further include an expert knowledge transfer task; The data processing of the first voiceprint training data based on the first original acoustic feature extraction module to obtain the acoustic feature representation associated with each self-supervised learning task, and determining the multiple error parameters corresponding to the multiple self-supervised learning tasks according to the acoustic feature representation further includes: Perform abnormal risk rating on the fourth acoustic feature representation according to the feature reconstruction module related to the expert knowledge transfer task to obtain the abnormal risk score corresponding to the first voiceprint training data; Obtain a pre-trained benchmark model related to the first original acoustic feature extraction module, and perform abnormal risk rating on the first voiceprint training data according to the pre-trained benchmark model to obtain the benchmark abnormal state score corresponding to the first voiceprint training data; Determine the error parameter corresponding to the expert knowledge transfer task according to the abnormal risk score and the benchmark abnormal state score; 6. The method according to claim 1, characterized in that, The abnormal feature vector is determined by performing abnormal feature extraction on the voiceprint embedding feature by the acoustic feature encoder in the first acoustic feature extraction module; the first acoustic feature extraction module belongs to a blade fault detection model for blade abnormal detection of a wind turbine; the blade fault detection model further includes a second acoustic feature extraction module with multi-level feature fusion after the acoustic feature encoder; The obtaining the abnormal state score corresponding to each fault type by respectively performing abnormal state recognition on the abnormal feature vector according to each fault type includes: Determine the second acoustic feature extraction module from the blade fault detection model; the second acoustic feature extraction module includes a fault diagnosis subnet corresponding to each fault type; Based on the fault diagnosis subnet corresponding to each fault type, perform abnormal state recognition on the abnormal feature vector to obtain the abnormal state score corresponding to each fault type; one fault diagnosis subnet is used to obtain the abnormal state score corresponding to one fault type; The method further includes: Obtain a second voiceprint sample set; the second voiceprint training data included in the second voiceprint sample set is associated with a fault type label; the fault type label includes at least two fault type encodings corresponding to the at least two fault types, one fault type corresponds to one fault type encoding, and the at least two fault type encodings are determined according to the fault type to which the second voiceprint training data belongs; Obtain a pre-trained first acoustic feature extraction module, perform voiceprint feature encoding on the second voiceprint training data based on the first acoustic feature extraction module to obtain a fifth training voiceprint embedding feature of the second voiceprint training data, and perform abnormal feature extraction on the fifth training voiceprint embedding feature to obtain a fifth acoustic feature representation corresponding to the fifth training voiceprint embedding feature; Obtain at least two original fault diagnosis subnets corresponding to the at least two fault types, and perform fault type discrimination on the fifth acoustic feature representation according to the at least two subnets respectively to obtain at least two fault type probabilities corresponding to the at least two original fault diagnosis subnets; one fault type corresponds to one original fault diagnosis subnet, and one original fault diagnosis subnet is used to obtain one fault type probability; Determine a second target error parameter according to the at least two fault type probabilities and the at least two fault type encodings, optimize the model parameters of the at least two original fault diagnosis subnets according to the second target error parameter, determine the optimized at least two original fault diagnosis subnets as the at least two fault diagnosis subnets, and determine the second acoustic feature extraction module according to the at least two fault diagnosis subnets.

7. The method according to claim 1, wherein The voiceprint embedding feature is determined by performing voiceprint feature encoding on time-frequency feature data by an acoustic feature encoding layer in the first acoustic feature extraction module; the time-frequency feature data is determined by performing time-frequency transformation processing on the blade voiceprint data by a voiceprint preprocessing layer in the first acoustic feature extraction module; the first acoustic feature extraction module belongs to a blade fault detection model for performing blade anomaly detection on the wind turbine; the blade fault detection model further includes a third acoustic feature extraction module with multi-level feature fusion after the acoustic feature encoding layer; The fault weight analysis of the voiceprint embedding feature to obtain the fault type weight corresponding to each fault type includes: Determine a third acoustic feature extraction module from the blade fault detection model; the third acoustic feature extraction module includes a multi-branch feature extractor and a dynamic weight allocator; Perform multi-modal acoustic feature extraction on the voiceprint embedding feature based on the multi-branch feature extractor to obtain a multi-modal feature mapping result corresponding to the voiceprint embedding feature; Perform fault weight calculation on the multi-modal feature mapping result based on the dynamic weight allocator to obtain the fault type weight corresponding to each fault type; The blade fault detection model further includes a second acoustic feature extraction module for obtaining an abnormal state score corresponding to each fault type; The method further includes: Obtain a third voiceprint sample set; the third voiceprint training data included in the third voiceprint sample set is associated with an abnormal type label; Obtain a pre-trained first acoustic feature extraction module, perform voiceprint feature encoding on the third voiceprint training data based on the first acoustic feature extraction module to obtain the sixth training voiceprint embedding feature of the third voiceprint training data, and perform abnormal feature extraction on the sixth training voiceprint embedding feature to obtain the sixth acoustic feature representation corresponding to the sixth training voiceprint embedding feature; Obtain a pre-trained second acoustic feature extraction module, and perform abnormal state recognition on the sixth acoustic feature representation based on at least two fault diagnosis subnets in the second acoustic feature extraction module to obtain the training abnormal state score corresponding to each fault type; one fault diagnosis subnet is used to obtain the training abnormal state score corresponding to one fault type; Perform fault weight analysis on the sixth training voiceprint embedding feature based on the third original acoustic feature extraction module to obtain the training fault type weight corresponding to each fault type; Determine the training blade anomaly detection result of the reference wind turbine associated with the third voiceprint training data according to the training abnormal state score corresponding to each fault type and the training fault type weight corresponding to each fault type; Determine the third target error parameter according to the anomaly type label and the training blade anomaly detection result, optimize the model parameters of the third original acoustic feature extraction module according to the third target error parameter, determine the optimized third original acoustic feature extraction module as the third acoustic feature extraction module, and determine the blade fault detection model according to the first acoustic feature extraction module, the second acoustic feature extraction module and the third acoustic feature extraction module.

8. The method according to claim 1, characterized in that, The at least two fault types include target fault types; The method further includes: Obtain a blade fault detection model for blade anomaly detection of a wind turbine, and obtain an edge computing optimization architecture for expert knowledge transfer from the blade fault detection model; Obtain the fault condition voiceprint training data under the target fault type; Determine the pending voiceprint training data from the voiceprint training data set for voiceprint feature screening; Perform acoustic fingerprint comparison on the pending voiceprint training data and the fault condition voiceprint training data based on the voiceprint sample expansion module to obtain the matching degree between the fault condition voiceprint training data and the pending voiceprint training data; If the matching degree reaches the data enhancement matching degree threshold, determine the pending voiceprint training data as the extended voiceprint training data related to the target fault type; Determine the extended voiceprint training data and the fault condition voiceprint training data as the fourth voiceprint training data related to the target fault type; the fourth voiceprint sample set includes the fourth voiceprint training data; the fourth voiceprint training data includes the fault condition voiceprint training data generated under the target fault type; Performing data processing on the fourth voiceprint training data based on the blade fault detection model to obtain a blade anomaly detection result corresponding to the fourth voiceprint training data, and determining the blade anomaly detection result corresponding to the fourth voiceprint training data as a simulation diagnosis label for training the edge computing optimization architecture; Performing data processing on the fourth voiceprint training data based on the edge computing optimization architecture to obtain an edge blade anomaly detection result corresponding to the fourth voiceprint training data; Optimizing the architecture parameters of the edge computing optimization architecture according to the edge blade anomaly detection result and the simulation diagnosis label, and determining the optimized edge computing optimization architecture as the lightweight diagnosis network of the blade fault detection model; the lightweight diagnosis network is used for performing blade anomaly detection on the blade voiceprint data under the target fault type.

9. A server system, characterized in that, It includes a server, and the server is used to execute the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • 500kV GIS multi-mode working condition abnormity comprehensive monitoring system and method based on voiceprint recognition

    CN119479690A

Cited By

  • A wind power blade intelligent diagnosis method and system and a storage medium

    CN122688103A