Wind driven generator blade anomaly detection method and system based on voiceprint collection
By encoding and extracting the soundprint data of wind turbine blades and detecting them in combination with the fault type weight, the problems of low efficiency and poor accuracy of wind turbine blade fault detection in the prior art are solved, and effective detection of blade abnormalities and early fault identification are achieved.
Patent Information
- Application Number
- CN202510505980.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-22
AI Technical Summary
The existing wind turbine blade fault detection methods are inefficient and have poor accuracy, making it difficult to detect potential faults early, and are susceptible to environmental interference.
Using a method based on voiceprint collection, the voiceprint feature encoding of the voiceprint data of the wind turbine blades is extracted, the voiceprint embedding features are performed, abnormal feature extraction and fault type identification are performed, and the target abnormal state detection is performed in combination with the fault type weight.
It realizes effective detection of wind turbine blade abnormalities, improves the accuracy and efficiency of detection, can detect potential faults in the early stage, and reduces safety risks and economic losses caused by blade failure.
Smart Images

Figure CN120012032A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of voiceprint data processing, and in particular to a method and system for detecting abnormalities in blades of wind turbines based on voiceprint collection. Background Art
[0002] Wind turbine blades are exposed to complex and harsh environments for a long time and are prone to failures, such as blade cracks and wear, which affect power generation efficiency and safety. Existing blade fault detection methods have certain limitations. Some rely on regular manual inspections, which are inefficient and difficult to detect early potential faults; some are based on vibration, temperature and other sensor detection, which are easily affected by environmental interference and have poor accuracy. Summary of the invention
[0003] The object of the present invention is to provide a method and system for detecting abnormalities in wind turbine blades based on voiceprint collection.
[0004] In a first aspect, an embodiment of the present invention provides a method for detecting abnormalities in blades of a wind turbine generator based on voiceprint collection, comprising: Encoding the voiceprint feature of the blade voiceprint data of the wind turbine to obtain the voiceprint embedding feature corresponding to the blade voiceprint data; Extracting abnormal features from the voiceprint embedded features to obtain abnormal feature vectors corresponding to the voiceprint embedded features; Determine at least two fault types associated with the blade soundprint data, perform abnormal state identification on the abnormal feature vector according to each fault type, and obtain an abnormal state score corresponding to each fault type; Performing fault weight analysis on the voiceprint embedded features to obtain a fault type weight corresponding to each fault type; Based on the fault type weight corresponding to each fault type, the abnormal state score weight corresponding to each fault type is allocated to obtain the target abnormal state detection result corresponding to each fault type; the weighted frequency domain response value corresponding to a fault type is used to allocate the abnormal state score weight corresponding to the corresponding fault type; The blade abnormality detection result of the wind turbine is determined according to the target abnormal state detection result corresponding to each fault type.
[0005] In a second aspect, an embodiment of the present invention provides a server system, including a server, wherein the server is used to execute the method described in the first aspect.
[0006] Compared with the prior art, the beneficial effects provided by the present invention include: using a wind turbine blade abnormality detection method and system based on voiceprint collection disclosed by the present invention, voiceprint feature encoding is performed on blade voiceprint data to obtain voiceprint embedding features, and its abnormal feature vector is extracted; at least two fault types are determined, and abnormal state scores are obtained by identifying abnormal feature vectors according to each fault type; the voiceprint embedding features are analyzed to obtain fault type weights; the abnormal state score weights are assigned in combination with the weights to obtain target abnormal state detection results; and finally the blade abnormality detection results are determined accordingly. The method realizes effective detection of wind turbine blade abnormalities. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative work.
[0008] Figure 1 A schematic diagram of the steps of a method for detecting abnormality of wind turbine blades based on voiceprint collection provided by an embodiment of the present invention; Figure 2 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0009] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0010] The specific implementation modes of the present invention are described in detail below in conjunction with the accompanying drawings.
[0011] In order to solve the technical problems in the aforementioned background technology, Figure 1 A flow chart of a method for detecting abnormality of wind turbine blades based on voiceprint collection provided in an embodiment of the present disclosure is provided below. The method for detecting abnormality of wind turbine blades based on voiceprint collection is introduced in detail.
[0012] Step S201, performing voiceprint feature encoding on the blade voiceprint data of the wind turbine to obtain voiceprint embedding features corresponding to the blade voiceprint data; Step S202, extracting abnormal features from the voiceprint embedded features to obtain abnormal feature vectors corresponding to the voiceprint embedded features; Step S203, determining at least two fault types associated with the blade soundprint data, and performing abnormal state identification on the abnormal feature vector according to each fault type to obtain an abnormal state score corresponding to each fault type; Step S204, performing fault weight analysis on the voiceprint embedded features to obtain a fault type weight corresponding to each fault type; Step S205, assigning a weight value of an abnormal state score corresponding to each fault type based on the fault type weight corresponding to each fault type, and obtaining a target abnormal state detection result corresponding to each fault type; a weighted frequency domain response value corresponding to a fault type is used to assign a weight value of an abnormal state score corresponding to a corresponding fault type; Step S206: determining a blade abnormality detection result of the wind turbine according to the target abnormal state detection result corresponding to each fault type.
[0013] In an embodiment of the present invention, for example, in a large wind farm, there are many wind turbines distributed. During the rotation process, the blades of each generator will generate unique sound signals due to their own state and surrounding environmental factors. These signals are collected by high-precision microphones arranged at appropriate positions in the tower or blade cavity of the wind turbine generator to form blade soundprint data and transmit them to the server. After the server obtains the blade soundprint data, it will call the blade fault detection model for blade abnormality detection, which includes the first acoustic feature extraction module. Taking a certain wind turbine as an example, the soundprint preprocessing layer in the first acoustic feature extraction module starts working. It will determine the acoustic signal characteristics associated with the blade soundprint data based on its own algorithm, such as the frequency range of the sound, amplitude change, etc. Then, according to these characteristics, the blade soundprint data is framed and processed to obtain multiple framed soundprint features, including the target framed soundprint feature. The server further determines the corresponding time-frequency analysis domain according to the acoustic signal characteristics to which the target framed soundprint feature belongs. For example, if the target frame voiceprint feature is mainly concentrated in the middle and high frequency bands, then the corresponding time-frequency analysis domain is the range related to the middle and high frequencies, and the time-frequency analysis domain contains multiple frequency domain response values, and each frequency domain response value has a corresponding frame feature range. Next, the voiceprint preprocessing layer determines the frame feature range to which the target frame voiceprint feature belongs, and determines the frequency domain response value corresponding to this range as the time-frequency feature graph corresponding to the target frame voiceprint feature. Subsequently, the first acoustic feature extraction module reconstructs the time-frequency feature of the time-frequency feature graphs corresponding to all frame voiceprint features. These time-frequency feature graphs are combined according to specific rules to obtain the time-frequency feature data of the blade voiceprint data. Finally, the time-frequency feature data is voiceprint feature encoded to obtain the voiceprint embedding feature of the blade voiceprint data. After the server completes the voiceprint embedding feature extraction, the acoustic feature encoder in the first acoustic feature extraction module will be used to extract the abnormal feature of the voiceprint embedding feature. When training this blade fault detection model, the server will obtain the first voiceprint sample set, in which the first voiceprint training data is not configured with an abnormal type label. At the same time, the server will also obtain multiple self-supervised learning tasks for the first original acoustic feature extraction module, such as the time-frequency graph mask reconstruction task, the time-frequency anomaly replacement detection task, the voiceprint denoising comparison task, and the expert knowledge transfer task. Taking the time-frequency graph mask reconstruction task as an example, the server performs time-frequency transformation processing on the first voiceprint training data based on the first original acoustic feature extraction module to obtain the first sample time-frequency feature data, which contains sample time-frequency feature graphs at multiple time frame positions. Then, the server selects the first time frame position to be masked in the time-frequency area from these time frame positions according to the preset rules. Assuming that the preset rule is to randomly select a certain proportion of time frame positions, the server may randomly select several time frame positions as the first time frame positions.Next, the server performs a time-frequency region masking process on the sample time-frequency feature graph at the first time frame position in the first sample time-frequency feature data according to the masking region identifier to obtain the masked time-frequency sample data. After that, the server performs voiceprint feature encoding on the masked time-frequency sample data based on the first original acoustic feature extraction module to obtain the first training voiceprint embedding feature, and performs abnormal feature extraction on it to obtain the first acoustic feature representation. Then, based on the feature reconstruction module associated with the task, the first acoustic feature representation is subjected to a masking time-frequency graph masking process to obtain the predicted time-frequency feature graph. Finally, according to the sample time-frequency feature graph at the first time frame position in the first sample time-frequency feature data and the predicted time-frequency feature graph, the error parameter corresponding to the task is determined. By continuously adjusting these error parameters, the model parameters of the first original acoustic feature extraction module are optimized, and finally a first acoustic feature extraction module with good performance is obtained for accurately extracting abnormal feature vectors. In the actual operation of a wind farm, blades may have a variety of fault types, such as blade cracks, surface wear, loose connection parts, etc. The server will determine at least two fault types related to the current blade voiceprint data. The second acoustic feature extraction module in the blade fault detection model contains a fault diagnosis subnet for each fault type. Taking the blade soundprint data of a certain wind turbine as an example, the server determines the second acoustic feature extraction module from the blade fault detection model, and then identifies the abnormal state of the previously extracted abnormal feature vector based on the fault diagnosis subnet corresponding to each fault type. When training these fault diagnosis subnets, the server obtains a second voiceprint sample set, in which the second voiceprint training data is associated with a fault type label, and these labels correspond to at least two fault type codes, and each fault type corresponds to a code. The server uses the pre-trained first acoustic feature extraction module to encode the second voiceprint training data with voiceprint features, obtains the fifth training voiceprint embedding feature, and performs abnormal feature extraction to obtain the fifth acoustic feature representation. Then, the server obtains at least two original fault diagnosis subnets, respectively distinguishes the fault type of the fifth acoustic feature representation, and obtains at least two fault type probabilities corresponding to at least two original fault diagnosis subnets. The second target error parameter is determined based on these probabilities and fault type codes, and the model parameters of the original fault diagnosis subnet are optimized to obtain the optimized fault diagnosis subnet, so that the abnormal feature vector can be accurately identified for each fault type, and the abnormal state score corresponding to each fault type can be obtained. For example, for the fault type of blade crack, the fault diagnosis subnet analyzes and calculates and gives an abnormal state score that reflects the possibility of blade crack fault corresponding to the current blade soundprint data. The server determines the third acoustic feature extraction module from the blade fault detection model, which includes a multi-branch feature extractor and a dynamic weight allocator. The server performs multimodal acoustic feature extraction on the soundprint embedded features based on the multi-branch feature extractor.For example, the multimodal feature mapping results corresponding to the voiceprint embedded features are obtained by analyzing the frequency, phase, amplitude and other dimensions of the voiceprint. Then, the dynamic weight allocator calculates the fault weights of these multimodal feature mapping results to obtain the fault type weights corresponding to each fault type. When training the third acoustic feature extraction module, the server obtains the third voiceprint sample set, in which the third voiceprint training data is associated with an abnormal type label. The server uses the pre-trained first acoustic feature extraction module to encode the voiceprint features of the third voiceprint training data, obtains the sixth training voiceprint embedded features, and extracts the abnormal features to obtain the sixth acoustic feature representation. Then, the fault diagnosis subnet in the pre-trained second acoustic feature extraction module is used to identify the abnormal state of the sixth acoustic feature representation, and obtains the training abnormal state score corresponding to each fault type. At the same time, the fault weight analysis is performed on the sixth training voiceprint embedded features based on the third original acoustic feature extraction module to obtain the training fault type weight corresponding to each fault type. According to the training abnormal state score and the training fault type weight, the training blade abnormality detection result of the reference wind turbine associated with the third voiceprint training data is determined. Then, the third target error parameter is determined according to the abnormal type label and the training blade abnormality detection result, and the model parameters of the third original acoustic feature extraction module are optimized to obtain a third acoustic feature extraction module with good performance, and the fault type weight corresponding to each fault type is accurately calculated. For example, after analysis and calculation, it is determined that the weight of the blade crack fault type is 0.6, the weight of the surface wear fault type is 0.4, and so on. After the server obtains the abnormal state score and fault type weight corresponding to each fault type, the weight is assigned. For example, for the blade crack fault type, it is assumed that its abnormal state score is 80 points and the fault type weight is 0.6; for the surface wear fault type, the abnormal state score is 70 points and the fault type weight is 0.4. Then the target abnormal state detection result corresponding to the blade crack fault type is 80×0.6=48 points, and the target abnormal state detection result corresponding to the surface wear fault type is 70×0.4=28 points. The weighted frequency domain response value here is used to perform weighted calculation on the abnormal state score of the corresponding fault type, so as to obtain a target abnormal state detection result that can more accurately reflect the actual abnormal degree of the fault type. The server combines the target abnormal state detection results corresponding to all fault types to determine the abnormal detection results of the blade. For example, if the target abnormal state detection result of the blade crack fault type exceeds the preset severe abnormal threshold, while the target abnormal state detection results of other fault types are relatively low, the server can determine that the wind turbine blade is likely to have a crack fault, and maintenance personnel need to be arranged in time to inspect and repair the blade. If the target abnormal state detection results of all fault types are within the normal range, it indicates that the current blade is in good operating condition.In this way, the server can provide accurate blade abnormality detection information to the operation and maintenance personnel of the wind farm, ensure the stable operation of the wind turbine, improve power generation efficiency, and reduce the safety risks and economic losses caused by blade failures.
[0014] In the embodiment of the present invention, the voiceprint feature encoding of the blade voiceprint data of the wind turbine to obtain the voiceprint embedding feature corresponding to the blade voiceprint data can be implemented through the following examples.
[0015] Performing time-frequency transformation processing on the blade soundprint data to obtain time-frequency feature data of the blade soundprint data; The time-frequency feature data is encoded with voiceprint features to obtain voiceprint embedding features of the time-frequency feature data.
[0016] In an embodiment of the present invention, exemplarily, after receiving the blade soundprint data of a certain wind turbine, the server starts the first step of soundprint feature encoding, namely, performing time-frequency transformation processing on the blade soundprint data to obtain time-frequency feature data. For example, the blades of this wind turbine have recently experienced some abnormal vibrations, causing the collected soundprint data to present fluctuations different from the normal state in the frequency and time dimensions. The server uses a specific time-frequency transformation algorithm, such as empirical mode decomposition, to convert these soundprint data from simple time domain signals to time-frequency feature data that simultaneously contains time and frequency information. For example, through the transformation, it can be seen that in a certain time period, the soundprint data has a phenomenon of energy concentration in the high frequency band, which may indicate that the blade encountered a specific abnormal condition at that moment. The server performs soundprint feature encoding on the obtained time-frequency feature data, thereby obtaining the soundprint embedding feature of the time-frequency feature data. This process is like creating a unique "digital identity" for the time-frequency feature data. The server uses the soundprint feature encoding algorithm to extract and encode the key features in the time-frequency feature data into a compact vector representation, namely the soundprint embedding feature. Taking the wind turbine as an example, the server extracts key information such as the frequency range and intensity change of abnormal vibration from the time-frequency feature data and converts it into a specific voiceprint embedding feature vector. This vector contains the core features of the blade voiceprint. The server can then perform a series of important operations such as abnormal feature extraction and fault type identification based on this vector to accurately determine the operating status of the blade and ensure the stable and efficient operation of the wind turbine.
[0017] In the embodiment of the present invention, the time-frequency transformation processing is performed on the blade voiceprint data to obtain the time-frequency feature data of the blade voiceprint data, which can be implemented through the following examples.
[0018] When the blade soundprint data of the wind turbine is obtained, a blade fault detection model for detecting blade abnormalities of the wind turbine is obtained; the blade fault detection model includes a first acoustic feature extraction module; Determining an acoustic signal characteristic associated with the blade voiceprint data based on the first acoustic feature extraction module, and extracting a framed voiceprint characteristic of the blade voiceprint data according to the acoustic signal characteristic; Based on the first acoustic feature extraction module, the framed voiceprint feature is converted into a time-frequency feature to obtain a time-frequency feature graph corresponding to the framed voiceprint feature; Based on the first acoustic feature extraction module, the time-frequency feature graph corresponding to the framed voiceprint feature is reconstructed to obtain the time-frequency feature data of the blade voiceprint data.
[0019] In an embodiment of the present invention, illustratively, in a large wind farm, the server continuously receives blade soundprint data from each wind turbine. When the server obtains the blade soundprint data of a wind turbine, it immediately calls the blade fault detection model for detecting abnormal blades of the wind turbine from the storage, and this model includes a first acoustic feature extraction module. The first acoustic feature extraction module starts to work, and it analyzes the incoming blade soundprint data to determine the acoustic signal characteristics associated with it. For example, it identifies the signal intensity change law of a specific frequency range in the blade soundprint data, the periodicity of the signal and other characteristics. Based on these characteristics, the module performs frame processing on the blade soundprint data and extracts the frame soundprint features. Assuming that the blade soundprint data has a duration of 10 seconds, the module divides it into one frame every 0.1 seconds, and obtains 100 frame soundprint features, each of which carries the soundprint characteristics within this 0.1 second. Then, the first acoustic feature extraction module performs time-frequency feature conversion on these frame soundprint features. It converts the frequency, amplitude and other information of each frame voiceprint feature into a time-frequency feature graph. For example, for a certain frame voiceprint feature, the module uses a specific algorithm to display the frequency information at different times within the frame in the form of a two-dimensional graph to form a time-frequency feature graph. The colors or grayscales at different positions in the graph represent different frequency intensities. Finally, the first acoustic feature extraction module reconstructs the time-frequency features of the time-frequency feature graphs corresponding to all frame voiceprint features. It arranges and combines the time-frequency feature graphs of each frame in chronological order to construct a multidimensional matrix, which is the time-frequency feature data of the blade voiceprint data.
[0020] In the embodiment of the present invention, the first acoustic feature extraction module includes a voiceprint preprocessing layer; the framed voiceprint feature includes a plurality of framed voiceprint features, and the plurality of framed voiceprint features include a target framed voiceprint feature; The time-frequency feature conversion of the framed voiceprint feature based on the first acoustic feature extraction module to obtain the time-frequency feature graph corresponding to the framed voiceprint feature can be implemented through the following examples.
[0021] According to the acoustic signal characteristics to which the target framed voiceprint feature belongs, determining the time-frequency analysis domain corresponding to the target framed voiceprint feature; the time-frequency analysis domain includes a plurality of frequency domain response values, each frequency domain response value having a corresponding framed feature range; The frame feature range to which the target frame voiceprint feature belongs is determined based on the voiceprint preprocessing layer, and the frequency domain response value corresponding to the frame feature range to which the target frame voiceprint feature belongs is determined as the time-frequency feature graph corresponding to the target frame voiceprint feature.
[0022] In an embodiment of the present invention, exemplarily, in a certain wind farm, after the server obtains the blade soundprint data of a wind turbine, it calls the first acoustic feature extraction module including the soundprint preprocessing layer. The module performs frame processing on the blade soundprint data to obtain multiple frame soundprint features, of which one is selected as the target frame soundprint feature. The server starts to process the target frame soundprint feature based on the first acoustic feature extraction module. First, according to the acoustic signal characteristics of the target frame soundprint feature itself, such as its frequency distribution range, energy concentration frequency band, etc., the corresponding time-frequency analysis domain is determined. For example, if the energy of the target frame soundprint feature is mainly concentrated in the 500Hz-1000Hz frequency band, the server determines the corresponding time-frequency analysis domain accordingly. This time-frequency analysis domain contains multiple frequency domain response values, each of which corresponds to a specific frame feature range. Then, the soundprint preprocessing layer comes into play. It further analyzes the target frame soundprint feature and accurately determines the frame feature range to which the target frame soundprint feature belongs. Assume that after analysis, it is determined that the target frame voiceprint feature has a frequency of 600Hz-800Hz and an amplitude in a certain range, which is the frame feature range to which it belongs. Then, the server extracts the frequency domain response value corresponding to this frame feature range, and then determines it as the time-frequency feature graph corresponding to the target frame voiceprint feature. This time-frequency feature graph presents the characteristics of the target frame voiceprint feature within a specific frequency and time range in an intuitive way, providing key information for the subsequent comprehensive analysis of the blade voiceprint data and anomaly detection. Through such sophisticated processing, the server can extract key features from complex blade voiceprint data, laying the foundation for accurately judging whether there are abnormalities in the wind turbine blades.
[0023] In the embodiment of the present invention, the abnormal feature vector is determined by extracting abnormal features from the voiceprint embedded features by an acoustic feature encoder in the first acoustic feature extraction module; the first acoustic feature extraction module belongs to a blade fault detection model for detecting blade abnormalities of the wind turbine; the embodiment of the present invention also provides the following implementation methods: Acquire a first voiceprint sample set; the first voiceprint training data included in the first voiceprint sample set is voiceprint training data that is not configured with an abnormal type label; Acquire a plurality of self-supervised learning tasks for the first original acoustic feature extraction module; Based on the first original acoustic feature extraction module, data processing is performed on the first voiceprint training data to obtain an acoustic feature representation associated with each self-supervised learning task, and a plurality of error parameters corresponding to the plurality of self-supervised learning tasks are determined according to the acoustic feature representation; one self-supervised learning task corresponds to one error parameter; A first target error parameter is determined according to the multiple error parameters, a model parameter of the first original acoustic feature extraction module is optimized according to the first target error parameter, and the optimized first original acoustic feature extraction module is determined as the first acoustic feature extraction module.
[0024] In an embodiment of the present invention, exemplarily, the server first obtains a first voiceprint sample set, in which the first voiceprint training data in the sample set are not configured with abnormal type labels. For example, these data may come from voiceprints collected when the blades of each wind turbine are operating normally under different time periods and different wind conditions in the power plant, and also contain a small amount of voiceprint data that may be abnormal but the abnormal type has not yet been clearly identified. Then, the server obtains multiple self-supervised learning tasks for the first original acoustic feature extraction module. These tasks are intended to allow the module to automatically learn useful features from unlabeled data. Taking the time-frequency map mask reconstruction task as an example, the server performs time-frequency transformation processing on the first voiceprint training data based on the first original acoustic feature extraction module, and converts the voiceprint data into the first sample time-frequency feature data containing sample time-frequency feature maps at multiple time frame positions. The server selects the first time frame position to be masked for the time-frequency region from a large number of time frame positions according to a predetermined preset rule. For example, the preset rule is to select one of every 10 time frame positions, and the server will select the corresponding position accordingly. Then, the server performs a time-frequency region masking process on the sample time-frequency feature graph at the first time frame position in the first sample time-frequency feature data according to the masked region identifier to obtain the masked time-frequency sample data. After that, the server again encodes the voiceprint feature of the masked time-frequency sample data based on the first original acoustic feature extraction module to obtain the first training voiceprint embedding feature, and then extracts the abnormal feature to obtain the first acoustic feature representation. The server uses the feature reconstruction module associated with the task to perform a masking process on the first acoustic feature representation to reconstruct the time-frequency graph to obtain the predicted time-frequency feature graph. Finally, by comparing the sample time-frequency feature graph at the first time frame position in the first sample time-frequency feature data with the predicted time-frequency feature graph, the error parameter corresponding to the task is determined. For other self-supervised learning tasks, such as the time-frequency abnormal replacement detection task, the voiceprint denoising comparison task, etc., the server also follows a similar process to obtain the corresponding error parameters respectively. After collecting multiple error parameters corresponding to all self-supervised learning tasks, the server determines the first target error parameter through a specific algorithm. This algorithm may be a weighted average of all error parameters. Based on the first target error parameter, the server optimizes the model parameters of the first original acoustic feature extraction module and adjusts the weights, biases and other parameters within the module. After a series of optimizations, the first original acoustic feature extraction module with improved performance is determined as the first acoustic feature extraction module, which is used to accurately extract abnormal feature vectors from the voiceprint embedded features in the subsequent process, so as to more accurately detect abnormal conditions of wind turbine blades.
[0025] In an embodiment of the present invention, the plurality of self-supervised learning tasks include a time-frequency graph mask reconstruction task; The first voiceprint training data is processed based on the first original acoustic feature extraction module to obtain the acoustic feature representation associated with each self-supervised learning task, and multiple error parameters corresponding to the multiple self-supervised learning tasks are determined according to the acoustic feature representation. The implementation can be performed through the following examples.
[0026] Based on the first original acoustic feature extraction module, the first voiceprint training data is subjected to time-frequency transformation processing to obtain first sample time-frequency feature data of the first voiceprint training data; the first sample time-frequency feature data includes sample time-frequency feature graphs at multiple time frame positions; Selecting a first time frame position to be masked for a time-frequency region from a plurality of time frame positions corresponding to the first sample time-frequency feature data according to a preset rule; Performing time-frequency region masking processing on the sample time-frequency feature graph at the first time frame position in the first sample time-frequency feature data according to the masking region identifier to obtain masked time-frequency sample data; Based on the first original acoustic feature extraction module, the masked time-frequency sample data is encoded with voiceprint features to obtain a first training voiceprint embedding feature of the masked time-frequency sample data, and an abnormal feature is extracted from the first training voiceprint embedding feature to obtain a first acoustic feature representation corresponding to the first training voiceprint embedding feature; the first acoustic feature representation belongs to the acoustic feature representation; Based on a feature reconstruction module associated with the time-frequency map mask reconstruction task, reconstructing the time-frequency map mask processing on the first acoustic feature representation to obtain a predicted time-frequency feature map corresponding to the masked area identifier in the masked time-frequency sample data; According to the sample time-frequency feature graph at the first time frame position in the first sample time-frequency feature data and the predicted time-frequency feature graph, an error parameter corresponding to the time-frequency graph mask reconstruction task is determined.
[0027] In an embodiment of the present invention, exemplarily, in a monitoring system of a wind farm, a server is responsible for processing a large amount of data related to wind turbine blades. At this time, the server is executing a time-frequency map mask reconstruction task based on the first original acoustic feature extraction module, which is one of multiple self-supervised learning tasks. The server first performs time-frequency transformation processing on the first voiceprint training data based on the first original acoustic feature extraction module. Assume that the first voiceprint training data is a voiceprint record of a wind turbine blade within a period of time. After processing, the first sample time-frequency feature data is obtained, and each time frame position records the sample time-frequency feature map at the corresponding moment, showing the characteristics of the voiceprint at the moment in the time and frequency dimensions. Then, the server selects the first time frame position to be masked in the time-frequency region from the multiple time frame positions according to the preset rules. The preset rule may be to randomly select 10% of the time frame positions. The server determines a specific first time frame position from a large number of time frame positions through a random algorithm. Afterwards, the server performs time-frequency region masking processing on the sample time-frequency feature map at the first time frame position in the first sample time-frequency feature data according to the masking region identifier. For example, if the masked region identifier specifies an area within a certain frequency range, the server masks the area corresponding to this frequency range in the sample time-frequency feature map to obtain the masked time-frequency sample data. Then, the server again encodes the voiceprint feature of the masked time-frequency sample data based on the first original acoustic feature extraction module, generates the first training voiceprint embedding feature, and then extracts the abnormal feature to obtain the first acoustic feature representation corresponding to the first training voiceprint embedding feature, which is part of the acoustic feature representation. Next, the server calls the feature reconstruction module associated with the task to perform mask processing on the first acoustic feature representation to reconstruct the time-frequency map. The feature reconstruction module attempts to restore the time-frequency features of the masked area based on the information in the first acoustic feature representation, and obtains the predicted time-frequency feature map corresponding to the masked region identifier in the masked time-frequency sample data. Finally, the server calculates the difference between the original sample time-frequency feature map at the first time frame position in the first sample time-frequency feature data and the reconstructed predicted time-frequency feature map, so as to determine the error parameter corresponding to the time-frequency map mask reconstruction task. This error parameter reflects the accuracy of the first original acoustic feature extraction module in processing the task, and provides an important basis for subsequent optimization modules.
[0028] In an embodiment of the present invention, the plurality of self-supervised learning tasks include a time-frequency anomaly replacement detection task; The first voiceprint training data is processed based on the first original acoustic feature extraction module to obtain the acoustic feature representation associated with each self-supervised learning task, and multiple error parameters corresponding to the multiple self-supervised learning tasks are determined according to the acoustic feature representation. The implementation can be performed through the following examples.
[0029] Based on the first original acoustic feature extraction module, the first voiceprint training data is subjected to time-frequency transformation processing to obtain first sample time-frequency feature data of the first voiceprint training data; the first sample time-frequency feature data includes sample time-frequency feature graphs at multiple time frame positions; Selecting a second time frame position to be subjected to time-frequency feature replacement according to a preset rule from a plurality of time frame positions corresponding to the first sample time-frequency feature data; Performing time-frequency feature replacement processing on the sample time-frequency feature graph at the second time frame position in the first sample time-frequency feature data according to preset time-frequency information to obtain replaced time-frequency sample data; the preset time-frequency information is different from the sample time-frequency feature graph at the second time frame position; the preset time-frequency information is an abnormal time-frequency feature fragment extracted from historical fault data, or an unnatural time-frequency feature generated by Gaussian noise; Based on the first original acoustic feature extraction module, the replaced time-frequency sample data is encoded with voiceprint features to obtain a second training voiceprint embedding feature of the replaced time-frequency sample data, and the second training voiceprint embedding feature is extracted with abnormal features to obtain a second acoustic feature representation corresponding to the second training voiceprint embedding feature; the second acoustic feature representation belongs to the acoustic feature representation; Based on a feature reconstruction module associated with the time-frequency anomaly replacement detection task, determining a target detection time frame position for the second acoustic feature representation, and obtaining time-frequency feature replacement prediction results corresponding to the multiple time frame positions; According to the time-frequency feature replacement prediction results corresponding to the second time frame position and the multiple time frame positions, the error parameter corresponding to the time-frequency anomaly replacement detection task is determined.
[0030] In an embodiment of the present invention, exemplarily, at the data processing server end of a large wind farm, a time-frequency anomaly replacement detection task for wind turbine blade soundprint data is being carried out, which is one of multiple self-supervised learning tasks. The server first uses the first original acoustic feature extraction module to perform time-frequency transformation processing on the first soundprint training data. For example, the training data comes from the blade soundprints of multiple wind turbines in different periods in a wind farm. After processing, the first sample time-frequency feature data is obtained, which is like an ordered set, and each element is a sample time-frequency feature graph at a time frame position, showing the frequency characteristics of the soundprint at different times. Then, the server selects the second time frame position to be replaced with the time-frequency feature from the many time frame positions according to the preset rules. Assuming that the preset rule is to select the position where the time frame position sequence number is an odd number and the frequency fluctuation exceeds a certain threshold, the server determines the second time frame position that meets the conditions by analyzing and screening the sample time-frequency feature data. Afterwards, the server performs time-frequency feature replacement processing on the sample time-frequency feature graph at the second time frame position in the first sample time-frequency feature data according to the preset time-frequency information. For example, a segment of abnormal time-frequency feature is extracted from the historical fault data as the preset time-frequency information, and this segment of information is different from the sample time-frequency feature map at the current second time frame position. The server replaces the corresponding part of the original sample time-frequency feature map with the abnormal time-frequency feature segment to obtain the replaced time-frequency sample data. Subsequently, the server again encodes the replaced time-frequency sample data with voiceprint features based on the first original acoustic feature extraction module to obtain the second training voiceprint embedding feature, and extracts the abnormal features, thereby obtaining the second acoustic feature representation corresponding to the second training voiceprint embedding feature, which belongs to a part of the overall acoustic feature representation. Next, the server determines the target detection time frame position for the second acoustic feature representation with the help of a feature reconstruction module associated with the task. The feature reconstruction module analyzes and determines the time-frequency feature replacement corresponding to multiple time frame positions based on the information in the second acoustic feature representation, and obtains the time-frequency feature replacement prediction result. Finally, the server calculates the difference between the first selected second time frame position and the obtained time-frequency feature replacement prediction results corresponding to the multiple time frame positions, so as to determine the error parameter corresponding to the time-frequency abnormal replacement detection task. This error parameter reflects the processing accuracy of the first original acoustic feature extraction module in this task, and provides a key basis for the subsequent optimization of module performance.
[0031] In an embodiment of the present invention, the plurality of self-supervised learning tasks include a voiceprint denoising comparison task; The first voiceprint training data is processed based on the first original acoustic feature extraction module to obtain the acoustic feature representation associated with each self-supervised learning task, and multiple error parameters corresponding to the multiple self-supervised learning tasks are determined according to the acoustic feature representation. The implementation can be performed through the following examples.
[0032] Injecting Gaussian white noise features into the framed voiceprint features of the first voiceprint training data according to the signal-to-noise ratio to obtain noisy voiceprint training data, and performing time-frequency transformation processing on the noisy voiceprint training data to obtain noise time-frequency feature data of the noisy voiceprint training data; Performing time-frequency transformation processing on the first voiceprint training data to obtain first sample time-frequency feature data of the first voiceprint training data; Based on the first original acoustic feature extraction module, the noise time-frequency feature data is encoded with voiceprint features to obtain a third training voiceprint embedding feature of the noise time-frequency feature data, and the third training voiceprint embedding feature is extracted with abnormal features to obtain a third acoustic feature representation corresponding to the third training voiceprint embedding feature; the third acoustic feature representation belongs to the acoustic feature representation; Based on the first original acoustic feature extraction module, the first sample time-frequency feature data is encoded with voiceprint features to obtain a fourth training voiceprint embedding feature of the first sample time-frequency feature data, and an abnormal feature is extracted from the fourth training voiceprint embedding feature to obtain a fourth acoustic feature representation corresponding to the fourth training voiceprint embedding feature; the fourth acoustic feature representation belongs to the acoustic feature representation; According to the third acoustic feature representation and the fourth acoustic feature representation, an error parameter corresponding to the voiceprint denoising comparison task is determined.
[0033] In an embodiment of the present invention, exemplarily, in the operation and maintenance data processing center of a large wind farm, the server is performing a voiceprint denoising and comparison task on the voiceprint data of wind turbine blades, which is an important link in multiple self-supervised learning tasks. The server first processes the first voiceprint training data. These training data come from the blade voiceprints collected by different wind turbines in different working conditions in the power plant. The server injects Gaussian white noise features into the framed voiceprint features of the first voiceprint training data according to a specific signal-to-noise ratio. For example, the signal-to-noise ratio is set to 20dB, and white noise that conforms to the Gaussian distribution is added to each framed voiceprint feature to obtain the noise voiceprint training data. Subsequently, the noise voiceprint training data is subjected to time-frequency transformation processing to obtain the noise time-frequency feature data of the noise voiceprint training data, which presents the characteristics of the voiceprint with noise interference in the time and frequency dimensions. At the same time, the server also performs time-frequency transformation processing on the first voiceprint training data without adding noise, and obtains the first sample time-frequency feature data of the first voiceprint training data, which represents the time-frequency characteristics of the original pure voiceprint. Next, the server further processes the noise time-frequency feature data and the first sample time-frequency feature data based on the first original acoustic feature extraction module. For the noise time-frequency feature data, the third training voiceprint embedding feature is obtained after voiceprint feature encoding, and then the abnormal feature extraction is performed to obtain the third acoustic feature representation corresponding to the third training voiceprint embedding feature, which is part of the overall acoustic feature representation. Similarly, the first sample time-frequency feature data is voiceprint feature encoded to obtain the fourth training voiceprint embedding feature, and then the abnormal feature extraction is performed to obtain the fourth acoustic feature representation corresponding to the fourth training voiceprint embedding feature. Finally, the server determines the error parameter corresponding to the voiceprint denoising comparison task based on the third acoustic feature representation and the fourth acoustic feature representation. The server quantifies the performance difference of the first original acoustic feature extraction module when processing noisy and pure voiceprint data by calculating the difference between the two acoustic feature representations, such as calculating the distance between the two in the feature vector space, the deviation of key feature parameters, and other indicators. This difference value is the error parameter corresponding to the voiceprint denoising comparison task. This error parameter will provide an important basis for the subsequent optimization of the first original acoustic feature extraction module, so as to improve its processing ability of blade soundprint data in a noisy environment and ensure the accuracy of abnormality detection of wind turbine blades.
[0034] In an embodiment of the present invention, the plurality of self-supervised learning tasks include an expert knowledge transfer task; The first voiceprint training data is processed based on the first original acoustic feature extraction module to obtain the acoustic feature representation associated with each self-supervised learning task, and multiple error parameters corresponding to the multiple self-supervised learning tasks are determined according to the acoustic feature representation. The implementation can be performed through the following examples.
[0035] According to the feature reconstruction module associated with the expert knowledge transfer task, the fourth acoustic feature representation is rated for abnormal risk to obtain an abnormal risk score corresponding to the first voiceprint training data; Acquire a pre-trained benchmark model associated with the first original acoustic feature extraction module, perform abnormal risk rating on the first voiceprint training data according to the pre-trained benchmark model, and obtain a benchmark abnormal state score corresponding to the first voiceprint training data; An error parameter corresponding to the expert knowledge transfer task is determined according to the abnormal risk score and the benchmark abnormal state score.
[0036] In an embodiment of the present invention, exemplarily, an expert knowledge transfer task is being performed on a data processing server of a wind farm, which is part of a plurality of self-supervised learning tasks. The server first performs time-frequency transformation processing on the first voiceprint training data. These first voiceprint training data are collected from multiple generator blades of a wind farm, covering voiceprint information under different operating conditions. After processing, the first sample time-frequency feature data of the first voiceprint training data is obtained, which clearly shows the distribution characteristics of the voiceprint in the time and frequency dimensions. Then, the server uses the first original acoustic feature extraction module to encode the first sample time-frequency feature data of the voiceprint feature, obtains the fourth training voiceprint embedding feature, and then extracts the abnormal feature to obtain the fourth acoustic feature representation corresponding to the fourth training voiceprint embedding feature, which is an important part of the overall acoustic feature representation. After that, the server calls the feature reconstruction module associated with the expert knowledge transfer task. This module performs abnormal risk rating on the fourth acoustic feature representation based on the knowledge and algorithm contained in itself. For example, the module will analyze various characteristic indicators in the fourth acoustic feature representation, such as energy changes in a specific frequency band, periodicity of the voiceprint, etc., and give a corresponding abnormal risk score for the first voiceprint training data based on these analysis results. Assuming that the score range is 0-100 points, the higher the score, the greater the abnormal risk. After evaluation, the first voiceprint training data scored 40 points. At the same time, the server obtains a pre-trained benchmark model associated with the first original acoustic feature extraction module. This pre-trained benchmark model has been trained on a large amount of data and has accumulated rich judgment experience. The server uses the pre-trained benchmark model to perform abnormal risk rating on the same first voiceprint training data and obtains the benchmark abnormal state score corresponding to the first voiceprint training data. For example, the score given by the benchmark model is 35 points. Finally, the server determines the error parameter corresponding to the expert knowledge transfer task based on the abnormal risk score and the benchmark abnormal state score. The server quantifies this error parameter by calculating the difference, ratio, etc. between the two. For example, by calculating the absolute value of the difference between the two (40-35=5), taking this value as part of the error parameter, and combining it with other relevant calculation rules, we can finally determine an error parameter that can reflect the gap between the first original acoustic feature extraction module and the expert knowledge (pre-trained benchmark model) in this task. This error parameter will guide the subsequent optimization of the first original acoustic feature extraction module, so that it can better utilize expert knowledge and improve the accuracy of abnormal detection of wind turbine blade soundprint data.
[0037] In the embodiment of the present invention, the abnormal feature vector is determined by extracting abnormal features from the voiceprint embedded features by an acoustic feature encoder in a first acoustic feature extraction module; the first acoustic feature extraction module belongs to a blade fault detection model for detecting blade abnormalities of the wind turbine; the blade fault detection model also includes a second acoustic feature extraction module after multi-level feature fusion of the acoustic feature encoder; The abnormal state identification is performed on the abnormal feature vector according to each fault type to obtain the abnormal state score corresponding to each fault type, which can be implemented through the following example.
[0038] Determining a second acoustic feature extraction module from the blade fault detection model; the second acoustic feature extraction module includes a fault diagnosis subnet corresponding to each fault type; Based on the fault diagnosis subnet corresponding to each fault type, the abnormal state of the abnormal feature vector is identified to obtain the abnormal state score corresponding to each fault type; one fault diagnosis subnet is used to obtain the abnormal state score corresponding to one fault type.
[0039] In an embodiment of the present invention, exemplarily, the server obtains the voiceprint embedding feature obtained after processing by the first acoustic feature extraction module, and the acoustic feature encoder in the module extracts the abnormal feature and determines the abnormal feature vector. The blade fault detection model also includes a second acoustic feature extraction module located after the acoustic feature encoder, which adopts a multi-level feature fusion technology to more accurately analyze the blade fault. When it is necessary to identify the abnormal state of the abnormal feature vector according to different fault types to obtain the abnormal state score corresponding to each fault type, the server first determines the second acoustic feature extraction module from the blade fault detection model. The module contains a fault diagnosis subnet for each fault type, such as a special fault diagnosis subnet for the blade crack fault type, and a corresponding subnet for the blade surface wear fault type. Taking a certain wind turbine as an example, the server obtains the abnormal feature vector of the wind turbine blade soundprint data after processing. For the blade crack fault type, the server calls the corresponding fault diagnosis subnet in the second acoustic feature extraction module. This subnet performs an in-depth analysis of the abnormal feature vector based on the pre-trained algorithm and parameters. It may examine specific characteristic indicators related to blade cracks in the abnormal feature vector, such as energy changes within a specific frequency range, mutations in the acoustic signal, etc. Through a comprehensive evaluation of these indicators, the fault diagnosis subnet outputs a value as the abnormal state score corresponding to the blade crack fault type. Assuming the full score is 100 points, after analysis and calculation, the fault diagnosis subnet gives an abnormal state score of 60 points for the blade crack fault type, indicating that the possibility of crack faults in the blade is at a medium level. Similarly, for the blade surface wear fault type, the server calls the corresponding fault diagnosis subnet, processes the same abnormal feature vector according to the subnet's unique analysis logic and algorithm, and obtains the abnormal state score corresponding to the blade surface wear fault type, such as 40 points, indicating that the possibility of surface wear faults in the blade is relatively low. In this way, the server uses different fault diagnosis subnets to perform abnormal state identification on the abnormal feature vector for each fault type, and obtains the abnormal state score corresponding to each fault type, providing a key basis for accurately judging the fault condition of the wind turbine blade.
[0040] In the embodiments of the present invention, the following implementation modes are also provided.
[0041] Acquire a second voiceprint sample set; the second voiceprint training data included in the second voiceprint sample set is associated with a fault type label; the fault type label includes at least two fault type codes corresponding to the at least two fault types, one fault type corresponds to one fault type code, and the at least two fault type codes are determined according to the fault type to which the second voiceprint training data belongs; Acquire a pre-trained first acoustic feature extraction module, perform voiceprint feature encoding on the second voiceprint training data based on the first acoustic feature extraction module, obtain a fifth training voiceprint embedding feature of the second voiceprint training data, perform abnormal feature extraction on the fifth training voiceprint embedding feature, and obtain a fifth acoustic feature representation corresponding to the fifth training voiceprint embedding feature; Obtain at least two original fault diagnosis subnets corresponding to the at least two fault types, and perform fault type discrimination on the fifth acoustic feature representations according to the at least two, respectively, to obtain at least two fault type probabilities corresponding to the at least two original fault diagnosis subnets; one fault type corresponds to one original fault diagnosis subnet, and one original fault diagnosis subnet is used to obtain one fault type probability; A second target error parameter is determined according to the at least two fault type probabilities and the at least two fault type codes, model parameters of the at least two original fault diagnosis subnets are optimized according to the second target error parameter, the optimized at least two original fault diagnosis subnets are determined as the at least two fault diagnosis subnets, and the second acoustic feature extraction module is determined according to the at least two fault diagnosis subnets.
[0042] In an embodiment of the present invention, exemplarily, in a data processing server of a wind farm, in order to more accurately construct a second acoustic feature extraction module for blade abnormality detection, the server performs the following series of operations. The server first obtains a second voiceprint sample set, in which the second voiceprint training data in the sample set are all associated with a fault type label. For example, during the long-term monitoring process of a large wind farm, a large number of voiceprint data of different wind turbine blades under various fault conditions are collected to form a second voiceprint sample set. These fault types include blade cracks, blade surface wear, blade imbalance, etc., and each fault type corresponds to a specific fault type code. For example, blade cracks correspond to the code "001", and blade surface wear corresponds to the code "002", which are determined based on the fault type to which the second voiceprint training data actually belongs. Next, the server calls the pre-trained first acoustic feature extraction module to process the second voiceprint training data. Taking one of the second voiceprint training data about blade surface wear as an example, the first acoustic feature extraction module first encodes the voiceprint feature to obtain the fifth training voiceprint embedding feature, and then further extracts the abnormal feature of the embedded feature to obtain the fifth acoustic feature representation. After that, the server obtains at least two original fault diagnosis subnets corresponding to at least two fault types. For example, there are corresponding original fault diagnosis subnets for the two fault types of blade crack and blade surface wear. The server uses these two original fault diagnosis subnets to distinguish the fault type of the fifth acoustic feature representation obtained above. Each original fault diagnosis subnet analyzes the feature information related to the fault type in the fifth acoustic feature representation based on its own algorithm and parameters, so as to obtain the corresponding fault type probability. For example, after analysis, the original fault diagnosis subnet corresponding to the blade crack obtains that the probability of the voiceprint data belonging to the blade crack fault type is 0.3, and the original fault diagnosis subnet corresponding to the blade surface wear obtains that the probability of belonging to the blade surface wear fault type is 0.6. Finally, the server determines the second target error parameter based on the at least two fault type probabilities and the corresponding at least two fault type codes. It uses a specific calculation method, such as comparing the difference between the fault type probability and the actual situation represented by the fault type code, to obtain a value that reflects the performance deviation of the original fault diagnosis subnet, namely the second target error parameter. Based on this error parameter, the server adjusts and optimizes the model parameters of at least two original fault diagnosis subnets. For example, modify the weight coefficients in the subnet, adjust the activation function of the neuron, etc. After optimization, the two original fault diagnosis subnets are determined as formal fault diagnosis subnets, and then the second acoustic feature extraction module is determined based on these optimized fault diagnosis subnets, so that it can more accurately identify abnormal states of different fault types in the subsequent abnormal detection of wind turbine blades.
[0043] In the embodiment of the present invention, the voiceprint embedding feature is determined by the acoustic feature encoding layer in the first acoustic feature extraction module performing voiceprint feature encoding on the time-frequency feature data; the time-frequency feature data is determined by the voiceprint preprocessing layer in the first acoustic feature extraction module performing time-frequency transformation processing on the blade voiceprint data; the first acoustic feature extraction module belongs to a blade fault detection model for detecting blade anomalies of the wind turbine; the blade fault detection model also includes a third acoustic feature extraction module in which multi-level features are fused after the acoustic feature encoding layer; The fault weight analysis of the voiceprint embedded features to obtain the fault type weight corresponding to each fault type may be performed through the following example.
[0044] determining a third acoustic feature extraction module from the blade fault detection model; Based on the third acoustic feature extraction module, a fault weight analysis is performed on the voiceprint embedding feature to obtain a fault type weight corresponding to each fault type.
[0045] In the embodiment of the present invention, the server first receives the voiceprint data from the blades of the wind turbine. Taking a wind turbine under a specific working condition as an example, the voiceprint data of its blades is transmitted to the server. The first acoustic feature extraction module starts to work, and the voiceprint preprocessing layer in the module performs time-frequency transformation processing on the blade voiceprint data. This is like converting the voiceprint data from one "language" to another "language" that is easier to analyze, converting the voiceprint signal that simply changes with time into time-frequency feature data, so that the characteristics of the voiceprint in the two dimensions of time and frequency are clearly presented. For example, the voiceprint preprocessing layer uses a specific algorithm to sort out the frequency information at different times in the voiceprint data, and obtains the time-frequency feature data containing the frequency distribution corresponding to each time point. Then, the acoustic feature encoding layer in the first acoustic feature extraction module encodes the time-frequency feature data for voiceprint features, thereby determining the voiceprint embedding feature. The acoustic feature encoding layer extracts the most representative key information from the time-frequency feature data and encodes it into a compact vector form, namely the voiceprint embedding feature. This voiceprint embedding feature contains the core features of the blade voiceprint data, providing an important basis for subsequent fault analysis. In the blade fault detection model, a third acoustic feature extraction module with multi-level feature fusion is also provided after the acoustic feature encoding layer. When it is necessary to perform fault weight analysis on the voiceprint embedding feature to obtain the fault type weight corresponding to each fault type, the server determines the third acoustic feature extraction module from the blade fault detection model. Then, the server performs fault weight analysis on the voiceprint embedding feature based on the third acoustic feature extraction module. For example, the module may use a multi-branch feature extractor to extract multimodal acoustic features from different dimensions of the voiceprint embedding feature, such as frequency, phase, amplitude, etc., to obtain the multimodal feature mapping results corresponding to the voiceprint embedding feature. Afterwards, the dynamic weight allocator in the module performs a comprehensive analysis and calculation on these multimodal feature mapping results. It assigns corresponding weights to each fault type based on pre-set algorithms and rules, combined with the degree of association between different fault types and these features. Assuming that the blades of a wind turbine may have three types of faults: blade crack, blade wear, and blade icing, after analysis and calculation by the third acoustic feature extraction module, it is determined that the weight of the blade crack fault type is 0.5, the weight of the blade wear fault type is 0.3, and the weight of the blade icing fault type is 0.2. These weights reflect the relative probability of each fault type under the current voiceprint embedding feature, providing a key quantitative basis for the subsequent accurate judgment of blade faults.
[0046] In an embodiment of the present invention, the blade fault detection model further includes a second acoustic feature extraction module for obtaining an abnormal state score corresponding to each fault type; The embodiment of the present invention also provides the following implementation mode: Acquire a third voiceprint sample set; the third voiceprint training data included in the third voiceprint sample set is associated with an abnormal type label; Acquire a pre-trained first acoustic feature extraction module, perform voiceprint feature encoding on the third voiceprint training data based on the first acoustic feature extraction module, obtain a sixth training voiceprint embedding feature of the third voiceprint training data, perform abnormal feature extraction on the sixth training voiceprint embedding feature, and obtain a sixth acoustic feature representation corresponding to the sixth training voiceprint embedding feature; Obtain a pre-trained second acoustic feature extraction module, perform abnormal state recognition on the sixth acoustic feature representation based on at least two fault diagnosis subnets in the second acoustic feature extraction module, and obtain a training abnormal state score corresponding to each fault type; one fault diagnosis subnet is used to obtain a training abnormal state score corresponding to one fault type; Performing a fault weight analysis on the sixth training voiceprint embedding feature based on the third original acoustic feature extraction module to obtain a training fault type weight corresponding to each fault type; Determine, according to the training abnormal state score corresponding to each fault type and the training fault type weight corresponding to each fault type, a training blade abnormality detection result of the reference wind turbine associated with the third voiceprint training data; A third target error parameter is determined according to the abnormality type label and the training blade abnormality detection result, and the model parameters of the third original acoustic feature extraction module are optimized according to the third target error parameter. The optimized third original acoustic feature extraction module is determined as the third acoustic feature extraction module, and the blade fault detection model is determined according to the first acoustic feature extraction module, the second acoustic feature extraction module and the third acoustic feature extraction module.
[0047] In an embodiment of the present invention, exemplarily, the server first obtains a third voiceprint sample set, and the third voiceprint training data in the sample set are all associated with an abnormal type label. For example, these data come from different wind turbine blade voiceprint records accumulated in power plants for a long time, covering various normal and abnormal operating conditions. The abnormal type label clearly marks the abnormal situation corresponding to each data, such as blade cracks, wear, deformation, etc. Then, the server calls the pre-trained first acoustic feature extraction module to process the third voiceprint training data. Taking one of the third voiceprint training data associated with the blade wear abnormal type label as an example, the first acoustic feature extraction module first encodes the voiceprint feature to obtain the sixth training voiceprint embedded feature, and then extracts the abnormal feature of the embedded feature to obtain the sixth acoustic feature representation. After that, the server obtains the pre-trained second acoustic feature extraction module. The module contains at least two fault diagnosis subnets, each for a fault type. Based on the fault diagnosis subnet in this module, the server identifies the abnormal state of the sixth acoustic feature representation. For example, for the blade crack fault diagnosis subnet and the blade wear fault diagnosis subnet, they analyze the feature information related to their respective fault types in the sixth acoustic feature representation according to their own algorithms and parameters, and obtain the training abnormal state score corresponding to each fault type. Assuming that the training abnormal state score given by the blade crack fault diagnosis subnet is 30 points, and the score given by the blade wear fault diagnosis subnet is 70 points, this indicates that under the current voiceprint data, the possibility of abnormal wear of the blade is relatively high. At the same time, the server performs a fault weight analysis on the sixth training voiceprint embedding feature based on the third original acoustic feature extraction module to obtain the training fault type weight corresponding to each fault type. For example, after analysis, it is determined that the training fault type weight of the blade crack fault type is 0.4, and the weight of the blade wear fault type is 0.6. Next, the server determines the training blade abnormality detection result of the reference wind turbine associated with the third voiceprint training data according to the training abnormal state score and training fault type weight corresponding to each fault type. For example, through a specific calculation method, the score of blade cracks (30 points) is multiplied by a weight of 0.4, and the score of blade wear (70 points) is multiplied by a weight of 0.6, and then the calculation results of other possible fault types are combined to obtain the training blade anomaly detection result of the reference wind turbine blade. Finally, the server determines the third target error parameter based on the anomaly type label and the training blade anomaly detection result. For example, the difference between the blade wear anomaly marked by the anomaly type label and the calculated training blade anomaly detection result is compared, and the third target error parameter is obtained through a specific algorithm. Based on this error parameter, the server optimizes the model parameters of the third original acoustic feature extraction module, such as adjusting the weights and thresholds of the internal algorithm. After the optimization is completed, it is determined as the third acoustic feature extraction module.By combining the first acoustic feature extraction module and the second acoustic feature extraction module, a blade fault detection model with better performance is finally determined to improve the accuracy of abnormality detection of wind turbine blades.
[0048] In an embodiment of the present invention, the third acoustic feature extraction module includes a multi-branch feature extractor and a dynamic weight allocator; The fault weight analysis of the voiceprint embedding feature based on the third acoustic feature extraction module to obtain the fault type weight corresponding to each fault type can be implemented through the following example.
[0049] Performing multimodal acoustic feature extraction on the voiceprint embedding feature based on the multi-branch feature extractor to obtain a multimodal feature mapping result corresponding to the voiceprint embedding feature; Based on the dynamic weight allocator, fault weight calculation is performed on the multimodal feature mapping result to obtain the fault type weight corresponding to each fault type.
[0050] In an embodiment of the present invention, exemplarily, in a central monitoring server of a large wind farm, a key link in wind turbine blade fault detection is to perform fault weight analysis on the voiceprint embedded features through a third acoustic feature extraction module to determine the weight of each fault type.
[0051] The server calls the third acoustic feature extraction module from the blade fault detection model, which consists of a multi-branch feature extractor and a dynamic weight allocator. Taking the soundprint embedded feature data of a wind turbine blade as an example, this data has been obtained by preliminary processing and transmitted to the third acoustic feature extraction module.
[0052] First, the multi-branch feature extractor starts working. It extracts multimodal acoustic features from the voiceprint embedded features from different dimensions. One branch focuses on frequency features, analyzing the distribution and changes of different frequency components in the voiceprint embedded features, such as identifying energy concentration areas in specific frequency bands, which may be related to specific faults. Another branch focuses on phase features, capturing the phase changes of voiceprint signals at different time points. Abnormal changes in phase may also indicate that there is a fault in the blade. There is also a branch that focuses on amplitude features to study the size and fluctuation law of the amplitude of the voiceprint signal. Through these multi-dimensional extractions, the multimodal feature mapping results corresponding to the voiceprint embedded features are obtained. This result integrates the detailed feature information of the voiceprint in multiple modes such as frequency, phase, and amplitude, providing a rich data basis for subsequent fault weight calculations.
[0053] Subsequently, the dynamic weight allocator takes over the multimodal feature mapping results. It comprehensively considers and calculates these multimodal features based on pre-set complex algorithms and rules. For example, for the three common fault types of blade cracks, blade wear and blade imbalance, the algorithm evaluates the degree of association between each feature in the multimodal feature mapping results and different fault types. If the energy concentration phenomenon in a certain high-frequency band in the frequency feature is closely related to the blade crack fault, then when calculating the blade crack fault type weight, the frequency feature will be given a higher weight. The dynamic weight allocator combines the influence of all relevant features, calculates the weight for each fault type separately, and finally obtains the fault type weight corresponding to each fault type. Assume that after calculation, the blade crack fault type weight is 0.5, indicating that under the current soundprint embedded feature, the possibility of blade crack fault is relatively high; the blade wear fault type weight is 0.3; and the blade imbalance fault type weight is 0.2. These weights provide a quantitative basis for the subsequent accurate judgment of the possibility of blade fault type, which helps operation and maintenance personnel to inspect and maintain wind turbine blades more targeted.
[0054] In an embodiment of the present invention, the at least two fault types include a target fault type; The embodiments of the present invention also provide the following implementation modes.
[0055] Obtaining a blade fault detection model for detecting blade anomalies of a wind turbine, and obtaining an edge computing optimization architecture for expert knowledge migration from the blade fault detection model; Acquire a fourth voiceprint sample set associated with the target fault type; the fourth voiceprint sample set includes fourth voiceprint training data; Performing data processing on the fourth voiceprint training data based on the blade fault detection model to obtain a blade abnormality detection result corresponding to the fourth voiceprint training data, and determining the blade abnormality detection result corresponding to the fourth voiceprint training data as a simulation diagnosis label for training the edge computing optimization architecture; Performing data processing on the fourth voiceprint training data based on the edge computing optimization architecture to obtain an edge blade anomaly detection result corresponding to the fourth voiceprint training data; The architectural parameters of the edge computing optimization architecture are optimized according to the edge blade anomaly detection result and the simulation diagnostic label, and the optimized edge computing optimization architecture is determined as the lightweight diagnostic network of the blade fault detection model; the lightweight diagnostic network is used to perform blade anomaly detection on the blade soundprint data under the target fault type.
[0056] In an embodiment of the present invention, it is assumed that among at least two fault types of a wind turbine blade, blade crack is set as a target fault type. The server first obtains a blade fault detection model for blade anomaly detection, and simultaneously obtains an edge computing optimization architecture for expert knowledge migration from the model. The blade fault detection model has been trained on a large amount of data and has accumulated rich experience in fault detection, while the edge computing optimization architecture aims to lightweight the model so that it can run efficiently on edge devices. Next, the server obtains a fourth voiceprint sample set associated with the target fault type (blade crack), and the sample set contains fourth voiceprint training data. These data are from voiceprints collected when crack faults are suspected to occur in wind turbine blades under different working conditions, covering voiceprint information corresponding to cracks of different degrees and different positions. Then, the server processes the fourth voiceprint training data based on the blade fault detection model. The various modules in the model work together, from voiceprint feature encoding to abnormal feature extraction, to fault type discrimination and other steps, and finally obtain the blade anomaly detection result corresponding to the fourth voiceprint training data. For example, after analysis and calculation, the model determines that the blades corresponding to some data have a high probability of crack faults, and the blades corresponding to some data have a low probability of crack faults. These results are determined as simulation diagnostic labels for training the edge computing optimization architecture, and serve as reference standards for subsequent evaluation of the performance of the edge computing optimization architecture. Afterwards, the server uses the edge computing optimization architecture to process the fourth voiceprint training data. The architecture analyzes the voiceprint data based on its own algorithm, and attempts to detect whether the blade has the target fault type (blade crack), thereby obtaining the edge blade anomaly detection results corresponding to the fourth voiceprint training data. For example, the edge computing optimization architecture may determine that the blades corresponding to some data have potential crack risks, but there is a certain difference from the judgment of the blade fault detection model. Finally, the server optimizes the architecture parameters of the edge computing optimization architecture based on the edge blade anomaly detection results and simulation diagnostic labels. By comparing the differences between the two, the error parameters are calculated, and the parameters such as the algorithm and weight within the architecture are adjusted based on these parameters. For example, if the edge computing optimization architecture misjudges some blade cracks, the server will adjust the relevant parameters to make the architecture more accurate in subsequent detections. After optimization, the edge computing optimization architecture is determined as a lightweight diagnostic network for the blade fault detection model. This lightweight diagnostic network is specifically used to quickly and accurately detect blade anomalies in blade soundprint data under the target fault type (blade crack), and can be deployed on edge devices close to wind turbines to improve the real-time and efficiency of fault detection. Based on the above, the corresponding causes of blade fault types in the embodiment of the present invention can be found in Table 1.
[0057] Table 1 In an embodiment of the present invention, the fourth voiceprint training data includes fault condition voiceprint training data generated under the target fault type; The obtaining of the fourth voiceprint sample set associated with the target fault type may be implemented through the following example.
[0058] Acquire fault condition voiceprint training data under the target fault type; Determine the pending voiceprint training data from the voiceprint training data set used for voiceprint feature screening; Based on the voiceprint sample expansion module, an acoustic fingerprint comparison is performed on the pending voiceprint training data and the fault condition voiceprint training data to obtain a matching degree between the fault condition voiceprint training data and the pending voiceprint training data; If the matching degree reaches the data enhancement matching degree threshold, the pending voiceprint training data is determined as the extended voiceprint training data associated with the target fault type; The extended voiceprint training data and the fault condition voiceprint training data are determined as fourth voiceprint training data associated with the target fault type.
[0059] In the embodiment of the present invention, exemplarily, assuming that the target fault type is blade surface corrosion, the server carries out the acquisition of the fourth voiceprint sample set around this fault type. First, the server obtains the fault condition voiceprint training data under the target fault type (blade surface corrosion). These data are derived from the actual operating conditions when the wind turbine blade has a surface corrosion fault. For example, in coastal areas, wind turbines have corrosion on the blade surface due to sea breeze erosion. The server collects the voiceprint data generated by blades with different corrosion degrees and different operating environments as fault condition voiceprint training data. Then, the server determines the pending voiceprint training data from the voiceprint training data set used for voiceprint feature screening. The set contains various voiceprint data accumulated by the power plant for a long time, covering different wind turbines and different operating states. The server selects voiceprint data that may be related to blade surface corrosion faults from the massive data as pending voiceprint training data through specific screening rules. For example, data that operates in a similar environment and has similar trends in voiceprint features in certain frequency ranges or energy distributions are screened out. Then, the server compares the acoustic fingerprints of the pending voiceprint training data with the fault condition voiceprint training data based on the voiceprint sample extension module. The voiceprint sample extension module compares the two sets of data in detail from the characteristics of various dimensions of the voiceprint data, such as frequency, phase, amplitude, etc., and calculates the matching degree between the fault condition voiceprint training data and the pending voiceprint training data. For example, through a complex algorithm, it is concluded that the matching degree between a certain pending voiceprint training data and the fault condition voiceprint training data in key frequency features is 80%. If the matching degree reaches the data enhancement matching degree threshold, assuming that the threshold is set to 70%, then the server determines this pending voiceprint training data as an extended voiceprint training data associated with the target fault type (blade surface corrosion). This means that although the data is not explicitly marked as blade surface corrosion fault data, judging from the voiceprint characteristics, it is very likely to be related to the fault. Finally, the server jointly determines the extended voiceprint training data and the fault condition voiceprint training data as the fourth voiceprint training data associated with the target fault type (blade surface corrosion). The fourth voiceprint training data composed of these data more comprehensively covers the voiceprint features related to blade surface corrosion failures, providing a rich and representative data basis for subsequent analysis based on blade fault detection models and training of edge computing optimization architectures, which helps to improve the accuracy and reliability of blade surface corrosion fault detection.
[0060] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the above-mentioned wind turbine blade abnormality detection method based on voiceprint collection. Figure 2 As shown, Figure 2The computer device 100 is a block diagram of a structure of a computer device 100 provided in an embodiment of the present invention. The computer device 100 includes a memory 111, a processor 112 and a communication unit 113. To achieve data transmission or interaction, the memory 111, the processor 112 and the communication unit 113 are directly or indirectly electrically connected to each other.
[0061] For illustrative purposes, the foregoing description is made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise form disclosed. Numerous modifications and variations are possible in accordance with the above teachings. These embodiments are selected and described in order to best illustrate the principles of the present disclosure and its practical application, so that those skilled in the art can best utilize the present disclosure and utilize various embodiments with different modifications to suit the intended specific application.
Claims
1. A method for detecting abnormality of wind turbine blades based on voiceprint collection, characterized in that: include: Encoding the voiceprint feature of the blade voiceprint data of the wind turbine to obtain the voiceprint embedding feature corresponding to the blade voiceprint data; Extracting abnormal features from the voiceprint embedded features to obtain abnormal feature vectors corresponding to the voiceprint embedded features; Determine at least two fault types associated with the blade soundprint data, perform abnormal state identification on the abnormal feature vector according to each fault type, and obtain an abnormal state score corresponding to each fault type; Performing fault weight analysis on the voiceprint embedded features to obtain a fault type weight corresponding to each fault type; Based on the fault type weight corresponding to each fault type, the abnormal state score weight corresponding to each fault type is allocated to obtain the target abnormal state detection result corresponding to each fault type; the weighted frequency domain response value corresponding to a fault type is used to allocate the abnormal state score weight corresponding to the corresponding fault type; The blade abnormality detection result of the wind turbine is determined according to the target abnormal state detection result corresponding to each fault type.
2. The method according to claim 1, characterized in that The step of encoding the voiceprint feature of the blade voiceprint data of the wind turbine to obtain the voiceprint embedding feature corresponding to the blade voiceprint data includes: When the blade soundprint data of the wind turbine is obtained, a blade fault detection model for detecting blade abnormalities of the wind turbine is obtained; the blade fault detection model includes a first acoustic feature extraction module; Based on the first acoustic feature extraction module, an acoustic signal characteristic associated with the blade voiceprint data is determined, and a framed voiceprint feature of the blade voiceprint data is extracted according to the acoustic signal characteristic; the first acoustic feature extraction module includes a voiceprint preprocessing layer; the framed voiceprint feature includes a plurality of framed voiceprint features, and the plurality of framed voiceprint features include a target framed voiceprint feature; According to the acoustic signal characteristics to which the target framed voiceprint feature belongs, determining the time-frequency analysis domain corresponding to the target framed voiceprint feature; the time-frequency analysis domain includes a plurality of frequency domain response values, each frequency domain response value having a corresponding framed feature range; Determine the frame feature range to which the target frame voiceprint feature belongs based on the voiceprint preprocessing layer, and determine the frequency domain response value corresponding to the frame feature range to which the target frame voiceprint feature belongs as the time-frequency feature graph corresponding to the target frame voiceprint feature; Based on the first acoustic feature extraction module, reconstructing the time-frequency feature graph corresponding to the framed voiceprint feature to obtain the time-frequency feature data of the blade voiceprint data; The time-frequency feature data is encoded with voiceprint features to obtain voiceprint embedding features of the time-frequency feature data.
3. The method according to claim 1, characterized in that The abnormal feature vector is determined by extracting abnormal features from the voiceprint embedding feature by an acoustic feature encoder in a first acoustic feature extraction module; the first acoustic feature extraction module belongs to a blade fault detection model for detecting blade abnormalities of the wind turbine; The method further comprises: Acquire a first voiceprint sample set; the first voiceprint training data included in the first voiceprint sample set is voiceprint training data that is not configured with an abnormal type label; Acquire a plurality of self-supervised learning tasks for the first original acoustic feature extraction module; Based on the first original acoustic feature extraction module, data processing is performed on the first voiceprint training data to obtain an acoustic feature representation associated with each self-supervised learning task, and a plurality of error parameters corresponding to the plurality of self-supervised learning tasks are determined according to the acoustic feature representation; one self-supervised learning task corresponds to one error parameter; A first target error parameter is determined according to the multiple error parameters, a model parameter of the first original acoustic feature extraction module is optimized according to the first target error parameter, and the optimized first original acoustic feature extraction module is determined as the first acoustic feature extraction module.
4. The method according to claim 3, characterized in that The multiple self-supervised learning tasks include a time-frequency map mask reconstruction task; The first voiceprint training data is processed based on the first original acoustic feature extraction module to obtain an acoustic feature representation associated with each self-supervised learning task, and a plurality of error parameters corresponding to the plurality of self-supervised learning tasks are determined according to the acoustic feature representation, including: Based on the first original acoustic feature extraction module, the first voiceprint training data is subjected to time-frequency transformation processing to obtain first sample time-frequency feature data of the first voiceprint training data; the first sample time-frequency feature data includes sample time-frequency feature graphs at multiple time frame positions; Selecting a first time frame position to be masked for a time-frequency region from a plurality of time frame positions corresponding to the first sample time-frequency feature data according to a preset rule; Performing time-frequency region masking processing on the sample time-frequency feature graph at the first time frame position in the first sample time-frequency feature data according to the masking region identifier to obtain masked time-frequency sample data; Based on the first original acoustic feature extraction module, the masked time-frequency sample data is encoded with voiceprint features to obtain a first training voiceprint embedding feature of the masked time-frequency sample data, and an abnormal feature is extracted from the first training voiceprint embedding feature to obtain a first acoustic feature representation corresponding to the first training voiceprint embedding feature; the first acoustic feature representation belongs to the acoustic feature representation; Based on a feature reconstruction module associated with the time-frequency map mask reconstruction task, reconstructing the time-frequency map mask processing on the first acoustic feature representation to obtain a predicted time-frequency feature map corresponding to the masked area identifier in the masked time-frequency sample data; According to the sample time-frequency feature graph at the first time frame position in the first sample time-frequency feature data and the predicted time-frequency feature graph, an error parameter corresponding to the time-frequency graph mask reconstruction task is determined.
5. The method according to claim 3, characterized in that: The multiple self-supervised learning tasks include a time-frequency anomaly replacement detection task; The first voiceprint training data is processed based on the first original acoustic feature extraction module to obtain an acoustic feature representation associated with each self-supervised learning task, and a plurality of error parameters corresponding to the plurality of self-supervised learning tasks are determined according to the acoustic feature representation, including: Based on the first original acoustic feature extraction module, the first voiceprint training data is subjected to time-frequency transformation processing to obtain first sample time-frequency feature data of the first voiceprint training data; the first sample time-frequency feature data includes sample time-frequency feature graphs at multiple time frame positions; Selecting a second time frame position to be subjected to time-frequency feature replacement according to a preset rule from a plurality of time frame positions corresponding to the first sample time-frequency feature data; Performing time-frequency feature replacement processing on the sample time-frequency feature graph at the second time frame position in the first sample time-frequency feature data according to preset time-frequency information to obtain replaced time-frequency sample data; the preset time-frequency information is different from the sample time-frequency feature graph at the second time frame position; the preset time-frequency information is an abnormal time-frequency feature fragment extracted from historical fault data, or an unnatural time-frequency feature generated by Gaussian noise; Based on the first original acoustic feature extraction module, the replaced time-frequency sample data is encoded with voiceprint features to obtain a second training voiceprint embedding feature of the replaced time-frequency sample data, and the second training voiceprint embedding feature is extracted with abnormal features to obtain a second acoustic feature representation corresponding to the second training voiceprint embedding feature; the second acoustic feature representation belongs to the acoustic feature representation; Based on a feature reconstruction module associated with the time-frequency anomaly replacement detection task, determining a target detection time frame position for the second acoustic feature representation, and obtaining time-frequency feature replacement prediction results corresponding to the multiple time frame positions; According to the time-frequency feature replacement prediction results corresponding to the second time frame position and the multiple time frame positions, the error parameter corresponding to the time-frequency anomaly replacement detection task is determined.
6. The method according to claim 3, characterized in that The multiple self-supervised learning tasks include a voiceprint denoising comparison task; The first voiceprint training data is processed based on the first original acoustic feature extraction module to obtain an acoustic feature representation associated with each self-supervised learning task, and a plurality of error parameters corresponding to the plurality of self-supervised learning tasks are determined according to the acoustic feature representation, including: Injecting Gaussian white noise features into the framed voiceprint features of the first voiceprint training data according to the signal-to-noise ratio to obtain noisy voiceprint training data, and performing time-frequency transformation processing on the noisy voiceprint training data to obtain noise time-frequency feature data of the noisy voiceprint training data; Performing time-frequency transformation processing on the first voiceprint training data to obtain first sample time-frequency feature data of the first voiceprint training data; Based on the first original acoustic feature extraction module, the noise time-frequency feature data is encoded with voiceprint features to obtain a third training voiceprint embedding feature of the noise time-frequency feature data, and the third training voiceprint embedding feature is extracted with abnormal features to obtain a third acoustic feature representation corresponding to the third training voiceprint embedding feature; the third acoustic feature representation belongs to the acoustic feature representation; Based on the first original acoustic feature extraction module, the first sample time-frequency feature data is encoded with voiceprint features to obtain a fourth training voiceprint embedding feature of the first sample time-frequency feature data, and an abnormal feature is extracted from the fourth training voiceprint embedding feature to obtain a fourth acoustic feature representation corresponding to the fourth training voiceprint embedding feature; the fourth acoustic feature representation belongs to the acoustic feature representation; Determining an error parameter corresponding to the voiceprint denoising comparison task according to the third acoustic feature representation and the fourth acoustic feature representation; The multiple self-supervised learning tasks also include an expert knowledge transfer task; The first voiceprint training data is processed based on the first original acoustic feature extraction module to obtain an acoustic feature representation associated with each self-supervised learning task, and a plurality of error parameters corresponding to the plurality of self-supervised learning tasks are determined according to the acoustic feature representation, further comprising: According to the feature reconstruction module associated with the expert knowledge transfer task, the fourth acoustic feature representation is rated for abnormal risk to obtain an abnormal risk score corresponding to the first voiceprint training data; Acquire a pre-trained benchmark model associated with the first original acoustic feature extraction module, perform abnormal risk rating on the first voiceprint training data according to the pre-trained benchmark model, and obtain a benchmark abnormal state score corresponding to the first voiceprint training data; An error parameter corresponding to the expert knowledge transfer task is determined according to the abnormal risk score and the benchmark abnormal state score.
7. The method according to claim 1, characterized in that The abnormal feature vector is determined by extracting abnormal features from the voiceprint embedded features by an acoustic feature encoder in a first acoustic feature extraction module; the first acoustic feature extraction module belongs to a blade fault detection model for detecting blade abnormalities of the wind turbine; the blade fault detection model also includes a second acoustic feature extraction module after multi-level feature fusion of the acoustic feature encoder; The performing abnormal state identification on the abnormal feature vector according to each fault type to obtain the abnormal state score corresponding to each fault type includes: Determining a second acoustic feature extraction module from the blade fault detection model; the second acoustic feature extraction module includes a fault diagnosis subnet corresponding to each fault type; Based on the fault diagnosis subnet corresponding to each fault type, the abnormal state of the abnormal feature vector is identified to obtain the abnormal state score corresponding to each fault type; one fault diagnosis subnet is used to obtain the abnormal state score corresponding to one fault type; The method further comprises: Acquire a second voiceprint sample set; the second voiceprint training data included in the second voiceprint sample set is associated with a fault type label; the fault type label includes at least two fault type codes corresponding to the at least two fault types, one fault type corresponds to one fault type code, and the at least two fault type codes are determined according to the fault type to which the second voiceprint training data belongs; Acquire a pre-trained first acoustic feature extraction module, perform voiceprint feature encoding on the second voiceprint training data based on the first acoustic feature extraction module, obtain a fifth training voiceprint embedding feature of the second voiceprint training data, perform abnormal feature extraction on the fifth training voiceprint embedding feature, and obtain a fifth acoustic feature representation corresponding to the fifth training voiceprint embedding feature; Obtain at least two original fault diagnosis subnets corresponding to the at least two fault types, and perform fault type discrimination on the fifth acoustic feature representations according to the at least two, respectively, to obtain at least two fault type probabilities corresponding to the at least two original fault diagnosis subnets; one fault type corresponds to one original fault diagnosis subnet, and one original fault diagnosis subnet is used to obtain one fault type probability; A second target error parameter is determined according to the at least two fault type probabilities and the at least two fault type codes, model parameters of the at least two original fault diagnosis subnets are optimized according to the second target error parameter, the optimized at least two original fault diagnosis subnets are determined as the at least two fault diagnosis subnets, and the second acoustic feature extraction module is determined according to the at least two fault diagnosis subnets.
8. The method according to claim 1, characterized in that: The voiceprint embedding feature is determined by the acoustic feature encoding layer in the first acoustic feature extraction module performing voiceprint feature encoding on the time-frequency feature data; the time-frequency feature data is determined by the voiceprint preprocessing layer in the first acoustic feature extraction module performing time-frequency transformation processing on the blade voiceprint data; the first acoustic feature extraction module belongs to a blade fault detection model for detecting blade anomalies of the wind turbine; the blade fault detection model also includes a third acoustic feature extraction module in which multi-level features are fused after the acoustic feature encoding layer; The performing fault weight analysis on the voiceprint embedded feature to obtain the fault type weight corresponding to each fault type includes: Determining a third acoustic feature extraction module from the blade fault detection model; the third acoustic feature extraction module includes a multi-branch feature extractor and a dynamic weight allocator; Performing multimodal acoustic feature extraction on the voiceprint embedding feature based on the multi-branch feature extractor to obtain a multimodal feature mapping result corresponding to the voiceprint embedding feature; Based on the dynamic weight allocator, fault weight calculation is performed on the multimodal feature mapping result to obtain the fault type weight corresponding to each fault type; The blade fault detection model also includes a second acoustic feature extraction module for obtaining an abnormal state score corresponding to each fault type; The method further comprises: Acquire a third voiceprint sample set; the third voiceprint training data included in the third voiceprint sample set is associated with an abnormal type label; Acquire a pre-trained first acoustic feature extraction module, perform voiceprint feature encoding on the third voiceprint training data based on the first acoustic feature extraction module, obtain a sixth training voiceprint embedding feature of the third voiceprint training data, perform abnormal feature extraction on the sixth training voiceprint embedding feature, and obtain a sixth acoustic feature representation corresponding to the sixth training voiceprint embedding feature; Obtain a pre-trained second acoustic feature extraction module, perform abnormal state recognition on the sixth acoustic feature representation based on at least two fault diagnosis subnets in the second acoustic feature extraction module, and obtain a training abnormal state score corresponding to each fault type; one fault diagnosis subnet is used to obtain a training abnormal state score corresponding to one fault type; Performing a fault weight analysis on the sixth training voiceprint embedding feature based on the third original acoustic feature extraction module to obtain a training fault type weight corresponding to each fault type; Determine, according to the training abnormal state score corresponding to each fault type and the training fault type weight corresponding to each fault type, a training blade abnormality detection result of the reference wind turbine associated with the third voiceprint training data; A third target error parameter is determined according to the abnormality type label and the training blade abnormality detection result, and the model parameters of the third original acoustic feature extraction module are optimized according to the third target error parameter. The optimized third original acoustic feature extraction module is determined as the third acoustic feature extraction module, and the blade fault detection model is determined according to the first acoustic feature extraction module, the second acoustic feature extraction module and the third acoustic feature extraction module.
9. The method according to claim 1, characterized in that: The at least two fault types include a target fault type; The method further comprises: Obtaining a blade fault detection model for detecting blade anomalies of a wind turbine, and obtaining an edge computing optimization architecture for expert knowledge migration from the blade fault detection model; Acquire fault condition voiceprint training data under the target fault type; Determine the pending voiceprint training data from the voiceprint training data set used for voiceprint feature screening; Based on the voiceprint sample expansion module, an acoustic fingerprint comparison is performed on the pending voiceprint training data and the fault condition voiceprint training data to obtain a matching degree between the fault condition voiceprint training data and the pending voiceprint training data; If the matching degree reaches the data enhancement matching degree threshold, the pending voiceprint training data is determined as the extended voiceprint training data associated with the target fault type; Determine the extended voiceprint training data and the fault condition voiceprint training data as fourth voiceprint training data associated with the target fault type; the fourth voiceprint sample set includes the fourth voiceprint training data; the fourth voiceprint training data includes the fault condition voiceprint training data generated under the target fault type; Performing data processing on the fourth voiceprint training data based on the blade fault detection model to obtain a blade abnormality detection result corresponding to the fourth voiceprint training data, and determining the blade abnormality detection result corresponding to the fourth voiceprint training data as a simulation diagnosis label for training the edge computing optimization architecture; Performing data processing on the fourth voiceprint training data based on the edge computing optimization architecture to obtain an edge blade anomaly detection result corresponding to the fourth voiceprint training data; The architectural parameters of the edge computing optimization architecture are optimized according to the edge blade anomaly detection result and the simulation diagnostic label, and the optimized edge computing optimization architecture is determined as the lightweight diagnostic network of the blade fault detection model; the lightweight diagnostic network is used to perform blade anomaly detection on the blade soundprint data under the target fault type.
10. A server system, characterized in that: The method comprises a server, wherein the server is used to execute the method described in any one of claims 1 to 9.
Citation Information
Patent Citations
Wind generating set blade sound monitoring system and method
CN117365872A
Fan multi-mode fault diagnosis method based on voiceprint feature and vibration data fusion
CN119393366A
500kV GIS multi-mode working condition abnormity comprehensive monitoring system and method based on voiceprint recognition
CN119479690A
Random voiceprint certification system, random voiceprint cipher lock and creating method therefor
US20100017209A1
Method and device for identifying near-bit lithology based on intelligent voiceprint identification
US20240344453A1
Cited By
Hydropower station unit equipment state detection method, device, equipment and storage medium
CN121600952A