Visual audio multi-modal monitoring system and method for identifying defects in wind turbine blades

By combining a visual and audio multimodal monitoring system with panoramic video and blade tip microphones, the accuracy and false alarm problems of wind turbine blade defect detection in low visibility conditions have been solved, achieving stable and efficient blade defect monitoring in all weather conditions.

CN121049272BActive Publication Date: 2026-02-03CHONGQING DUCHEN IND TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511562807.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-02-03
Estimated Expiration
2045-10-30

AI Technical Summary

Technical Problem

Existing wind turbine blade defect detection systems have low detection accuracy and high false alarm rate in low visibility conditions, and are also costly and complex to maintain.

Method used

A visual and audio multimodal monitoring system is adopted, which combines a panoramic video monitoring mechanism and a blade tip microphone. The system makes intelligent decisions under different visibility conditions through image and audio recognition algorithms. The panoramic video monitoring mechanism acquires images of the back of the blades and the blade tip microphone collects audio, thereby realizing environmental perception and intelligent decision-making and reducing the risk of false alarms.

Benefits of technology

It improves the accuracy and reliability of blade defect detection under various weather conditions, reduces equipment costs and maintenance complexity, and achieves stable defect monitoring around the clock.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121049272B_ABST
    Figure CN121049272B_ABST
Patent Text Reader

Abstract

The application discloses a kind of visual audio multimodal monitoring system and method for identifying fan blade defects, panoramic video monitoring mechanism can obtain the back panoramic image of the blade of current wind driven generator and the image of the surrounding environment of wind driven generator;Root pickup is set on panoramic video monitoring mechanism and tip pickup is set on the tower of wind driven generator, which can obtain the audio of the tip section of wind driven generator and the audio of the root section.First, the environmental visibility under the current weather condition is identified by panoramic video monitoring mechanism, and then different real-time monitoring strategies are executed according to the environmental visibility, which not only makes the blade defect monitoring no longer restricted by weather environment, improves the application scenario, but also greatly reduces the risk of false alarm through visual and audio mutual verification, greatly improves the accuracy of blade surface defect identification.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image communication and sound processing, and particularly relates to a visual and audio multi-modal monitoring system and method for identifying defects of fan blades. BACKGROUND

[0002] Fan blades are one of the most important components of wind turbines. During the daily operation of wind turbines, damage may occur due to bird strikes and natural disasters, etc. Therefore, in order to ensure the safe operation of wind turbines, it is necessary to detect the damage of fan blades as soon as possible and to facilitate timely maintenance. More and more wind turbines begin to install cameras on the nacelle to monitor the fan blades.

[0003] At present, multiple cameras are usually used to segmentally and jointly monitor the fan blades. When the number of cameras is small, this segmental and joint monitoring system cannot achieve high-definition shooting of the discovered damage points, which often leads to inaccurate judgment of the damage degree by the backend. If the shooting resolution is to be improved, a large number of cameras are needed for segmental and joint monitoring, which not only has high cost, but also requires high processing and storage capabilities of the backend server due to the large amount of data generated by the all-day uninterrupted monitoring. Moreover, due to the large number of cameras, the existing segmental and joint monitoring system is inefficient to install, and the risk of damage to the cameras by natural disasters is higher due to the large installation area of multiple cameras. Once one camera stops working, the entire monitoring system will be paralyzed.

[0004] Therefore, the applicant of the present application has designed a fan blade panoramic monitoring system, an installation structure and a panoramic monitoring and shooting method (see Chinese patent for invention with publication number CN117596365A), which integrates sensors, wide-angle cameras, long-focus micro cameras and blade monitoring cameras into one, with extremely high integration. This not only simplifies the installation and improves the installation efficiency, but also makes the overall structure extremely compact due to the layered installation method, greatly reducing the risk of being affected by natural disasters and improving the stability of operation. First, the sensor is used to detect whether the blade has reached the detection position. If the blade has reached the detection position, the wide-angle camera is used to perform overall visual detection on the back of the blade. If a damage point is found, the long-focus micro camera is started to perform local high-definition shooting of the damage point. Through one sensor and two cameras, high-resolution monitoring of the back of the blade can be achieved, and the damage degree can be accurately identified. This not only greatly reduces the cost, but also reduces the amount of data by not requiring the two cameras to monitor all day long, thereby greatly reducing the requirements for the processing and storage capabilities of the backend server.

[0005] However, the applicant of this invention found in practical application that it is very limited by weather conditions. It can only accurately identify defects on the blade surface when visibility is good. In weather conditions such as wind, sand, rain, snow, lightning, and fog, it is not only impossible to accurately identify defects on the blade surface, but false alarms often occur. Even under clear night conditions, the accuracy of identifying defects on the blade surface is not ideal.

[0006] Solving these problems is now a top priority. Summary of the Invention

[0007] To address the technical problem of low detection accuracy in existing wind turbine blade defect detection under low visibility conditions, this invention provides a visual and audio multimodal monitoring system and method for identifying wind turbine blade defects.

[0008] The technical solution is as follows:

[0009] The first aspect of this application relates to a visual-audio multimodal monitoring system for identifying defects in wind turbine blades, comprising a dual-purpose monitoring device installed on the nacelle of a wind turbine and multiple blade tip microphones evenly distributed at the same height on the tower of the wind turbine. The pickup direction of each blade tip microphone is horizontal and directed away from the center line of the tower, thereby forming a blade tip annular pickup surface in space. The dual-purpose monitoring device includes a device base fixedly installed on the nacelle and a panoramic video monitoring mechanism and blade root microphones installed on the device base. The panoramic video monitoring mechanism is used to acquire panoramic images of the back of the wind turbine blade from the blade root to the blade tip. The pickup direction of the blade root microphones is parallel to the orientation of the nacelle, thereby forming a blade root pickup line in the horizontal direction.

[0010] When all the blades of a wind turbine rotate synchronously, the tips of each blade pass through the blade tip annular pickup surface, and the roots of each blade pass through the blade root pickup line.

[0011] The above-mentioned visual and audio multimodal monitoring system for identifying wind turbine blade defects not only acquires panoramic images of the back of the wind turbine blades, enabling panoramic monitoring of the back of the blades, but also acquires images of the surrounding environment, thus obtaining visibility and weather information. Simultaneously, by installing blade root microphones on the panoramic video monitoring system and blade tip microphones on the wind turbine tower, the system can acquire audio from the blade tip and root segments. Furthermore, since the wind turbine blades rotate with the nacelle, the evenly distributed blade tip microphones on the tower ensure that the sound emitted by the blades is accurately collected by the corresponding blade tip microphones regardless of their orientation.

[0012] The second aspect of this application relates to a visual-audio multimodal monitoring method performed by the aforementioned visual-audio multimodal monitoring system, comprising the following steps:

[0013] S1. The panoramic video monitoring agency identifies whether the environmental visibility under the current weather conditions is higher than the set value: if yes, proceed to step S2; if no, proceed to step S3.

[0014] S2. Perform real-time monitoring of high visibility, following these steps:

[0015] S21. The panoramic video monitoring agency collects panoramic images of the back of the leaf from the leaf root to the leaf tip.

[0016] S22. Process the panoramic image using an image recognition algorithm to determine if there are defects in the blade: if yes, determine the type and location of the defect and proceed to step S23; if no, return to step S1.

[0017] S23. Determine whether the defect is located near the blade tip annular pickup surface or the blade root pickup line: If it is near the blade tip annular pickup surface, proceed to step S24; if it is near the blade root pickup line, proceed to step S25.

[0018] S24. Based on the nacelle orientation, determine the blade tip microphone closest to the blade. Process the audio collected by the blade tip microphone using a voiceprint recognition algorithm and determine whether there is a defect in the section of the blade near the blade tip: if yes, determine the type of defect and proceed to step S26; if no, record the log and wait for manual review, then return to step S1.

[0019] S25. Process the audio collected by the leaf root microphone using the voiceprint recognition algorithm to determine whether there is a defect in the section of the leaf near the leaf root: if yes, determine the type of defect and proceed to step S26; if no, record the log and wait for manual review, then return to step S1.

[0020] S26. Determine whether the defect type identified by audio is consistent with the defect type identified by panoramic image in step S22: If yes, send an alarm message containing the defect type and location to the maintenance center, and then return to step S1; if no, record the log and wait for manual review, and then return to step S1.

[0021] S3. Perform real-time monitoring of low visibility, following these steps:

[0022] S31. Determine the blade tip pickup closest to the blade based on the nacelle orientation, and collect the blade tip audio through the blade tip pickup, while simultaneously collecting the blade root audio through the blade root pickup.

[0023] S32. The visual and audio multimodal monitoring system identifies the current weather type: if it is sandstorm or rain / snow, proceed to step S33; if it is thunderstorm, proceed to step S34; if it is other weather, proceed to step S3.

[0024] S33. Process the leaf tip segment audio and leaf root segment audio using spectral subtraction and adaptive filtering to suppress wind noise and / or rain noise in the leaf tip segment audio and leaf root segment audio, and then proceed to step S35.

[0025] S34. Use the elimination algorithm to process the leaf tip segment audio and leaf root segment audio, and after eliminating the thunder sound, proceed to step S35.

[0026] S35. Use the voiceprint recognition algorithm to process the audio of the leaf tip segment and the audio of the leaf root segment to determine whether there is a defect in the leaf: If yes, after determining the type and location of the defect, send a risk warning message containing the defect type and location to the maintenance center, mark the leaf as high risk and require re-inspection, and then proceed to step S36; if no, return to step S1.

[0027] S36. The panoramic video monitoring agency identifies whether the environmental visibility under the current weather conditions is higher than the set value: if yes, proceed to step S2; if no, repeat step S36.

[0028] The above-mentioned visual and audio multimodal monitoring method for identifying wind turbine blade defects first identifies the environmental visibility under the current weather conditions through a panoramic video monitoring system. Then, different real-time monitoring strategies are executed according to the environmental visibility. In high visibility environments, visual monitoring is the primary method, with audio monitoring used for auxiliary verification. In low visibility environments, audio monitoring is the primary method, with visual monitoring used for verification after the weather improves. This not only removes the limitations of weather conditions on blade defect monitoring, expanding the application scenarios, but also greatly reduces the risk of false alarms by using visual and audio verification methods in both high and low visibility environments, significantly improving the accuracy of blade surface defect identification. Therefore, this invention realizes a complete engineering solution for environmental perception, intelligent decision-making, and action closed loop, which can significantly improve the all-weather working capability and engineering practical value of wind turbine blade defect detection. Attached Figure Description

[0029] Figure 1 A schematic diagram of a visual and audio multimodal monitoring system installed on a wind turbine.

[0030] Figure 2 This is a schematic diagram of the dual-purpose monitoring equipment;

[0031] Figure 3 This is a cross-sectional view of the dual-purpose monitoring equipment;

[0032] Figure 4 forFigure 3 Enlarged view of point C in the middle;

[0033] Figure 5 for Figure 3 Enlarged view of point D in the middle;

[0034] Figure 6 This is a schematic diagram of the internal structure of the segmented monitoring turntable located at the top.

[0035] Figure 7 This is a schematic diagram of the internal structure of the electrical installation compartment;

[0036] Figure 8 A schematic diagram of the internal structure of the equipment mounting slot;

[0037] Figure 9 This is a schematic diagram of the CBAM module.

[0038] Figure 10 This is a schematic diagram of the encoder and decoder. Detailed Implementation

[0039] The present invention will be further described below with reference to the embodiments and accompanying drawings.

[0040] Example 1:

[0041] like Figure 1 As shown, a visual and audio multimodal monitoring system for identifying defects in wind turbine blades mainly includes a dual-purpose monitoring device 2 and multiple blade tip microphones 3.

[0042] The dual-purpose monitoring device 2 is preferably installed on the top of the nacelle 1a of the wind turbine generator 1, which ensures the reliable installation of the dual-purpose monitoring device 2.

[0043] Each blade tip microphone 3 is mounted at the same height on the tower 1b of the wind turbine 1, and each blade tip microphone 3 is evenly distributed along the circumference of the tower 1b. The pickup direction of each blade tip microphone 3 is horizontal and directed away from the center line of the tower 1b, thus forming a blade tip annular pickup surface A in space. When each blade 1c of the wind turbine 1 rotates synchronously, the blade tip 1c1 of each blade 1c passes through the blade tip annular pickup surface A. Therefore, no matter which direction the blade 1c of the wind turbine 1 rotates with the nacelle 1a, the sound emitted near the blade tip 1c1 of the blade 1c can be accurately collected by the corresponding blade tip microphone 3.

[0044] Please see Figures 1-3The dual-purpose monitoring device 2 includes an equipment base 2a fixedly mounted on the nacelle 1a, and a panoramic video monitoring mechanism and a blade root microphone 2e both mounted on the equipment base 2a. The equipment base 2a is generally reliably fixed to the top of the nacelle 1a with multiple bolts. The panoramic video monitoring mechanism is used to acquire panoramic images of the wind turbine blades 1c from the blade root 1c2 to the blade tip 1c1. The pickup direction of the blade root microphone 2e is parallel to the orientation of the nacelle 1a, thus forming a blade root pickup line B in the horizontal direction. Since the dual-purpose monitoring device 2 rotates synchronously with the nacelle 1a, when each blade 1c of the wind turbine 1 rotates synchronously, the blade root 1c2 of each blade 1c passes through the blade root pickup line B. Therefore, the sounds emitted near the blade root 1c2 of the blade 1c can be accurately collected by the blade root microphone 2e.

[0045] Please see Figures 2-8 The equipment base 2a has an open upper equipment mounting slot 2a1. The center of the bottom of the equipment mounting slot 2a1 has an upwardly extending central shaft 2a2. The blade root pickup 2e is mounted on the outer wall of the equipment mounting slot 2a1 near the blade 1c, ensuring the reliable installation of the blade root pickup 2e.

[0046] Multiple lower electromagnets 2a4 are evenly distributed circumferentially along the central axis 2a2 at the bottom of the equipment mounting slot 2a1. Correspondingly, upper electromagnets 2b2 are installed at the bottom of the panoramic video monitoring mechanism, one above each of the lower electromagnets 2a4. When energized, each upper electromagnet 2b2 has the opposite magnetic pole to the corresponding lower electromagnet 2a4, thereby maintaining a gap between the panoramic video monitoring mechanism and the equipment mounting slot 2a1. This allows the panoramic video monitoring mechanism to be suspended under the action of magnetic force. When the cabin 1a vibrates, the panoramic video monitoring mechanism can maintain a relatively stable state, thus effectively improving the shooting quality of the panoramic video monitoring mechanism.

[0047] Further, please see Figure 3 and Figure 8 The bottom of the equipment mounting slot 2a1 is provided with multiple lower shielding sleeves 2a3 that are fitted one-to-one with each lower electromagnet 2a4. The bottom of the panoramic video monitoring mechanism is provided with multiple upper shielding sleeves 2b1 that are fitted one-to-one with each upper electromagnet 2b2. Both the upper shielding sleeves 2b1 and the lower shielding sleeves 2a3 are made of non-magnetic materials, which can reduce the influence of the magnetic field on other electronic components.

[0048] Further, please see Figure 5 Each upper shielding sleeve 2b1 has an enlarged lower end to form a lifting guide section 2b11 that fits outside the corresponding lower shielding sleeve 2a3, thereby achieving the function of lifting and guiding, and thus effectively improving the lifting stability of the panoramic video monitoring mechanism.

[0049] Please see Figures 2-8 The panoramic video monitoring mechanism includes an electrical installation chamber 2b, at least one segmented monitoring turntable 2c, and a multi-functional detection turntable 2d, which are sequentially mounted on the central axis 2a2 from bottom to top. In this embodiment, the bottom of the electrical installation chamber 2b is equipped with an upper electromagnet 2b2 that is correspondingly arranged above each lower electromagnet 2a4. At the same time, the bottom of the electrical installation chamber 2b is provided with a plurality of upper shielding sleeves 2b1 that are correspondingly mounted outside each upper electromagnet 2b2.

[0050] The multi-functional detection turntable 2d and the segmented monitoring turntable 2c can rotate along the central axis 2a2 under the control of the corresponding rotation control component 2f, thereby flexibly adjusting the shooting angle.

[0051] Specifically, the portion of the central shaft 2a2 located in the multi-functional detection turntable 2d and each segmented monitoring turntable 2c is integrally formed with a ring-shaped boss 2a22. Each rotation control component 2f includes a motor 2f1 fixedly installed in the corresponding multi-functional detection turntable 2d or segmented monitoring turntable 2c, and rollers 2f2 synchronously mounted on the motor shaft of the motor 2f1. The circumferential outer wall of each roller 2f2 frictionally engages with the end face of the corresponding ring-shaped boss 2a22. By driving the rollers 2f2 to rotate forward and backward using the motor 2f1, the rotation angle of the multi-functional detection turntable 2d and the segmented monitoring turntable 2c can be precisely controlled.

[0052] Furthermore, the roller 2f2 is made of high-friction rubber material, which can avoid slippage and further improve the control accuracy of the rotation angle of the multi-functional detection turntable 2d and the segmented monitoring turntable 2c.

[0053] Wide-angle cameras 2m and thermal imaging cameras 2h are installed on both the electrical installation bay 2b and the segmented monitoring turntable 2c. Since the blade 1c of the wind turbine 1 is relatively long, this embodiment achieves a larger shooting range by using wide-angle cameras 2m, thereby reducing the number of layers on the segmented monitoring turntable 2c and lowering equipment costs. In this embodiment, three wide-angle cameras 2m are preferably used. The three wide-angle cameras 2m respectively capture images of the front, middle, and rear sections of the blade 1c. After stitching the three images together, a panoramic image of the back of the blade 1c can be obtained. Simultaneously, by configuring thermal imaging cameras 2h, the corresponding positions of the blade 1c can be clearly observed under weather conditions or at night. In this embodiment, the three thermal imaging cameras 2h also respectively capture thermal images of the front, middle, and rear sections of the blade 1c.

[0054] In this embodiment, a CCD camera 2g and an infrared rangefinder 2i are mounted on the multi-functional detection turntable 2d. The CCD camera 2g captures images of the surrounding scene, and through analysis, it can identify visibility information and weather conditions. The infrared rangefinder 2i not only detects the arrival of the blade 1c at the detection position, but also, if it detects the blade 1c approaching the shooting position, activates three wide-angle cameras 2m to capture images of the back of the blade 1c, and can measure the distance to the observed target in real time.

[0055] Furthermore, the wide-angle camera 2m is mounted on the corresponding electrical mounting compartment 2b or segmented monitoring turntable 2c via a stabilizer bracket 2m1. Specifically, the stabilizer bracket 2m1 has a common mobile phone gimbal stabilizer structure, which can achieve excellent image stabilization for the wide-angle camera 2m.

[0056] Furthermore, ambient lighting 2o is also installed on both sides of the CCD camera 2g camera.

[0057] Since the wind turbine 1 is at high risk of being struck by lightning, seven surge protectors 2j are installed in the electrical installation compartment 2b of this embodiment to provide lightning protection for all cameras.

[0058] Furthermore, electrical installation compartment 2b also contains electrical components such as PCB circuit boards and switches. Meanwhile, the topmost segmented monitoring turntable 2c contains a battery 2n to power the entire dual-purpose monitoring equipment 2.

[0059] Meanwhile, in order to facilitate wiring, the central shaft 2a2 is a hollow shaft structure, and several wire-passing holes 2a21 are opened on the central shaft 2a2, which are respectively connected to the electrical installation compartment 2b, the multi-functional detection turntable 2d and each segment monitoring turntable 2c.

[0060] Please see Figure 2 The outer walls of the electrical installation compartment 2b, the multi-functional detection turntable 2d, and the segmented monitoring turntables 2c together form an egg-shaped structure, which makes the overall appearance streamlined, effectively reducing wind resistance and wind noise, and improving the stability of video and audio acquisition of the dual-purpose monitoring equipment 2.

[0061] Example 2:

[0062] A visual-audio multimodal monitoring method using a visual-audio multimodal monitoring system according to Embodiment 1 is performed according to the following steps:

[0063] S1. The panoramic video monitoring agency identifies whether the environmental visibility under the current weather conditions is higher than the set value: if yes, proceed to step S2; if no, proceed to step S3.

[0064] Specifically, the panoramic video monitoring agency uses a CCD camera to capture photos of the surrounding environment. By analyzing these photos, environmental visibility information can be obtained. For details on this method, please refer to Chinese Invention Patent Publication No. CN113408415A. The environmental visibility is then compared with a set value. If the environmental visibility is higher than the set value, the process proceeds to step S2; if the environmental visibility is lower than the set value, the process proceeds to step S3.

[0065] S2. Perform real-time monitoring of high visibility, following these steps:

[0066] S21. The infrared rangefinder 2i of the panoramic video monitoring agency detects that a blade 1c has arrived at the shooting position. Then, three wide-angle cameras 2m are activated to shoot the front, middle and rear sections of the back of the blade 1c respectively. Then, by stitching and combining, a panoramic image of the back of the blade 1c from the root to the tip is obtained.

[0067] S22. Process the panoramic image using an image recognition algorithm to determine whether there is a defect in blade 1c: if yes, determine the type and location of the defect and proceed to step S23; if no, return to step S1.

[0068] Specifically, step S22 is performed according to the following steps:

[0069] S221. Using a dataset of known blade defects, train a YOLOv12 model embedding a CBAM module using a machine learning algorithm until the YOLOv12 model embedding a CBAM module can accurately classify defects using feature values. Then, use this YOLOv12 model embedding a CBAM module as the target detection model. Specifically, follow these steps:

[0070] S2211. Data Acquisition. Following the method in step S21, acquire multiple sets of panoramic images of defective and normal leaves from different angles, distances, and environments. Typically, at least 20,000 panoramic images of defective and normal leaves should be acquired.

[0071] S2212, Data Annotation. Using a data annotation tool (e.g., Label-Studio), the feature-oriented bounding box information of all panoramic images of defective and normal blades from step S2211 is annotated to obtain an image dataset of known blade defect types. The annotation information includes: damage, cracks, fractures, detachment, and corrosion.

[0072] S2213. Data Partitioning. The known leaf image dataset is divided into training and validation sets. The preferred ratio of training to validation sets is 8:2 to ensure the trained model has good generalization ability and avoids overfitting.

[0073] S2214. Model Training. First, the training set is input into the YOLOv12 model embedded with the CBAM module for training, outputting labeled target images. Then, the labeled target images are evaluated using a validation set. If the evaluation fails, the next round of training begins; if it succeeds, training terminates, and the YOLOv12 model embedded with the CBAM module is used as the object detection model. The parameters used to evaluate the labeled target images using the validation set include: precision, recall, and mean precision.

[0074] S222. Perform cropping, scaling, noise reduction, normalization, and binarization on the panoramic image to obtain the preprocessed image. Specifically, follow these steps:

[0075] S2221. Crop the panoramic image. Remove irrelevant areas (sky, ground, etc.) from the panoramic image, focusing on the area containing the target blade to reduce computational redundancy. Typically, extend the boundary area by 50-100 pixels centered on the blade to ensure complete target coverage.

[0076] S2222. Scale the cropped panoramic image. Unify the panoramic image size to the model input size to meet network structure requirements. Typically, bilinear interpolation is used to scale the image, with an input size of 640×640, maintaining a consistent target aspect ratio and avoiding stretching distortion.

[0077] S2223. Denoise the scaled panoramic image. Use a Gaussian filtering algorithm to eliminate random noise in the image caused by imaging system or environmental factors, thereby improving image clarity and feature recognition.

[0078] S2224. Normalize the denoised panoramic image. Linear normalization is typically used to normalize the denoised panoramic image, mapping each pixel value in the panoramic image from the original range of [0, 255] to the interval [0, 1]. Image normalization is used to standardize the pixel values ​​of the input image to a uniform numerical range, improving the stability of model training and the consistency of inference.

[0079] S2225. Binarize the normalized panoramic image. By setting a fixed threshold, the image is converted into a black-and-white binary image for subsequent edge extraction and angle estimation. Binarization enhances the edge features of the target bolt in the image, and by extracting the edge contour of the target bolt, the displacement and rotation angle of the target bolt can be accurately estimated.

[0080] S223. Use the target detection model to extract the features of the blade in the preprocessed image and determine whether there is a defect in blade 1c: if yes, determine the type and location of the defect and proceed to step S23; if no, return to step S1.

[0081] Step S23 is performed in the following steps:

[0082] S2231. Overall network structure design.

[0083] The overall network structure consists of three parts: the backbone network, the neck network, and the head network.

[0084] The backbone network is responsible for feature extraction and consists of Conv convolutional layers, C3k2 modules, and A2C2f modules. To reduce network size and improve performance, the backbone network employs residual connections and a bottleneck structure. The C3k2 module extracts multi-scale features, and the A2C2f module further enhances the feature representation capability.

[0085] Neck Network: Located between the backbone network and the head network, the neck network includes a concat layer, an upsample layer, and an A2C2f module. It fuses and adjusts the features extracted by the backbone network, and integrates feature information from different levels through upsampling and concatenation operations to enhance feature expression.

[0086] Head Network: The head network is the decision-making part of the object detection model. It consists of multiple detection layers and is responsible for generating the final detection results, including bounding boxes, confidence scores, and class labels.

[0087] Through these designs and improvements, it performs excellently in feature extraction, fusion, and detection, and can effectively handle target detection tasks at different scales and in complex scenes.

[0088] S2232, Backbone Network Design.

[0089] The YOLOv12 model is introduced into the A2C2f module. The A2C2f module is an improved feature extraction module proposed in the YOLOv12 model, which combines area attention and residual connections, mainly to improve the efficiency and accuracy of feature extraction.

[0090] Integrating the CBAM module into the output layer of the backbone network enables channel and spatial attention adjustment of the extracted features, thereby making the model more focused on features related to wind turbine blade defects, suppressing irrelevant information, improving feature representation capabilities, and thus improving detection accuracy.

[0091] S2233, Path aggregation network design.

[0092] The path aggregation network design is mainly reflected in the neck network, used for feature fusion and enhancement. The YOLOv12 model's neck network adopts the FPN and PAN structures, but with optimizations and improvements to enhance model efficiency and performance. Upsampling is performed first, followed by downsampling, with two branches connected by two cross-layer fusion connections, enhancing the overall semantic and positional information extraction of the model. The main function of the FPN is to fuse feature maps of different scales to obtain multi-scale feature information. A top-down path is used to pass high-level semantic information to low-level feature maps, enhancing the semantic information of low-level features. Each level of the FPN adjusts the size of the high-level feature map to be the same as the low-level feature map through upsampling and convolution operations, and then the two are added together to obtain the fused feature map. PAN adds a bottom-up path to the FPN to further enhance feature aggregation and propagation. PAN passes low-level detailed information to high-level feature maps through a bottom-up path, enhancing the detailed information of high-level features. Each layer of the PAN performs convolution and downsampling operations to resize the lower-level feature maps to the same size as the higher-level feature maps, and then adds the two together to obtain the final fused feature map. By improving the FPN and PAN structures, the YOLOv12 model's ability to detect defects at different scales is enhanced.

[0093] S2234, Output of test results.

[0094] The detection results are mainly output in the Head section of the network. It is responsible for receiving features extracted by the backbone network and transforming these features into the final detection results, thus achieving the detection of main defects.

[0095] Please refer to Table 1. The YOLOv12 model in this embodiment consists of two parts: a backbone for feature extraction and a head module for target detection. The backbone includes 5 convolutional layers (Conv), 2 C3k2 modules, 2 A2C2f modules, and 1 CBAM attention module. The first convolutional layer divides the 3×640×640 image into 320×320×12 pixels, converts it into a 64×320×320 feature map through 64 3×3 convolutions, and then multiplies the output channel by a coefficient of 0.25 to convert it into a 16×320×320 feature map, achieving downsampling and feature extraction while preserving spatial information. Each dynamic convolutional layer enhances feature diversity by adaptively adjusting the convolution kernel parameters. The C3k2 module uses a 4-layer cascaded 1×1 and 3×3 convolutional structure, combined with skip connections to improve feature representation. The A2C2f module enhances the model's robustness to target deformation through pooling operations, while the CBAM module focuses on key region features through channel and spatial attention mechanisms. The head section includes two convolutional layers (Conv), one C3k2 module, three A2C2f modules, two upsampling layers (Upsample), four concatenation units (Concat), and one detection unit (Detect). Upsampling and concatenation fuse features at different levels; the detection unit outputs detection results at three scales (20×20, 40×40, and 80×80), achieving efficient end-to-end inference.

[0096] Table 1. Specific network parameters of the YOLOv12 model with embedded CBAM module.

[0097]

[0098] Please see Figure 9 The feature map output by the backbone network is ,in, Indicates the input feature map, Represents the set of real numbers. Indicates the number of channels. Indicates altitude, Indicates the width.

[0099] (1) Channel attention mechanism (CAM).

[0100] The channel attention module captures inter-channel dependencies using global average pooling and global max pooling. (Input feature map) After global average pooling (AvgPool) and max pooling (MaxPool), two channel description vectors are generated. and :

[0101] ;

[0102] ;

[0103] in , Represents the input feature map The Middle The first channel, the first line, number The feature value at the column position.

[0104] Channel relationships are extracted using a shared multilayer perceptron (MLP), and then summed and a sigmoid activation function is applied to generate channel attention weights. :

[0105] ;

[0106] in, ( ) represents the Sigmoid activation function. ( ) represents a multilayer perceptron function, which typically consists of two fully connected layers and a ReLU activation function. Its expression is:

[0107] ;

[0108] in, , Indicates the reduction ratio. ( ) represents the ReLU activation function.

[0109] Channel attention-weighted feature map Input feature map With channel attention Multiply by channel:

[0110] ;

[0111] in, This indicates element-wise multiplication.

[0112] (2) Spatial attention mechanism (SAM).

[0113] The spatial attention module uses the feature map after channel attention weighting. To capture spatial information. Average pooling and max pooling are performed along the channel dimension to obtain two spatial feature maps. and :

[0114] ;

[0115] ;

[0116] Spatial attention weights are generated through concatenation and convolution. :

[0117] ;

[0118] in, It is a 7×7 convolution.

[0119] (3) The final output of the CBAM module.

[0120] Input feature map Output after passing through the CBAM module:

[0121] ;

[0122] in, Represents the input feature map The output feature map after processing by the CBAM module.

[0123] In summary, the image data of the wind turbine blades to be inspected is preprocessed according to the steps described above. The processed images are then input into a trained network model to obtain feature vectors. Defect type prediction results are output through each task branch, thus achieving the detection of wind turbine blade defects. The detection results are then visualized, and a detection report is generated. When a serious defect is detected, an alarm is triggered on the backend management server, and relevant personnel are notified.

[0124] For example, the dataset is imported into a YOLOv12 model with an embedded CBAM module for image conversion and training to label blade defects. Blade defects typically include features such as protective film damage, rainproof ring detachment, delamination damage, resin-rich cracks at the trailing edge corner, structural cracks, lightning strike damage, drain blockage, leading edge corrosion, coating damage, and adhesive cracking. The training sample set in this paper is labeled with five categories: damage, crack, fracture, detachment, and corrosion.

[0125] After image annotation was completed, the training and testing datasets were split in an 8:2 ratio. The entire training process was conducted in PyCharm. During training, 16 images were used per iteration, for a total of 300 training iterations. The training image size was 640 pixels, the data loading threads were 8, the weight decay coefficient was set to 0.0005, and the momentum parameter was set to 0.9. Precision, recall, and mean precision (MAP) were used to evaluate the effectiveness of the wind turbine blade defect detection method.

[0126] Table 2 Comparison of detection performance between YOLOv12 and YOLOv11

[0127]

[0128] As shown in Table 2, the YOLOv12 model with embedded CBAM modules exhibits improved precision, recall, and mean precision compared to the YOLOv11 model with embedded CBAM modules. Experiments demonstrate that the YOLOv12 model with embedded CBAM modules effectively learns information between feature maps, enhancing its ability to detect multi-scale defects and improving the accuracy of defect type detection. This allows the defect detection model to more accurately capture the most critical wind turbine defect features, significantly improving the detection results. Specifically, the model training output selects samples with less prominent features and more complex backgrounds for the wind turbine blades, while the detection device output accurately identifies defect features. This image recognition algorithm can accurately identify wind turbine blade defects even in environments with strong feature variations and complex environments.

[0129] S23. Determine whether the defect is closer to the blade tip annular pickup surface A or the blade root pickup line B: If it is closer to the blade tip annular pickup surface A, proceed to step S24; if it is closer to the blade root pickup line B, proceed to step S25.

[0130] S24. Based on the orientation of the nacelle 1a, determine the blade tip microphone 3 that is closest to the blade 1c. Use the voiceprint recognition algorithm to process the audio collected by the blade tip microphone 3 and determine whether there is a defect in the section of the blade 1c near the blade tip: if yes, determine the type of defect and proceed to step S26; if no, record the log and wait for manual review, then return to step S1.

[0131] S25. Process the audio collected by the leaf root microphone 2e using the voiceprint recognition algorithm to determine whether there is a defect in the section of the leaf 1c near the leaf root: if yes, determine the type of defect and proceed to step S26; if no, record the log and wait for manual review, then return to step S1.

[0132] S26. Determine whether the defect type identified by audio is consistent with the defect type identified by panoramic image in step S22: If yes, send an alarm message containing the defect type and location to the maintenance center, and then return to step S1; if no, record the log and wait for manual review, and then return to step S1.

[0133] S3. Perform real-time monitoring of low visibility, following these steps:

[0134] S31. Determine the blade tip pickup 3 that is closest to the blade 1c based on the orientation of the nacelle 1a, and collect the blade tip audio through the blade tip pickup 3, while collecting the blade root audio through the blade root pickup 2e.

[0135] S32. The visual and audio multimodal monitoring system identifies the current weather type: if it is sandstorm or rain / snow, proceed to step S33; if it is thunderstorm, proceed to step S34; if it is other weather, proceed to step S3.

[0136] S33. Process the leaf tip segment audio and leaf root segment audio using spectral subtraction and adaptive filtering to suppress wind noise and / or rain noise in the leaf tip segment audio and leaf root segment audio, and then proceed to step S35.

[0137] S34. Use a rejection algorithm to process the audio of the leaf tip segment and the audio of the leaf root segment. After rejecting the thunder sound, proceed to step S35.

[0138] S35. Use the voiceprint recognition algorithm to process the audio of the leaf tip segment and the audio of the leaf root segment to determine whether there is a defect in the blade 1c: If yes, after determining the type and location of the defect, send a risk warning message containing the defect type and location to the maintenance center, and mark the blade 1c as high risk and await re-inspection, and then proceed to step S36; if no, return to step S1.

[0139] S36. The panoramic video monitoring agency identifies whether the environmental visibility under the current weather conditions is higher than the set value: if yes, proceed to step S2; if no, repeat step S36.

[0140] The above-mentioned audio processing using voiceprint recognition algorithms shall be performed according to the following steps:

[0141] a. Normal feature extraction.

[0142] First, acoustic features are extracted from normal audio files acquired from defect-free wind turbine blades (1c). Each audio signal is converted into a Mel spectrogram, a time-frequency feature representing the signal's frequency content on a Mel scale. Subsequently, the spectrogram is converted to a logarithmic scale to better represent amplitude variations. Simultaneously, to capture temporal context information, a sliding window method is used to concatenate multiple consecutive frames into a high-dimensional feature vector, thus forming the training dataset.

[0143] Specifically, the short-time Fourier transform technique is used to convert the time-domain waveform signal into a frequency-domain representation. The Fourier transform window length is set to 1024 sampling points, and the Mel-scale spectrogram is calculated using a sliding window method. The Mel-scale filter bank contains 64 frequency bands and, based on the characteristics of human auditory perception, can better capture key frequency components in the audio.

[0144] Given a speech signal x[n] with a sampling rate of sr, perform a short-time Fourier transform:

[0145] ;

[0146] in, Represents the two-dimensional DFT coefficients in the frequency domain. For FFT points, For frame shift, For windowing functions, For frequency index, For frame index, Represents the time-domain sample index. Represents the imaginary unit. ;

[0147] Then, the Mel frequency spectrum is obtained by projecting it onto the Mel frequency axis through the Mel filter bank. :

[0148] ;

[0149] in, Represents the weights of the Mel filter bank. Indicates the index of the Mel filter. .

[0150] The linear-scale Mel spectrum amplitude values ​​are converted to logarithmic-scale decibel values ​​and then normalized using the maximum amplitude reference value.

[0151] ;

[0152] in, For the first Frame, First The sound pressure level of a Mel filter, Indicates a reference value.

[0153] The conversion process enhances sensitivity to low-amplitude signals while compressing the dynamic range of high-amplitude signals, thereby improving the ability to characterize features.

[0154] A sliding window method was used to construct the temporal feature vector. The time window length was set to 5 consecutive frames, and the 64 Mel-band features within each window were concatenated to form a 320-dimensional feature vector (64 bands × 5 frames). This operation preserves the temporal context information of the audio signal, providing rich feature input for subsequent deep learning models.

[0155] Let the window size be Each time step is spliced ​​continuously frame:

[0156] ;

[0157] in, Indicates time eigenvectors.

[0158] That is to put Frames 3D features are concatenated into:

[0159] ;

[0160] After expanding into a vector:

[0161] ;

[0162] in, Representing the eigenvector Dimensions.

[0163] For all audio files After generating features, stack them:

[0164] ;

[0165] in, Represents the entire dataset. Indicates the first The audio file is in time eigenvectors.

[0166] Therefore, the total dimension is:

[0167] ;

[0168] in, Indicates the first The duration of each audio file.

[0169] b. Model training.

[0170] After feature extraction, a deep autoencoder model is constructed to learn feature representations of normal audio. This model employs methods such as... Figure 10 The illustrated symmetrical encoder and decoder structure works as follows: the encoder compresses input features into a low-dimensional latent representation through progressively decreasing fully connected layers, capturing the essential features of normal audio; the decoder then attempts to reconstruct the original input features from the latent representation. The model uses mean squared error as the loss function, and minimizes the reconstruction error through an optimization algorithm, enabling the model to learn to accurately reconstruct normal audio features. During the model training phase, normal audio data is used as input and the target output, and the network parameters are optimized through multiple iterations.

[0171] Specifically, as shown in Table 3, the encoder module consists of three fully connected layers cascaded sequentially. The first encoding layer receives a 320-dimensional feature vector as input, performs a linear transformation through 128 neurons, and then applies a ReLU activation function for non-linear mapping. The second encoding layer compresses the 128-dimensional features to a 64-dimensional space, also using the ReLU activation function to enhance feature representation. The third encoding layer, as the latent space representation layer, further compresses the features to 32 dimensions; the low-dimensional representation output by this layer contains the essential information of the input features. The decoder module has a symmetrical structure with the encoder module and contains three fully connected reconstruction layers. The first decoding layer expands the 32-dimensional latent features to 64 dimensions and reconstructs the features using the ReLU activation function. The second decoding layer restores the feature dimension to 128 dimensions, gradually recovering the original feature details. The final output layer uses a linear activation function to accurately reconstruct the 128-dimensional features into a 320-dimensional output vector, consistent with the input dimension.

[0172] Table 3. Self-encoder structural parameters

[0173]

[0174] During model training, the input features are compressed by the encoder during forward propagation and then reconstructed by the decoder. During backpropagation, the mean squared error loss function is used to calculate the difference between the reconstructed output and the original input. The network weight parameters are adjusted by the optimization algorithm so that the model can minimize the reconstruction error of normal samples.

[0175] c. Defect detection.

[0176] During the detection phase, the same features are extracted from new audio samples and input into the trained autoencoder to calculate its reconstruction error. Since the model is trained only on normal data, abnormal audio often produces a high reconstruction error. By setting a threshold, abnormal audio that deviates significantly from the normal pattern can be detected, thereby achieving audio defect detection in industrial equipment.

[0177] d. Preprocessing and segmentation of abnormal audio signals.

[0178] Abnormal audio detected by the defect detection model is collected and resampled to the sampling rate required by the Wav2Vec 2.0 pre-trained model. The continuous audio stream is divided into overlapping segments of fixed duration using a sliding window technique. The window size and sliding step are adaptively configured according to the characteristics of the abnormal duration: a 2.0-second window with a 1.0-second step is used for mixed abnormalities. Multi-channel audio is fused by mean fusion and converted into a mono signal, and then audio normalization is performed.

[0179] Specifically, a pre-trained Wav2Vec 2.0 model is used as the anomaly feature extractor to obtain the raw waveform data and metadata of the anomalous audio. For audio inputs with different sampling rates, resampling technology is used to uniformly convert them to the 16kHz sampling rate required by the pre-trained model, ensuring consistency in feature extraction. An anti-aliasing filter is used in the resampling process to avoid frequency aliasing.

[0180] ;

[0181] in, This represents the resampled audio signal. Represents the resampling function. The original waveform. The number of sampling points. This indicates the number of sampling points in the resampled audio signal. For the number of vocal tracts, Sampling rate, The target sampling rate.

[0182] For multi-channel data, a channel mean fusion algorithm is used.

[0183] ;

[0184] in, Indicates the first At the time point of the audio channel The sampled values, To output a mono signal.

[0185] A suitable sliding window segmentation strategy is adopted, with a window length of 2.0 seconds, a sliding step size of 1.0 second, and an overlap ratio of 50%. This parameter setting can balance temporal resolution and computational efficiency, and is suitable for most industrial equipment anomaly detection scenarios.

[0186] Define the window size and step size, and the number of samples:

[0187] ;

[0188] in, For window duration, The duration of the step. The number of sampling points in the window. This represents the number of sampling points for the step size.

[0189] Set up a segment validity verification mechanism: exclude end segments whose length is less than 50% of the window size and verify the minimum length; amplitude threshold detection: remove invalid segments that are silent or have too low energy (RMS value below -40dBFS) and detect the amplitude threshold; ensure that each segment contains complete sampling points and check the sampling integrity.

[0190] Establish metadata records for each audio segment: source file identification information, segment start timestamp and duration, segment position index in the original audio, audio energy level and signal-to-noise ratio estimate.

[0191] From mono Segment Extraction part:

[0192] ;

[0193] Effective scope:

[0194] ;

[0195] The shape of each segment: ;

[0196] in For the first Each segment (tensor, shape) ).

[0197] Amplitude normalization, DC component elimination, and format unification are performed on all generated audio segments to ensure that all segments have the same tensor dimension [batch_size, 32000].

[0198] For fragments Amplitude normalization:

[0199] ;

[0200] in, Indicates the first The maximum normalized amplitude of each audio segment Indicates the first The audio clip at the time point The signal amplitude at that location.

[0201] ;

[0202] in, Indicates the first The audio clip at the time point The normalized signal amplitude at the location.

[0203] DC to:

[0204] ;

[0205] in, Indicates the first The audio clip at the time point The signal after mean normalization at the specified location.

[0206] Combined and written as:

[0207] ;

[0208] in, Indicates the first The audio clip at the time point The original signal amplitude at that location.

[0209] After standardization ,in, For standardized fragments, the shape is .

[0210] e. Extraction of abnormal depth features.

[0211] The pre-trained Wav2Vec 2.0 model is used as the feature extractor. This model is pre-trained on large-scale audio data and has powerful audio representation capabilities. The processed audio segments are input into the model to extract high-dimensional deep feature representations. Temporal average pooling is performed on the feature sequences of each segment to generate fixed-dimensional segment-level embedding vectors.

[0212] Specifically, the Wav2Vec 2.0 model is trained using a self-supervised learning paradigm and possesses powerful audio representation capabilities. After loading the model, it is set to evaluation mode, and all weight parameters are frozen to ensure consistency in feature extraction. The standardized audio segments are then input into the pre-trained model to complete the following calculations:

[0213] ;

[0214] in, The acoustic model returns a feature map to the input. For this is the first Spectral features of an audio segment after feature extraction As auxiliary feature vectors, For frame length, For feature dimensions.

[0215] Perform time averaging (for frame length) (Calculate the mean) to obtain the fragment embedding :

[0216] ;

[0217] in, This is average pooling over the time dimension.

[0218] Then, the extracted feature vectors are normalized using the L2 norm to ensure that all feature vectors lie on the unit hypersphere, thereby improving the clustering effect.

[0219] Embed the unnormalized fragment as ,but:

[0220] ;

[0221] in:

[0222] , Indicates the first Embedding vectors of fragments In the Components in each dimension;

[0223] result satisfy .

[0224] f. Density cluster analysis.

[0225] The density-based DBSCAN clustering algorithm is used to perform unsupervised clustering of feature embedding vectors. The optimal combination of clustering parameters is found through a grid search strategy, and the evaluation indicators include the number of clusters, the proportion of noise points, and the stability of clustering. The parameter configuration with a noise point ratio lower than the set ratio and forming effective clusters is selected as the final model parameters.

[0226] Specifically, the DBSCAN clustering algorithm is based on the following key definitions:

[0227] Let the sample set be .

[0228] ε-neighborhood: A circular region with radius ε centered at the sample point. Its ε-neighborhood is defined as .

[0229] Key point: The ε-neighborhood must contain at least For each sample point, if the point... The ε-neighborhood contains at least If there are 1 point, then The core point: .

[0230] Boundary point: A point located in the neighborhood of the core point but which does not satisfy the core point condition.

[0231] Noise points: Points that are neither core points nor boundary points.

[0232] Direct density can be achieved: .

[0233] Density attainable: If a sequence of points exists satisfy , .

[0234] Density connected: If there exists a point , making and The density can reach , then call and Density connected.

[0235] Cluster definition: A cluster Defined as all core points and the points whose density is connected to them;

[0236] ;

[0237] By traversing all sample points, the number of samples in the ε-neighborhood of each point is calculated, and points that meet the conditions are marked as core points. Starting from any unvisited core point, all density-reachable points are found (connected through core point chains), and density-connected points are grouped into the same cluster. This process is expanded until no new points can be added, and all points not assigned to any cluster are marked as noise points.

[0238] Prepare a large number of audio signal samples of wind turbine blades, including normal blades and blades with different types of defects (such as cracks, pinholes, corrosion, etc.). Divide these samples into training set and test set. The training set uses only normal samples to train the autoencoder, while the test set contains both normal and abnormal samples, with the ratio of abnormal samples to normal samples in the test set being 1:1. All remaining normal samples are used in the training set.

[0239] Feature vectors extracted from all normal audio samples are combined to form a training dataset. The feature vectors are arranged in order of sample origin, forming a two-dimensional feature matrix. The row dimension of the matrix equals the total number of feature vectors, and the column dimension is fixed at 320. This dataset serves as input to an autoencoder model to learn the feature distribution of normal audio. A signal length verification mechanism is implemented during feature extraction. For audio signals with excessively short durations (to which at least one complete feature vector cannot be extracted), the sample is excluded and logged to ensure the integrity and validity of the training data.

[0240] An autoencoder model was written in Python. Data from the training set was preprocessed and features extracted before being imported into the model for training. The mean squared error loss function was used as the optimization objective during training, with the Adam optimizer selected and a learning rate of 0.001. A total of 300 training iterations were performed. Precision, recall, and F1 score were used to evaluate the effectiveness of the wind turbine blade defect detection method. The model was considered complete after multiple iterations until its accuracy on the test set no longer improved.

[0241] g. Defect identification.

[0242] Each cluster in the clustering results is considered a potential anomaly type to identify whether blade 1c has defects and the type of defects. Noise points are considered as difficult-to-classify anomalies or novel anomaly patterns. By combining the temporal information of audio clips and source file information, a temporal distribution map of anomaly events is established, generating detailed feature descriptions and statistical information for each anomaly type.

[0243] The anomaly detection mechanism is based on the following principle: the trained autoencoder has the best reconstruction capability for normal audio features with a small reconstruction error; however, for abnormal audio features, because they deviate from the training data distribution, the decoder cannot accurately reconstruct them, resulting in a significant increase in reconstruction error. By setting an error threshold, effective identification of abnormal audio can be achieved.

[0244] The defect detection component learns the feature representation of normal audio through a deep autoencoder and leverages the significantly increased reconstruction error of abnormal audio to achieve accurate detection. This method eliminates the need for manual annotation of abnormal samples, adaptively learns the normal audio features of different devices, and boasts advantages such as high detection accuracy, strong adaptability, and ease of implementation. It can be widely applied to preventive maintenance and fault diagnosis of industrial equipment.

[0245] After anomaly detection is completed, the preprocessed anomalous audio is input into the Wav2Vec 2.0 model to extract anomalous features. The DBSCAN algorithm is then used to cluster the anomalous audio features. After density clustering analysis, the clustering results are presented to domain experts for manual review. This involves listening to representative audio samples from each cluster and reviewing the cluster distribution using t-SNE dimensionality reduction visualization results, examining the spectral characteristics, temporal characteristics, and other acoustic properties of each cluster. Based on the review results, each confirmed anomalous cluster is labeled, given a meaningful name based on its anomalous acoustic characteristics (e.g., "bearing wear noise," "blade collision noise," etc.), and the severity and urgency of the anomaly type are indicated. The manually labeled results are then structured and stored in an anomaly knowledge base.

[0246] For new audio samples to be detected, the same Wav2Vec 2.0 model is used to extract deep features. The feature similarity between the new sample and existing anomaly types in the knowledge base is calculated. If the similarity exceeds a threshold, it is labeled as the corresponding anomaly type. When a new audio sample cannot match any known anomaly type, it is marked as "unidentified anomaly". After accumulating a certain number of such anomalies, cluster analysis is performed again, and then domain experts are invited to label the new clusters.

[0247] This embodiment utilizes a two-stage processing architecture, learning feature representations of normal audio through a deep autoencoder and leveraging the significantly increased reconstruction error of abnormal audio to achieve accurate detection. This method eliminates the need for manual annotation of abnormal samples, adaptively learns normal audio features from different devices, and boasts advantages such as high detection accuracy, strong adaptability, and ease of implementation. It can be widely applied to preventative maintenance and fault diagnosis of industrial equipment. Furthermore, the powerful representational capabilities of the pre-trained model overcome the problem of scarce labeled data in industrial scenarios; adaptive parameter configuration adapts to different types of abnormal patterns; and unsupervised clustering is used to achieve automatic discovery and classification of abnormal types, providing fine-grained analytical tools for industrial equipment fault diagnosis. This method offers advantages such as strong adaptability, high accuracy, and ease of deployment.

[0248] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention. Those skilled in the art, under the guidance of the present invention, can make various similar representations without departing from the spirit and claims of the present invention, and such modifications all fall within the protection scope of the present invention.

Claims

1. A visual-audio multimodal monitoring method using a visual-audio multimodal monitoring system, characterized in that, The visual and audio multimodal monitoring system includes a dual-purpose monitoring device installed on the nacelle of the wind turbine and multiple blade tip microphones evenly distributed at the same height on the tower of the wind turbine. The pickup direction of each blade tip microphone is horizontal and directed away from the center line of the tower, thus forming a blade tip annular pickup surface in space. The dual-purpose monitoring device includes an equipment base fixedly installed on the nacelle and a panoramic video monitoring mechanism and blade root microphones installed on the equipment base. The panoramic video monitoring mechanism is used to acquire panoramic images of the back of the wind turbine blades from the blade root to the blade tip. The pickup direction of the blade root microphones is parallel to the orientation of the nacelle, thus forming a blade root pickup line in the horizontal direction. When all the blades of the wind turbine rotate synchronously, the blade tip of each blade passes through the blade tip annular pickup surface, and the blade root of each blade passes through the blade root pickup line. The visual-audio multimodal monitoring method is performed according to the following steps: S1. The panoramic video monitoring agency identifies whether the environmental visibility under the current weather conditions is higher than the set value: if yes, proceed to step S2; if no, proceed to step S3. S2. Perform real-time monitoring of high visibility, following these steps: S21. The panoramic video monitoring agency collects panoramic images of the back of the leaf from the leaf root to the leaf tip. S22. Use image recognition algorithm to process panoramic image and determine whether there is a defect in the blade: if yes, after determining the type and location of the defect, proceed to step S23. No, return to step S1; S23. Determine whether the defect is located near the blade tip annular pickup surface or the blade root pickup line: If it is near the blade tip annular pickup surface, proceed to step S24; if it is near the blade root pickup line, proceed to step S25. S24. Based on the nacelle orientation, determine the blade tip microphone closest to the blade. Process the audio collected by the blade tip microphone using a voiceprint recognition algorithm and determine whether there is a defect in the section of the blade near the blade tip: if yes, determine the type of defect and proceed to step S26; if no, record the log and wait for manual review, then return to step S1. S25. Process the audio collected by the leaf root microphone using the voiceprint recognition algorithm to determine whether there is a defect in the section of the leaf near the leaf root: if yes, determine the type of defect and proceed to step S26; if no, record the log and wait for manual review, then return to step S1. S26. Determine whether the defect type identified by audio is consistent with the defect type identified by panoramic image in step S22: If yes, send an alarm message containing the defect type and location to the maintenance center, and then return to step S1; if no, record the log and wait for manual review, and then return to step S1. S3. Perform real-time monitoring of low visibility, following these steps: S31. Determine the blade tip pickup closest to the blade based on the nacelle orientation, and collect the blade tip audio through the blade tip pickup, while simultaneously collecting the blade root audio through the blade root pickup. S32. The visual and audio multimodal monitoring system identifies the current weather type: if it is sandstorm or rain / snow weather, proceed to step S33. If it is thunderstorm weather, proceed to step S34; if it is other weather, proceed to step S3. S33. Process the leaf tip segment audio and leaf root segment audio using spectral subtraction and adaptive filtering to suppress wind noise and / or rain noise in the leaf tip segment audio and leaf root segment audio, and then proceed to step S35. S34. Use the elimination algorithm to process the leaf tip segment audio and leaf root segment audio, and after eliminating the thunder sound, proceed to step S35. S35. Use the voiceprint recognition algorithm to process the audio of the leaf tip segment and the audio of the leaf root segment to determine whether there is a defect in the leaf: If yes, after determining the type and location of the defect, send a risk warning message containing the defect type and location to the maintenance center, mark the leaf as high risk and await re-inspection, and then proceed to step S36. No, return to step S1; S36. The panoramic video monitoring agency identifies whether the environmental visibility under the current weather conditions is higher than the set value: if yes, proceed to step S2; if no, repeat step S36.

2. The visual-audio multimodal monitoring method according to claim 1, characterized in that, Step S22 is performed according to the following steps: S221. Using a known image dataset of blade defects, train a YOLOv12 model with an embedded CBAM module using a machine learning algorithm until the YOLOv12 model with an embedded CBAM module can accurately classify defects using feature values. Then, use the YOLOv12 model with an embedded CBAM module as the target detection model. S222. Perform cropping, scaling, noise reduction, normalization, and binarization on the panoramic image to obtain the preprocessed image; S223. Use the target detection model to extract the features of the blade in the preprocessed image and determine whether there is a defect in the blade: If yes, proceed to step S23 after determining the type and location of the defect. No, return to step S1.

3. The visual-audio multimodal monitoring method according to claim 1, characterized in that, The voiceprint recognition algorithm is performed according to the following steps: a. Normal feature extraction: Acoustic features are extracted from normal audio files collected from defect-free wind turbine blades. Each audio signal is converted into a Mel spectrogram, and then the spectrogram is converted into a logarithmic scale. The sliding window method is used to stitch multiple consecutive frames into a high-dimensional feature vector to form a training dataset. b. Model Training: A deep autoencoder model is constructed to learn the feature representation of normal audio. The model adopts a symmetrical encoder and decoder structure. The encoder compresses the input features into a low-dimensional latent representation through progressively decreasing fully connected layers, capturing the essential features of normal audio. The decoder reconstructs the original input features from the latent representation. The model uses mean squared error as the loss function and minimizes the reconstruction error through an optimization algorithm, enabling the model to reconstruct the features of normal audio. During the model training phase, normal audio data is used as input and target output, and the network parameters are optimized through multiple rounds of iteration. c. Defect detection: Extract the same features from new audio samples and input them into the trained autoencoder. Calculate the reconstruction error and, by setting a threshold, detect abnormal audio that deviates significantly from the normal pattern. d. Abnormal audio signal preprocessing and segmentation: Abnormal audio signals identified by the defect detection model are collected and resampled to the sampling rate required by the Wav2Vec 2.0 pre-trained model. The continuous audio stream is segmented into overlapping segments of fixed duration using a sliding window technique. The window size and sliding step size are adaptively configured according to the characteristics of the abnormal duration. For mixed abnormalities, a 2.0-second window with a 1.0-second step size is used. Multi-channel audio is fused by mean fusion and converted into a mono signal. Finally, audio normalization is performed. e. Abnormal depth feature extraction: Using the pre-trained Wav2Vec 2.0 model as the feature extractor, the processed audio segments are input into the model, and temporal average pooling is performed on the feature sequence of each segment to generate segment-level embedding vectors of fixed dimensions. f. Density clustering analysis: The density-based DBSCAN clustering algorithm is used to perform unsupervised clustering of the feature embedding vectors. The optimal combination of clustering parameters is found through a grid search strategy. The evaluation indicators include the number of clusters, the proportion of noise points, and the clustering stability. The parameter configuration with a noise point ratio lower than the set ratio and forming effective clusters is selected as the final model parameters. g. Defect identification: Each cluster in the clustering results is regarded as a potential anomaly type to identify whether there are defects in the blades and the type of defects.

4. The visual-audio multimodal monitoring method according to claim 1, characterized in that, The equipment base has an open upper equipment mounting slot. The center of the bottom of the equipment mounting slot has an upwardly extending central axis. Multiple lower electromagnets are installed at the bottom of the equipment mounting slot and are evenly distributed circumferentially along the central axis. The blade root pickup is installed on the outer wall of the equipment mounting slot near the blade. The panoramic video monitoring mechanism includes an electrical installation chamber, at least one segmented monitoring turntable, and a multi-functional detection turntable, which are sequentially mounted on a central axis from bottom to top. The bottom of the electrical installation chamber has upper electromagnets positioned above each lower electromagnet, with each upper electromagnet having opposite magnetic poles to its corresponding lower electromagnet when energized, thus maintaining a gap between the electrical installation chamber and the equipment mounting slot. The multi-functional detection turntable and the segmented monitoring turntable can rotate along the central axis under the control of corresponding rotation control components. Both the electrical installation chamber and the segmented monitoring turntable are equipped with wide-angle cameras and thermal imaging cameras, while the multi-functional detection turntable is equipped with a CCD camera and an infrared rangefinder.

5. The visual-audio multimodal monitoring method according to claim 4, characterized in that, The bottom of the equipment installation slot is provided with multiple lower shielding sleeves that are fitted onto each lower electromagnet in a one-to-one correspondence. The bottom of the electrical installation compartment is provided with multiple upper shielding sleeves that are fitted onto each upper electromagnet in a one-to-one correspondence. Both the upper and lower shielding sleeves are made of non-magnetic material. The lower end of each upper shielding sleeve is enlarged to form a lifting guide section that fits onto the corresponding lower shielding sleeve.

6. The visual-audio multimodal monitoring method according to claim 4, characterized in that, The central shaft is a hollow shaft structure and has several wire holes that connect to the electrical installation compartment, the multi-functional testing turntable, and each segment monitoring turntable.

7. The visual-audio multimodal monitoring method according to claim 4 or 6, characterized in that, The central shaft located in the multi-functional detection turntable and each segmented monitoring turntable is integrally formed with a ring-shaped boss. The rotation control components include a motor fixedly installed in the corresponding multi-functional detection turntable or segmented monitoring turntable, and rollers synchronously rotated and mounted on the motor shaft. The circumferential outer wall of each roller is in frictional engagement with the end face of the corresponding ring-shaped boss.

8. The visual-audio multimodal monitoring method according to claim 4, characterized in that, The outer walls of the electrical installation compartment, the multi-functional testing turntable, and the monitoring turntables in each segment together form an egg-shaped structure.

9. The visual-audio multimodal monitoring method according to claim 4, characterized in that, The electrical installation compartment is equipped with multiple surge protectors.

Citation Information

Patent Citations

  • Airport visibility and runway visual range detection and display system based on image recognition technology

    CN113408415A

  • Blade status monitoring system based on video image processing for wind turbine generator and detection method

    CN109322796A

  • Fan blade panoramic monitoring system, mounting structure and panoramic monitoring shooting method

    CN117596365A

  • Fan blade anomaly detection method, device and system and medium

    CN118998003A