Image detection method for wind turbine generator blade cracks in strong light environment

By using an overexposure correction network and defuzzing processing module in the image detection of wind turbine blades, the exposure whitening and motion blur problems of image detection in strong light environments are solved, the detection accuracy and reliability are improved, and the detection speed is improved.

CN119991645APending Publication Date: 2025-05-13CHONGQING LEIRUN TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510163980.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In a strong light environment, the image detection of wind turbine blades has problems with whitening exposure and motion blur, resulting in low crack detection accuracy and reliability.

Method used

The overexposure correction network and defuzzing processing module are used to perform exposure correction and image defuzzing processing through the combination of an encoder, backbone network and decoder to improve detection accuracy.

Benefits of technology

The accuracy and reliability of blade crack detection of wind turbine sets are improved, the adaptability of detection methods is enhanced, and the detection speed is improved through the improvement of network structure and the embedding of defuzzing processing modules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991645A_ABST
    Figure CN119991645A_ABST
Patent Text Reader

Abstract

The invention discloses an image detection method for wind turbine generator blade cracks in a strong light environment, which comprises the following steps: inputting a wind turbine generator blade image into an overexposure correction network, and generating a correction image by the overexposure correction network; the overexposure correction network is composed of an encoder, a backbone network and a decoder. And adding a deblurring processing module at the front end of the original YOLOv3 network to obtain an improved YOLOv3 network, inputting a corrected image generated by the overexposure correction network into the improved YOLOv3 network, and detecting crack defects in the image through the YOLOv3 network. According to the method, crack detection is carried out on the basis of the fan blade image collected in the overexposure and fan blade high-speed rotation environment, the exposure correction method and the deblurring detection network are fused, the fan blade crack detection precision and reliability are improved, and the adaptability of the wind turbine generator blade crack method is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical fields of artificial intelligence, image restoration, target detection and the like, and in particular proposes an image detection method for cracks in blades of a wind turbine generator set. Background Art

[0002] As energy problems become increasingly severe, clean and renewable new energy has become an important direction for solving energy problems. Among them, wind power generation, as an important part of new energy technology, has received more and more attention. However, wind turbine blades are prone to cracks and structural damage due to long-term exposure to high wind loads, rain and snow erosion, lightning strikes, extreme temperature changes and other environments. Blade cracks not only affect the power generation efficiency of wind turbines, but may also cause serious safety accidents. Therefore, the detection and health monitoring of wind turbine blade cracks have become an important research direction in the wind power industry. When using drones to detect cracks in wind turbine blades, drones face many challenges in the complex environment of strong exposure and high-speed rotation of blades. In a strong exposure environment, the reflection on the surface of the blade will cause the camera to be oversensitive, blurring the details of the crack area or even completely losing them. Although the existing high dynamic range imaging technology can improve the exposure problem, the computational cost is high and it is difficult to use it efficiently in real-time detection. At the same time, due to the high-speed rotation of the blades, motion blur will occur during the shooting process, resulting in unclear crack boundaries and reduced detection accuracy. Traditional deblurring algorithms are difficult to restore clear edges in the case of large-scale blur.

[0003] At present, there are many relevant literatures and patents at home and abroad that study over-exposed image restoration and target detection in high-speed motion environments, and some effective solutions have been proposed.

[0004] 1. In the article titled "Research on Image Recognition Methods of Substation Insulator Strings Based on Deep Learning", the authors Feng Han et al. proposed an image restoration method based on contour features, which combines image preprocessing, edge detection and gradient information to improve the recognition accuracy in overexposed environments. However, this method may not perform well in complex backgrounds or high-contrast lighting scenes. Under backlight conditions, the shadow part of the insulator string may lose details, making the restoration effect unstable. Restoration based solely on contour features may not be able to supplement complex texture information, affecting detection accuracy.

[0005] 2. In the article titled "Research on Motion Image Deblurring Algorithm Based on Generative Adversarial Network", the authors Qi Haocheng et al. proposed a deblurring method that combines adversarial generative networks, residual networks, super-resolution reconstruction and other technologies. However, the deblurring effect depends on the stability of GAN training. GAN training has a mode collapse problem, which may lead to missing details of the deblurred image or unrealistic artifacts.

[0006] 3. In the article titled “Choosing Smartly: Adaptive Multimodal Fusion for Object Detection in Changing Environments”, the authors OierMees et al. proposed that CNN expert networks process RGB, depth, motion and other modal data separately, and use a gating network to dynamically adjust the weights of different expert networks according to high-level features. However, the learning ability of the gating network depends on the distribution of training data. If the training data is insufficient, it may lead to incorrect modal weight distribution and affect the accuracy of target detection. At the same time, this method is mainly based on single-frame target detection and is difficult to handle continuous frame target detection. Summary of the invention

[0007] In view of the problems that pictures taken by drones when cruising to inspect wind turbine blades in an overexposed environment have whitened images and some information is missing, as well as the problem that the high-speed rotation of the wind turbine blades and the shaking of the drone during flight cause motion blur in the taken pictures, the purpose of the present invention is to provide an image detection method for cracks in wind turbine blades in a strong light environment to solve the aforementioned problems and improve the accuracy and reliability of wind turbine blade crack detection.

[0008] The image detection method of cracks in wind turbine blades under strong light environment of the present invention comprises the following steps:

[0009] 1) Input the wind turbine blade image into the overexposure correction network, and the overexposure correction network generates a corrected image;

[0010] The overexposure correction network consists of three parts: an encoder, a backbone network and a decoder;

[0011] The encoder includes a convolution layer for extracting features of an input image, a first detail enhancement encoding block for enhancing detail information of a feature image output by the convolution layer, a first downsampling layer for downsampling an output of the first detail enhancement encoding block, a second detail enhancement encoding block for enhancing detail information of an output of the first downsampling layer, and a second downsampling layer for downsampling an output of the second detail enhancement encoding block;

[0012] The first detail enhancement coding block and the second detail enhancement coding block are respectively composed of several layers of DEM connected in series, wherein the DEM includes a convolution layer, a BN layer, a ReLU activation function layer, a depth-separable convolution layer and a SENet connected in series in sequence, and a feature image of the input DEM is multiplied by a feature image output by the SENet as the output of the DEM;

[0013] The backbone network is composed of several layers of CBA networks connected in series, and the CBA network is composed of a Transformer module and a ResNet network connected in series, and the Transformer module includes a first normalization layer, a linear attention gating mechanism module, a second normalization layer and a feedforward layer connected in series, and the linear attention gating mechanism module includes: a first convolutional layer for convolution processing the output of the first normalization layer, a linear attention network for processing the output of the first normalization layer, an activation function layer for processing the output of the first convolutional layer, an element-by-element multiplication operation block for performing element-by-element multiplication of the activation function layer and the output of the linear attention network, a second convolutional layer for convolution processing the output of the element-by-element multiplication operation block, and an element-by-element addition operation block for performing element-by-element addition of the output of the second convolutional layer and the output of the first normalization layer;

[0014] The decoder first uses the attention-guided feature fusion module to fuse the feature map output by the backbone network and the output by the second detail enhancement coding block, then performs a first upsampling on the fused feature map to halve the number of channels of the feature map, and then processes the feature map obtained by the first upsampling through the first DCBlock to reduce excessive details in the high-exposure area, and then performs feature fusion on the output of the first DCBlock and the output of the first detail enhancement coding block, and then performs a second upsampling on the fused features to reduce the number of channels, and then processes the feature map obtained by the second upsampling through the second DCBlock, and finally processes the output of the second DCBlock through the convolution layer to obtain the final corrected image;

[0015] 2) A deblurring processing module is added to the front end of the original YOLOv3 network to obtain an improved YOLOv3 network, the corrected image generated by the overexposure correction network is input into the improved YOLOv3 network, and the crack defects in the image are detected by the YOLOv3 network; the deblurring processing module is composed of 9 residual modules connected in series, and the residual module is composed of a first convolutional layer, a first instance normalization layer, a ReLU activation function layer, a second convolutional layer and a second instance normalization layer connected in series in sequence.

[0016] Furthermore, the training data set of the overexposure correction network includes simulated overexposed images obtained by the following method:

[0017] For a given image I c , the simulated overexposure effect is added through formula (1) to obtain the simulated overexposure image:

[0018]

[0019] Where x, y are the center coordinates of the overexposure range, i, j are the coordinates of the affected pixel points, s is the exposure intensity, which takes a random value between 255 and 400, and r is the exposure radius; c (i, j) is the original image, I c (i, j) are images after simulated overexposure enhancement.

[0020] The value of r is determined by formula (2), and rand represents a random number;

[0021]

[0022] Where w is the image width; h is the image height; σ takes the smaller value of the image width and height.

[0023] Furthermore, the feature extraction channels of the convolutional residual modules from layers 60 to 80 in the improved YOLOv3 network are reduced to half of the original ones.

[0024] Beneficial effects of the present invention:

[0025] 1. The image detection method for cracks in wind turbine blades under strong light environment of the present invention performs crack detection based on the images of wind turbine blades collected under overexposure and high-speed rotation of wind turbine blades. By integrating the exposure correction method and the deblurring detection network, the accuracy and reliability of wind turbine blade crack detection are improved, and the adaptability of the wind turbine blade crack detection method is increased.

[0026] 2. The present invention improves the exposure correction backbone network by introducing the Transformer module, so that the network can effectively capture the global context information in the input feature map and model long-distance dependencies, thereby improving the network's repair accuracy for overexposed images.

[0027] 3. The present invention embeds the defuzzification processing module into the YOLO detection network, thereby improving the detection accuracy of the model.

[0028] 4. The present invention compresses the structure of the YOLO detection network, making the network more lightweight and improving the detection speed of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 Schematic diagram of the structure of the exposure correction network.

[0030] Figure 2 This is the structure diagram of DEM.

[0031] Figure 3 This is a structural diagram of the CAB network.

[0032] Figure 4 This is the network structure diagram of the defuzzification processing module.

[0033] Figure 5 Flowchart for improving YOLOv3’s detection of image cracks. DETAILED DESCRIPTION

[0034] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0035] In this embodiment, the image detection method for cracks in wind turbine blades under strong light environment includes the following steps:

[0036] 1) The wind turbine blade image is input into the overexposure correction network, and the overexposure correction network generates a corrected image.

[0037] The overexposure correction network consists of three parts: an encoder, a backbone network and a decoder, and its structure is shown in Figure 1.

[0038] The encoder includes a 3×3 convolutional layer for extracting features of an input image, a first detail enhancement coding block for enhancing detail information of a feature image output by the convolutional layer, a first downsampling layer for downsampling the output of the first detail enhancement coding block, a second detail enhancement coding block for enhancing detail information of the output of the first downsampling layer, and a second downsampling layer for downsampling the output of the second detail enhancement coding block.

[0039] In this embodiment, the first detail enhancement coding block and the second detail enhancement coding block are respectively composed of four layers of DEM in series. The structure of the DEM is as follows: Figure 2As shown in the figure, it includes a 3×3 convolutional layer, a BN layer (BatchNorm2d), a ReLU activation function layer, a depth-separable convolutional layer (DepthConv) and a SENet connected in series. The feature image of the input DEM is multiplied with the feature image output by the SENet as the output of the DEM. In the DEM structure, the input over-exposed image is first subjected to a 3x3 convolution operation to extract local features of the image; the BN layer normalizes the convolution output to increase the stability and convergence speed of the network; the ReLU activation function performs a nonlinear transformation on the normalized features; a depth-separable convolution operation is performed after the nonlinear transformation to reduce the number of network parameters and improve computational efficiency; the input feature map is then processed by SENet: first, global average pooling (GAP) is performed, and then the pooled result is sent to two fully connected layers, and then the feature map is nonlinearly mapped through the activation function to output the feature map. SENet is used to automatically learn and adjust the weights of detail features to achieve the purpose of enhancing the detail information of the over-exposed area; finally, the feature map output by SENet is point-multiplied with the original feature map of the input DEM to achieve weighted features of different channels. DEM can retain the detail information in the original feature map, and can capture texture feature information after adding SENet. DEM can learn the relationship between channels, enhance image detail information, and avoid over-enhancing the image.

[0040] In this embodiment, the first downsampling layer and the second downsampling layer of the encoder are downsampling layers with a convolution kernel size of 3×3 and a step size of 2.

[0041] In this embodiment, the backbone network is composed of 6 layers of CBA networks connected in series. The structure of the CBA network is as follows: Figure 3 As shown. The CBA network is composed of a Transformer module and a ResNet network connected in series. The Transformer module includes a first normalization layer, a linear attention gating mechanism module, a second normalization layer and a feedforward layer connected in series in sequence. The linear attention gating mechanism module includes: a first convolutional layer for convolution processing the output of the first normalization layer, a linear attention network for processing the output of the first normalization layer, an activation function layer for processing the output of the first convolutional layer, an element-by-element multiplication operation block for performing element-by-element multiplication of the activation function layer and the output of the linear attention network, a second convolutional layer for convolution processing the output of the element-by-element multiplication operation block, and an element-by-element addition operation block for performing element-by-element addition of the output of the second convolutional layer and the output of the first normalization layer. The first convolutional layer and the second convolutional layer in the linear attention gating mechanism module perform 1x1 convolution.

[0042] Existing overexposure correction algorithms often use ResNet as the backbone network for feature extraction, but the modeling ability of global information is not strong. This embodiment introduces the Transformer module to enable the backbone network to effectively capture the global context information in the input feature map and model long-distance dependencies. However, the computational complexity of the standard Transformer is quadratic with the spatial size of the input feature map; in addition, the standard Transformer is not suitable for high-resolution images. To solve this problem, this embodiment designs a linear attention gating mechanism module to replace the multi-head attention module in the standard Transformer block. The computational complexity of the linear attention gating mechanism module is linearly related to the spatial size, which reduces the computational complexity compared to the multi-head attention module; and the linear attention gating mechanism module also enables the network to focus on the overexposed area of ​​the image at different positions and different scales, and can extract important feature information and highlight the correlation of specific dimensions, while retaining the original information by fusing the original features.

[0043] In this embodiment, the decoder first uses the attention guided feature fusion module (AGFF) to fuse the output of the backbone network and the feature map output by the second detail enhancement coding block, and then upsamples the fused feature map for the first time to halve the number of channels of the feature map, and then processes the feature map obtained by the first upsampling through the first DCBlock to reduce excessive details in the high exposure area, and then fuses the output of the first DCBlock and the output of the first detail enhancement coding block, and then upsamples the fused features for a second time to reduce the number of channels, and then processes the feature map obtained by the second upsampling through the second DCBlock, and finally processes the output of the second DCBlock through a 3×3 convolutional layer to obtain the final corrected image.

[0044] 2) A deblurring processing module is added to the front end of the original YOLOv3 network to obtain an improved YOLOv3 network, and the corrected image generated by the overexposure correction network is input into the improved YOLOv3 network, and the crack defects in the image are detected by the YOLOv3 network; the deblurring processing module is composed of 9 residual modules connected in series, and the residual module is composed of a first convolutional layer, a first instance normalization layer, a ReLU activation function layer, a second convolutional layer, and a second instance normalization layer connected in series in sequence. The network structure of the deblurring processing module is as follows: Figure 4 The improved YOLOv3 detection process for image cracks is shown in Figure 5As shown in the figure, when the improved YOLOv3 network receives the exposure-corrected image, it generates a 416×416 frame image through adaptive processing, and then the deblurring module performs a deblurring operation on it. After the deblurring is completed, the image size is restored to 416×416 through the deconvolution network layer, and then the feature extraction network resnet network of YOLOv3 extracts features to obtain the detection result.

[0045] As an improvement to the above embodiment, the training data set of the overexposure correction network includes simulated overexposed images obtained by the following method:

[0046] For a given image I C , the simulated overexposure effect is added through formula (1) to obtain the simulated overexposure image:

[0047]

[0048] Where x, y are the center coordinates of the overexposure range, i, j are the coordinates of the affected pixel points, s is the exposure intensity, which takes a random value between 255 and 400, and r is the exposure radius; C (i, j) is the original image, I c (i, j) are images after simulated overexposure enhancement.

[0049] The value of r is determined by formula (2), and rand represents a random number;

[0050]

[0051] Where w is the image width; h is the image height; σ takes the smaller value of the image width and height, and * is the product sign.

[0052] When a drone inspects wind turbine blades, it needs to shoot 360 degrees around the blades, so it will inevitably be affected by light. When light shines directly on the camera, the image will be partially overexposed, causing the image to be white and some information to be missing, affecting the detection results of the neural network. Since there are fewer samples of blade crack images in backlight environments, the directly trained detection model is more prone to errors when dealing with this situation, and the network can only detect the parts that are not missing. In order to solve this problem, this embodiment designs the above-mentioned simulated overexposure enhancement algorithm. By adding simulated overexposed images to the training data set as part of data enhancement, the robustness of the network in strong exposure environments can be improved.

[0053] As a further improvement to the above embodiment, this embodiment reduces the feature extraction channels of the convolutional residual modules of the 60th to 80th layers in the improved YOLOv3 network to half of the original ones.

[0054] Adding a deblurring processing module to the YOLOv3 network will inevitably increase the overall computational load of the network and greatly reduce its detection speed. To this end, by analyzing the network structure of YOLO v3, its feature extraction network for the target is mainly composed of 23 residual structures, each of which is composed of 1×1 and 3×3 convolutional layers, and then the detection of the target is completed by three YOLO layers of different scales. The number of convolutional layer channels used by the YOLOv3 network layer for feature maps of scales 13 and 26 is large, which contains more invalid connections. Therefore, for the feature extraction modules of these two scales, this embodiment performs dimensionality reduction operations on its convolutional layer. Through experimental verification, removing half of the 52-scale feature extraction channel detection, the mAP obtained by the final training can reach 76.48%, which is only 2.66 percentage points lower than the mAP measured before the network modification, which is 79.14%, indicating that the fine-grained scene detection of 52 scales has no significant detection enhancement effect on the detection object. Based on this, this embodiment reduces the feature extraction channels of the convolutional residual modules of the lower 60 to 80 layers in the YOLOv3 network to half of the original ones, making the network more lightweight and improving the speed of the model.

[0055] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution of the present invention, which should be included in the scope of the claims of the present invention.

Claims

1. An image detection method for cracks in wind turbine blades under strong light environment, characterized by: The following steps are involved: 1) Input the wind turbine blade image into the overexposure correction network, and the overexposure correction network generates a corrected image; The overexposure correction network consists of three parts: an encoder, a backbone network and a decoder; The encoder includes a convolution layer for extracting features of an input image, a first detail enhancement encoding block for enhancing detail information of a feature image output by the convolution layer, a first downsampling layer for downsampling an output of the first detail enhancement encoding block, a second detail enhancement encoding block for enhancing detail information of an output of the first downsampling layer, and a second downsampling layer for downsampling an output of the second detail enhancement encoding block; The first detail enhancement coding block and the second detail enhancement coding block are respectively composed of several layers of DEM connected in series, wherein the DEM includes a convolution layer, a BN layer, a ReLU activation function layer, a depth-separable convolution layer and a SENet connected in series in sequence, and a feature image of the input DEM is multiplied by a feature image output by the SENet as the output of the DEM; The backbone network is composed of several layers of CBA networks connected in series, and the CBA network is composed of a Transformer module and a ResNet network connected in series, and the Transformer module includes a first normalization layer, a linear attention gating mechanism module, a second normalization layer and a feedforward layer connected in series, and the linear attention gating mechanism module includes: a first convolutional layer for convolution processing the output of the first normalization layer, a linear attention network for processing the output of the first normalization layer, an activation function layer for processing the output of the first convolutional layer, an element-by-element multiplication operation block for performing element-by-element multiplication of the activation function layer and the output of the linear attention network, a second convolutional layer for convolution processing the output of the element-by-element multiplication operation block, and an element-by-element addition operation block for performing element-by-element addition of the output of the second convolutional layer and the output of the first normalization layer; The decoder first uses the attention-guided feature fusion module to fuse the feature map output by the backbone network and the output by the second detail enhancement coding block, then performs a first upsampling on the fused feature map to halve the number of channels of the feature map, and then processes the feature map obtained by the first upsampling through the first DCBlock to reduce excessive details in the high-exposure area, and then performs feature fusion on the output of the first DCBlock and the output of the first detail enhancement coding block, and then performs a second upsampling on the fused features to reduce the number of channels, and then processes the feature map obtained by the second upsampling through the second DCBlock, and finally processes the output of the second DCBlock through the convolution layer to obtain the final corrected image; 2) A deblurring processing module is added to the front end of the original YOLOv3 network to obtain an improved YOLOv3 network, the corrected image generated by the overexposure correction network is input into the improved YOLOv3 network, and the crack defects in the image are detected by the YOLOv3 network; the deblurring processing module is composed of 9 residual modules connected in series, and the residual module is composed of a first convolutional layer, a first instance normalization layer, a ReLU activation function layer, a second convolutional layer and a second instance normalization layer connected in series in sequence.

2. The image detection method for cracks in wind turbine blades under strong light environment according to claim 1 is characterized by: The training data set of the overexposure correction network includes simulated overexposed images obtained by the following method: For a given image I C , the simulated overexposure effect is added through formula (1) to obtain the simulated overexposure image: Where x, y are the center coordinates of the overexposure range, i, j are the coordinates of the affected pixel points, s is the exposure intensity, which takes a random value between 255 and 400, and r is the exposure radius; C (i, j) is the original image, I c (i, j) are images after simulated overexposure enhancement. The value of r is determined by formula (2), and rand represents a random number; Where w is the image width; h is the image height; σ takes the smaller value of the image width and height.

3. The image detection method for cracks in wind turbine blades under strong light environment according to claim 1 is characterized by: The feature extraction channels of the convolutional residual modules in the 60th to 80th layers of the improved YOLOv3 network are reduced to half of the original ones.

Citation Information

Cited By

  • Tunnel crack automatic identification method based on deep learning

    CN121259586A