Unmanned aerial vehicle aerial photography fan blade defect detection method based on computer vision
By adopting the GAN-based image defuzzing method and the improved YOLOv5s model in fan blade defect detection, the problem of reducing detection accuracy caused by image motion blur is solved, and higher detection accuracy and more accurate positioning effect are achieved, which significantly improves the reliability and maintenance efficiency of wind turbine blade defect detection.
Patent Information
- Application Number
- CN202510144385.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-13
AI Technical Summary
In fan blade defect detection, image motion blur leads to a decrease in detection accuracy, and it is difficult for the prior art to effectively deal with image deblurring and small object detection in environments where feature limitations are encountered.
The image defuzzing method based on Generative Adversarial Network (GAN) is adopted, combined with the improved YOLOv5s model, and the full-scale, binary-scale and quarter-scale defuzzing images are generated through the collaborative work of the generator and the discriminator, and the accuracy of small-objective detection is improved through the improved convolutional structure and loss function.
It effectively reduces the interference of image blur on fan blade detection, improves detection accuracy and generalization capabilities, achieves higher accuracy and more accurate positioning effects, and significantly improves the reliability and maintenance efficiency of wind turbine blade defect detection.
Smart Images

Figure CN119991638A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence, computer vision, generative adversarial networks, target detection, and the like, and in particular to a fan blade defect detection method based on computer vision. Background Art
[0002] Among the many renewable energy sources, wind power generation has become a popular energy technology worldwide due to its economic and environmental advantages. Wind turbines are usually installed in places with frequent and harsh environmental changes, such as deserts, mountains, and Gobi. However, these environmental conditions often lead to many defects on the surface of wind turbine blades, affecting the overall operation and power generation quality of wind turbines. Therefore, it is crucial to detect surface defects of wind turbine blades in a timely and accurate manner. In order to reduce expensive costs and downtime, drone-based visual inspection methods have been widely used in wind turbine blade surface defect detection.
[0003] When drones collect images of wind turbine blades, motion blur is likely to occur due to the influence of various factors such as airflow disturbance, rotor vibration, and relative motion between the camera and the blades. In the image detection of wind turbine blade defects, defects such as cracks, corrosion, and trachoma often face problems such as motion blur, smear deformation, loss of texture details, and unclear edges because they occupy a small proportion in the image and lack feature information. These problems seriously reduce the detection accuracy of small targets of wind turbine blade defects. Therefore, it is a technical problem that needs to be solved urgently to restore the motion blurred wind turbine blade defect small target image with high quality and improve the detection ability of the target detection network for small targets of wind turbine blade defects.
[0004] Although existing image deblurring methods perform well in feature-rich real-world scenarios, their effectiveness tends to weaken in feature-constrained environments. And although target detection algorithms have been widely used in wind blade defect detection, due to the variety of target types, large scale variations, and large size differences in wind blade surface defect detection targets, existing algorithms still have a lot of room for improvement in wind blade surface defect small target detection.
[0005] At present, there are many relevant literatures and patents at home and abroad that study motion image blur and small target detection, and some effective solutions have been proposed.
[0006] 1. In the article titled "Transmission Line Small Target Detection Algorithm Based on Motion Blurred Image Restoration", the author Tang Xunhao uses the conditional generative adversarial network (ViT-GAN) to restore small target motion blurred images, strengthen its feature extraction backbone's ability to perceive information on the global and regional aspects of the image, and improve the image restoration quality to facilitate subsequent target detection; by introducing a multi-head self-attention mechanism, adding a small target detection layer, and optimizing the bounding box loss function to improve the YOLOv8 network, the network's ability to detect small targets in a transmission line environment with complex backgrounds and large changes in target scale is improved. However, the processing time of this method algorithm is longer than that of other algorithms, and it requires more computing resources.
[0007] 2. In the article titled "Multi-stage progressive image restoration", the author Zamir proposed a multi-stage progressive image restoration architecture (MPRNet), which uses a coding and decoding structure to learn image context information, and then integrates it with local information and branch information. In each stage, a full-pixel adaptive mechanism is designed to weight local features, thereby effectively fusing multi-stage feature information to obtain higher restoration quality. However, the network structure of this method still has limitations in feature extraction capabilities, which makes it difficult for the algorithm to fully capture the overall information of the image, especially in the restoration of global and regional details, where there are significant differences.
[0008] 3. In the article titled "Research on Precise Detection Algorithm for Small Target Defects on Wind Blades", the author Zhang Teming proposed an improved small target detection algorithm for YOLOv5, adding the Gather-and-Distribute (GD) mechanism to the YOLOv5 network. This mechanism improves the multi-scale fusion capability by improving the convolution and self-attention operations, and is used to capture the pixel-level relationship between different scales, achieving an ideal balance between latency and accuracy. However, this method does not take into account the impact of motion image blur on small target defect detection on wind blades, and only detects trachoma defects, with weak generalization ability and overfitting problems in the data set. Summary of the invention
[0009] In view of this, the purpose of the present invention is to provide a method for detecting defects in wind turbine blades using drone aerial photography based on computer vision, so as to solve the technical problem of interference of image blur on wind turbine blade detection and improve detection accuracy.
[0010] The method for detecting defects of wind turbine blades by drone aerial photography based on computer vision of the present invention comprises the following steps:
[0011] 1) The wind turbine blade image taken by the drone is pre-processed and then input into the generative adversarial network for deblurring. The generative adversarial network includes a generator and a discriminator. The network architecture of the generator adopts the MIMO-UNet network. The discriminator consists of a global discriminator and a local discriminator. The MIMO-UNet network generates deblurred images of three scales: full scale, half scale and quarter scale. The total loss function of the generative adversarial network for optimizing the quality of the deblurred image generated by the generator is as follows:
[0012] L=λ1L Adversarial +λ2L MSE +λ3L MSFR +λ4L Yolo (1)
[0013] Among them, L Adversarial is the adversarial loss, L MSE is the mean square error loss, L MSFR is the multi-scale structural similarity loss, L Yolo is the perceptual loss, λ1, λ2, λ3 and λ4 are the weight coefficients of each loss;
[0014] The full-scale deblurred image is input into the perceptual loss function module to calculate the perceptual loss to measure the difference between the CNN feature maps of the generated image and the target image; the perceptual loss function is defined as follows:
[0015]
[0016] Among them, I B is the input blurred image, I S To represent the generated deblurred image, φ is the feature map obtained by the first layer CSP1_3 in the YOLOv5s backbone network, and W and H are the dimensions of the feature map; Represents a G Generator function, x, y represent the indexes used to traverse the width and height dimensions of the feature map;
[0017] The deblurred images of three scales are simultaneously input into the MSFR loss function module to calculate the MSFR loss. The MSFR loss is used to measure the distance between the multi-scale true image and the deblurred image in the frequency domain to measure the similarity of the image structure. The MSFR loss function is defined as follows:
[0018]
[0019] Where: F represents the fast Fourier transform that transfers the image signal to the frequency domain, K is the number of scales, k is the specific scale, W k and H k Respectively represent the image width and height of the k-th scale, represents the generated deblurred image at the kth scale, represents the clear image of the kth scale;
[0020] The deblurred images of three scales are simultaneously input into the MSE loss function module to calculate the MSE loss to measure the pixel-level error between the generated image and the real image. The MSE loss function is defined as follows:
[0021]
[0022] in, and Represent the generated RGB image and the real RGB image of the kth scale respectively;
[0023] The deblurred images of three scales are input into the global discriminator and the local discriminator respectively, and the adversarial loss is calculated to judge the deblurring effect at the overall and detail levels respectively. The adversarial loss function L Adversarial The definition is as follows:
[0024]
[0025] Among them, G represents the generator, D represents the discriminator, and p data (x) is the real data distribution, p z is the generated data distribution; x is the sample data in the real data set, D(x) is the output of the discriminator D for the real sample x, and z is the generator from the input noise distribution p z The noise vector sampled in is D(z), and D(z) is the output of the discriminator D for the samples generated by the generator based on the noise z. It is expectation; D is the Lipschitz constant of the discriminator D, L D The definition is as follows:
[0026]
[0027] That is L D is the smallest real number among the following:
[0028]
[0029] 2) The dataset is composed of deblurred images of wind turbine blades processed by generative adversarial networks;
[0030] 3) Improve the YOLOv5s model, including:
[0031] 1) Add an upsampling layer after the second upsampling operation in the Neck part of the YOLOv5s model;
[0032] 2) Use SPD convolutional building blocks to replace the convolutional structure in the original YOLOv5s model;
[0033] 3) Using loss function L EIoU Instead of the CIoU loss function in the original YOLOv5s model, the EIoU loss function is defined as follows:
[0034]
[0035] where h w and h c are the width and height of the smallest external bounding box of the predicted bounding box and the target bounding box, respectively; p represents the Euclidean distance between two points, b represents the coordinates of the center point of the predicted box, w represents the width of the predicted box, h represents the height of the predicted box, and b gt Represents the center coordinates of the ground-truth bounding box, w gt Indicates the width of the real box, h gt Represents the height of the real box, IoU represents the intersection-over-union ratio of the predicted box and the real box, and the calculation formula of IoU is as follows:
[0036]
[0037] Where: intersection area A intersection =max(0,x2-x1)×max(0,y2-y1),
[0038] Prediction box area
[0039] Real frame area
[0040] The coordinates of the upper left corner of the prediction box are The coordinates of the lower right corner are The coordinates of the upper left corner of the real box are The coordinates of the lower right corner are
[0041]
[0042] 4) Use the data set obtained in step 2) to train and test the improved YOLOv5s model;
[0043] 5) Inputting the wind turbine blade image taken by the drone in real time into the generative adversarial network described in step 1), inputting the deblurred image output by the generative adversarial network into the YOLOv5s improved model that has passed the test in step 4), and detecting the defects of the wind turbine blades through the YOLOv5s improved model.
[0044] Furthermore, the computer vision-based UAV aerial photography wind turbine blade defect detection method also includes replacing the residual block in the generator network architecture with a fast Fourier transform convolution residual block, the structure of the fast Fourier transform convolution residual block includes two convolution residual streams, the first convolution residual stream includes four residual modules connected in series, each residual module is composed of two 3×3 convolution blocks and a ReLU activation function connected between the two 3×3 convolution blocks; the second convolution residual stream includes a RealFFT2d module for performing a 2D real fast Fourier transform, four residual modules connected in series and an InvRealFFT2d module for calculating an inverse 2D real fast Fourier transform, each residual module is composed of two 1×1 convolution blocks and a ReLU activation function connected between the two 1×1 convolution blocks; after the image feature data is processed by the first convolution residual stream and the second convolution residual stream respectively, the outputs of the two convolution residual streams are processed element by element as the final output of the fast Fourier transform convolution residual block.
[0045] Beneficial effects of the present invention:
[0046] 1. The present invention deeply integrates the image deblurring method based on the generative adversarial network with the small target detection method based on the improved YOLOv5s, which greatly reduces the interference of image blur on the detection of fan blades and effectively improves the detection generalization ability of the method of the present invention in a fuzzy image environment. At the same time, for the detection of small targets with fan blade defects, the present invention can better extract small target feature information, accelerate the model convergence speed and improve the regression accuracy by improving the YOLOv5s algorithm, and achieves higher accuracy and more precise positioning effect in the detection of small targets with fan blade defects.
[0047] 2. In terms of data processing, the defective images of wind turbine blades collected by drones are blurred due to vibration and relative motion and have high labeling costs. Based on some field-collected and labeled images, the present invention uses GAN to perform image deblurring, effectively improving image quality while creating a high-quality wind turbine blade defect detection data set, providing solid data support for subsequent detection tasks. In terms of model optimization, compared with the prior art, the target detection model obtained by the present invention has fewer parameters, which not only reduces the computing resources required for model training and reasoning, greatly improves the operating efficiency, but also shows stronger robustness when processing motion-blurred images, and can better adapt to complex and changeable practical application scenarios.
[0048] 3. In practical applications, the method of the present invention can significantly reduce the detection error caused by image blur, greatly improving the reliability and maintenance efficiency of wind turbine blade defect detection. This helps to detect wind turbine blade defects in a timely manner, thereby reducing expensive costs and downtime, and plays a vital role in maintaining the efficient and reliable operation of wind turbines and increasing economic benefits. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 Schematic diagram of the training process of generating adversarial networks.
[0050] Figure 2 This is the generator network architecture diagram.
[0051] Figure 3 This is the network structure diagram of the global discriminator.
[0052] Figure 4 This is the network structure diagram of the local discriminator.
[0053] Figure 5 The network structure after adding an upsampling layer to the Neck part of the YOLOv5s model.
[0054] Figure 6 Schematic diagram of the working principle of the building block for SPD convolution.
[0055] Figure 7 Figure 2 is the structural diagram of the fast Fourier transform convolution residual block.
[0056] Figure 8 Flow chart of the fan blade defect detection method. DETAILED DESCRIPTION
[0057] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0058] In this embodiment, the method for detecting defects in wind turbine blades by drone aerial photography based on computer vision includes the following steps:
[0059] 1) The wind turbine blade image taken by the drone is pre-processed and then input into the generative adversarial network for deblurring. The generative adversarial network includes a generator and a discriminator. The network architecture of the generator adopts the MIMO-UNet network. The discriminator consists of a global discriminator and a local discriminator. The MIMO-UNet network generates deblurred images at three scales: full scale (i.e., original image scale), half-scale, and quarter-scale. The training process of the generative adversarial network is as follows: Figure 1 As shown. The total loss function used by the generative adversarial network to optimize the quality of the deblurred image generated by the generator is as follows:
[0060] L=λ1L Adversarial +λ2LMSE +λ3L MSFR +λ4L Yolo (1)
[0061] Among them, L Adversarial is the adversarial loss, L MSE is the mean square error loss, L MSFR is the multi-scale structural similarity loss, L Yolo is the perceptual loss, and λ1, λ2, λ3, and λ4 are the weight coefficients of each loss.
[0062] The full-scale deblurred image is input into the perceptual loss function module to calculate the perceptual loss to measure the difference between the CNN feature maps of the generated image and the target image; the perceptual loss function is defined as follows:
[0063]
[0064] Among them, I B is the input blurred image, I S To represent the generated deblurred image, φ is the feature map obtained by the first layer CSP1_3 in the YOLOv5s backbone network, and W and H are the dimensions of the feature map; Represents a G The generator function of , x, y represents the index used to traverse the width and height dimensions of the feature map. Perceptual loss focuses on restoring general content, while adversarial loss focuses on restoring texture details. Using perceptual loss training can improve the deblurring effect.
[0065] The deblurred images of three scales are simultaneously input into the MSFR loss function module to calculate the MSFR loss. The MSFR loss is used to measure the distance between the multi-scale true image and the deblurred image in the frequency domain to measure the similarity of the image structure. The MSFR loss function is defined as follows:
[0066]
[0067] Where: F represents the fast Fourier transform that transfers the image signal to the frequency domain, K is the number of scales, k is the specific scale, and W k and H k Respectively represent the image width and height of the k-th scale, represents the generated deblurred image at the kth scale, Represents the clear image at the kth scale. Adding the MSFR loss to the total loss function can reduce the difference in frequency space and allow the network to learn to recover the lost high-frequency components.
[0068] The deblurred images of three scales are simultaneously input into the MSE loss function module to calculate the MSE loss to measure the pixel-level error between the generated image and the real image. The MSE loss function is defined as follows:
[0069]
[0070] in, and Represent the generated RGB image and the real RGB image of the kth scale respectively;
[0071] The deblurred images of the three scales are input into the global discriminator and the local discriminator respectively (the structures of the global discriminator and the local discriminator are as follows: Figure 3 and Figure 4 As shown in Figure 2), by calculating the adversarial loss to judge the deblurring effect at the overall and detail levels, the adversarial loss function L Adversarial The definition is as follows:
[0072]
[0073] Among them, G represents the generator, D represents the discriminator, and p data (x) is the real data distribution, p z is the generated data distribution; x is the sample data in the real data set, D(x) is the output of the discriminator D for the real sample x, and z is the generator from the input noise distribution p z The noise vector sampled in is D(z), and D(z) is the output of the discriminator D for the samples generated by the generator based on the noise z. It is expectation; D is the Lipschitz constant of the discriminator D, L D The definition is as follows:
[0074]
[0075] That is L D is the smallest real number among the following:
[0076]
[0077] A sufficient condition for the Lipschitz constant is Using gradient normalization, D(x) is transformed into Make it automatically satisfied In the prior art, ReLU or LeakyReLU is usually used as the activation function. Under this condition, D(x) is actually a "piecewise linear function", which means that, except for the boundary, D(x) is a linear function in the local continuous region. Accordingly, is a constant vector. So gradient normalization tries to make So we have:
[0078]
[0079] In order to avoid the denominator being zero, |D(x)| is directly added to the denominator, which also ensures the boundedness of the function.
[0080]
[0081] The gradient normalization step can stabilize the training process of the generative adversarial network and improve the model performance.
[0082] For blurred images in complex scenes, using only the “patch” scale discriminator will lead to uneven distortion in spatial background restoration and cause offset of object detection positioning frames. Unlike traditional generative adversarial networks, the multi-scale dual discriminator used in this embodiment is a discriminator structure that provides multiple supervisions at different scales. The outputs of three different scales are all input into the discriminator. This structure enables the discriminator to obtain more data for training in each training round, making it converge faster, more efficient and more stable, while prompting the generator to obtain better training performance, which can greatly improve the accuracy of target detection.
[0083] 2) The dataset is composed of deblurred images of wind blades processed by a generative adversarial network.
[0084] 3) Improve the YOLOv5s model, including:
[0085] 1) Add an upsampling layer after the second upsampling operation in the Neck part of the YOLOv5s model.
[0086] The feature fusion network of YOLOv5s is composed of a feature pyramid network (FPN) and a path aggregation network (PANet). In the process of feature extraction from shallow to deep, shallow features have higher resolution and richer geometric information, while deep features have stronger receptive fields and richer semantic information. When the drone takes images of defects in wind turbine blades, the defects are small targets in the entire image, and these small targets are mainly represented by shallow features. After multiple convolution and merging operations, the original YOLOv5s model has a low resolution of the output feature map and lacks the expression of shallow information, making it difficult for the original model to learn the features of small targets, thereby affecting the detection accuracy of small targets.
[0087] In order to improve the utilization rate of shallow features, this embodiment reconfigures the feature fusion network of the original model by adding an upsampling layer (layer Upsample) after the second upsampling operation of the original YOLOv5s model. The structure is as follows Figure 5As shown in the figure. This change enables the feature map to go through convolution and upsampling operations again, thereby adjusting the output scale of the feature map, and finally obtaining feature maps of size 160*160, 80*80, and 40*40. Among them, the feature map of size 160*160 is specifically used to handle the detection of smaller objects in the image. It has a higher resolution and can better preserve shallow feature information, making it easier for the network to learn the features of small objects, thereby improving the detection accuracy of small objects.
[0088] 2) The SPD convolutional building block is used to replace the convolutional structure in the original YOLOv5s model. The background of drone aerial images is complex, and the feature information of small targets is also fuzzy. The original YOLOv5s model uses a strided convolutional layer with a step size of 2 to downsample the feature map, which easily leads to a lack of discriminant information in the feature extraction process, thereby affecting the accuracy of detection. In order to better extract the feature information of small targets, this embodiment uses the SPD convolutional building block to replace the original convolutional structure. The SPD convolutional building module consists of a space-to-depth layer and a non-stride convolutional layer. Its working principle is as follows: Figure 6 As shown, specifically including:
[0089] (1) Splitting: The original feature map is split into multiple sub-feature maps according to the scale factor. For example, when the scale factor is 2, 4 sub-feature maps can be obtained. The size of each sub-feature map is
[0090] (2) Concatenation: The sub-feature maps are concatenated according to the channel dimension to obtain the intermediate feature map X′. The spatial dimension of the intermediate feature map is reduced by the scale factor, while the channel dimension is enlarged by the scale factor. This will produce a final size
[0091] (3) Non-strided convolution: The intermediate feature map X’ is further transformed into the final feature map X” by using a non-strided (stride = 1) convolution layer with C2 filters. The purpose of using a non-strided convolution layer is to retain all feature discriminative information as much as possible.
[0092] The original image is first sliced into sub-feature maps after the SPD convolution building block; then, the sub-feature maps are connected and features are extracted; finally, the extracted feature information is filtered so that the feature information in the image is retained to the greatest extent while downsampling the feature map. After this process, the feature extraction capability of the network is improved, thus solving the problem of losing feature information due to downsampling.
[0093] 3) Using loss function L EIoU Instead of the CIoU loss function in the original YOLOv5s model, the EIoU loss function is defined as follows:
[0094]
[0095] Where: h w and h c are the width and height of the smallest external bounding box of the predicted bounding box and the target bounding box, respectively; p represents the Euclidean distance between two points, b represents the coordinates of the center point of the predicted box, w represents the width of the predicted box, h represents the height of the predicted box, and b gt Represents the center coordinates of the ground-truth bounding box, w gt Indicates the width of the real box, h gt Represents the height of the real box, IoU represents the intersection-over-union ratio of the predicted box and the real box, and the calculation formula of IoU is as follows:
[0096]
[0097] Where: intersection area A intersection =max(0,x2-x1)×max(0,y2-y1),
[0098] Prediction box area
[0099] Real frame area
[0100] The coordinates of the upper left corner of the prediction box are The coordinates of the lower right corner are The coordinates of the upper left corner of the real box are The coordinates of the lower right corner are
[0101]
[0102] The three parts of the EIoU loss function are the IoU loss L IoU , distance loss L dis The aspect ratio loss Lasp considers the real difference between the overlapping area, the distance between the center points, and the length and width of the edge, respectively, solving the problem of the fuzzy definition of the aspect ratio based on CIoU. Separating the difference between the width and height of the predicted bounding box and the width and height of the minimum enclosing bounding box speeds up the convergence of the model and improves the regression accuracy.
[0103] 4) Use the dataset obtained in step 2) to train and test the improved YOLOv5s model.
[0104] 5) Inputting the wind turbine blade image taken by the drone in real time into the generative adversarial network described in step 1), inputting the deblurred image output by the generative adversarial network into the YOLOv5s improved model that has passed the test in step 4), and detecting the defects of the wind turbine blades through the YOLOv5s improved model.
[0105] As an improvement to the above embodiment, the method for detecting defects in wind turbine blades by drone aerial photography based on computer vision further includes replacing the residual block in the generator network architecture with a fast Fourier transform convolution residual block (FFTRes), and the structure of the fast Fourier transform convolution residual block (FFTRes) is as follows: Figure 7 As shown, it includes two convolution residual streams, the first convolution residual stream includes four residual modules connected in series, each of which is composed of two 3×3 convolution blocks and a ReLU activation function connected between the two 3×3 convolution blocks; the second convolution residual stream includes a RealFFT2d module for 2D real fast Fourier transform, four residual modules connected in series, and an InvRealFFT2d module for calculating the inverse 2D real fast Fourier transform, each of which is composed of two 1×1 convolution blocks and a ReLU activation function connected between the two 1×1 convolution blocks; after the image feature data are processed by the first convolution residual stream and the second convolution residual stream respectively, the outputs of the two convolution residual streams are processed element by element as the final output of the fast Fourier transform convolution residual block FFTRes.
[0106] The processing of the second convolution residual stream is as follows:
[0107] The feature map from the previous layer Input the RealFFT2d module, which performs a 2D real fast Fourier transform on the input to obtain Will The input is the residual concatenation module, and the InvRealFFT2d module performs an inverse 2D real fast Fourier transform on the output of the residual concatenation module, converts it back to the spatial domain, and obtains the output of the second convolution residual stream;
[0108] The output of the first convolution residual stream and the output of the second convolution residual stream are added element by element as the final output of the fast Fourier transform convolution residual block FFTRes.
[0109] The purpose of deblurring is to restore the lost high-frequency components, so the difference in frequency space must be reduced. The first convolution residual stream of the fast Fourier transform convolution residual block (FFTRes) of this improved design learns the pixel spatial domain features of the blurred image; the second convolution residual stream performs fast Fourier transform on the feature map in the spatial domain, and then learns features in the frequency domain through the convolution residual part on the feature map. The FFTRes block can achieve better performance and improve inference efficiency while using fewer residual blocks.
[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution of the present invention, which should be included in the scope of the claims of the present invention.
Claims
1. A method for detecting defects in wind turbine blades by drone aerial photography based on computer vision, characterized in that: The following steps are involved: 1) The wind turbine blade image taken by the drone is pre-processed and then input into the generative adversarial network for deblurring. The generative adversarial network includes a generator and a discriminator. The network architecture of the generator adopts the MIMO-UNet network. The discriminator consists of a global discriminator and a local discriminator. The MIMO-UNet network generates deblurred images of three scales: full scale, half scale and quarter scale. The total loss function of the generative adversarial network for optimizing the quality of the deblurred image generated by the generator is as follows: L=λ1L Adversarial +λ2L MSE +λ3L MSFR +λ4L Yolo (1) Among them, L Adversarial is the adversarial loss, L MSE is the mean square error loss, L MSFR is the multi-scale structural similarity loss, L Yolo is the perceptual loss, λ1, λ2, λ3 and λ4 are the weight coefficients of each loss; The full-scale deblurred image is input into the perceptual loss function module to calculate the perceptual loss to measure the difference between the CNN feature maps of the generated image and the target image; the perceptual loss function is defined as follows: Among them, I B is the input blurred image, I S To represent the generated deblurred image, φ is the feature map obtained by the first layer CSP1_3 in the YOLOv5s backbone network, and W and H are the dimensions of the feature map; Represents a G Generator function, x, y represent the indexes used to traverse the width and height dimensions of the feature map; The deblurred images of three scales are simultaneously input into the MSFR loss function module to calculate the MSFR loss. The MSFR loss is used to measure the distance between the multi-scale true image and the deblurred image in the frequency domain to measure the similarity of the image structure. The MSFR loss function is defined as follows: Where: F represents the fast Fourier transform that transfers the image signal to the frequency domain, K is the number of scales, k is the specific scale, and W k and H k Respectively represent the image width and height of the k-th scale, represents the generated deblurred image at the kth scale, represents the clear image of the kth scale; The deblurred images of three scales are simultaneously input into the MSE loss function module to calculate the MSE loss to measure the pixel-level error between the generated image and the real image. The MSE loss function is defined as follows: in, and Represent the generated RGB image and the real RGB image of the kth scale respectively; The deblurred images of three scales are input into the global discriminator and the local discriminator respectively, and the adversarial loss is calculated to judge the deblurring effect at the overall and detail levels respectively. The adversarial loss function L Adversarial The definition is as follows: Among them, G represents the generator, D represents the discriminator, and p data (x) is the real data distribution, p z is the generated data distribution; x is the sample data in the real data set, D(x) is the output of the discriminator D for the real sample x, and z is the generator from the input noise distribution p z The noise vector sampled in is D(z), and D(z) is the output of the discriminator D for the samples generated by the generator based on the noise z. It is expectation; D is the Lipschitz constant of the discriminator D, L D The definition is as follows: That is L D is the smallest real number among the following: 2) The dataset is composed of deblurred images of wind turbine blades processed by generative adversarial networks; 3) Improve the YOLOv5s model, including: 1) Add an upsampling layer after the second upsampling operation in the Neck part of the YOLOv5s model; 2) Use SPD convolutional building blocks to replace the convolutional structure in the original YOLOv5s model; 3) Using loss function L EIoU Instead of the CIoU loss function in the original YOLOv5s model, the EIoU loss function is defined as follows: Where: h w and h c are the width and height of the smallest external bounding box of the predicted bounding box and the target bounding box, respectively; p represents the Euclidean distance between two points, b represents the coordinates of the center point of the predicted box, w represents the width of the predicted box, h represents the height of the predicted box, and b gt Represents the coordinates of the center point of the real box, w gt Indicates the width of the real box, h gt Represents the height of the real box, IoU represents the intersection-over-union ratio of the predicted box and the real box, and the calculation formula of IoU is as follows: Where: intersection area A intersection =max(0,x2-x1)×max(0,y2-y1), Prediction box area Real frame area The coordinates of the upper left corner of the prediction box are The coordinates of the lower right corner are The coordinates of the upper left corner of the real box are The coordinates of the lower right corner are 4) Use the data set obtained in step 2) to train and test the improved YOLOv5s model; 5) Inputting the wind turbine blade image taken by the drone in real time into the generative adversarial network described in step 1), inputting the deblurred image output by the generative adversarial network into the YOLOv5s improved model that has passed the test in step 4), and detecting the defects of the wind turbine blades through the YOLOv5s improved model.
2. The method for detecting defects in wind turbine blades by drone aerial photography based on computer vision according to claim 1 is characterized in that: It also includes replacing the residual block in the generator network architecture with a fast Fourier transform convolution residual block, the structure of the fast Fourier transform convolution residual block includes two convolution residual streams, the first convolution residual stream includes four residual modules connected in series, each residual module consists of two 3×3 convolution blocks and a ReLU activation function connected between the two 3×3 convolution blocks; the second convolution residual stream includes a RealFFT2d module for performing a 2D real fast Fourier transform, four residual modules connected in series, and an InvRealFFT2d module for calculating an inverse 2D real fast Fourier transform, each residual module consists of two 1×1 convolution blocks and a ReLU activation function connected between the two 1×1 convolution blocks; after the image feature data is processed by the first convolution residual stream and the second convolution residual stream respectively, the outputs of the two convolution residual streams are processed element by element as the final output of the fast Fourier transform convolution residual block.
Citation Information
Cited By
Fan blade surface defect detection method and system based on visual image of unmanned aerial vehicle
CN120259302A
Fan defect detection method and system based on machine vision
CN120672669A
A fan defect detection method and system based on machine vision
CN120672669B