Fast Crack Detection Method Based on the Combination of Improved YOLOv5 Neural Network and UAV Video

By improving the YOLOv5 neural network combined with drone video and SSA sparrow search algorithm, the problems of traditional low crack detection efficiency and high hardware requirements are solved, and fast and accurate crack recognition is achieved, which is suitable for general computer and drone video detection.

CN115223060BActive Publication Date: 2025-07-04ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210608946.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-31
Publication Date
2025-07-04
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

The existing crack detection methods are inefficient and have poor results, making it difficult to quickly and accurately collect crack information within the skylight time, and the hardware equipment requirements are high. The traditional deep learning model has a lot of work in image acquisition, and the initial clustering center needs to be manually set, which may lead to inaccurate classification.

Method used

The improved YOLOv5 neural network is used to combine drone video, and the SSA sparrow search algorithm is used to replace the K-means algorithm for feature point clustering. The bridge surface image data set is collected through the drone, image processing and training is performed, and video frame-by-frame recognition is performed by combining OpenCV.

Benefits of technology

It realizes fast and accurate crack detection, reduces hardware requirements, improves identification efficiency and accuracy, is suitable for general computers, overcomes space limitations and is easy to operate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115223060B_ABST
    Figure CN115223060B_ABST
Patent Text Reader

Abstract

A fast crack detection method based on the combination of an improved YOLOv5 neural network and drone video, comprising: using a drone to capture a video of the surface of a target structure, extracting images in the video frame by frame, and obtaining a training and validation data set through labeling; based on the YOLOv5 algorithm model, introducing the SSA (Sparrow Search Algorithm) to replace and improve the original K-means algorithm in the model, and training the improved model with the completed data set; using the trained model to perform frame-by-frame image recognition on the new video captured by the drone, and finally detecting the position of cracks in the video. The present invention combines the YOLOv5 and SSA algorithms to accurately identify target cracks, and realizes crack detection through video input, having the advantages of high efficiency, high precision, low cost, strong robustness, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of structural detection, and particularly to a rapid crack detection method based on the combination of improved YOLOv5 and drone-shot videos. Background Art

[0002] The construction of world road and bridge traffic facilities has been developing year by year. Bridges are an important part of transportation. Most of their structures are composed of concrete, and they will crack under the action of heavy vehicle rolling, rain and snow surface corrosion, and the aging of concrete materials themselves during operation. Cracking is a macroscopic manifestation of the stress state at the mesoscopic level of the structure. The continuous development of cracks may lead to the loss of bearing capacity of the structure or even collapse. Especially for prestressed concrete bridges and reinforced concrete bridges with a higher safety level, the regular detection of crack damage is more important.

[0003] Traditional crack detection methods include ultrasonic detection method, acoustic emission detection method, fiber optic sensing detection method, etc. These methods have low detection efficiency, poor effect, and it is difficult to quickly and accurately collect crack information during the skylight time. In recent years, digital image processing technology has become the mainstream research direction of crack detection. Among them, deep learning is a relatively advanced method for automatic detection of some targets. By extracting features from a large amount of real image data, the machine can be made to have the ability of analysis and learning. Deep learning classification neural network models for overall targets include AlexNet, GoogLeNet, VGGNet, etc., and those for image segmentation include SegNet, UNet, etc. The cracks detected by the above neural network models are more objective and reliable compared with traditional detection. However, these methods require a large amount of picture data, have a large workload in image acquisition, take a long time in the pre-processing process of images, and have high requirements for hardware devices.

[0004] YOLOv5 is a relatively novel deep learning network for image recognition at present. It has the advantages of small memory, high recognition accuracy, fast running speed, can accept the input of pictures of any size and perform adaptive gray filling into a consistent size, and can accept video input. It is a brand-new lightweight network, which can run on a computer with a lower configuration, and can speed up the target detection speed on the basis of ensuring accuracy, which fully meets the requirements of rapid crack detection. The detection head adopts the K-means clustering algorithm. This algorithm has a clear physical meaning and can achieve accurate classification with a small number of samples. However, the initial clustering center needs to be manually selected, and inaccurate classification may be caused by ineffective clustering due to the improper selection of the initial value. The SSA algorithm can effectively avoid these problems.

[0005] The SSA Sparrow Search Algorithm is an innovative algorithm inspired by the foraging and searching habits of sparrows. It does not require manual setting of initial values, automatically searches for feature points and updates the comfort value, and has the advantages of strong search ability, excellent search results, fast convergence, and good clustering effect. Compared with the original K-means clustering algorithm in YOLOv5, it has higher classification accuracy, faster speed, and stronger robustness. Combining SSA with YOLOv5 for recognition can identify small target objects such as cracks faster and more accurately. Summary of the Invention

[0006] The present invention aims to overcome the above problems and disadvantages of the background art, and proposes a fast crack detection method based on an improved YOLOv5 neural network combined with drone video.

[0007] The present invention first uses a drone to collect pictures of surface cracks of the existing bridge superstructure to form a training and validation data set; trains the image data set obtained in the previous step with the YOLOv5 neural network improved by SSA; finally, uses the computer vision software library Open CV to process each frame of the newly collected bridge surface video to identify cracks in the video.

[0008] The fast crack detection method based on an improved YOLOv5 neural network combined with drone video of the present invention includes the following steps:

[0009] A. Construction of a structural crack image data set based on a drone.

[0010] A1: Use a drone to photograph the surface of the target structure. The drone adopts a cruise route along the bridge longitudinal direction, and the ground workstation connects to a wireless video receiver to receive the collected video.

[0011] A2: Extract images from the video frame by frame at a certain time interval to form an image library for subsequent operations.

[0012] A3: Modify the original image to an image with a smaller resolution to relieve the pressure on hardware devices. Since bridge cracks do not occur explosively and have a long cracking period, there are not many observable images containing cracks. Therefore, perform operations such as flipping, brightness transformation, and zooming on the collected images to expand the data set.

[0013] A4: Use the labelImg software to create labels for each image of the collected pictures, store the crack labels and non-crack labels separately, and randomly generate a training set and a validation set.

[0014] B. Improve the YOLOv5 network based on SSA.

[0015] B1: Replace the K-means algorithm in the YOLOv5 network structure with the SSA algorithm that has stronger feature search capabilities. The SSA first initializes the number of clusters, the number of iterations, and the ratios between different feature points; calculates the fitness values in sequence to provide regions and directions for subsequent classification of feature points; and introduces the following formula to update the positions of homogeneous feature points among heterogeneous feature points:

[0016]

[0017] t represents the current iteration number, j = 1, 2, 3, …; itermax is a constant representing the maximum iteration number. X i,j is the position information of the i-th pixel point in the j-th dimension. α ∈ (0, 1] is a random number R; R2 ∈ [0, 1], ST ∈ [0.5, 1] are the warning value and the safety value respectively; Q is a random number following a normal distribution. L is a 1×d matrix with all elements being 1.

[0018] Introduce the following formula to update the positions of newly added homogeneous feature points:

[0019]

[0020] X P is the optimal position occupied by the discoverer, and X worst is the worst position. A is a 1×d matrix with elements being 1 or -1, and A + = A T (AA T ) -1 .

[0021] Introduce the following formula to update the positions of marginal feature points:

[0022]

[0023] X best is the global optimal position. β is the step size control parameter, β ∼ N(0, 1), K ∈ [-1, 1], and f i is the adaptive value of each feature point. f g and f w are the best and worst fitness values, and ε is a constant to prevent the denominator from being 0.

[0024] Calculate the fitness values and update the positions of feature points through the above steps. If the stopping condition is met, exit the process and output the results; otherwise, repeat the steps after initialization.

[0025] B2: Based on the B1 feature point update algorithm, first perform random sampling on the data points to be classified to obtain the initial classification candidate points of the SSA. Secondly, through the iteration of the above-mentioned sparrow search algorithm, update the positions of the feature points to move each feature point in the global space towards the optimal value and converge quickly, and classify the feature points of the same type into one category.

[0026] B3: Through the above calculations, the classification of different feature points can be achieved, that is, the classification between different color blocks, and finally the classification and recognition of all features in the image can be realized.

[0027] C. Input the expanded dataset into the improved network for training.

[0028] C1: Use a script to convert the image file label format to TXT, adopt the improved YOLOv5 network to speed up the recognition speed of the target object, use the Mosica algorithm for data augmentation at the input end, cut off a part of the image content and fill it with pixel values in other areas of the dataset, and detect the object background composed of four pictures to adaptively adjust the anchor box.

[0029] C2: Continue to input the pictures processed in step C1 into the improved YOLOv5 backbone network CSP1-X to slice the image. The neck network uses FPN+PAN for upsampling and downsampling, and uses the CSP2-X structure to minimize the information loss of the original image and strengthen the network feature fusion.

[0030] C3: Output the accuracy rate obtained after each training, and observe whether the required accuracy rate is reached after the training of the specified number of times. The detection head uses the sparrow search algorithm in B to generate the prediction box, and the loss function uses GIoU. The GIoU loss propagates backward in the network to gradually reduce its loss until the model converges.

[0031] C4: After the model training is completed, output the accuracy rate, observe whether the required precision is met, and finally obtain the recognition result of the model network.

[0032] D. Collect videos and input them into the trained network to achieve fast detection.

[0033] D1: Reshoot the surface of the upper structure of the bridge where images have not been collected, and input the new crack video collected by the drone on site into the network trained in C to achieve fast crack detection.

[0034] D2: Use the Video read function in the computer vision software library OpenCV to read the original video collected by the drone, extract images frame by frame at a certain time interval and input them into the network in C for crack recognition and store the detection results.

[0035] D3: After specifying the output address, video size, frame rate and other information for the already recognized video frames through the Video write function, rewrite them to generate a new video with crack recognition completed, and finally realize an efficient and practical crack fast detection method based on the combination of the improved YOLOv5 neural network and the drone video.

[0036] In step A1, a drone is used to photograph the surface of the target structure. Try to select an area with better light for image acquisition to reduce the impact of overexposure or too weak light on image clarity.

[0037] Compared with the prior art, the present invention has the following advantages:

[0038] 1. The present invention overcomes the spatial limitation of crack acquisition. The drone can be used to obtain crack images at any height, and the operation is convenient and efficient.

[0039] 2. The YOLOv5 network is good at identifying small target objects, and cracks are exactly small target objects, which well meet the requirements for identifying structural cracks. Moreover, the input pictures can be directly input without additional cropping processing. Four picture data can be processed at one time during the standardization process, with high efficiency, saving time for subsequent structural repair and preventing greater damage.

[0040] 3. The improved YOLOv5 model used in the present invention does not need to run on a computer with high configuration, and general computers can be used for running and recognition, with relatively lower hardware requirements.

[0041] 4. The present invention improves the YOLOv5 neural network. The original K-means for collecting clustering points is replaced by the SSA sparrow search algorithm. After replacement, the algorithm has stronger optimization ability, faster convergence speed, higher accuracy in identifying cracks, and strong robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 is a schematic diagram of crack image acquisition of the present invention.

[0043] Figure 2 is a flowchart of the implementation of the method of the present invention.

[0044] Specific implementation steps:

[0045] The following combines Figure 1 the schematic diagram of crack acquisition by the drone in Figure 1 and Figure 2 the flowchart of implementation in Figure 1 to illustrate the specific implementation manner of the present invention by taking the crack detection of a bridge as an example: Legend description:

[0046] 1 - The upper structure of the bridge

[0047] 2 - The image acquisition area of the upper structure

[0048] 3 - The drone

[0049] A rapid crack detection method based on the combination of an improved YOLOv5 neural network and drone video according to the present invention comprises the following specific steps:

[0050] A. Construction of an Unmanned Aerial Vehicle (UAV)-Based Structural Crack Image Dataset

[0051] A1: Use the UAV 3 to photograph the surface of the upper structure 1 of the bridge. Try to select an image acquisition area 2 with better light for image acquisition to reduce the impact of overexposure or too weak light on image clarity. The UAV 3 adopts a cruising route along the bridge longitudinal direction, and the ground workstation connects to the wireless video transmitter receiver to receive the collected video.

[0052] A2: Extract images from the video frame by frame at a certain time interval to form an image library for subsequent operations.

[0053] A3: Modify the original images to images with a smaller resolution to relieve the pressure on hardware devices. Since bridge cracks do not occur explosively and have a long cracking period, there are not many observable images containing cracks. Therefore, perform operations such as flipping, brightness transformation, zooming in and out on the collected images to expand the dataset.

[0054] A4: Use the labelImg software to create labels for each image in the collected pictures, including crack labels and non-crack labels, which are stored separately and randomly generate a training set and a validation set.

[0055] B. Improvement of YOLOv5 Network Based on SSA

[0056] B1: Replace the K-means algorithm in the YOLOv5 network structure with the SSA algorithm with stronger search feature ability. SSA first initializes the number of categories, the number of iterations, and the ratio between different feature points; calculates the fitness value in sequence to provide the area and direction for the subsequent classification of feature points; introduces the following formula to update the position of the same type of feature points among different types of feature points:

[0057]

[0058] t represents the current iteration number, j = 1, 2, 3, …; itermax is a constant representing the maximum iteration number. X i,j is the position information of the i-th pixel point in the j-th dimension. α ∈ (0, 1] is a random number R; R2 ∈ [0, 1], ST ∈ [0.5, 1] are the early warning value and the safety value respectively; Q is a random number obeying the normal distribution. L is a 1×d matrix with all elements being 1.

[0059] Introduce the following formula to update the position of the newly added same type of feature points:

[0060]

[0061] X P is the optimal position occupied by the discoverer, X worstis the worst position. A is a 1×d matrix with elements 1 or -1, A + = A T (AA T ) -1 .

[0062] Introduce the following formula to update the position of edge feature points:

[0063]

[0064] X best is the global optimal position. β is the step size control parameter, β~N(0,1), K∈[-1,1], f i is the adaptive value of each feature point. f g and f w are the best and worst fitness values, and ε is a constant to prevent the denominator from being 0.

[0065] Calculate the fitness value and update the feature point position through the above steps. If the stop condition is met, exit the process and output the result; otherwise, repeat the steps after initialization.

[0066] B2: Based on the B1 feature point update algorithm, first perform random sampling on the data points to be classified to obtain the initial classification candidate points of SSA. Secondly, through the iteration of the above sparrow search algorithm, update the feature point position to make each feature point in the global move towards the optimal value and converge quickly, and classify the feature points of the same type into one category.

[0067] B3: Through the above calculations, the classification of different feature points can be realized, that is, the classification between different color blocks, and finally the classification recognition of all features in the image can be realized.

[0068] C. Input the expanded dataset into the improved network for training.

[0069] C1: Use a script to convert the image file label format to TXT, adopt the improved YOLOv5 network to speed up the recognition speed of target objects, use the Mosica algorithm for data augmentation at the input end, cut off a part of the image content and fill it with pixel values in other areas of the dataset, and detect the object background composed of four pictures to adaptively adjust the anchor box.

[0070] C2: Continue to input the pictures processed in step C1 into the improved YOLOv5 backbone network CSP1-X to slice the images. The neck network uses FPN+PAN for upsampling and downsampling, and uses the CSP2-X structure to minimize the information loss of the original image and strengthen the network feature fusion.

[0071] C3: Output the accuracy rate obtained after each training, and observe whether the required accuracy rate is achieved after the training of the specified number of times. The detection head uses the sparrow search algorithm in B to generate prediction boxes, and the loss function uses GIoU. The GIoU loss is backpropagated in the network to gradually reduce its loss until the model converges.

[0072] C4: After the model training is completed, output the accuracy rate, observe whether the required accuracy is met, and finally obtain the recognition result of the model network.

[0073] D. Collect videos and input them into the trained network to achieve rapid detection.

[0074] D1: Reshoot the surface of the upper structure 1 of the bridge where images have not been collected, and input the new crack video collected by the unmanned aerial vehicle 3 on-site into the network trained in C to achieve rapid crack detection.

[0075] D2: Use the Video read function in the computer vision software library OpenCV to read the original video collected by the unmanned aerial vehicle 3, extract images frame by frame at a certain time interval, input them into the network in C for crack recognition, and store and retrieve the detection results.

[0076] D3: After specifying the output address, video size, frame rate and other information for the already recognized video frames through the Video write function, rewrite them to generate a new video with crack recognition completed, and finally realize an efficient and practical crack rapid detection method based on the combination of the improved YOLOv5 neural network and the unmanned aerial vehicle video.

Claims

1. A rapid crack detection method based on the combination of an improved YOLOv5 neural network and drone video, comprising the following steps: A. Construct a structural crack image dataset based on a drone; A1: Use a drone to photograph the surface of the target structure; the drone adopts a cruising route along the bridge direction, and the ground workstation connects to a wireless video transmitter receiver to receive the collected video; A2: Extract images from the video frame by frame at a certain time interval to form an image library for subsequent operations; A3: Modify the original image to an image with a smaller resolution to relieve the pressure on hardware devices. Since bridge cracks do not occur explosively and have a long cracking period, there are not many observable images containing cracks. Therefore, the collected images are flipped, brightness-transformed, enlarged and reduced, etc. to augment the dataset; A4: Use the labelImg software to create labels for each image in the collected pictures, store the crack labels and non-crack labels separately, and randomly generate a training set and a validation set; B. Improve the YOLOv5 neural network based on SSA; B1: Replace the K-means algorithm in the YOLOv5 network structure with the SSA algorithm with stronger feature search ability; SSA first initializes the types, the number of iterations, and the ratio between different feature points; calculates the fitness value in order to provide the area and direction for the subsequent classification of feature points; introduces the following formula to update the positions of the same-type feature points among different-type feature points: Let \(t\) denote the current iteration number, \(j = 1, 2, 3, \cdots\); \(itermax\) is a constant representing the maximum number of iterations; \(X\) i,j is the position information of the \(i\)-th pixel point in the \(j\)-th dimension; \(\alpha\in(0, 1]\) is a random number \(R\); \(R_2\in[0, 1]\), \(ST\in[0.5, 1]\) are the warning value and the safety value respectively; \(Q\) is a random number following a normal distribution; \(L\) is a \(1\times d\) matrix with all elements being \(1\); Introduce the following formula to update the positions of the newly added same-type feature points: X P is the optimal position occupied by the discoverer, X worst is the worst position; A is a 1×d matrix with elements 1 or -1, A + = A T (AA T ) -1 ; Introduce the following formula to update the positions of the edge feature points: X best is the global optimal position; β is the step size control parameter, β ∼ N(0,1), K ∈ [-1,1], f i is the adaptive value of each feature point; f g and f w are the best and worst fitness values, ε is a constant to prevent the denominator from being zero; Calculate the fitness value and update the feature point positions through the above steps. If the stop condition is met, exit the process and output the result; otherwise, repeat the steps after initialization; B2: Based on the B1 feature point update algorithm, first initialize random sampling for the data points to be classified to obtain the initial classification candidate points of SSA. Secondly, through the iteration of the above-mentioned sparrow search algorithm SSA, update the feature point positions to make each feature point in the global move towards the optimal value and converge quickly, and classify the feature points of the same type into one category; B3: Through the above calculations, the classification of different feature points can be achieved, that is, the classification between different color blocks, and finally the classification recognition of all features in the image can be realized; C. Train the improved YOLOv5 neural network based on the augmented dataset; C1: Use a script to convert the image file label format to TXT, adopt the improved YOLOv5 network to speed up the recognition speed of the target object, use the Mosica algorithm for data augmentation at the input end, cut off a part of the image content and fill it with pixel values in other areas of the dataset, and detect the object background composed of four pictures to adaptively adjust the anchor box; C2: Continue to input the pictures processed in step C1 into the improved YOLOv5 backbone network CSP1-X to slice the images. The neck network uses FPN+PAN for upsampling and downsampling, and uses the CSP2-X structure to minimize the information loss of the original image and strengthen the network feature fusion; C3: Output the accuracy rate obtained after each training, and observe whether the required accuracy rate is achieved after the training of the specified number of times; the detection head uses the sparrow search algorithm in B to generate prediction boxes, and the loss function uses GIoU; the GIoU loss is backpropagated in the network to gradually reduce its loss until the model converges. C4: After the model training is completed, output the accuracy rate, observe whether the required accuracy is met, and finally obtain the recognition result of the model network. D. Rapid crack detection based on the improved YOLOv5 neural network. D1: Re-take the surface of the upper structure of the bridge where images have not been collected, and input the new crack video collected by the drone on-site into the network trained in C to achieve rapid crack detection. D2: Use the Video read function in the computer vision software library OpenCV to read the original video collected by the drone, extract images frame by frame at a certain time interval, input them into the network in C for crack recognition, and store the detection results. D3: After specifying the output address, video size, frame rate and other information for the already recognized video frames through the Video write function, rewrite them to generate a new video with completed crack recognition, and finally realize an efficient and practical crack rapid detection method based on the combination of the improved YOLOv5 neural network and the drone video.

2. The rapid crack detection method based on the combination of the improved YOLOv5 neural network and UAV video according to claim 1, characterized in that: In step A1, use the drone to take pictures of the surface of the target structure, and try to select areas with better light for image collection to reduce the impact of overexposure or too weak light on the image clarity.