Steel box girder crack detection method based on improved YOLOv7
By adding Dynamic Head dynamic detection head to the YOLOv7 neural network and combining the MPDIoU loss function, the steel box girder crack detection method is improved, solving the shortcomings in complex backgrounds and fine crack detection, and achieving high-precision and efficient detection effects.
Patent Information
- Application Number
- CN202510181782.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-13
AI Technical Summary
The existing steel box girder crack detection methods have shortcomings in complex backgrounds and fine crack detection, which are difficult to meet the needs of steel box girder crack detection.
Using the steel box girder crack detection method based on improved YOLOv7, the model's detection accuracy and fine crack recognition ability in complex backgrounds are improved by adding Dynamic Head dynamic detection head to the YOLOv7 neural network and combining the MPDIoU loss function.
High-precision crack detection and small crack identification in complex backgrounds are achieved, which reduces detection costs and improves the safety of inspectors and the objectivity and accuracy of detection results.
Smart Images

Figure CN120147818A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of steel box girder disease detection, and particularly relates to a steel box girder crack detection method based on improved YOLOv7. Background Technique
[0002] With the increase in the number of steel bridges and the past development model of "emphasizing reconstruction over maintenance", more and more steel bridges have entered the high-incidence period of diseases. Among them, due to the influence of environmental factors and vehicle traffic, etc., the cracking of steel box girders has become one of the main problems affecting the safety of steel bridges, and the importance of steel bridge crack detection has become increasingly prominent. At present, the detection of steel bridge cracks mainly adopts the method of manual detection. However, the internal environment of steel box girders is complex, and some places that are difficult for some inspectors to reach need to rely on tools such as ladders for crack detection, and the detection results may also vary due to different subjective judgments of inspectors. Therefore, there are problems of high inspection difficulty and poor accuracy, and the method based on manual detection has gradually been unable to meet the requirements of steel box girder crack detection.
[0003] With the development of computer vision and deep learning, disease detection technologies based on deep learning have been widely applied in the field of civil engineering. The detection technologies based on deep learning mainly include image recognition and image segmentation. Among them, image recognition can output one or more rectangular boxes (bounding boxes) with class labels in an image, so as to give the position and class of the target in the image. Image segmentation can not only label the position and class of the target in the image, but also generate a pixel-level mask for each target object, so as to give the exact boundary and shape of the target in the image. In addition, for image recognition and image segmentation technologies, their common algorithms are divided into two categories: two-stage and one-stage. Typical two-stage algorithms are Faster R-CNN and Mask R-CNN. First, in the first stage, the Region Proposal Network (RPN) is used to extract candidate regions that may contain target objects. Subsequently, in the second stage, a classification network is used to determine the target class of each candidate region, and a regression network is used to adjust the position and size of the candidate region bounding box, so as to achieve accurate positioning and segmentation of the target. Since the two-stage algorithm processes the target twice, it can achieve better performance in terms of accuracy. However, its computational complexity is large and the processing speed is relatively slow, which is not suitable for real-time detection. Compared with the two-stage algorithm, the one-stage detection algorithm regards object detection and segmentation as a regression problem, and directly predicts the position, class and shape of the target in the image based on a deep convolutional neural network. Typical one-stage algorithms include YOLO (You Only Look Once), SSD (Single Shot MultiBox Detector) and RetinaNet. They skip the step of generating candidate regions, directly extract features from the image, and output the class, position and shape of the target through a neural network. Since the one-stage detection algorithm does not need to process the image multiple times, it has the advantage of fast detection speed and can be applied to real-time detection application scenarios. However, its accuracy may be slightly lower than that of the two-stage algorithm.
[0004] In summary, the steel box girder crack detection method based on computer vision has not been optimized for the characteristics of steel bridge cracks and cannot be applied to crack detection under complex backgrounds and detection of fine cracks. Therefore, the current steel box girder crack detection method needs to be further improved in terms of detection under complex backgrounds and detection of fine cracks. It is necessary to develop a model optimized for the characteristics of steel box girder cracks, so as to provide effective technical support for the condition assessment of steel box girder cracks and help build an all-round and full-process intelligent bridge detection system in the future. Summary of the Invention
[0005] The present invention discloses a steel box girder crack detection method based on improved YOLOv7, aiming to provide a computer vision-based detection method for steel box girder crack detection that is low-cost, highly secure for personnel, applicable to complex backgrounds, and can improve the detection effect of fine cracks, making the detection results more objective, accurate, fast, convenient, intuitive, and effective.
[0006] To achieve the above object, the technical solution of the present invention is as follows:
[0007] A steel box girder crack detection method based on improved YOLOv7, comprising the following steps:
[0008] (1) Construct a steel box girder crack data set;
[0009] (2) Based on the YOLOv7 detection model, add a Dynamic Head dynamic detection head after the backbone layer in its neural network;
[0010] (3) Implement object detection and instance segmentation through bounding box regression.
[0011] Preferably, the step (1) includes the following specific steps: According to the methods of providing by bridge maintenance units and personal shooting and collection, collect steel box girder disease images, and use the annotation method of a paintbrush to annotate all images for diseases. Finally, establish a steel box girder crack data set containing 4,418 disease images; the steel box girder crack data set includes photos with different devices, angles, environments, and resolutions, and increases the robustness of the detection model by ensuring the diversity of the data set.
[0012] Preferably, in the step (2), the Dynamic Head dynamic detection head unifies different object detection heads by using an attention mechanism; the attention mechanism between feature levels is used for scale perception, the attention mechanism between spatial positions is used for spatial perception, and the attention mechanism within the output channel is used for task perception, improving the expression ability of the model's object detection head without increasing the amount of calculation.
[0013] Preferably, in the step (2), the general self-attention formula is:
[0014] W(F) = π(F)·F F ∈ R L×S×C (1)
[0015] Since directly learning the attention function in all dimensions by this method will lead to excessive calculation amount and there are high-dimensional problems; therefore, the attention function is converted into 3 sequential attentions, and each attention only focuses on one dimension:
[0016] W(F) = π C (π S (π L(F)·F)·F)·F (2)
[0017] In formula (2), π L , π S , π C are three attention functions applied to scale perception, spatial perception, and task perception, respectively;
[0018] Scale perception attention π L : Fuses different scales based on semantic importance:
[0019]
[0020] In formula (3), f(·) is a linear function approximated by a 1×1 convolutional layer, is the hard-sigmoid function;
[0021] Spatial perception attention π S : First, uses deformable convolution to sparsify attention learning, and then performs cross-scale feature integration;
[0022]
[0023] In formula (4), K is the number of sparse sampling positions, p k +Δp k is the displacement position focused on a distinguishing area by the self-learned spatial offset Δp k , and Δm k is the self-learned important scalar at position p k ; both are learned from the input features at the median level of F;
[0024] Task perception attention π C : Is used to dynamically switch feature channels to assist different tasks, promoting the generalization of joint learning and target expression capabilities:
[0025] π C (F)·F = max(α 1 (F)·F c +β 1 (F), α 2 (F)·F c +β 2 (F)) (5)
[0026] In formula (5), F c is the feature slice of the c-th channel, [α 1 , α 2 , β 1 , β 2 T = θ(·) is a hyperparameter used to control the activation threshold, and θ(·) is similar to DyReLU;
[0027] Any type of backbone network is used to extract the feature pyramid and further scale it to a unified scale to construct a 3D tensor, which is then used as the input of the Dynamic Head; Next, multiple DyHead modules including scale awareness, spatial position awareness, and task awareness are stacked serially. Finally, various types of predictions are attached to the head layer, and the output of the dynamic detection head is used for image recognition and segmentation tasks.
[0028] Preferably, step (3) includes: combining the loss functions IoU, GIoU, DIoU, and CIoU to propose MPDIoU to directly minimize the distance between the upper left and lower right corner points of the predicted bounding box and the ground truth bounding box. The calculation method is as follows:
[0029] For A and B, define (x 1 A , y 1 A ), (x 2 A , y 2 A ) to represent the upper left corner coordinate point and the lower right corner coordinate point of A, and (x 1 B , y 1 B ), (x 2 B , y 2 B ) to represent the upper left corner coordinate point and the lower right corner coordinate point of B, as shown in equations (7)-(9):
[0030]
[0031] On this basis, a bounding box regression loss function based on MPDIoU is proposed, and the loss function is defined as follows:
[0032]
[0033]
[0034] The beneficial effects of the steel box girder crack detection method based on the improved YOLOv7 of the present invention are:
[0035] 1. The present invention can help bridge inspection departments save costs. Compared with traditional steel box girder crack detection, the present invention requires less manpower. At the same time, for steel box girders with relatively harsh inspection environments, inspectors can complete the inspection at a safe location, improving the safety of inspectors. In terms of inspection results, it also reduces the subjectivity of inspectors and improves the accuracy of crack detection. Therefore, with the aid of the present invention, rapid, objective, and accurate steel box girder crack detection can be achieved.
[0036] 2. The present invention makes up for the disadvantages of poor applicability of steel box girder crack detection under complex backgrounds and poor detection effects for fine cracks at the present stage. The present invention combines the specification requirements and actual engineering needs, and establishes a steel box girder crack dataset based on 4418 images. By adding a DyHead dynamic detection head to the YOLOv7 neural network, it can better adapt to the detection tasks under complex backgrounds and improve the recognition effect of fine cracks. And the MPDIoU loss function is introduced to further improve the detection accuracy of the YOLOv7 detection model.
[0037] 3. The present invention realizes the visualization of inspection results based on computer vision technology. The YOLOv7 deep learning model used in the present invention can visualize the inspection results based on computer vision technology. In the image, it can automatically mark the detected crack positions, categories, and shapes, making the inspection results look more convenient and intuitive for statistics.
[0038] 4. The present invention has a wide range of applications. The web-based inspection platform developed using the trained YOLOv7 detection model in a local / cloud server can provide an image upload and inspection channel for inspectors. In addition, the web-based inspection platform developed in the cloud server can also provide API interfaces for third-party applications and platforms to call. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a structural diagram of the improved YOLOv7 model.
[0040] Figure 2 It is a detection framework diagram of Dynamic Head.
[0041] Figure 3 It is a schematic diagram of the principle of DyHead.
[0042] Figure 4 It is a schematic diagram of bounding box regression calculation ((a) is a schematic diagram of IoU bounding box calculation; (b) is a schematic diagram of MPDIoU bounding box calculation).
[0043] Figure 5 It is an effect diagram of steel box girder crack detection. DETAILED DESCRIPTION OF THE INVENTION
[0044] The following is only the preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
[0045] The following embodiments can be understood as separately expressing a part of the local structure or method of the present invention, or can also be understood as the embodiments being combined with each other to explain the connotation of the structure or method of a larger scope of the present invention.
[0046] In the initial embodiment, a steel box girder crack detection method based on improved YOLOv7 of the present invention includes the following steps:
[0047] (4) Construct a steel box girder crack dataset;
[0048] (5) Based on the YOLOv7 detection model, add a Dynamic Head dynamic detection head after the backbone layer in its neural network. The improved YOLOv7 model structure is as Figure 1 shown;
[0049] (6) Achieve object detection and instance segmentation through bounding box regression.
[0050] In a further embodiment, the step (1) includes the following specific steps: collect steel box girder disease images according to the methods provided by bridge maintenance units and personal shooting, and use the annotation method of a paintbrush to annotate diseases for all images. Finally, establish a steel box girder crack dataset containing 4418 disease images; the steel box girder crack dataset includes photos with different devices, angles, environments, and resolutions, and the robustness of the detection model is increased by ensuring the diversity of the dataset.
[0051] In a further embodiment, in the step (2), the Dynamic Head dynamic detection head unifies different object detection heads by using the attention mechanism; the attention mechanism between feature levels is used for scale perception, the attention mechanism between spatial positions is used for spatial perception, and the attention mechanism within the output channel is used for task perception, significantly improving the expression ability of the model's object detection head without increasing the computational amount. The detection framework of Dynamic Head is as Figure 2 shown, where n is the number of Dynamic Head modules.
[0052] In a further embodiment, in the step (2), the general self-attention formula is:
[0053] W(F) = π(F)·F F ∈ R L×S×C (1)
[0054] Since directly learning the attention function in all dimensions by this method leads to excessive computational complexity and there are high-dimensional problems; therefore, the attention function is converted into three sequential attentions, each of which focuses on only one dimension:
[0055] W(F) = π C (π S (π L (F)·F)·F)·F (2)
[0056] In formula (2), π L , π S , π C are three attention functions applied to scale perception, spatial perception, and task perception respectively;
[0057] Scale perception attention π L : Fuses different scales based on semantic importance:
[0058]
[0059] In formula (3), f(·) is a linear function approximated by a 1×1 convolutional layer, is the hard-sigmoid function;
[0060] Spatial perception attention π S : First, uses deformable convolution to sparsify attention learning, and then performs cross-scale feature integration;
[0061]
[0062] In formula (4), K is the number of sparse sampling positions, p k +Δp k is the displacement position focused on a distinguishing area by the self-learned spatial offset Δp k , Δm k is the self-learned important scalar at position p k ; both are learned from the input features at the median level of F;
[0063] Task perception attention π C : It can dynamically switch feature channels to assist different tasks, promoting the generalization of joint learning and target expression capabilities:
[0064] π C (F)·F = max(α 1 (F)·f c +β 1 (F), α 2 (F)·f c +β 2 (F)) (5)
[0065] In formula (5), F c is the feature patch of the c-th channel, [α 1 , α 2 , β 1 , β 2 T = θ(·) is a hyperparameter used to control the activation threshold, and θ(·) is similar to DyReLU;
[0066] Backbone networks of any type are used to extract feature pyramids and further scale them to a unified scale to construct 3D tensors, which are then used as the input of the Dynamic Head; next, multiple DyHead modules including scale perception, spatial position perception, and task perception are stacked serially. Finally, various types of predictions are attached to the head layer, and the output of the dynamic detection head is used for image recognition and segmentation tasks.
[0067] After adding the DyHead dynamic detection head to the YOLOv7 neural network, the detection effect in complex backgrounds and the detection effect for small cracks are better than those of the original model. The reason is that the DyHead dynamic detection head combines the three dimensions of scale perception attention, spatial perception attention, and task perception attention to construct a unified head. As Figure 3 shown, the scale perception attention module uses the subsequent non-linear activation function to greatly increase the non-linear characteristics, increase the network depth, and enhance the expression ability of the neural network, making the model more sensitive to cracks of different scales; the spatial position perception attention module makes the crack features sparser and focuses on cracks in different positions; the task perception attention module will form different activations based on different downstream tasks and can be applied to different tasks (classification, recognition, segmentation).
[0068] In a further embodiment, step (3) includes: initially, the I n -norm loss function is applied to bounding box regression, but it is very sensitive to each scale; in YOLOv1, the square root of w and h is used to reduce the sensitivity; in YOLOv3, 2wh is used; to better calculate the difference between the true bounding box and the predicted bounding box, the IoU loss is used. As Figure 4 (a) shows, the initial IoU represents the ratio of the intersection area of the predicted bounding box and the true bounding box to the union area, which is expressed as:
[0069]
[0070] In formula (6), represents the true bounding box, represents the predicted bounding box;
[0071] Traditional IoU metrics have limitations when dealing with bounding boxes with similar width and height values, thus hindering the effectiveness of convergence speed and accuracy. GIoU loses its effectiveness when the predicted bounding box is completely surrounded by the actual bounding box. DIoU degenerates into the original IoU when the center point of the predicted bounding box coincides with the center point of the true bounding box. CIoU takes into account both the center point distance and aspect ratio, but it is defined using relative values rather than absolute values. When the predicted bounding box and the true bounding box have the same aspect ratio under different width and height values, the loss function for bounding box regression will lose its effectiveness. Therefore, MPDIoU is proposed by combining the loss functions IoU, GIoU, DIoU, and CIoU. As shown in Figure 4 (b), to directly minimize the distance between the top-left and bottom-right points of the predicted bounding box and the true bounding box, the calculation method is as follows:
[0072] For A and B, define (x 1 A , y 1 A ), (x 2 A , y 2 A ) to represent the top-left and bottom-right coordinate points of A, and (x 1 B , y 1 B ), (x 2 B , y 2 B ) to represent the top-left and bottom-right coordinate points of B, as shown in Eqs. (7)-(9):
[0073]
[0074] On this basis, a bounding box regression loss function based on MPDIoU is proposed, and the loss function is defined as follows:
[0075]
[0076] The working principle of the present invention:
[0077] (1) The experimental model deployment environment of the present invention is the Pytorch 2.1.0 deep learning framework. The computer operating system is Ubuntu 20.04 LTS, with python 3.11 and CUDA 12.1. The platform hardware configuration is as follows: CPU: AMD RyzenTM 9 7900x processor, 32g of memory, and the graphics card is NVIDIA RTX4090d with 24g of video memory. The environment for model deployment is installed through Anaconda, and Tensorboard is used to supervise the training process during the experiment. The network structure parameters are shown in Table 1.
[0078] Table 1 Network Structure Parameter Table
[0079]
[0080] (2) First, the present invention conducted experiments on the original YOLOv7 model. The mAP for crack recognition was 67.3%, and the mAP for crack segmentation was 58.4% respectively. Since the detection effect of the original YOLOv7 model is poor in complex backgrounds, the present invention added the DyHead dynamic detection head after the backbone layer of the YOLOv7 neural network to obtain the improved model YOLOv7-D, and set the number of dynamic detection heads to 2, 4, 6, and 8 respectively, and conducted experiments on the above four cases. As can be seen from Table 2, when the number of dynamic detection heads is 6, the crack recognition and segmentation effects are better than those of the original model and the optimized models with the number of 2, 4, and 8. The mAP for crack recognition is 69%, which is 2.5% higher than that of the original YOLOv7 model, and the mAP for crack segmentation is 58.8% respectively, which is 0.7% higher than that of the original YOLOv7 model.
[0081] Table 2 Comparison Table of Detection Performance of v7+D(n)
[0082]
[0083] (3) The present invention replaced the CIoU loss function used in the original YOLOv7 model with the MPDIoU loss function to obtain the improved model YOLOv7-M. As can be seen from Table 3, the mAP for crack recognition is 70.1%, which is 4.1% higher than that of the original YOLOv7 model. The mAP for crack segmentation is 63.8% respectively, which is 9.2% higher than that of the original model.
[0084] Table 3 Performance Comparison Table of v7, v7+MP, v7+MP+D(6)
[0085]
[0086] The present invention integrates the dynamic detection head DyHead(6) with the MPDIoU loss function to obtain the fusion model YOLOv7-DM, and then conducts experiments on it. As can be seen from Table 3, after fusion, the mAP for crack recognition is 70.3%, which is 4.5% higher than that of the original YOLOv7 model, and the mAP for crack segmentation is 61.9%, which is 6% higher than that of the original YOLOv7 model.
[0087] From the above results, it can be seen that the YOLOv7-DM fusion model has improved in terms of recognition accuracy and the mAP evaluation index compared to the original model, and the mAP is 0.3% higher than that of the YOLOv7-M model and 1.9% higher than that of the YOLOv7-D model. Therefore, in the crack recognition task, the YOLOv7-DM fusion model has the best comprehensive effect. In the crack segmentation task, all evaluation indexes of the YOLOv7-M optimized model are better than those of the fusion model, and the mAP is 8.5% higher than that of the YOLOv7-D optimized model. Therefore, the YOLOv7-M optimized model has the best comprehensive effect in the crack segmentation task.
[0088] The trained YOLOv7 detection model is deployed on a local / cloud server, and a web-based detection platform is developed on the local / cloud server to provide an upload and recognition channel for users. Detection personnel can take pictures of the steel box girder cracks at the detection site using shooting devices such as mobile phones and cameras, and upload the captured images to the cloud server through websites and other means for steel box girder crack detection. In addition, the web-based detection platform developed on the cloud server can also provide API interfaces for third-party applications and platforms to call. The web-based detection platform developed on the local / cloud server will perform intelligent detection of steel box girder cracks on the uploaded images and automatically mark the position, type, and shape of the cracks in the images. The detection results are as Figure 5 shown.
Claims
1. A steel box girder crack detection method based on improved YOLOv7, characterized by comprising the following steps: (1) Construct a steel box girder crack dataset; (2) Based on the YOLOv7 detection model, a Dynamic Head dynamic detection head is added after the backbone layer in its neural network; (3) Object detection and instance segmentation are achieved through bounding box regression.
2. A steel box girder crack detection method based on improved YOLOv7 as claimed in claim 1, characterized in that: The step (1) comprises the following specific steps: collecting steel box girder defect images according to the methods provided by the bridge maintenance unit and taken by individuals, and annotating all images with defects using a brush annotation method, and finally establishing a steel box girder crack dataset containing 4418 defect images; the steel box girder crack dataset includes photos of different equipment, angles, environments and resolutions, and the robustness of the detection model is increased by ensuring the diversity of the dataset.
3. A steel box girder crack detection method based on improved YOLOv7 as claimed in claim 2, characterized in that: In the step (2), the Dynamic Head dynamic detection head uses an attention mechanism to unify different target detection heads; the attention mechanism between feature levels is used for scale perception, the attention mechanism between spatial positions is used for spatial perception, and the attention mechanism within the output channel is used for task perception, thereby improving the expression ability of the model target detection head without increasing the amount of calculation.
4. A steel box girder crack detection method based on improved YOLOv7 as claimed in claim 3, characterized in that: In step (2), the general self-attention formula is: W(F)=π(F)-F F∈R L×S×C (1) Since this method directly learns the attention function in all dimensions, it will lead to excessive computation and high dimensionality problem; therefore, the attention function is converted into 3 sequence attentions, each focusing on only one dimension: W(F)=πC(πS(πL(F)·F)·F)·F (2) In formula (2), π L , π S , π C There are three attention functions applied to scale perception, space perception, and task perception respectively; Scale-aware attention π L : Fusion of different scales based on semantic importance: In formula (3), f(·) is a linear function approximated by a 1×1 convolutional layer. is the hard-sigmoid function; Spatial Perception Attention π S : First, deformable convolution is used to sparse attention learning, and then features are integrated across scales; In formula (4), K is the number of sparse sampling positions, p k +Δp k is the self-learning space offset Δp k Focus on the displacement position of a distinguishing area, Δm k is the position p k The self-learned important scalar at ; both are learned from the input features at the median level of F; Task-aware attention π C :Used to dynamically switch feature channels to assist different tasks and promote the generalization of joint learning and target expression capabilities: p C (F)-F=max(α I (F)·F c +b l (F),a 2 (F)·F c +b 2 (F)) (5) In formula (5), F c is the feature slice of the cth channel, [α 1 , α 2 , β 1 , β 2 ] T =θ(·) is a hyperparameter used to control the activation threshold, θ(·) is similar to DyReLU; Any type of backbone network is used to extract feature pyramids and further scale them to a uniform scale to construct a 3D tensor, which is then used as the input of DynamicHead; next, multiple DyHead modules containing scale awareness, spatial position awareness, and task awareness are stacked in series. Finally, various types of predictions are attached to the head layer, and the output of the dynamic detection head is used for image recognition and segmentation tasks.
5. A steel box girder crack detection method based on improved YOLOv7 as claimed in claim 4, characterized in that: The step (3) includes: combining the loss functions IoU, GIoU, DIoU, and CIoU to propose MPDIoU to directly minimize the distance between the upper left corner and the lower right corner of the predicted bounding box and the true bounding box. The calculation method is as follows: For A and B, define (x1 A ,y1 A ),(x2 A ,y2 A ) represents the coordinates of the upper left corner and the lower right corner of A, (x1 B , y1 B ),(x2 B ,y2 B ) represents the coordinates of the upper left corner and the lower right corner of B, as shown in equations (7)-(9): On this basis, a bounding box regression loss function based on MPDIoU is proposed, and the loss function is defined as follows:
Citation Information
Cited By
Road surface crack evaluation method and system based on improved deep learning model
CN121392483A