An intelligent identification and measurement method for bridge cracks

The integration of YOLOv8 with a super-resolution image network and UNet, along with optimized loss functions, addresses the challenge of accurately reflecting crack features and measuring widths, achieving precise bridge crack detection and segmentation.

CN117911371BActive Publication Date: 2025-07-15SOUTHWEST JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410082136.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-19
Publication Date
2025-07-15
Estimated Expiration
2044-01-19

AI Technical Summary

Technical Problem

Although most existing image semantic segmentation algorithms can achieve high-precision image segmentation, they lack improvements to the input images, making it difficult to effectively reflect the characteristics of bridge cracks, and bridge crack identification and measurement lack of full process intelligence.

Method used

Transfer learning is performed using YOLOv8 network, combined with super-resolution image segmentation model and UNet network, through loss function optimization, high-precision identification and positioning of bridge cracks and pixel width measurement, and pixel width measurement is performed using the mixing method of the shortest distance method and the orthogonal framework method.

Benefits of technology

It realizes intelligent recognition of bridge cracks throughout the process, improves the accuracy of crack recognition and the accuracy of pixel width measurement, breaks down the information barrier between identification positioning, segmentation and measurement, and reduces misjudgment in complex backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117911371B_ABST
    Figure CN117911371B_ABST
Patent Text Reader

Abstract

The present invention proposes an intelligent identification and measurement method for bridge cracks, belonging to the technical field of bridge structure detection, to solve the problem that most existing image semantic segmentation algorithms fuse low-level and high-level semantic features through skip connections to achieve high-precision image segmentation, but few algorithms improve the input images. First, the bridge image to be processed is input into the target recognition and positioning model to identify and position the cracks in the image, and the bridge crack image is output. Then, the bridge crack image output in step 1 is input into the ultra-high definition image segmentation model to obtain the binary image of the bridge cracks. Finally, the pixel width of the binary image of the cracks obtained in step 2 is measured based on the hybrid method of the shortest distance method and the orthogonal skeleton method. The method of the present invention can more quickly and accurately identify and position the bridge cracks in a complex environment; and improves the problem of insufficient crack segmentation accuracy by optimizing the semantic segmentation model from the input end.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bridge structure detection, and particularly relates to a method for intelligent identification and measurement of bridge cracks. Background Art

[0002] Bridge cracks will reduce the resistance of the bridge structure and pose a major threat to the safety and durability of the bridge. Therefore, regularly inspecting bridge cracks, including their locations, widths, and development trends, is crucial for understanding the damage condition of the bridge.

[0003] With the rapid development of new technologies such as unmanned aerial vehicles and artificial intelligence, bridge detection is moving towards digitization and intelligence. Using computer vision technology to perform intelligent identification on bridge photos, locate cracks, and extract information has advantages such as significantly reducing the labor burden, improving accuracy and objectivity compared to traditional manual detection. The application of deep learning enables intelligent identification to handle various crack backgrounds and morphologies, which can be divided into two categories: object detection algorithms and semantic segmentation algorithms. As the latest algorithm in the YOLO series of object detection algorithms, YOLOv8 can achieve high-precision and fast identification and location of bridge cracks. However, its output results are difficult to reflect the characteristics of the cracks. Therefore, it is necessary to use a semantic segmentation algorithm to perform pixel labeling on the crack images obtained by YOLOv8 and extract the pixels belonging to the cracks. However, most image semantic segmentation algorithms achieve high-precision image segmentation by skipping and fusing low-level and high-level semantic features, but few algorithms improve the input images. Summary of the Invention

[0004] In view of this, the present invention provides a method for intelligent identification and measurement of bridge cracks to solve the problem that most existing image semantic segmentation algorithms achieve high-precision image segmentation by skipping and fusing low-level and high-level semantic features, but few algorithms improve the input images.

[0005] The technical solution adopted by the present invention is as follows:

[0006] A method for intelligent identification and measurement of bridge cracks includes the following steps:

[0007] Step 1: Input the bridge crack dataset into the YOLOv8 network for transfer learning to obtain a target recognition and location model. Input the bridge image to be processed into the target recognition and location model to identify and locate the cracks in the image, and output the bridge crack image;

[0008] The specific steps of Step 1 are as follows:

[0009] Step 1.1: Divide the bridge crack dataset into a training set and a validation set, preprocess the bridge crack data in the training set and the validation set, and then input them into the YOLOv8 network for transfer learning;

[0010] Step 1.1 specifically includes the following steps:

[0011] Step 1.11: Obtain open-source concrete surface crack images and crack images collected by a crack magnifier to obtain a bridge crack dataset. Preprocess the bridge crack dataset through image scaling and image normalization, and enhance the bridge crack dataset using image flipping and image rotation;

[0012] Step 1.12: Divide the bridge crack dataset processed in Step 1.11 into a training and validation set and a test set in a ratio of 9:1. Subsequently, divide the training and validation set into a training set and a validation set again in a ratio of 9:1, and then input the training set and the validation set into the YOLOv8 network for transfer learning.

[0013] Step 1.2: The YOLOv8 network extracts image features from the bridge crack data input in Step 1.1 based on the backbone network, neck network, and prediction end, performs object detection, obtains the recognition and localization results of the cracks, and outputs the bridge crack image.

[0014] The backbone network is the core of feature extraction. It performs downsampling on the image using convolution, constructs the feature extraction structure of the model using the C2f module to obtain the key feature information of the image, and uses the SPPF module to achieve fast extraction and splicing of multi-scale feature information. The neck network utilizes the C2f module to strengthen the feature extraction ability of the model to achieve better feature fusion. The prediction end is the output part of the model, and it adopts the Decoupled-Head mechanism that decouples localization and classification, so as to more flexibly process semantic information of different scales.

[0015] Step 2: Establish an ultra-high-definition image segmentation model, input the bridge crack image output in Step 1 into the ultra-high-definition image segmentation model to obtain a binary image of the bridge crack;

[0016] The establishment of the ultra-high-definition image segmentation model in Step 2 specifically includes the following steps:

[0017] Step 2.1: Construct a super-resolution image network that enhances image features. The super-resolution image network extracts features from the bridge crack image using 3×3 convolution, reduces the dimension of the bridge crack image through 1×1 convolution, and then inputs the features after dimension reduction into the projection unit, and continuously alternates upward and downward projections to output high-definition and low-definition feature maps. Fuse all the high-definition feature maps generated by the upward projection units, and reconstruct it through 3×3 convolution to obtain the final high-definition crack image;

[0018] Among them, the up-projection unit receives the low-resolution feature maps output by all previous down-projection units, fuses and deconvolves them to obtain high-resolution feature maps; while the down-projection unit receives the high-resolution feature maps output by all previous up-projection units, extracts deep-level image features, and outputs low-resolution feature maps through convolution. In the up-projection unit, using e l obtains the residual value between low-resolution images, deconvolves it, and finally passes through Figure 4 the "+", in it, adds it to the backed-up high-resolution features to output the final high-resolution image; similarly, in the down-projection unit, eh represents the residual value between high-resolution images, and finally outputs a low-resolution image.

[0019] Step 2.2: Construct a semantic segmentation network UNet, and perform binary segmentation on the high-resolution crack images in Step 2.1 through the encoder, decoder and skip connection structure of the UNet network to obtain crack binary images.

[0020] In Step 2.2, the encoder is composed of a VGG16 network. The 5 convolutional layers of the encoder respectively use 64, 128, 256, 512, and 512 convolutional kernels, the size of the convolutional kernels is 3×3, and the activation function is ReLU;

[0021] The decoder is composed of 5 deconvolutional layers. The size of the convolutional kernels of each deconvolutional layer is 3×3, and the activation function is ReLU;

[0022] The skip connection structure connects the corresponding layers of the encoder and the decoder to fuse the high-level semantic information extracted by the encoder and the low-level texture information extracted by the decoder.

[0023] The 5 convolutional layers of the encoder are respectively responsible for extracting different levels of features of the image. Among them, the first layer extracts the edge and texture features of the image; the second layer extracts the local structural features of the image; the third layer extracts the semantic features of the image; the fourth layer extracts the semantic detail features of the image; the fifth layer extracts the global semantic features of the image. The 5 deconvolutional layers of the decoder are responsible for reconstructing the features extracted by the encoder.

[0024] The ultra-high-resolution image segmentation model is jointly composed of a super-resolution image network and an image semantic segmentation network through a loss function. The ultra-high-resolution image segmentation model continuously improves the weight coefficients through the backpropagation of the loss. Since the positive and negative samples in the crack images are extremely unbalanced, taking the distance between the predicted boundary contour of the crack and the boundary contour of the manually annotated label as the boundary loss helps to weaken the impact of the positive and negative sample imbalance on the model performance during the semantic segmentation process. Among them, the change in the distance between the predicted crack boundary and the label boundary can be expressed as:

[0025]

[0026] Wherein, and respectively represent the contours of the label image and the predicted image, p is a point on , represents the point on corresponding to the normal direction of point p on and The enclosed area is denoted as ΔS, q represents an arbitrary point in space, and D G (q) is the distance from point q to the nearest point on and can be expressed as:

[0027]

[0028] According to the change of the boundary distance, φ G (q) = -D G (q), s(q) and g(q) are contour indicator functions. When q falls in the spatial domain S, s(q) is 1, otherwise it is 0; similarly, when q falls in the spatial domain G, g(q) is 1, otherwise it is 0. Accordingly, let S = S θ , s(q) = s θ (q) to represent the probability of pixel classification, and the boundary loss function is obtained:

[0029]

[0030] The boundary loss may fall into a local minimum due to class imbalance. Therefore, the regional loss Generalized Dice Loss (GDL) is introduced and combined to form the loss function:

[0031]

[0032] In the above formula, the initial value of α is 0.01, but its value will increase by 0.01 for each training round. When training the model, the influence of the super-resolution technology on the crack input image needs to be considered. Here, the L1 loss is introduced to describe this error:

[0033]

[0034] Here, I SR and I HR respectively represent the reconstructed super-resolution image and the original high-definition image, and I is the set of high-definition image pixel points. Through the weight parameter, the semantic segmentation loss function and the super-resolution loss function are combined by the weight coefficient β ∈ [0, 1] to obtain the overall loss function of the model:

[0035]

[0036] Step 2.4: Collect the concrete semantic segmentation crack images and their labels shared on the network. Divide the training and validation sets into a training set and a validation set at a ratio of 9:1. Input the training set and the validation set into the super-resolution image segmentation model, evaluate the quality of the super-resolution image by the peak signal-to-noise ratio and the structural similarity index, and evaluate the accuracy of the image segmentation result by the intersection over union of the images.

[0037] Step 3: Measure the pixel width of the binary image of the crack obtained in Step 2 based on a hybrid method of the shortest distance method and the orthogonal skeleton method.

[0038] The specific steps of Step 3 are as follows:

[0039] Step 3.1: Denoise the connected regions of the binary image of the crack obtained in Step 2, eliminate the non-crack connected regions, then use the medial axis transformation method to extract the single-pixel skeleton of the denoised binary image of the crack, and perform pruning to remove the incorrect skeleton information. Finally, obtain the crack edge information through the Canny operator and the morphological medial axis transformation.

[0040] Step 3.2: Based on the crack edge and the single-pixel skeleton obtained in Step 3.1, use a constant kernel to find the orthogonal vector of the single-pixel skeleton, project the edge points of the crack edge onto the orthogonal vector, find the edge points with the projection coefficient greater than a given threshold, and divide the edge points greater than the given threshold into two groups, positive and negative, according to the orientation. Screen the two points with the smallest distance in the positive and negative groups. The Euclidean distance between the two points with the smallest distance is the pixel width of the crack.

[0041] The specific steps of Step 3.2 are as follows:

[0042] Step 3.21: Based on the obtained crack edge and single-pixel skeleton, the normal vector between the crack edge and the single-pixel skeleton is obtained as:

[0043] Bn = B - skel

[0044] where B is the local boundary point, skel is the adjacent skeleton point, and Bn is the vector group of the local boundary point relative to the skeleton point;

[0045] Step 3.32: Obtain the orthogonal vector y of the crack skeleton through the principal component analysis method, and deduce the projection coefficient of the local boundary point on the orthogonal vector y: P = Bn·y T , and introduce a balance factor γ∈(0,1) to divide the required boundary points into two candidate groups, positive and negative, as follows:

[0046]

[0047] Step 3.33: Finally, through the shortest distance method, screen the two points with the smallest distance in the candidate group points C1 and C2 to obtain the pixel width of the crack:

[0048] Step 3.34: After obtaining the pixel width of the crack, through pixel calibration, that is, obtaining the true width information of the crack, the physical size corresponding to each pixel in the image is obtained from the following formula:

[0049] W pp = 10W d / W f P c

[0050] where W pp is the physical size corresponding to a single pixel, W d is the shooting distance of the camera, W f is the focal length of the camera, and P c is the number of pixels contained in 1 cm of the camera's photosensitive element.

[0051] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are as follows:

[0052] 1. The present invention makes full use of the YOLOv8 algorithm to achieve high-precision and fast bridge crack recognition and positioning, solves the problem that the YOLOv8 algorithm is difficult to reflect the characteristics of cracks through the super-resolution image segmentation model, and finally measures the pixel width of bridge cracks based on the hybrid method of the shortest distance method and the orthogonal skeleton method, thus breaking the information barrier among crack recognition and positioning, crack segmentation, and pixel width measurement, and realizing the full-process intelligent recognition of bridge cracks.

[0053] 2. The present invention uses the trained YOLOv8 model to extract the features of the input image and perform object detection to obtain the recognition and positioning results.

[0054] 3. The present invention uses super-resolution image technology to enhance the features of the crack image, extracts the high-level semantics of the crack image through the VGG16 network, uses the shallow feature extraction network to extract the position contour information of the crack, and then fuses through skip connections, reducing the misjudgment of the model for complex background pixels and improving the accuracy of crack semantic segmentation. In addition, the present invention improves the loss function, optimizes and trains the model network parameters using boundary loss, reducing the problem of imbalance between positive and negative samples in the crack image. Through the joint training of the super-resolution network and the UNet network, the semantic segmentation performance of the model is improved from the input end.

[0055] 4. The present invention combines two crack pixel width measurement methods, the shortest distance method and the orthogonal skeleton method. While considering the crack propagation direction, it avoids the incorrect selection of edge points to positions far from the crack, thus realizing more accurate measurement of the crack pixel width. Description of the Drawings

[0056] The present invention will be described by way of examples and with reference to the accompanying drawings, where:

[0057] Figure 1 is a schematic flowchart of the intelligent identification and measurement method for bridge cracks based on YOLOv8 and ultra-high-definition image segmentation of the present invention;

[0058] Figure 2 is a schematic diagram of the YOLOv8 network structure;

[0059] Figure 3 is a schematic diagram of the ultra-high-definition image segmentation network structure;

[0060] Figure 4 is a schematic diagram of the upper projection unit and the lower projection unit structure;

[0061] Figure 5 is a schematic diagram of the UNet network structure;

[0062] Figure 6 is a flowchart for measuring the pixel width of cracks. Detailed implementation manners

[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated in the drawings here can be arranged and designed in various different configurations.

[0064] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents the selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0065] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0066] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0067] In the present invention, unless otherwise clearly specified and defined, the first feature being "above" or "below" the second feature may include the direct contact between the first and second features, or may include the situation where the first and second features are not in direct contact but in contact through additional features therebetween. Moreover, the first feature being "above", "over" and "on top of" the second feature includes that the first feature is directly above and obliquely above the second feature, or merely indicates that the horizontal height of the first feature is higher than that of the second feature. The first feature being "below", "beneath" and "underneath" the second feature includes that the first feature is directly below and obliquely below the second feature, or merely indicates that the horizontal height of the first feature is lower than that of the second feature.

[0068] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0069] Embodiment 1

[0070] As Figures 1-6 shown, an intelligent identification and measurement method for bridge cracks is disclosed in the embodiment of the present invention, including the following steps:

[0071] Step 1: Input the bridge crack data set into the YOLOv8 network for transfer learning to obtain a target recognition and positioning model, input the bridge image to be processed into the target recognition and positioning model, identify and position the cracks in the image, and output the bridge crack image;

[0072] The specific steps of the said Step 1 include the following steps:

[0073] Step 1.1: Divide the bridge crack data set into a training set and a validation set, preprocess the bridge crack data in the training set and the validation set, and then input them into the YOLOv8 network for transfer learning;

[0074] The specific steps of the said Step 1.1 include the following steps:

[0075] Step 1.11: Obtain the open-source concrete surface crack images and the crack images collected by a crack magnifier to get the bridge crack data set, preprocess the bridge crack data set through image scaling and image normalization, and enhance the bridge crack data set by image flipping and image rotation;

[0076] Step 1.12: Divide the bridge crack data set processed in Step 1.11 into a training and validation set and a test set at a ratio of 9:1, then divide the training and validation set into a training set and a validation set again at a ratio of 9:1, and then input the training set and the validation set into the YOLOv8 network for transfer learning.

[0077] Step 1.2: The YOLOv8 network extracts image features from the bridge crack data input in Step 1.1 based on the backbone network, neck network, and prediction head, performs object detection, obtains the recognition and localization results of the cracks, and outputs the bridge crack image.

[0078] The backbone network is the core of feature extraction. It downsamples the image using convolution, constructs the feature extraction structure of the model using the C2f module to obtain the key feature information of the image, and uses the SPPF module to achieve fast extraction and splicing of multi-scale feature information. The neck network utilizes the C2f module to enhance the feature extraction ability of the model to achieve better feature fusion. The prediction head is the output part of the model, which adopts the Decoupled-Head mechanism that decouples localization and classification, thus handling semantic information at different scales more flexibly.

[0079] Step 2: Establish an ultra-high-definition image segmentation model, input the bridge crack image output in Step 1 into the ultra-high-definition image segmentation model to obtain the binary image of the bridge crack;

[0080] The establishment of the ultra-high-definition image segmentation model described in Step 2 specifically includes the following steps:

[0081] Step 2.1: Construct a super-resolution image network that enhances image features. The super-resolution image network extracts features from the bridge crack image using 3×3 convolution, reduces the dimension of the bridge crack image through 1×1 convolution, and then inputs the dimension-reduced features into the projection unit, continuously alternating upward and downward projections to output high-definition and low-definition feature maps. Among them, the up-projection unit accepts all previous low-definition feature maps, fuses and deconvolves them to obtain high-definition feature maps; the down-projection unit accepts all high-definition feature maps, obtains deep image features, and outputs low-definition feature maps through convolution. In the up-projection unit, use e l to obtain the residual value between low-definition images, perform deconvolution on it, and finally through Figure 4 the “+” in, add it to the backed-up high-definition features to output the final high-definition image; similarly, in the down-projection unit, e h represents the residual value between high-definition images, and finally outputs the low-definition image. Fuse all the high-definition feature maps generated by the up-projection units, and reconstruct them through 3×3 convolution to obtain the final high-definition crack image.

[0082] Step 2.2: Construct a semantic segmentation network UNet, and perform binary segmentation on the high-definition crack image in Step 2.1 through the encoder, decoder, and skip connection structure of the UNet network to obtain the crack binary image.

[0083] In step 2.2, the encoder is composed of a VGG16 network. The five convolutional layers of the encoder respectively use 64, 128, 256, 512, and 512 convolutional kernels, the size of the convolutional kernels is 3×3, and the activation function is ReLU;

[0084] The decoder is composed of five transposed convolutional layers. The size of the convolutional kernels of each transposed convolutional layer is 3×3, and the activation function is ReLU;

[0085] The skip connection structure connects the corresponding layers of the encoder and the decoder to fuse the high-level semantic information extracted by the encoder and the low-level texture information extracted by the decoder.

[0086] The five convolutional layers of the encoder are respectively responsible for extracting different levels of features of the image. Among them, the first layer extracts the edge and texture features of the image; the second layer extracts the local structural features of the image; the third layer extracts the semantic features of the image; the fourth layer extracts the semantic detail features of the image; the fifth layer extracts the global semantic features of the image. The five transposed convolutional layers of the decoder are responsible for reconstructing the features extracted by the encoder.

[0087] Step 2.3: The ultra-high definition image segmentation model is jointly composed of a super-resolution image network and an image semantic segmentation network through a loss function. The ultra-high definition image segmentation model continuously improves the weight coefficients through the backpropagation of the loss. Since the positive and negative samples in the crack image are extremely unbalanced, taking the distance between the predicted crack boundary contour and the boundary contour of the manually annotated crack label as the boundary loss helps to weaken the impact of the positive and negative sample imbalance on the model performance during the semantic segmentation process. Among them, the change in the distance between the predicted crack boundary and the label boundary can be expressed as:

[0088]

[0089] In the formula, and respectively represent the contours of the label image and the predicted image, p is a point on , represents the point on corresponding to the normal direction of the p point on and The enclosed area, q represents an arbitrary point in space, D G (q) is the distance from the q point to the nearest point on , which can be expressed as:

[0090]

[0091] According to the change in the boundary distance, φ G (q) = -D G(q), s(q) and g(q) are contour indicator functions. When q falls within the spatial domain S, s(q) is 1, otherwise it is 0; similarly, when q falls within the spatial domain G, g(q) is 1, otherwise it is 0. Accordingly, let S = S θ , s(q) = s θ (q) to represent the probability of pixel classification, and obtain the boundary loss function:

[0092]

[0093] The boundary loss may fall into a local minimum due to class imbalance. Therefore, the regional loss generalized Dice loss (GDL) is introduced and combined to form the loss function:

[0094]

[0095] In the above formula, the initial value of α is 0.01, but its value will increase by 0.01 for each training round. When training the model, the impact of super-resolution technology on the input crack images needs to be considered. Here, the L1 loss is introduced to describe this error:

[0096]

[0097] Here, I SR and I HR represent the reconstructed super-resolution image and the original high-definition image respectively, and I is the set of high-definition image pixel points. Through the weight parameter, the semantic segmentation loss function and the super-resolution loss function are combined by the weight coefficient β ∈ [0, 1] to obtain the overall loss function of the model:

[0098]

[0099] Step 2.4: Collect the concrete semantic segmentation crack images and their labels shared by the network. And divide the training and validation sets into a training set and a validation set at a ratio of 9:1. Input the training set and the validation set into the super-resolution image segmentation model, evaluate the quality of the super-resolution image by the peak signal-to-noise ratio and the structural similarity index, and evaluate the accuracy of the image segmentation result by the intersection over union of the images.

[0100] Step 3: Measure the pixel width of the binary image of the crack obtained in Step 2 based on the hybrid method of the shortest distance method and the orthogonal skeleton method.

[0101] The specific steps of Step 3 include the following steps:

[0102] Step 3.1: Denoise the binary crack image obtained in Step 2 to eliminate non-crack connected regions. Then, use the medial axis transformation method to extract the single-pixel skeleton of the denoised binary crack image, and perform pruning to remove incorrect skeleton information. Finally, obtain the crack edge information through the Canny operator and morphological medial axis transformation;

[0103] Step 3.2: Based on the crack edge and single-pixel skeleton obtained in Step 3.1, use a constant kernel to find the orthogonal vector of the single-pixel skeleton, project the edge points of the crack edge onto the orthogonal vector, find the edge points with projection coefficients greater than a given threshold, and divide the edge points greater than the given threshold into positive and negative groups according to the orientation. Screen the two points with the smallest distance in the positive and negative groups. The Euclidean distance between the two points with the smallest distance is the pixel width of the crack.

[0104] The specific steps of Step 3.2 are as follows:

[0105] Step 3.21: Based on the obtained crack edge and single-pixel skeleton, the normal vector between the crack edge and the single-pixel skeleton is obtained as:

[0106] Bn = B - skel

[0107] where B is the local boundary point, skel is the adjacent skeleton point, and Bn is the vector group of the local boundary point relative to the skeleton point;

[0108] Step 3.32: Obtain the orthogonal vector y of the crack skeleton through the principal component analysis method, and deduce the projection coefficient of the local boundary point on the orthogonal vector y: P = Bn · y T , and introduce a balance factor γ ∈ (0, 1) to divide the required boundary points into positive and negative candidate groups, as follows:

[0109]

[0110] Step 3.33: Finally, through the shortest distance method, screen the two points with the smallest distance in the candidate group points C1 and C2 to obtain the pixel width of the crack:

[0111] Step 3.34: After obtaining the pixel width of the crack, through pixel calibration, that is, obtaining the true width information of the crack, the physical size corresponding to each pixel in the image is obtained from the following formula:

[0112] W pp = 10W d / W f P c

[0113] where W pp is the physical size corresponding to a single pixel, W d is the shooting distance of the camera, Wf is the camera focal length, P c is the number of pixels per 1 cm of the camera's photosensitive element.

[0114] The circuits, electronic components, and modules involved are all prior art, which can be fully implemented by those skilled in the art without further elaboration. The content protected by the present invention does not involve improvements to software and methods either.

[0115] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.

[0116] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. An intelligent identification and measurement method for bridge cracks, characterized in that, It includes the following steps: Step 1: Input the bridge crack dataset into the YOLOv8 network for transfer learning to obtain a target recognition and localization model. Input the bridge image to be processed into the target recognition and localization model to identify and locate the cracks in the image, and output the bridge crack image; Step 2: Establish an ultra-high-definition image segmentation model. Input the bridge crack image output in Step 1 into the ultra-high-definition image segmentation model to obtain a binary image of the bridge cracks; The ultra-high-definition image segmentation model in Step 2 is jointly composed of a super-resolution image network and an image semantic segmentation network through a loss function, including: Step A: Use the L1 loss function to calculate the error generated by the super-resolution network; Step B: Introduce the boundary loss function to optimize the problem of unbalanced positive and negative samples in the semantic segmentation process; Step C: Use the generalized Dice loss to avoid the problem of falling into local minima due to class imbalance; The establishment of the ultra-high-definition image segmentation model in Step 2 specifically includes the following steps: Step 2.1: Construct a super-resolution image network that enhances image features. The super-resolution image network uses 3×3 convolution to extract features from the bridge crack image, reduces the dimension of the bridge crack image through 1×1 convolution, and then inputs the dimension-reduced features into the projection unit. High-definition and low-definition feature maps are continuously alternately output through the up-projection unit and the down-projection unit. The high-definition feature maps generated by all up-projection units are fused and reconstructed through 3×3 convolution to obtain the final high-definition crack image; Among them, the up-projection unit fuses the low-definition feature maps output by all previous down-projection units and performs deconvolution on them to obtain high-definition feature maps; the down-projection unit fuses the high-definition feature maps output by all previous up-projection units, obtains deep-level image features, and outputs low-definition feature maps through convolution; Step 2.2: Construct a semantic segmentation network UNet. Perform binary segmentation on the high-definition crack image in Step 2.1 through the encoder, decoder, and skip connection structure of the UNet network to obtain a crack binary image. Step 3: Measure the pixel width of the binary image of the cracks obtained in Step 2 based on a hybrid method of the shortest distance method and the orthogonal skeleton method. After obtaining the pixel width of the cracks, obtain the true width information of the cracks through pixel calibration. The physical size corresponding to each pixel in the image is obtained from the following formula: W pp = 10W d / W f P c Among them, W pp is the physical size corresponding to a single pixel, W d is the shooting distance of the camera, W f is the focal length of the camera, P c is the number of pixels contained in 1 cm of the camera's photosensitive element.

2. The intelligent identification and measurement method for bridge cracks according to claim 1, characterized in that The specific steps of Step 1 include the following steps: Step 1.1: Divide the bridge crack dataset into a training set and a validation set. Preprocess the bridge crack data in the training set and the validation set and input them into the YOLOv8 network for transfer learning; Step 1.2: The YOLOv8 network extracts image features from the bridge crack data input in Step 1.1 based on the backbone network, neck network, and prediction end, performs target detection, obtains the recognition and localization results of the cracks, and outputs the bridge crack image.

3. A method for intelligent identification and measurement of bridge cracks according to claim 2, characterized in that, The specific steps of Step 1.1 include the following steps: Step 1.11: Obtain the open-source concrete surface crack images and the crack images collected by a crack magnifier to obtain a bridge crack dataset. Preprocess the bridge crack dataset through image scaling and image normalization, and enhance the bridge crack dataset using image flipping and image rotation. Step 1.12: Divide the bridge crack dataset processed in Step 1.11 into a training and validation set and a test set at a ratio of 9:

1. Subsequently, divide the training and validation set into a training set and a validation set again at a ratio of 9:1, and then input the training set and the validation set into the YOLOv8 network for transfer learning.

4. A method for intelligent identification and measurement of bridge cracks according to claim 1, characterized in that, In Step 2.2, the encoder is composed of a VGG16 network. The 5 convolutional layers of the encoder use 64, 128, 256, 512, and 512 convolutional kernels respectively, the size of the convolutional kernel is 3×3, and the activation function is ReLU. The decoder is composed of 5 transposed convolutional layers. The size of the convolutional kernel of each transposed convolutional layer is 3×3, and the activation function is ReLU. The skip connection structure connects the corresponding layers of the encoder and the decoder to fuse the high-level semantic information extracted by the encoder and the low-level texture information extracted by the decoder.

5. A method for intelligent identification and measurement of bridge cracks according to claim 1, characterized in that The specific steps of Step 3 are as follows: Step 3.1: Perform connected component denoising on the binary crack image obtained in Step 2 to eliminate non-crack connected components. Then, use the medial axis transformation method to extract the single-pixel skeleton of the denoised binary crack image and perform pruning to remove incorrect skeleton information. Finally, obtain the crack edge information through the Canny operator and morphological medial axis transformation. Step 3.2: Based on the crack edge and the single-pixel skeleton obtained in Step 3.1, use a constant kernel to find the orthogonal vector of the single-pixel skeleton, project the edge points of the crack edge onto the orthogonal vector, find the edge points with a projection coefficient greater than a given threshold, and divide the edge points greater than the given threshold into positive and negative groups according to the orientation. Screen the two points with the smallest distance in the positive and negative groups. The Euclidean distance between the two points with the smallest distance is the pixel width of the crack.

6. The intelligent identification and measurement method for bridge cracks according to claim 5, characterized in that, The specific steps of Step 3.2 are as follows: Step 3.21: Based on the obtained crack edge and single-pixel skeleton, the normal vector between the crack edge and the single-pixel skeleton is expressed as: Bn = B - skel where B is the local boundary point, skel is the adjacent skeleton point, and Bn is the vector group of the local boundary point relative to the skeleton point. Step 3.32: Obtain the orthogonal vector y of the crack skeleton through the principal component analysis method, and deduce the projection coefficient of the local boundary point on the orthogonal vector y: P = Bn·y T , and introduce a balance factor γ∈(0,1) to divide the required boundary points into two candidate groups of positive and negative, as follows: Step 3.33: By the shortest distance method, screen the two points with the smallest distance among the candidate group points C1 and C2 to obtain the pixel width of the crack: After obtaining the pixel width of the crack, through pixel calibration, obtain the true width information of the crack. The physical size corresponding to each pixel in the image is obtained by the following formula: W pp = 10W d / W f P c Among them, W pp is the physical size corresponding to a single pixel, W d is the shooting distance of the camera, W f is the focal length of the camera, P c is the number of pixels contained in 1 cm of the camera's photosensitive element.

Citation Information

Patent Citations

  • Point cloud pre-training method based on shielding modeling

    CN116109763A

  • Pavement crack detection method for improving UNet network structure

    CN116309485A