Crack detection method based on double-branch network structure

By adopting a dual-branch network structure and a cross-layer attention fusion module in road surface crack detection, the existing algorithm ignores the error detection and missed detection problems caused by crack pixel correlation, achieving higher robustness and generalization.

CN120182237APending Publication Date: 2025-06-20CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510328860.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing automated pavement crack detection algorithm ignores the correlation between crack pixels, resulting in problems such as mis-checking, missed detection and poor robustness.

Method used

The fracture detection method based on the dual-branch network structure is adopted, through the combination of deep network branches and shallow network branches, the cross-layer attention fusion module, bar convolution layer and feature fusion module are used to extract the edge information, detail information and semantic features of the cracks, and the loss function is constructed for model training.

Benefits of technology

It improves the generalization and robustness of crack detection, reduces missed detection, and enhances the robust detection ability of crack characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182237A_ABST
    Figure CN120182237A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of pavement detection, in particular to a crack detection method based on a double-branch network structure, and the method comprises the steps: obtaining an original pavement crack image training data set; training the crack detection model by using the original pavement crack image training data set; and inputting a to-be-detected pavement crack image into the trained crack detection model to obtain a crack detection graph, using a double-branch structure as a backbone network to extract crack features, using a cross-layer attention fusion module to extract detail information of a low-layer network and semantic information of a high-layer network, and obtaining the crack detection graph. An existing network model used for crack detection is effectively improved, generalization and robustness of crack detection are improved, and the problems that according to an existing automatic pavement crack detection algorithm, correlation between crack pixels is ignored, wrong detection and missing detection exist, and robustness is not high are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of road surface detection, and particularly to a crack detection method based on a dual-branch network structure. Background Art

[0002] The mileage of highways in China shows an increasing trend year by year, and the maintenance work is becoming more and more arduous. In the past, highway maintenance work completely relied on manual inspections. This method requires closing the road, causing traffic jams on the highway, and may increase the unsafe factors in the maintenance work. Traditional manual inspections mainly rely on technologies such as laser scanning, radar detection, acoustic detection, and computed tomography. These technologies require professional personnel to operate complex instruments to achieve. Traditional manual inspections have a large workload, low detection efficiency, and rely on the subjectivity of the participating personnel, which is not conducive to objective evaluation and large-scale intelligent and efficient applications. Therefore, it is necessary to implement a safer, more convenient, more efficient, greener, and more economical automated road surface crack detection algorithm.

[0003] With the development of digital image technology, some automated road surface crack detection algorithms have emerged, such as algorithms based on wavelet transform, image threshold, minimum path, edge detection, and percolation model. However, these algorithms are usually applicable to simple background scenarios with uniform illumination, no occlusion, no shadow, etc., and ignore the correlation between crack pixels, resulting in problems such as false detection, missed detection, and weak robustness. Therefore, road surface crack detection algorithms still need to be further studied and improved to be applicable to more complex scenarios. Summary of the Invention

[0004] The purpose of the present invention is to provide a crack detection method based on a dual-branch network structure, aiming to solve the problems of the current automated road surface crack detection algorithms, which ignore the correlation between crack pixels, have false detection, missed detection, and weak robustness.

[0005] To achieve the above purpose, the present invention provides a crack detection method based on a dual-branch network structure, including the following steps:

[0006] Obtain the original road surface crack image training data set;

[0007] Use the original road surface crack image training data set to train the crack detection model;

[0008] Input the road surface crack image to be detected into the trained crack detection model to obtain a crack detection map.

[0009] Among them, the original road surface crack image training data set includes the original road surface crack image and the corresponding road surface crack detection map label.

[0010] Among them, the crack detection model includes a dual-branch network structure, a cross-layer attention fusion module, a strip convolution layer, and a feature fusion module.

[0011] Among them, the specific method for training the crack detection model using the original road surface crack image training dataset is as follows:

[0012] Input the original road surface crack image into both the deep network branch and the shallow network branch of the dual-branch network simultaneously;

[0013] The low-level network of the deep network branch retains the edge information of the crack and obtains 4 low-level features;

[0014] The high-level network of the deep network branch extracts semantic features and obtains 4 high-level features;

[0015] Input the 4 low-level features and the 4 high-level features into the cross-layer attention fusion module for fusion to obtain 4 side outputs of the deep network branch;

[0016] The strip convolution layer of the shallow network branch retains the crack detail information and obtains 4 side outputs of the shallow network branch;

[0017] Input the 4 side outputs of the deep network branch and the 4 side outputs of the shallow network branch into the feature fusion module together to obtain a crack detection prediction map;

[0018] Construct a loss function for the crack detection model based on the crack detection prediction map, the original road surface crack image, and the corresponding crack detection map label, and update the parameters of the crack detection model with the minimum loss function value as the optimization goal to complete the training of the crack detection model.

[0019] Among them, in the cross-layer attention fusion module, the input original features first perform a 1×1 convolution operation, then are processed by an activation function to obtain attention weights, the attention weights are multiplied by the original features for fusion to obtain a feature matrix with attention weights, the feature matrix with attention weights then performs a 1×1 convolution operation, normalization, activation function, and 1×1 convolution operation to capture the mutual relationship between feature channels, and finally is added to the original features for fusion to obtain features with detailed information.

[0020] A crack detection method based on a dual-branch network structure of the present invention acquires a training dataset of original pavement crack images; trains a crack detection model using the training dataset of original pavement crack images; inputs a pavement crack image to be detected into the trained crack detection model to obtain a crack detection map. This method adopts an end-to-end trainable network architecture for pavement crack detection, and the backbone of this network architecture is a dual-branch network. The deep branch network has four layers of encoders, changes the direct mapping from the encoder to the decoder of the U-shaped network structure, and through a layer-by-layer feature fusion strategy based on a cross-layer attention mechanism, gives full play to the characteristics of each network layer, reduces the feature mapping gap between network layers, and improves the utilization of the obtained features. Fuses the four side outputs of the deep branch network and the four side outputs of the shallow branch network to obtain richer features, and realizes a more robust detection of crack features. This method uses a dual-branch structure as the backbone network to extract crack features, and uses a cross-layer attention fusion module to extract the detailed information of the low-level network and the semantic information of the high-level network, effectively improving the existing network model for crack detection, improving the generalization and robustness of crack detection, and solving the problems of the current automated pavement crack detection algorithm, such as ignoring the correlation between crack pixels, having misdetection, missed detection, and weak robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0022] Figure 1 It is a schematic structural diagram of the crack detection model of the present invention.

[0023] Figure 2 It is a schematic structural diagram of the cross-layer attention fusion module of the present invention.

[0024] Figure 3 It is a schematic structural diagram of the strip convolution layer of the present invention.

[0025] Figure 4 It is a schematic structural diagram of the feature fusion module of the present invention.

[0026] Figure 5 It is a flowchart of a crack detection method based on a dual-branch network structure provided by the present invention.

[0027] Figure 6 It is a flowchart of the specific method for training the crack detection model using the training dataset of original pavement crack images. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as a limitation of the present invention.

[0029] Please refer to Figures 1 to 6 , in a first aspect, the present invention provides a crack detection method based on a dual-branch network structure, including the following steps:

[0030] S1 Obtain the original road surface crack image training dataset;

[0031] In an embodiment of the present invention, the original road surface crack image training dataset includes original road surface crack images and corresponding road surface crack detection map labels.

[0032] In specific implementation, the original road surface crack image training datasets used are DeepCrack, Crack500, and EdmCrack600. The DeepCrack dataset includes 537 road surface crack images, with the resolution of each image being 384×544 pixels. This dataset includes different types of cracks, such as exposed cracks, thick cracks, and dirty cracks. Crack500 includes 3368 complex road surface cracks, with the resolution of each image being 600×360. EdmCrack600 includes 600 crack images with a resolution of 1920×1080. This dataset covers factors such as different weather conditions, different lighting conditions, shadows of other objects, and texture differences between different road surfaces. When implementing the present invention in network training experiments, the sizes of all sample images and ground truth label mask maps are adjusted to 256×256. Among them, 300 images in the DeepCrack dataset are used for model training, and the remaining 237 images are used for testing. 1896 images in the Crack500 dataset are used for model training, 348 images are used for model validation, and the remaining 1124 images are used for testing. 420 images in the EdmCrack600 dataset are used for model training, 60 images are used for model validation, and the remaining 120 images are used for testing.

[0033] S2 Use the original road surface crack image training dataset to train the crack detection model;

[0034] In an embodiment of the present invention, the crack detection model includes a dual-branch network structure, a cross-layer attention fusion module, a strip convolutional layer, and a feature fusion module.

[0035] Specific method:

[0036] S21 Input the original road surface crack image into both the deep network branch and the shallow network branch of the dual-branch network simultaneously;

[0037] S22 The low-level network of the deep network branch retains the edge information of the crack and obtains 4 low-level features;

[0038] In the embodiment of the present invention, the low-level network of the deep network branch retains the edge information of the crack and obtains 4 low-level features L1, L2, L3, and L4.

[0039] S23 The high-level network of the deep network branch extracts semantic features and obtains 4 high-level features;

[0040] In the embodiment of the present invention, the high-level network of the deep network branch extracts semantic features and obtains 4 high-level features H1, H2, H3, and H4.

[0041] S24 Input the 4 low-level features and the 4 high-level features into the cross-layer attention fusion module for fusion to obtain 4 side outputs of the deep network branch;

[0042] In the embodiment of the present invention, input the 4 low-level features L1, L2, L3, L4 and the 4 high-level features H1, H2, H3, H4 into the cross-layer attention fusion module for fusion to obtain four side outputs S1, S2, S3, S4 of the deep network branch; please refer to Figure 2 , the cross-layer attention fusion module extracts the rich spatial information of the low-level network and the channel information of the high-level network, improves the full utilization of the characteristics of different network layers by fusing the two kinds of information, reduces the loss and redundancy of crack feature information, and thus realizes more robust detection of crack features. The low-level network of the deep network has rich spatial information. The low-level network features first perform a 1×1 convolution operation, and then are processed by a Sigmoid activation function to obtain attention weights. The attention weights are multiplied by the original feature x for fusion to obtain a feature matrix with attention weights. The feature matrix with attention weights then performs a 1×1 convolution operation, BatchNormalization normalization, ReLU activation function, and a 1×1 convolution operation to capture the mutual relationship between feature channels, and finally is added and fused with the original input feature to obtain a feature with detailed information. The high-level network has rich channel information. The high-level network features first perform an average pooling operation along the channel, obtain channel attention through a 3×3 convolution operation, dimension reshaping, and a Sigmoid activation function, multiply the channel attention by the original feature for fusion to obtain a weighted channel matrix, and finally add and fuse the channel feature map with the original feature to obtain a feature with semantic information. Add and fuse the feature with detailed information and the feature with semantic information to realize more refined extraction of crack features.

[0043] S25: the strip convolution layer of the shallow network branch retains crack detail information to obtain four side outputs of the shallow network branch;

[0044] In the embodiment of the present invention, please refer to Figure 3 The strip convolution layer uses 3×1, 1×3 strip convolution operations instead of ordinary 3×3 convolution operations, and connects three strip convolutions with kernels of 3, 5, and 7 in parallel to extract multi-scale features. In order to retain the details and semantic information in the original input data and make it easier for information to propagate to the following layers, residual connections are used to avoid information loss. The multi-scale information extracted by the strip convolution layer and the residual jump information are added to obtain richer feature information. The nonlinear ability of the network is increased by the ReLU activation function, and finally the number of channels is increased by the 1×1 convolution layer.

[0045] S26: inputting the four side outputs of the deep network branch and the four side outputs of the shallow network branch into the feature fusion module to obtain a crack detection prediction map;

[0046] In the embodiment of the present invention, please refer to Figure 4 The feature fusion module reduces the number of channels of the four side output features of the shallow network branch through a 1×1 convolution operation to retain complete feature information as much as possible; the four side output features of the deep network branch are first upsampled to restore the feature resolution, and then the weights of the deep network output features are obtained by splicing along the channels and Softmax normalization operations. Finally, the weights are multiplied and fused with the corresponding features, and the number of channels is reduced through a 1×1 convolution operation to obtain features with richer semantics. The features extracted by the two branch networks are fused layer by layer, and then spliced ​​along the channels and 1×1 convolution operations are performed to achieve better detection results.

[0047] S27 constructs the loss function of the crack detection model based on the crack detection prediction map, the original pavement crack image and the corresponding crack detection map label, and updates the parameters of the crack detection model with the minimum loss function value as the optimization goal to complete the training of the crack detection model.

[0048] In the embodiment of the present invention, according to the crack detection prediction feature map P pre , the original pavement crack image P and the crack detection map label corresponding to the original pavement crack image are used to construct the loss function of the crack detection model. The parameters of the crack detection model are updated with the minimum loss function value as the optimization goal to complete the training of the crack detection model.

[0049] S3 inputs the road surface crack image to be detected into the trained crack detection model to obtain a crack detection map.

[0050] In the embodiment of the present invention, in the cross-layer attention fusion module, the input feature x first undergoes a 1×1 convolution operation, and then is processed by a Sigmoid activation function to obtain an attention weight. The attention weight is multiplied by the original feature x for fusion to obtain a feature matrix with attention weight. The feature matrix with attention weight then undergoes a 1×1 convolution operation, BatchNormalization normalization, a ReLU activation function, and a 1×1 convolution operation to capture the mutual relationship between feature channels, and finally is added to the original feature x for fusion to obtain a feature x1 with detailed information.

[0051] The input feature x undergoes average pooling operation along the channel, and through a 3×3 convolution operation, dimension reshaping, and a Sigmoid activation function to obtain channel attention. The channel attention is multiplied by the original feature for fusion to obtain a channel matrix with weight, and finally the channel feature map is added to the original feature for fusion to obtain a feature x2 with semantic information.

[0052] The feature x1 with detailed information and the feature x2 with semantic information are added for fusion to obtain a more refined feature out. out undergoes a 1×1 convolution operation to obtain a side output S i (i = 1, 2, 3, 4). out undergoes a 1×1 convolution operation and upsampling to obtain the input output of the next cross-layer attention fusion module i (i = 1, 2, 3, 4).

[0053] The loss function of the crack detection model includes:

[0054]

[0055] Among them, the training set is (X n , Y n ), n = 1,..., N, X n is a crack image, Y n is a label image, and all parameters of the network are represented as W. m represents the m-th side output layer, δ m represents the weight hyperparameter of the m-th output layer, Pre = {Pre i , i = 1,..., |X|}, and Pre i ∈{0, 1} represents the prediction result of the m-th output layer.

[0056] Δ is the cross-entropy loss function, which is expressed as:

[0057]

[0058] Among them, is the weight of crack pixels, W1 = 1 is the weight of non-crack pixels, G + is the total number of crack pixels in the input image, G -is the total number of non-crack pixels in the input image, and P is the probability that a pixel is predicted as a crack or a non-crack pixel. P i = 1 represents the probability that pixel ii is a crack, and P i = 0 represents the probability that pixel ii is a non-crack.

[0059] After splicing and fusing the outputs of each side, a fused prediction is obtained, and the fused loss function is expressed as:

[0060]

[0061] The dual-branch network consists of a deep network branch and a shallow network branch. In order to make the semantic information extracted by the deep network richer, the total loss only calculates the side output loss function of the deep network branch. The total loss function is expressed as:

[0062] L = L fuse (X, Y, W) + L fuse1 (X, Y, W) + L fuse2 (X, Y, W) + L side (X, Y, W)

[0063] Among them, L fuse (·) represents the feature fusion loss function of the deep network branch and the shallow network branch, and L fuse1 (·) represents the side output fusion loss function of the deep network branch, and L fuse2 (·) represents the side output fusion loss function of the shallow network branch, and L side (·) represents the side output loss function of the deep network branch.

[0064] Furthermore, the crack detection model is implemented using the Pytorch framework and trained on a Tesla V100 GPU with 32GB of video memory. The crack detection model requires 200 epochs to complete training, and the batchsize for each training is set to 1 (set to 4 for the Crack500 dataset). The learning rate policy is reduced by 10 times after every 40 epochs. The SGD optimizer with a momentum of 0.9 and a weight decay of 0.0002 is used as the optimization algorithm to update the network parameters. In the training phase and the testing phase, the size of all input images is adjusted to 256×256. The data augmentation strategy adopted is random rotation, vertical flipping, and horizontal flipping.

[0065] The method adopts an end-to-end trainable network architecture for pavement crack detection, and the backbone of this network architecture is a dual-branch network. The deep branch network has four layers of encoders, which changes the direct mapping from the encoder to the decoder of the U-shaped network structure. Through a layer-by-layer feature fusion strategy based on a cross-layer attention mechanism, the characteristics of each network layer are fully utilized, the feature mapping gap between network layers is reduced, and the utilization of the obtained features is improved. The four side outputs of the deep branch network and the four side outputs of the shallow branch network are fused to obtain richer features, realizing a more robust detection of crack features. This method uses a dual-branch structure as the backbone network to extract crack features, and uses a cross-layer attention fusion module to extract the detailed information of the low-level network and the semantic information of the high-level network, effectively improving the existing network model for crack detection, enhancing the generalization and robustness of crack detection, and solving the problems of the current automated pavement crack detection algorithm, such as ignoring the correlation between crack pixels, having false detections, missed detections, and weak robustness.

[0066] In a second aspect, the present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the crack detection method based on a dual-branch network structure as described above.

[0067] In a third aspect, the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it causes the processor to execute the crack detection method based on a dual-branch network structure as described above.

[0068] The above-disclosed is only a preferred embodiment of the crack detection method based on a dual-branch network structure of the present invention. Of course, it cannot be used to limit the scope of the rights of the present invention. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.

Claims

1. A crack detection method based on a double-branch network structure, characterized in that: The following steps are involved: Obtaining an original pavement crack image training dataset; Training a crack detection model using the original pavement crack image training data set; The pavement crack image to be detected is input into the trained crack detection model to obtain a crack detection map.

2. The crack detection method based on the double-branch network structure according to claim 1, characterized in that ; The original pavement crack image training data set includes original pavement crack images and corresponding pavement crack detection image labels.

3. The crack detection method based on the double-branch network structure as claimed in claim 2, characterized in that; The crack detection model includes a dual-branch network structure, a cross-layer attention fusion module, a strip convolution layer and a feature fusion module.

4. The crack detection method based on the double-branch network structure as claimed in claim 3, It is characterized by: The specific method of training the crack detection model using the original pavement crack image training data set is: Inputting the original pavement crack image into the deep network branch and the shallow network branch of the dual-branch network simultaneously; The lower layer network of the deep network branch retains the edge information of the crack and obtains 4 low-level features; The high-level network of the deep network branch extracts semantic features to obtain 4 high-level features; Inputting the four low-level features and the four high-level features into the cross-layer attention fusion module for fusion to obtain four side outputs of the deep network branch; The strip convolution layer of the shallow network branch retains crack detail information to obtain four side outputs of the shallow network branch; Inputting the four side outputs of the deep network branch and the four side outputs of the shallow network branch into the feature fusion module to obtain a crack detection prediction map; A loss function of the crack detection model is constructed based on the crack detection prediction map, the original pavement crack image and the corresponding crack detection map label, and the parameters of the crack detection model are updated with the minimum value of the loss function as the optimization goal to complete the training of the crack detection model.

5. The crack detection method based on the double-branch network structure as claimed in claim 2, It is characterized by: The original features input into the cross-layer attention fusion module are first subjected to a 1×1 convolution operation, and then processed through an activation function to obtain attention weights. The attention weights are multiplied and fused with the original features to obtain a feature matrix with attention weights. The feature matrix with attention weights is then subjected to a 1×1 convolution operation, normalization, activation function, and 1×1 convolution operation to capture the relationship between feature channels, and finally added and fused with the original features to obtain features with detailed information.