Pavement crack segmentation algorithm based on dynamic feature fusion and deformable space focusing

By using dynamic feature fusion and deformable spatial focus algorithms in road surface crack detection, the problems of low efficiency and poor accuracy of traditional detection methods are solved, and more efficient and accurate crack detection and recognition are achieved.

CN120070894APending Publication Date: 2025-05-30JIANGXI HIGHWAY RES & DESIGN INST CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510144905.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional pavement crack detection relies on manual inspection, which has problems such as low efficiency, poor accuracy and difficulty in dealing with complex backgrounds. In addition, traditional image processing methods are sensitive to noise and have poor robustness, making it difficult to achieve high-precision crack segmentation in complex environments.

Method used

The pavement crack segmentation algorithm based on dynamic feature fusion and deformable spatial focus is adopted to capture common features through pre-training branch transfer learning, and the specific features focused on the task are extracted using self-training, combined with deformable convolution for spatial focus, dynamically adjust the sampling position, and enhance generalization and recognition capabilities.

Benefits of technology

It realizes more precise positioning and identification of cracks in complex environments, reduces labor costs, improves detection efficiency and accuracy, and enhances the robustness and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070894A_ABST
    Figure CN120070894A_ABST
Patent Text Reader

Abstract

The invention provides a pavement crack segmentation algorithm based on dynamic feature fusion and deformable space focusing, and the algorithm comprises the steps: obtaining a crack image from a data set, constructing a segmentation network comprising an encoder and decoder architecture, dynamically adjusting the position of a sampling point through introducing deformable convolution, designing a dual-branch feature coding network, and carrying out the segmentation of a pavement crack through a dual-branch feature coding network. Wherein one branch is pre-trained on a large-scale data set, the other branch is self-trained, the dynamic feature fusion technology is utilized to integrate the features of the two branches, a jump connection mechanism is adopted to gradually recover the resolution of a feature map, a binary cross entropy loss and Dice loss joint optimization model is used, an Adam optimizer is used in the training process, and the resolution of the feature map is optimized. The effectiveness of the final algorithm model is verified through index evaluation and ablation experiments, and experimental results show that compared with other algorithms, the algorithm shows higher accuracy and robustness in a crack segmentation task, and is particularly suitable for pavement crack detection in a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of pavement crack segmentation and recognition, and in particular to a pavement crack segmentation algorithm based on dynamic feature fusion and deformable space focusing. Background Art

[0002] With the acceleration of urbanization, road traffic volume continues to increase, and pavement disease problems are becoming increasingly serious. Pavement cracks are one of the most common road diseases. If pavement cracks are not repaired in time in the early stage of formation, it may cause the situation to deteriorate further or even cause more serious damage, thereby triggering a series of complex chain reactions such as damage to the base layer and instability of the pavement structure. This not only greatly reduces the strength and durability of the road, but also increases the risk of traffic accidents, seriously endangering driving safety and affecting the traffic efficiency of the road. Regular crack detection can identify and deal with potential cracks in time, effectively prevent their further deterioration, thereby significantly improving the safety of the structure and reducing subsequent maintenance costs. It is of great significance to ensure the safety of public transportation and reduce social and economic costs.

[0003] Traditional crack detection mainly relies on manual inspection, which is easily affected by human factors, resulting in deviations in detection results and making it difficult to ensure its accuracy and consistency. Manual crack detection still has obvious shortcomings in efficiency, especially in large-scale road surface inspections, which often require a lot of manpower and time costs. In addition, when dealing with tiny cracks and certain dangerous or inaccessible areas, the limitations of manual crack detection become more obvious, and it is often unable to provide a comprehensive assessment of the crack condition. These shortcomings make it difficult for traditional manual crack detection to meet the actual application needs of modern traffic management.

[0004] Traditional image processing methods occupy an important position in the crack segmentation task. These methods usually rely on low-level features of images, such as edge information, texture patterns and grayscale distribution. By combining classic image processing techniques such as edge detection, threshold segmentation and region growing, the crack area can be extracted and separated. Although traditional segmentation methods have been widely used in crack detection, they still have certain limitations in actual scenarios. First, these methods are highly sensitive to noise, which may lead to unstable segmentation results. Second, in complex backgrounds or changing environments, traditional methods have poor robustness and it is difficult to effectively deal with the confusion between cracks and backgrounds. In addition, traditional methods perform poorly when dealing with large differences in crack morphology and size, and it is often difficult to achieve high-precision segmentation in various scenarios, especially when the crack features are fuzzy or irregular, the segmentation effect is even worse. Summary of the invention

[0005] The purpose of the present invention is to provide a pavement crack segmentation algorithm based on dynamic feature fusion and deformable space focusing.

[0006] The problems to be solved by the present invention are as follows: capturing general features through pre-trained branch transfer learning, extracting task-focused specific features using self-training, adjusting the weights of the features extracted by the pre-trained branch and the self-training branch through dynamic feature fusion adjustment, proposing deformable spatial focusing by combining deformable convolution, dynamically adjusting the sampling position by introducing learnable offset parameters, enhancing the generalization ability, balancing the contributions of general features and specific features, and improving the recognition and localization capabilities of various shaped cracks, and more accurately locating and identifying cracks.

[0007] The technical solution adopted by the pavement crack segmentation algorithm based on dynamic feature fusion and deformable spatial focusing is as follows: S1: Obtain crack image detection data, obtain crack images from the DeepCrack dataset, and divide the dataset into a training set and a test set; S2: Build a pavement crack segmentation network with dynamic feature fusion and deformable spatial focusing based on ResNet, adopt an encoder-decoder architecture, where the encoder is a dual-branch feature encoding network for pre-training and self-training, and the encoder-decoder architecture configures a deep learning development environment based on the PyTorch framework, including the PyTorch library, CUDA support, and numerical calculation and data processing tools, including NumPy and OpenCV; S3: Introduce an offset to adaptively adjust the position of the sampling point for deformable spatial focusing, use max pooling and average pooling to compress the channel dimension of the feature 1 obtained by self-training, concatenate the pooled feature 1 channel by channel, and perform convolution to obtain feature D, , where AvgPool is average pooling, MaxPool is max pooling, Conv represents convolution, and Concat represents the concatenation operation; Perform deformable convolution on feature 1 to extract features, concatenate feature D with the features extracted by deformable convolution, and calculate the attention weight through the Sigmoid activation function for the concatenated result , , , where represents the Sigmoid activation function, is , Def is deformable convolution, represents matrix multiplication, is the feature 2 obtained by self-training; S4: Process the feature 2 obtained by self-training according to the deformable spatial focusing described in S3 to finally obtain the self-training feature 3 ; S5: The pre-training and self-training dual-branch feature encoding network performs feature extraction to obtain the highest-level features extracted by the pre-training network and the self-training network, which are the pre-training feature 5 and the self-training feature 5 , the pre-training feature 5 is trained based on ResNet, and the self-training feature 5 is the self-training feature 3 obtained in S4 and is trained based on ResNet; S6: Based on the features extracted by the dual-branch, dynamic feature fusion is performed. For and a concatenation operation is carried out, and the result after the concatenation operation is added to and . Then, the number of channels of the added features is adjusted through a convolution operation to obtain the intermediate feature , , where represents matrix addition; Input into the Sigmoid function to generate an adaptive weight coefficient, and the weight coefficient is weighted and fused with and respectively to obtain the adaptive fusion feature . The fusion process is represented by the formula ; S7: Adopt the skip connection mechanism, and gradually upsample the feature map through the transposed convolution operation to map the extracted fusion feature into the segmentation mask of the road crack for crack segmentation; S8: Use the sum of the binary cross-entropy loss and the Dice loss as the overall loss function L of the algorithm model, , , , where N is the total number of pixels, is the true label of the i-th pixel, is the probability that the algorithm model predicts the i-th pixel as a positive sample; S9: Train the algorithm model. During the training process, use the Adam optimizer, set the initial learning rate to 1e-3 and decay to 1e-5, and train for a total of 100 epochs with a batch size of 16; S10: After the algorithm model training is completed, save the algorithm model weight file, read the algorithm model weight file and the test images, and use the evaluation metrics of P, R, F1, and MIoU to evaluate the segmentation results of different algorithm models on the test set. Here, P is the proportion of pixels predicted as cracks that are actually cracks, R is the proportion of crack pixels successfully predicted, F1 is a comprehensive evaluation metric that balances P and R, and MIoU measures the spatial overlap degree between the prediction result and the true label. The value range of MIoU is 0 - 1, and the closer MIoU is to 1, the better the overlap between the prediction result and the true label. S11: After the evaluation, design ablation experiments and comparative experiments using the control variable method to determine the contributions of the dynamic double-branch encoder and deformable spatial focusing to the final performance. S12: Read the captured crack images, and load the trained algorithm model weight file to perform crack segmentation on the images. Calculate the length, width, and area of the cracks by analyzing the number of pixels in the crack regions, and judge the severity, expansion trend, and potential impact range on the structural safety of the cracks. S13: Introduce an artificial verification mechanism to continuously optimize the crack segmentation algorithm model. Based on the segmentation results output by the algorithm model, conduct quantitative evaluation and feedback analysis on the recognition accuracy of different types of cracks. S14: Install the Python and Conda environments, and install the LabelMe tool through the pip command in the activated Conda virtual environment. Use LabelMe to annotate the crack regions in the collected images, and automatically save the annotation results as JSON format files to construct a dynamically iterative dataset. Continuously optimize the recognition ability of the algorithm model for crack types by continuously iterating the annotated data.

[0008] Further, obtaining the crack image detection data in S1 includes: The dataset of DeepCrack is a dataset of 537 annotated images, including crack images of different scales and scenes, which are JPG images in RGB format with an image resolution of 544 * 384. Among them, 377 images are used as the training set and 160 images are used as the test set.

[0009] Further, splicing the feature D with the feature extracted by deformable convolution in S3 includes: Deformable convolution dynamically adjusts the sampling positions according to the input data, learns the offsets in the network, moves the sampling points on the input feature map by the convolution kernel, and focuses on the regions and objects of interest. Deformable convolution introduces an offset in the regular lattice sampling , , where is the center position of the convolutional window on the input feature map, is the receptive field size and dilation rate, is the relative coordinate of the n-th position of the convolutional window with respect to the relative coordinate, is the convolutional kernel weight, is the pixel at the n-th position within the convolutional window.

[0010] Furthermore, the pre-training and self-training dual-branch feature encoding network in S5 performs feature extraction, including: The pre-trained network branch is used to extract general features including crack color, edges, and texture. The self-training branch integrated with deformable spatial focusing is used to extract detailed features of crack shapes. Dynamic feature fusion is employed to integrate the features of different branches, enabling information collaboration and optimization among features.

[0011] Furthermore, in S10, the evaluation metrics of P, R, F1, and MIoU are used to evaluate the segmentation results of different algorithm models on the test set, including: , , , ; where TP is the true positive, representing the number of pixels correctly predicted as cracks; FP is the false positive, representing the number of pixels wrongly predicted as cracks; FN is the false negative, representing the number of pixels wrongly predicted as the background; and mean represents taking the average.

[0012] Furthermore, in S11, the ablation experiment and the comparative experiment are designed using the control variable method, including: The dual-branch feature encoding network and the deformable spatial focusing are embedded into the baseline algorithm model and combined with the evaluation metrics to quantitatively analyze the impact of the dual-branch feature encoding network and the deformable spatial focusing on the performance of the algorithm model. All experiments are conducted in the same hardware environment, and the algorithm model parameters and training strategies are consistent during the ablation experiment; This algorithm model is compared with image segmentation methods such as UNet, UNet++, Attention UNet, DeepLab v3+, CE-Net, MET-Net, and DeepCrack. The comparative experiment uses P, R, F1, and MIoU as evaluation metrics. Based on the evaluation metrics of P, R, F1, and MIoU, the performance of this algorithm model and other image segmentation methods on the DeepCrack dataset is analyzed, and the performance of this algorithm model in pavement crack segmentation is evaluated based on the evaluation metrics of P, R, F1, and MIoU; Set a classification threshold to determine whether each pixel belongs to a crack or the background. Different threshold settings may lead to different segmentation results. Five groups of thresholds are set, and the robustness of the algorithm model is evaluated by analyzing the MIoU of the segmentation results obtained by different methods under these five groups of thresholds.

[0013] The beneficial effects of the present invention are as follows: Traditional crack detection usually relies on manual inspections, which is time-consuming and laborious, and is susceptible to factors such as weather and light. By using crack segmentation technology, all-weather and all-time monitoring can be achieved in an automated and intelligent manner, reducing labor costs. If pavement cracks are not repaired in a timely manner, they may deteriorate over time, forming potholes or other more serious damages, seriously affecting vehicle driving safety. The crack segmentation technology can detect cracks at the very beginning, timely discover potential safety hazards, and avoid accidents. Capture general features through pre-trained branch transfer learning, and use self-training to extract task-focused specific features, improving the comprehensiveness and richness of feature extraction, and enhancing the generalization ability of the algorithm model. Adjust the weights of the features extracted by the pre-trained branch and the self-trained branch through dynamic feature fusion, dynamically optimize the representation ability of crack features, effectively balance the contributions of general features and specific features, and thus improve the overall performance of the algorithm model. Combine deformable convolution to propose deformable spatial focusing. By introducing learnable offset parameters to dynamically adjust the sampling positions, the algorithm model can effectively represent the complex morphological features of cracks, and further enhance the model's recognition and localization abilities for various shaped cracks. The proposed method achieves good detection results on the DeepCrack dataset. Compared with other mainstream methods, the algorithm model can more accurately locate and identify cracks, demonstrating its potential and advantages in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a structural diagram of a pavement crack segmentation algorithm based on dynamic feature fusion and deformable spatial focusing; Figure 2 It is a structural diagram of dynamic feature fusion; Figure 3 It is a sample diagram of the DeepCrack dataset; Figure 4 It is a schematic diagram of deformable convolution; Figure 5 It is a structural diagram of deformable spatial focusing; Figure 6 It is a visualization display diagram of ablation experiments; Figure 7 It is a schematic diagram of a confusion matrix; Figure 8Visualization diagrams of pavement crack segmentation results for different methods. Detailed implementation manners

[0015] The present invention will be further described clearly and completely below, but the protection scope of the present invention is not limited thereto.

[0016] The technical solutions adopted by the pavement crack segmentation algorithm based on dynamic feature fusion and deformable spatial focusing are as follows: S1: Obtain crack image detection data, obtain crack images from the DeepCrack dataset, and divide the dataset into a training set and a test set; S2: Build a pavement crack segmentation network based on ResNet for dynamic feature fusion and deformable spatial focusing, adopt an encoder-decoder architecture, where the encoder is a dual-branch feature encoding network of pre-training and self-training, and the encoder-decoder architecture is configured with a deep learning development environment based on the PyTorch framework, including the PyTorch library, CUDA support, and numerical calculation and data processing tools, including NumPy and OpenCV; S3: Introduce an offset to adaptively adjust the positions of sampling points for deformable spatial focusing, use max pooling and average pooling to compress the channel dimension of the self-trained feature 1 of, splice the pooled feature 1 by channels, and perform convolution to obtain feature D, , where AvgPool is average pooling, MaxPool is max pooling, Conv represents convolution, and Concat represents the splicing operation; Perform deformable convolution on feature 1 to extract features, splice feature D with the features extracted by deformable convolution, and calculate the attention weights through the Sigmoid activation function for the splicing result , , , where represents the Sigmoid activation function, is , Def is deformable convolution, represents matrix multiplication, is the self-trained feature 2; S4: Process the self-trained feature 2 according to the deformable spatial focusing described in S3 to finally obtain the self-trained feature 3 ; S5: Perform feature extraction on the dual-branch feature encoding network of pre-training and self-training to obtain the highest-level features extracted by the pre-training network and the self-training network, which are the pre-trained feature 5 and the self-trained feature 5 , the pre-trained feature 5 Obtained by training based on ResNet, self-training feature 5 Is the self-training feature 3 obtained in S4 Obtained by training based on ResNet; S6: Based on the features extracted by the dual branches, perform dynamic feature fusion, and and Perform a concatenation operation, and add the result after the concatenation operation to and Then adjust the number of channels of the added features through a convolution operation to obtain the intermediate feature , , where Represents matrix addition; Input Into the Sigmoid function to generate an adaptive weight coefficient, and multiply the weight coefficient with and respectively for weighted fusion to obtain the adaptive fusion feature , and the fusion process is represented by the formula ; S7: Adopt the skip connection mechanism, gradually upsample the feature map through transposed convolution operation, map the extracted fusion feature into the segmentation mask of road cracks, and perform crack segmentation; S8: Use the sum of binary cross-entropy loss and Dice loss As the overall loss function L of the algorithm model, , , , where N is the total number of pixels, Is the true label of the i-th pixel, Is the probability that the algorithm model predicts the i-th pixel as a positive sample; S9: Train the algorithm model. During the training process, use the Adam optimizer, set the initial learning rate to 1e-3, and decay to 1e-5. Train for a total of 100 epochs, and the batch size is 16; S10: After the algorithm model is trained, save the algorithm model weight file, read the algorithm model weight file and test images, and use the evaluation metrics of P, R, F1, and MIoU to evaluate the segmentation results of different algorithm models on the test set. Among them, P is the proportion of pixels actually being cracks among all pixels predicted as cracks, R is the proportion of successfully predicted crack pixels, F1 is a comprehensive evaluation metric that balances P and R, and MIoU measures the spatial overlap degree between the prediction result and the true label. The value range of MIoU is 0-1, and the closer MIoU is to 1, the better the overlap between the prediction result and the true label; S11: After the evaluation, use the method of controlling variables to design ablation experiments and comparative experiments to determine the contributions of the dynamic double-branch encoder and deformable spatial focusing to the final performance; S12: Read the captured crack images, load the trained algorithm model weight file to segment the cracks in the images, calculate the length, width, and area of the cracks by analyzing the number of pixels in the crack regions, and judge the severity, expansion trend, and potential impact range on the structural safety of the cracks; S13: Introduce an artificial verification mechanism to continuously optimize the crack segmentation algorithm model, and conduct quantitative evaluation and feedback analysis on the recognition accuracy of different types of cracks based on the segmentation results output by the algorithm model; S14: Install the Python and Conda environments, install the LabelMe tool through the pip command in the activated Conda virtual environment, use LabelMe to annotate the crack regions in the collected images, automatically save the annotation results as JSON format files, construct a dynamically iterative dataset, and continuously optimize the recognition ability of the algorithm model for crack types by continuously iterating the annotated data.

[0017] Reference Figure 1 As shown, it is the structural diagram of the pavement crack segmentation algorithm based on dynamic feature fusion and deformable spatial focusing. Reference Figure 2 As shown, it is the structural diagram of dynamic feature fusion.

[0018] Furthermore, the acquisition of crack image detection data in S1 includes: The DeepCrack dataset is a dataset of 537 annotated images, including crack images of different scales and scenes, which are JPG images in RGB format, with an image resolution of 544*384. Among them, 377 images are used as the training set and 160 images are used as the test set.

[0019] Reference Figure 3 As shown, it is a sample diagram of the DeepCrack dataset.

[0020] Furthermore, the splicing of feature D and the feature extracted by deformable convolution in S3 includes: The deformable convolution dynamically adjusts the sampling positions according to the input data, learns the offsets in the network, moves the sampling points of the convolution kernel on the input feature map, and focuses on the regions and objects of interest; Deformable convolution introduces an offset in the regular lattice sampling , , where is the center position of the convolution window on the input feature map, is the receptive field size and dilation rate, is the relative coordinate of the nth position of the convolutional window with respect to , is the convolutional kernel weight, is the pixel at the nth position within the convolutional window.

[0021] Refer to Figure 4 as shown, it is a schematic diagram of deformable convolution. Refer to Figure 5 as shown, it is a structural diagram of deformable spatial focusing.

[0022] Furthermore, the pre-training and self-training dual-branch feature encoding network performs feature extraction, including: The pre-trained network branch is used to extract general features including crack color, edges, and texture. The self-training branch integrated with deformable spatial focusing is used to extract detailed features of crack shapes. Dynamic feature fusion is adopted to integrate the features of different branches, and information collaboration and optimization between features are carried out.

[0023] Furthermore, in the S10, evaluation metrics such as P, R, F1, and MIoU are used to evaluate the segmentation results of different algorithm models on the test set, including: , , , ; where TP is the true positive, representing the number of pixels correctly predicted as cracks; FP is the false positive, representing the number of pixels wrongly predicted as cracks; FN is the false negative, representing the number of pixels wrongly predicted as the background; mean represents taking the average value.

[0024] Furthermore, in the S11, the control variable method is used to design ablation experiments and comparative experiments, including: The dual-branch feature encoding network and deformable spatial focusing are embedded into the baseline algorithm model and combined with the evaluation metrics to quantitatively analyze the impact of the dual-branch feature encoding network and deformable spatial focusing on the performance of the algorithm model. All experiments are carried out in the same hardware environment, and the algorithm model parameters and training strategies are consistent during the ablation experiment process.

[0025] Table 1 Results of ablation experiments Dual-branch Feature Encoding Network Deformable Spatial Focus P R F1 MIoU w / o w / o 82.99 83.37 83.18 71.19 w w / o 82.73 86.14 84.40 73.05 w w 85.34 86.16 85.75 75.23 The ablation experiment results are shown in Table 1, where "w" means embedded into the algorithm model and "w / o" means not embedded into the algorithm model. By comparing the experimental results of the first and second rows, it can be found that the algorithm model combined with the dual-branch feature encoding network has certain improvements in all four metrics. Among them, the R metric has increased by 2.77%. By comparing the experimental results of the second and third rows, it can be found that the model embedded with deformable spatial focusing has further improvements in all four metrics. Among them, P and MIoU have increased by 2.61% and 2.18% respectively. This is because the dual-branch feature encoding network and deformable spatial focusing enhance the model's ability to represent pavement crack features and effectively extract the key features suitable for pavement crack segmentation, thus significantly improving the crack segmentation performance of the model. The ablation experiment shows that the proposed dual-branch feature encoding network and deformable spatial focusing can effectively improve the crack segmentation accuracy of the model, verifying the effectiveness of the dual-branch feature encoding network and deformable spatial focusing; Reference Figure 6 As shown, it is a visualization display diagram of the ablation experiment. The color and texture of the stone inside the crack are similar to the background, resulting in the baseline model misidentifying it as the background. In the segmentation mask, this misidentification is manifested as the black area inside the crack. Analyzing the segmentation results after combining the dual-branch feature encoding network and deformable spatial focusing, it can be seen that the area of the black area inside the crack in the segmentation mask has significantly decreased, and the segmentation performance has been significantly improved. Combining the confusion matrix for quantitative analysis of the segmentation results, reference Figure 7 As shown, it is a schematic diagram of the confusion matrix. It can be seen from the figure that after sequentially applying the dual-branch feature encoding network and deformable spatial focusing, the proportion of correctly classified pixels gradually increases, further verifying the effectiveness of the dual-branch feature encoding network and deformable spatial focusing.

[0026] Compare this algorithm model with image segmentation methods such as UNet, UNet++, Attention UNet, DeepLab v3+, CE-Net, MET-Net, and DeepCrack. The comparative experiment uses P, R, F1, and MIoU as evaluation metrics. Analyze the performance of this algorithm model and other image segmentation methods on the DeepCrack dataset based on the P, R, F1, and MIoU evaluation metrics, and evaluate the performance of this algorithm model in pavement crack segmentation based on the P, R, F1, and MIoU evaluation metrics.

[0027] Table 2 Comparative experiment results of different methods P R F1 MIoU UNet 82.04 82.05 82.04 69.39 UNet++ 80.93 85.48 83.14 70.64 Attention UNet 80.15 83.73 81.90 69.49 DeepLab v3+ 80.04 85.10 82.49 70.32 CE-Net 82.99 83.37 83.18 71.19 MET-Net 80.61 85.91 83.18 71.48 DeepCrack 83.09 83.98 83.53 72.41 This algorithm model 85.34 86.16 85.75 75.23 Table 2 shows the comparison results between the proposed algorithm model and other methods. The proposed algorithm model exhibits the best segmentation performance in the single evaluation metrics P and R. The P metric is 2.25% higher than the sub-optimal DeepCrack. This is because the dual-branch feature encoding network and deformable spatial focusing can enhance the model's ability to extract and represent pavement crack features, effectively improving the model's segmentation performance. F1 and MIoU can comprehensively evaluate the model's segmentation performance, and the proposed algorithm model also achieves the best results in these two comprehensive metrics. Compared with the sub-optimal DeepCrack, the proposed algorithm model improves by 2.22% and 2.83% in the F1 and MIoU metrics, respectively. The experimental results show that the proposed algorithm model has the best segmentation effect in the four metrics, verifying the effectiveness of the proposed algorithm model for segmenting pavement cracks.

[0028] Set the classification threshold to determine whether each pixel belongs to a crack or the background. Different threshold settings may lead to different segmentation results. Five groups of thresholds are set, and the robustness of the algorithm model is evaluated by analyzing the MIoU of the segmentation results obtained by different methods under these five groups of thresholds.

[0029] Table 3 MIoU of Different Methods under Different Thresholds 0.1 0.3 0.5 0.7 0.9 Avg SD UNet 69.38 69.48 69.39 69.20 68.62 69.21 0.3106 UNet++ 70.22 70.51 70.64 70.69 70.63 70.54 0.1696 Attention UNet 68.78 69.29 69.49 69.58 69.39 69.31 0.2803 DeepLab v3+ 69.74 70.23 70.32 70.19 69.54 70.00 0.3068 CE-Net 71.18 71.22 71.19 71.15 71.06 71.16 0.0548 MET-Net 70.46 71.19 71.48 71.64 71.39 71.23 0.4125 DeepCrack 72.43 72.48 72.41 72.22 71.84 72.28 0.2352 This algorithm model 75.19 75.26 75.23 75.23 75.14 75.21 0.0415 Table 3 presents the MIoU of the segmentation results of different methods under different thresholds. It can be found that as the threshold increases, the MIoU of each method shows a trend of first increasing and then decreasing. Avg represents the average value of MIoU under different thresholds, and SD represents the standard deviation of MIoU under different thresholds. The smaller the SD value, the more stable the model. The SD value of the proposed algorithm model is 0.0415, which is the lowest among all methods. This is because its dual-branch structure can extract stable and high-quality features from crack images, enabling the model to maintain high stability against threshold changes.

[0030] Reference Figure 8As shown in the figure, the pavement crack segmentation results of different methods are visualized. In the analysis of the first column of images, except for the algorithm model and MET-Net, all methods mistake the ruler in the background as a crack. Although MET-Net has a certain ability to resist background interference, there is a phenomenon of missed detection in some local areas. In the images in the second and third columns, the cracks are relatively small and have a low contrast with the background, resulting in poor crack segmentation performance of UNet, UNet++, Attention UNet, DeepLabv3+, CE-Net and MET-Net. These methods cannot accurately identify the complete crack contour, and some crack edges cannot even be segmented correctly. DeepCrack cannot accurately segment the crack area in the third column in the upper right corner of the image. The segmentation effect of this algorithm model in this area is also slightly insufficient. The crack image in the fourth column is blurred and the crack width is large. UNet mistakenly identifies the background area as a crack, resulting in unnecessary white spots in the segmentation results. DeepLab v3+, CE-Net, MET-Net and DeepCrack performed well in the task of segmenting wider cracks, but were still inferior to this algorithm model in the precise segmentation of crack edges. The above results show that this algorithm model can effectively and accurately segment cracks of different shapes in complex environments, further verifying the effectiveness and advancement of this algorithm model.

[0031] The present invention provides a pavement crack segmentation algorithm based on dynamic feature fusion and deformable spatial focusing. Crack images are obtained from a data set, and a segmentation network including an encoder and a decoder architecture is constructed. The sampling point positions are dynamically adjusted by introducing deformable convolution. A dual-branch feature encoding network is designed, in which one branch is pre-trained on a large-scale data set and the other branch is self-trained. The dynamic feature fusion technology is used to integrate the features of the two branches, and a skip connection mechanism is used to gradually restore the feature map resolution. The binary cross entropy loss and Dice loss are used to jointly optimize the model. The Adam optimizer is used in the training process. The final algorithm model is evaluated by indicators and its effectiveness is verified by ablation experiments. Experimental results show that compared with other algorithms, the algorithm shows higher accuracy and robustness in crack segmentation tasks, and is particularly suitable for pavement crack detection in complex scenarios.

Claims

1. A pavement crack segmentation algorithm based on dynamic feature fusion and deformable space focusing, characterized by: include: S1: Obtain crack image detection data, obtain crack images from the DeepCrack dataset, and divide the dataset into a training set and a test set; S2: A pavement crack segmentation network with dynamic feature fusion and deformable spatial focusing is built based on ResNet, using an encoder-decoder architecture. The encoder is a pre-trained and self-trained dual-branch feature encoding network. The encoder-decoder architecture is configured with a deep learning development environment based on the PyTorch framework, including the PyTorch library, CUDA support, and numerical computing and data processing tools, including NumPy and OpenCV. S3: Introduce offset to adaptively adjust the position of sampling points for deformable spatial focusing, and use maximum pooling and average pooling to self-trained feature 1 The channel dimension is compressed to make the pooled feature 1 Concatenate by channel and perform convolution to obtain feature D. , where AvgPool is average pooling, MaxPool is maximum pooling, Conv represents convolution, and Concat represents concatenation operation; For feature 1 Perform deformable convolution to extract features, concatenate feature D with the features extracted by deformable convolution, and calculate the attention weight of the concatenated result through the Sigmoid activation function , , ,in Represents the Sigmoid activation function, which is , Def is a deformable convolution, represents matrix multiplication, Feature 2 obtained from self-training; S4: Feature 2 obtained from self-training According to the deformable spatial focusing described in S3, the self-training feature 3 is finally obtained. ; S5: Pre-trained and self-trained dual-branch feature encoding networks are used for feature extraction to obtain the highest-level features extracted by the pre-trained network and the self-trained network, which are pre-trained feature 5 and self-training features 5 , pre-trained features 5 Based on ResNet training, self-training feature 5 is the self-training feature 3 obtained in S4 Based on ResNet training; S6: Dynamic feature fusion based on the features extracted by the two branches. and Perform a splicing operation and compare the result of the splicing operation with and Add, and then adjust the number of channels of the added features through convolution operation to obtain the intermediate features , ,in, represents matrix addition; Will Input to the Sigmoid function to generate adaptive weight coefficients, and the weight coefficients are respectively and Perform weighted fusion to obtain adaptive fusion features , the fusion process is expressed as express; S7: Using the skip connection mechanism, the feature map is gradually upsampled through the transposed convolution operation, and the extracted fusion features are mapped into the segmentation mask of the road cracks to perform crack segmentation; S8: Using binary cross entropy loss and Dice loss The sum is used as the overall loss function L of the algorithm model, , , , where N is the total number of pixels, is the true label of the i-th pixel, The algorithm model predicts the probability that the i-th pixel is a positive sample; S9: Train the algorithm model. The Adam optimizer is used during the training process. The initial learning rate is set to 1e-3 and decays to 1e-5. The total training period is 100 cycles, and the batch size is 16. S10: After the algorithm model training is completed, save the algorithm model weight file, read the algorithm model weight file and the test image, and use the P, R, F1 and MIoU evaluation indicators to evaluate the segmentation results of different algorithm models on the test set, where P is the proportion of all pixels predicted as cracks that are actually cracks, R is the proportion of crack pixels predicted successfully, F1 is a comprehensive evaluation indicator of the balance between P and R, and MIoU is a measure of the spatial overlap between the predicted result and the true label. The MIoU value range is 0-1. The closer the MIoU is to 1, the better the overlap between the predicted result and the true label. S11: After the evaluation is completed, the control variable method is used to design ablation experiments and comparative experiments to determine the contribution of the dynamic dual-branch encoder and deformable spatial focusing to the final performance; S12: Read the captured crack image and load the trained algorithm model weight file to segment the image into cracks. Calculate the length, width and area of ​​the crack by analyzing the number of pixels in the crack area to determine the severity of the crack, its expansion trend and its potential impact on the structural safety. S13: Introduce a manual verification mechanism to continuously optimize the crack segmentation algorithm model. Based on the segmentation results output by the algorithm model, conduct quantitative evaluation and feedback analysis on the recognition accuracy of different types of cracks. S14: Install Python and Conda environment, and install LabelMe tool through pip command in the activated Conda virtual environment. Use LabelMe to annotate the crack areas of the collected images, and automatically save the annotation results as JSON format files to build a dynamically iterative data set. By continuously iterating the annotated data, the algorithm model's ability to identify crack types is continuously optimized.

2. The pavement crack segmentation algorithm based on dynamic feature fusion and deformable space focusing as claimed in claim 1, characterized in that: The step of obtaining crack image detection data in S1 includes: The DeepCrack dataset is a dataset of 537 annotated images, including crack images of different scales and scenes. It is a JPG image in RGB format with an image resolution of 544*384, of which 377 images are used as training sets and 160 images are used as test sets.

3. The pavement crack segmentation algorithm based on dynamic feature fusion and deformable space focusing as claimed in claim 1, characterized in that: In S3, feature D is concatenated with features extracted by deformable convolution, including: Deformable convolution dynamically adjusts the sampling position according to the input data, learns the offset in the network, and moves the convolution kernel to the sampling point on the input feature map to focus on the area and object of interest; Deformable Convolution To introduce an offset in the regular lattice sampling , ,in is the center position of the convolution window on the input feature map, is the receptive field size and dilation rate, For the nth position of the convolution window The relative coordinates of is the convolution kernel weight, is the pixel at the nth position in the convolution window.

4. The pavement crack segmentation algorithm based on dynamic feature fusion and deformable space focusing as claimed in claim 1, characterized in that: The pre-trained and self-trained dual-branch feature encoding network in S5 performs feature extraction, including: The pre-trained network branch is used to extract general features including crack color, edge and texture. The self-training branch with integrated deformable spatial focusing is used to extract detailed features including crack shape. Dynamic feature fusion is used to integrate the features of different branches to coordinate and optimize information between features.

5. The pavement crack segmentation algorithm based on dynamic feature fusion and deformable space focusing as claimed in claim 1, characterized in that: In S10, P, R, F1 and MIoU evaluation indicators are used to evaluate the segmentation results of different algorithm models on the test set, including: , , , ; Among them, TP is true positive, which means the number of pixels correctly predicted as cracks, FP is false positive, which means the number of pixels incorrectly predicted as cracks, FN is false negative, which means the number of pixels incorrectly predicted as background, and mean means the average value.

6. The pavement crack segmentation algorithm based on dynamic feature fusion and deformable space focusing as claimed in claim 1, characterized in that: In S11, the control variable method is used to design ablation experiments and comparative experiments, including: The dual-branch feature encoding network and deformable space focusing are embedded in the baseline algorithm model and combined with evaluation indicators to quantitatively analyze the impact of the dual-branch feature encoding network and deformable space focusing on the performance of the algorithm model. All experiments are conducted under the same hardware environment, and the algorithm model parameters and training strategies are consistent during the ablation experiment. The algorithm model is compared with UNet, UNet++, Attention UNet, DeepLab v3+, CE-Net, MET-Net, and DeepCrack image segmentation methods. The comparative experiment uses P, R, F1, and MIoU as evaluation indicators. The performance of the algorithm model and other image segmentation methods on the DeepCrack dataset is analyzed based on the P, R, F1, and MIoU evaluation indicators. The performance of the algorithm model in pavement crack segmentation is evaluated based on the P, R, F1, and MIoU evaluation indicators. The classification threshold is set to determine whether each pixel belongs to the crack or the background. Five groups of thresholds are set. The robustness of the algorithm model is evaluated by analyzing the MIoU of the segmentation results obtained by different methods under the five groups of thresholds.

Citation Information

Cited By

  • Crack identification method and system based on image processing

    CN121190932A

  • Crack identification method and system based on image processing

    CN121190932B

  • Fixed ice linear deformation detection method based on satellite-borne InSAR and deep learning

    CN122313296A

  • Small crack detection and segmentation method based on sampling and geometric constraint

    CN122391244A