Lightweight and efficient pavement crack segmentation method
Through the improved U-Net architecture, combined with deep separation convolution and multi-branch grouping convolution, the problems of bulky model and blurred boundaries in road surface crack detection are solved, and lightweight and high-precision crack segmentation is achieved, suitable for mobile devices and complex environments.
Patent Information
- Application Number
- CN202510475784.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art has problems such as bulky model, missing multi-scale features and blurred boundaries in pavement crack detection, making it difficult to efficiently and accurately segment fractures on mobile devices.
Using an improved U-Net architecture, combining depth separation convolution, Hadamard product attention, multi-branch grouping convolution and channel attention, a lightweight multi-scale feature fusion and boundary optimization module is designed to optimize crack segmentation through joint loss functions.
It realizes efficient and accurate pavement crack segmentation on mobile devices, adapts to complex lighting and background conditions, and improves the positioning accuracy and segmentation effect of crack edges.
Smart Images

Figure CN120298701A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing and computer vision, and specifically to a pavement crack semantic segmentation method based on a lightweight neural network. It realizes efficient crack detection and precise segmentation through an improved U-Net architecture, and is applicable to mobile devices and real-time detection scenarios. Background Art
[0002] Pavement cracks are important hidden dangers affecting road safety. Traditional detection methods (such as manual inspection, threshold segmentation) have problems such as low efficiency and poor accuracy. Although semantic segmentation technologies based on deep learning (such as U-Net) have improved the detection ability, they have the following defects: 1. Bulky model: The stacking of traditional convolutional layers leads to a large number of parameters, making it difficult to deploy on mobile devices; 2. Lack of multi-scale features: A single convolutional kernel cannot capture cracks of different widths and shapes, and it is easy to miss fine cracks; 3. Blurred boundaries: The lack of an efficient feature aggregation mechanism results in unclear crack edge segmentation, especially in low-contrast or shadow scenarios where the performance drops significantly.
[0003] Existing lightweight models generally reduce parameters through depthwise separable convolutions, but sacrifice the feature expression ability and have poor crack detection effects in complex backgrounds. Therefore, there is an urgent need for a pavement crack segmentation method that combines lightweight and high precision. Summary of the Invention
[0004] The present invention aims to provide a lightweight, efficient and robust pavement crack segmentation method, which is realized through innovative module design: 1. Lightweight: Significantly reduce the model parameters and computational amount while ensuring accuracy; 2. Multi-scale feature fusion: Effectively capture the structural information of cracks of different scales; 3. Boundary optimization: Enhance the positioning accuracy of crack edges and adapt to complex lighting and background conditions.
[0005] This method is designed based on the U-Net architecture, Figure 1 The overall architecture design is given, Figures 2 to 4 The detailed design of the internal modules is given. The present invention realizes end-to-end crack segmentation through three core modules, and the specific steps are as follows:
[0006] Step 1: Image preprocessing and feature extraction (FEM module)
[0007] Input: 3-channel color pavement image ( is the batch size, is the image size).
[0008] Depthwise Separable Convolution (DWConv): Perform depthwise convolution on the input image and pointwise convolution to extract initial features , reducing the computational cost.
[0009] Hadamard Product Attention (HPA): 1. Initialize the learnable spatial weight tensor ; 2. Perform DWConv and ReLU activation on to generate the adjusted weight ; 3. Calculate (element-wise product) to highlight the features in the crack area.
[0010] Feature Fusion: Concatenate with the original input , output after DWConv, and finally fuse through residual connection as the encoder input.
[0011] Step 2: Multi-scale Feature Extraction and Downsampling (Encoder - MBM Module)
[0012] Multi-branch Group Convolution: Use 4 branches for the input feature map, and perform group convolution with dilation rates (the number of groups is of the number of input channels) to extract multi-scale features .
[0013] Channel Attention with Shared Parameters: Perform global average pooling (GAP) on each branch feature , generate channel weights through two-layer convolution (dimensionality reduction + dimensionality increase) and Sigmoid activation, and return the original features by element-wise product to enhance key channels.
[0014] Feature Fusion: Concatenate the 4 branch features, adjust the dimension to through pointwise convolution to generate ; at the same time, process the input feature through convolution as the skip connection , and finally output .
[0015] Downsampling: Implement 4-level downsampling through a max pooling layer with a stride of 2 and a pooling window size of . The size of the feature map is halved successively, and the number of channels is doubled.
[0016] Step 3: Feature Upsampling and Cross-Layer Fusion (Decoder - MBM Module)
[0017] Upsampling: The high-level features output by the encoder are magnified by a factor of 2 through bilinear interpolation and concatenated along the channel dimension with the shallow features (after being processed by MBM) at the corresponding encoder stage.
[0018] Multi-branch Feature Extraction: Similar to the MBM structure of the encoder, through grouped convolutions with different dilation rates and channel attention, the multi-scale features after upsampling are fused to gradually restore the spatial resolution.
[0019] Step 4: Feature Aggregation and Boundary Optimization (FAM Module)
[0020] Spatial Weight Generation: For the feature map output by the decoder, initialize a learnable tensor , through DWConv and ReLU activation, and then generate weights through Sigmoid .
[0021] Feature Weighting: Perform point convolution to adjust the number of channels to 1 for the output of the decoder, and perform element-wise multiplication with to generate the final segmentation probability map , highlighting the crack area and refining the boundary.
[0022] Step 5: Optimization of the Joint Loss Function
[0023] Simultaneously calculate the Binary Cross Entropy Loss (BCE Loss), Dice Loss, and mIoU Loss, and sum all the losses as the total loss of the model. The formula is as follows:
[0024]
[0025] where distinguish the pixel-level differences between cracks and the background, reduce the breaks in the crack area and improve the recall rate, optimize the overlap between the prediction and the ground truth label to ensure boundary accuracy. Description of the Drawings
[0026] Figure 1 is the complete architecture of the method, showing the complete process from image input to segmentation output, including the connection relationships of the FEM, MBM, and FAM modules.
[0027] Figure 2 is the detailed diagram of the FEM module, showing the specific operation steps of the Hadamard product attention and skip connections.
[0028] Figure 3It is the detailed diagram of the MBM module, showing the specific operation steps of multi-dilation rate grouped convolution, channel attention, and skip connection.
[0029] Figure 4 It is the feature weighting process of the FAM module, demonstrating the key steps of spatial weight tensor generation and element-wise multiplication.
[0030] Figure 5 It is the inference result of this method, representing the original image, ground truth, and crack segmentation result from left to right respectively. Specific implementation manners
[0031] In a specific example of applying the present invention, the dataset consists of road surface crack images taken by a mobile phone, including shadow parts under different lighting conditions. These data cover different scenarios, road surface backgrounds, crack types, and lighting conditions, which can more powerfully illustrate the method of the present invention.
[0032] I. Model parameter configuration
[0033] The input dimension and output dimension of FEM are set to 3 and 16 respectively, and the input dimension and output dimension of FAM are set to 8 and 1 respectively. The model altogether includes an encoder with four downsampling modules and a decoder with four upsampling modules. The input dimension and output dimension of the encoder are [16, 32, 64, 128] and [32, 64, 128, 256] respectively, and the input dimension and output dimension of the decoder are [256, 128, 64, 32] and [64, 32, 16, 8] respectively. Four evaluation metrics, namely Precision, Recall, F1-score, and mean Intersection over Union (mIOU), are used.
[0034] II. Training and inference steps
[0035] Data preprocessing: The input image is resized to 224×224, and data is augmented by random cropping, flipping, and brightness / contrast adjustment. The label is binarized, with the crack area being 1 and the background being 0.
[0036] Training process: The AdamW optimizer is used, with a learning rate of 0.01, a weight decay of 0.01, a batch size of 8, and training for 120 epochs. When the mIoU of the validation set does not improve for 10 consecutive epochs, the learning rate is decayed to 0.1 times.
[0037] Inference process: The input single image is processed by FEM, encoder, decoder, and FAM, and a 1-channel probability map is output. A threshold of 0.5 is applied to the probability map to generate a binary segmentation result, visualizing the crack area.
[0038] III. Experimental verification
[0039] The road surface crack images are split into a training set, a validation set, and a test set according to a ratio of 7:2:1. After obtaining the trained model through the above steps, the inference results on the test set are shown in Table 1, and the crack segmentation results are as Figure 5 shown.
[0040] The method described in the present invention can be extended to other scenarios such as building crack, medical image lesion segmentation, and industrial defect detection by replacing training data, appropriately adjusting the dilation rate of the MBM branch or the FAM weight tensor parameters, and reusing the backbone network through transfer learning and only adaptively training the end classification layer.
Claims
1. A lightweight and efficient pavement crack segmentation method, characterized in that It includes the following steps: (1) Feature extraction: Preprocess the input image through depthwise separable convolution and Hadamard product attention mechanism (HPA) to generate an enhanced feature map; (2) Multi-scale feature extraction: Use multi-branch grouped convolution (dilation rate 1-4) and shared channel attention to extract multi-scale context features in the encoder and decoder; (3) Feature aggregation and optimization: Weight the decoder output through a learnable spatial weight tensor to refine the crack boundary and generate the final segmentation map; (4) Joint loss optimization: Simultaneously use BCE loss, Dice loss, and mIoU loss to train the model to balance class imbalance and boundary accuracy.
2. The method according to claim 1, characterized in that In the multi-branch grouped convolution step, the number of groups in each branch is 1 / 4 of the number of input channels, and the multi-scale structure of the crack is captured through different dilation rates.
3. The method according to claim 1, characterized in that In the feature aggregation step, the spatial weight tensor is generated through depthwise separable convolution and Sigmoid activation to dynamically weight the features in the crack region.
4. The method according to claim 1, characterized in that The encoder and decoder adopt a symmetric structure, and fuse shallow and deep features through bilinear interpolation upsampling and cross-layer splicing.