A real-time road surface strip repair detection method based on an improved lightweight MANet
Patent Information
- Application Number
- CN202610990561.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-11
AI Technical Summary
[0013]本发明的目的在于提供一种基于改进的轻量化MANet的实时路面条状修补检测方法,旨在解决现有技术在条状修补的语义分割任务上存在参数量和计算量大、推理速度较慢且分割精度较低的技术问题
Smart Images

Figure CN122737620A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a real-time road surface patch detection method based on an improved lightweight MANet, belonging to the field of road surface patch detection technology. Background Technology
[0002] Asphalt pavement has a vast mileage and wide range of applications in highway construction. However, due to changing environmental factors and the duration of traffic, pavement defects such as cracks and potholes continuously develop over time. This not only significantly degrades the quality of the asphalt pavement itself but also poses challenges to road traffic safety. Strip repair is one of the most common types of defects on asphalt pavements, characterized by its elongated shape and irregular size, which significantly affects the overall consistency of the pavement. Given the limited length of road mileage in my country, manual inspection would be extremely resource-intensive. Therefore, research on automated target detection and semantic segmentation for strip repair is currently a hot topic.
[0003] Existing technology and existing problems:
[0004] Currently, automatic segmentation methods for asphalt pavement patch repair mainly include methods based on a combination of traditional image processing and machine learning, instance segmentation methods based on deep learning, and semantic segmentation methods based on deep learning.
[0005] The main ideas of each method are as follows:
[0006] (1) A method combining traditional image processing and machine learning: This method first uses traditional image processing methods such as clustering and segmentation to handle the complex lighting and background interference in road surface damage images, and then uses machine learning methods to further segment the optimized images. This method combines traditional and advanced technical solutions and can achieve relatively good results in specific scenarios.
[0007] (2) Instance segmentation method based on deep learning: A typical example of this method is Mask R-CNN, which improves upon the object detection algorithm Faster R-CNN to perform pixel-level segmentation of the object in the original object detection result box. This method can not only segment road surface defects, but also further distinguish the individual defects and their quantities.
[0008] (3) Semantic segmentation methods based on deep learning: Typical examples of this method include DeepLabV3+ and SegNet. They generally consist of an encoder and a decoder, which perform feature extraction, downsampling, feature fusion, and upsampling respectively. This method can segment targets belonging to the same category in an image pixel by pixel. This method can simultaneously identify the type of road surface defects and segment the target.
[0009] Problems with existing detection methods:
[0010] (1) Existing problems of combining traditional image processing with machine learning: Although traditional image processing methods have a long history of development, their generalization ability is relatively average. They may perform well in some scenarios but perform extremely poorly in others. In addition, this method is not an end-to-end integrated technical solution. It usually requires step-by-step processing, which is relatively complex and cumbersome.
[0011] (2) Existing problems of deep learning-based instance segmentation methods: Since instance segmentation needs to solve the tasks of segmentation and differentiation of individuals at the same time, the model needs to be responsible for and implement many functions, so its accuracy and precision are generally low and it is difficult to meet the requirements of engineering applications.
[0012] (3) Existing problems of deep learning-based semantic segmentation methods: This method requires pixel-by-pixel segmentation of the target. However, existing public semantic segmentation models are usually difficult to balance segmentation accuracy and inference speed in the segmentation task of strip patching. Some models even have low segmentation accuracy while having a huge number of parameters and computational load, and the network structure lacks targeted optimization. Summary of the Invention
[0013] The purpose of this invention is to provide a real-time road surface patch detection method based on an improved lightweight MANet, which aims to solve the technical problems of large number of parameters and computation, slow inference speed and low segmentation accuracy in the semantic segmentation task of patch.
[0014] To achieve the above objectives, the technical solution of this invention is: a real-time road surface patch detection method based on an improved lightweight MANet. This method improves upon and proposes a lightweight MANet, which is superior to the benchmark model in all aspects. The lightweight MANet can achieve ideal lightweighting and significantly improve inference speed while maintaining high segmentation accuracy, thereby achieving an optimized balance of multiple indicators. This is beneficial for achieving the goal of real-time and high-precision detection of patch surface repairs. The method includes the following steps:
[0015] Step 1: Construct a semantic segmentation dataset for strip repair of road surface defects;
[0016] Step 2: Improve the baseline MANet to obtain a lightweight MANet; wherein, the improvement is to replace the backbone, increase the number of channels and add depthwise convolution;
[0017] Step 3: Train the lightweight MANet;
[0018] Step 4: Input the semantic segmentation dataset into the trained lightweight MANet to obtain the road surface patch detection results.
[0019] Optionally, Step 1 specifically includes:
[0020] Step 1.1: Obtain grayscale images of strip repairs for road surface defects with uniform dimensions;
[0021] Step 1.2: Use semantic segmentation and annotation software to draw polygonal labels for the strip repair of road surface defects, and obtain a mask file corresponding to the strip repair image;
[0022] Step 1.3: Merge the striped patch images and the corresponding mask files into the same dataset, and divide them into training set, validation set and test set according to the preset ratio.
[0023] Optionally, the replacement trunk specifically refers to:
[0024] The original MANet backbone ResNet-50 was discarded and replaced with a ResNet-34 backbone. The ResNet-34 backbone adopts a basic residual block structure, which contains two consecutive 3×3 convolutions with two convolution kernels.
[0025] Optionally, the increased number of channels specifically refers to:
[0026] In the Decoder Block, a constituent unit of the baseline MANet decoder, the two 1×1 convolutions at the beginning and end remain unchanged, while the number of channels of the 3×3 transposed convolution in the middle, which is responsible for upsampling the feature map, is increased by a preset factor.
[0027] Optionally, the addition of depthwise convolution specifically refers to:
[0028] In the baseline MANet decoder, an additional 3×3 depthwise convolution is added after each feature fusion stage, with the number of groups equal to the number of channels. The expression for the feature fusion stage of the depthwise convolution is as follows:
[0029]
[0030] In the formula, output is the output feature map, DWConv3×3 is the additional 3×3 depthwise convolution, and FusionFeat is the fused feature map.
[0031] Optionally, Step 3 specifically includes:
[0032] Step 3.1: Set the hyperparameters for training the lightweight MANet, including the number of iterations, batch size, number of threads, floating-point precision selection, and number of segmentation categories;
[0033] Step 3.2: Select the road surface patching image prediction mask of the benchmark MANet and the lightweight MANet for comparison, and train the lightweight MANet model according to the hyperparameters.
[0034] The beneficial effects of this invention are as follows: This invention improves the MANet model network and performs semantic segmentation on road surface patch repairs in three aspects: replacing the feature extraction backbone with a lighter one, increasing the number of channels in the transposed convolution in the Decoder Block, and adding depthwise convolutions to the feature fusion part of the decoder to further extract features, resulting in a lightweight MANet. Compared with the benchmark model, the lightweight MANet achieves significant improvements in parameter quantity, computational cost, inference speed, and segmentation accuracy, which helps to deploy it on devices with limited computing resources, such as mobile phones or drones. Simultaneously, the faster inference speed enables the lightweight MANet to perform real-time semantic segmentation. Finally, the lightweight MANet has higher segmentation accuracy; in the task of semantic segmentation for road surface patch repairs, there are fewer omissions and missegments in the visualized prediction results, and the overall inference effect and correctness are better. It can provide more accurate and comprehensive semantic and contour information of patch repairs, which is beneficial to the efficient advancement of subsequent asphalt pavement maintenance work. Attached Figure Description
[0035] Figure 1 This is an overall flowchart of the present invention;
[0036] Figure 2 A comparative diagram showing the overall network structure of the MANet of the present invention and the lightweight MANet;
[0037] Figure 3 This is a comparative illustration of the Bottleneck and Basic Block network structures of the present invention;
[0038] Figure 4 This is a bar chart comparing the inference speed and segmentation accuracy of the lightweight MANet model of the present invention and the contrasting model.
[0039] Figure 5 This is a visual comparison of the predicted strip repair results of MANet and lightweight MANet according to the present invention. Detailed Implementation
[0040] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0041] Example 1: As Figure 1 As shown, a real-time road surface patch detection method based on an improved lightweight MANet includes the following steps:
[0042] Step 1: Construct a semantic segmentation dataset for strip repair of road surface defects;
[0043] Step 1.1: Using a road image acquisition vehicle, grayscale images of strip repairs for road defects are automatically acquired during the driving process. The images are uniformly 2048*2048 pixels in size and in .jpg format.
[0044] Step 1.2: Use the semantic segmentation and annotation software LabelMe to draw fine polygon labels for the strip repair of road surface defects, and obtain a mask file corresponding to the strip repair image in .png format. In the mask file, the pixel value of the background road surface is 0, the pixel value of the foreground repair is 1, and the number of segmentation categories is 2.
[0045] Step 1.3: Integrate the strip-patched images and their corresponding mask files into a single dataset. The initial dataset contains 500 images of road strip-patched surfaces. After a 90° rotation for data augmentation, the dataset size is expanded to 1000 images, which are then divided into training, validation, and test sets in an 8:1:1 ratio. The original and augmented images will only appear in the same subset simultaneously to ensure no data leakage occurs, making the dataset's construction more scientific, reasonable, and reliable in terms of both quality and quantity.
[0046] Step 2: As Figure 2 As shown, the baseline MANet (Multi-Scale Aware-Relation Network) is improved to obtain a lightweight MANet; wherein, the improvement is to replace the backbone, increase the number of channels and add depthwise convolution;
[0047] Optionally, the replacement trunk specifically refers to:
[0048] The original MANet backbone ResNet-50 was discarded and replaced with a ResNet-34 backbone. The ResNet-34 backbone adopts a basic residual block structure, which contains two consecutive 3×3 convolutions with two convolution kernels.
[0049] Understandably, although ResNet-50 and ResNet-34 have the same depth, following the same "3+4+6+3" four-stage feature extraction design, and both employ residual structures, their basic feature extraction modules differ in network structure. ResNet-50 uses a Bottleneck residual block structure, containing two 1×1 convolutions and one 3×3 convolution, while ResNet-34 uses a Basic Block structure, containing two consecutive 3×3 convolutions. The number of convolutional kernels is reduced from three to two, making the backbone more lightweight and efficient. A comparison of the Bottleneck and Basic Block network structures is shown below. Figure 3 As shown;
[0050] Optionally, the increased number of channels specifically refers to:
[0051] In the Decoder Block, a constituent unit of the baseline MANet decoder, the two 1×1 convolutions at the beginning and end remain unchanged, while the number of channels of the 3×3 transposed convolution in the middle, which is responsible for upsampling the feature map, is increased by a preset factor.
[0052] Understandably, within the Decoder Block, a constituent unit of the decoder, the network structure acts as a bottleneck, consisting of two 1×1 convolutions and one 3×3 transposed convolution. This embodiment keeps the two 1×1 convolutions at the beginning and end unchanged, but increases the number of channels in the transposed convolution responsible for feature map upsampling by a preset factor. In this embodiment, the preset factor is set to 2, and the number of channels in the output feature map also doubles. Due to this significant increase in the number of channels, feature extraction and representation capabilities are considerably enhanced, contributing to a better understanding of the image's structure and semantic information.
[0053] Optionally, the addition of depthwise convolution specifically refers to:
[0054] In the baseline MANet decoder, an additional 3×3 depthwise convolution is added after each feature fusion stage, with the number of groups equal to the number of channels. The expression for the feature fusion stage of the depthwise convolution is as follows:
[0055]
[0056] In the formula, output is the output feature map, DWConv3×3 is the additional 3×3 depthwise convolution, and FusionFeat is the fused feature map.
[0057] Understandably, in the decoder, a 3×3 depthwise convolution is added after each feature fusion stage, with the number of groups equal to the number of channels. This means that additional feature extraction is performed on the feature map obtained from each feature fusion. In this embodiment, there are four feature fusion stages, so four depthwise convolutions are added respectively. This helps to eliminate residual background noise after feature map fusion and further enhances the ability to extract contextual semantic information from strip patching. At the same time, it only causes a very limited increase in the computational cost and parameter count of the model. This is mainly because conventional convolution requires a lot of computation across channels, while depthwise convolution is only responsible for the independent computation of a single channel, so it can greatly improve computational efficiency and reduce the model size.
[0058] It is understood that after three improvements, this embodiment yields a lightweight MANet, which has fewer parameters, less computation, and faster inference speed, and is used for subsequent training and prediction tasks to calculate segmentation accuracy.
[0059] Step 3: Train the lightweight MANet;
[0060] Step 3.1: Set the hyperparameters for training the lightweight MANet, including the number of iterations, batch size, number of threads, floating-point precision selection, and number of segmentation categories;
[0061] Optionally, in this embodiment, the number of iterations is 40, the batch size is 2 or 4, the number of threads is 4, the floating-point precision is FP32, the number of segmentation classes is 2, the optimizer is Adam, the input size is 512*512*3, the random seed is uniformly set to 11, the initial learning rate is 1e-4, and the learning rate descent method is cosine. Furthermore, the system platform for training the model is Windows 11, the CPU is an Intel 12800HX, the GPU is an RTX 4070 Laptop, and the memory size is 32 GB.
[0062] Step 3.2: Select the road surface patching image prediction mask of the benchmark MANet and the lightweight MANet for comparison, and train the lightweight MANet model according to the hyperparameters.
[0063] Step 4: Input the semantic segmentation dataset into the trained lightweight MANet to obtain the road surface patch detection results.
[0064] The effectiveness of the present invention will be further demonstrated through the following experiments.
[0065] Specifically, this experiment measured various metrics of the model, such as segmentation accuracy, parameter count, computational cost, and inference speed, and compared the results in detail with publicly available mainstream and lightweight semantic segmentation models. The mainstream non-lightweight semantic segmentation models selected included DeepLabV3+, SegNet, PSPNet, FCN, and MANet. These models typically demonstrate strong feature extraction performance and generalization ability in various semantic segmentation scenarios. Lightweight models selected included ICNet and BiSeNet V1. These models are state-of-the-art (SOTA) technologies from previous years, possessing a lightweight structure and good feature extraction performance, achieving a certain balance between the two. The metric used to measure the model's segmentation accuracy in this invention is mIoU (mean interaction ratio), expressed as:
[0066]
[0067] In the formula, k is the number of foreground categories, TP is a true positive (i.e., a pixel predicted as the correct valid category), FN is a false negative (i.e., a pixel predicted as the incorrect invalid category), and FP is a false positive (i.e., a pixel predicted as the correct invalid category).
[0068] Furthermore, Table 1 shows a comparison of various metrics between lightweight MANet and other semantic segmentation models on the stripe patching dataset.
[0069] Table 1. Comparison of metrics between lightweight MANet and other semantic segmentation models
[0070]
[0071] As shown in Table 1, the lightweight MANet not only outperforms the benchmark MANet in all aspects, but also, compared with the comparison models, in terms of parameter count, the lightweight MANet reduces the number of parameters by about 16% compared to the second smallest ICNet, and significantly reduces it by about 51% compared to the largest PSPNet; in terms of computational cost, the lightweight MANet reduces the number of parameters by about 8% compared to the second smallest ICNet, and significantly reduces it by about 81% compared to the largest DeepLabV3+; in terms of inference speed, the lightweight MANet is 6 fps faster than the second fastest BiSeNet V1, and significantly faster than the slowest SegNet by 45 fps; finally, in terms of segmentation accuracy, the lightweight MANet is 0.81 percentage points higher than the benchmark model, and significantly higher than the lowest accurate BiSeNet V1 by 15.02 percentage points.
[0072] In summary, compared with the comparative model, the lightweight MANet proposed in this embodiment not only has the fewest parameters and computational cost, but also the fastest inference speed and the highest segmentation accuracy, achieving a very ideal lightweight effect. This facilitates efficient and accurate detection and segmentation of strip repairs and subsequent road maintenance work. A bar chart comparing the lightweight MANet and the comparative model in terms of inference speed and segmentation accuracy is shown below. Figure 4 As shown.
[0073] Furthermore, to demonstrate the rationality and effectiveness of the improvements to MANet made by this invention, an ablation experiment was conducted in this embodiment to analyze and demonstrate the effects of the three improvements from the baseline model to the lightweight MANet. The ablation experiment analysis of the lightweight MANet is shown in Table 2.
[0074] Table 2 Ablation test results of lightweight MANet
[0075]
[0076] As shown in Table 2, when the original backbone ResNet-50 was replaced with ResNet-34, although the segmentation accuracy mIoU decreased by 2.2 percentage points, the number of parameters decreased by about 38.3%, the computational cost (FLOPs) decreased by about 59.2%, and the inference speed (FPS) increased significantly by about 78.5%. After adding depthwise convolutions in the fusion stage, the mIoU improved to 87.55%, which is better than the accuracy of the original MANet, with only minimal changes in the number of parameters and computational cost, and a decrease of only 2 fps. Finally, after doubling the number of channels in the transposed convolution, the mIoU further improved to 88.07%, with only a slight increase in the number of parameters, a mere increase in computational cost of about 6.5%, and a decrease in inference speed of only 1 fps. Therefore, the lightweight MANet achieved the ideal lightweight effect while also possessing significantly higher segmentation accuracy, and comprehensively outperformed the benchmark MANet in all four of the above indicators.
[0077] Furthermore, this embodiment also selects the road surface patching image prediction result masks of the benchmark MANet and the lightweight MANet for separate display and visual comparison analysis. The visual comparison of some patching prediction results of MANet and lightweight MANet is shown below. Figure 5 As shown.
[0078] from Figure 5As can be seen from the results, in the first row, the lightweight MANet performs significantly better than the baseline MANet in the strip patching segmentation in the lower right corner of the image, while the latter shows a large number of missing pixels in the lower right corner. In the second row, the lightweight MANet performs significantly more complete strip patching segmentation in the lower half of the image, while the baseline MANet misses most of the main strip patching part. In the third row, the lightweight MANet can almost completely segment the ends of the strip patching in the input image, while the baseline MANet has obvious missing segmentation problems in the lower right corner.
[0079] Therefore, after the three improvements proposed in this invention, the resulting lightweight MANet has significant advantages in terms of parameter quantity, computational load, inference speed, and segmentation accuracy. It not only surpasses the benchmark MANet in all aspects, but also exceeds several mainstream non-lightweight and lightweight models to varying degrees. It achieves an ideal balance between lightweight model and strip repair segmentation performance, which helps to provide higher quality road surface strip repair information more efficiently and accurately, so that subsequent road maintenance work can be carried out more smoothly, quickly, and stably.
[0080] In summary, existing semantic segmentation models struggle to balance segmentation accuracy, inference efficiency, and the number of model parameters for patching road surfaces. This invention improves MANet in three aspects: First, the original backbone ResNet-50 is replaced with the lighter ResNet-34; second, the number of upsampled transposed convolution channels in the Decoder Block is doubled; and finally, an additional depthwise convolution is added after each feature fusion stage in the decoder to enhance feature extraction capabilities, resulting in a lightweight MANet. Experiments show that the proposed lightweight MANet, compared to benchmark models, not only boasts higher segmentation accuracy but also significantly fewer parameters and computational cost, along with faster inference speed. It can more efficiently and accurately complete the semantic segmentation task of patching road surfaces with less computational resources.
[0081] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.
Claims
1. A real-time road surface patch repair detection method based on an improved lightweight MANet, characterized in that, The method includes the following steps: Step 1: Construct a semantic segmentation dataset for strip repair of road surface defects; Step 2: Improve the baseline MANet to obtain a lightweight MANet; wherein, the improvement is to replace the backbone, increase the number of channels and add depthwise convolution; Step 3: Train the lightweight MANet; Step 4: Input the semantic segmentation dataset into the trained lightweight MANet to obtain the road surface patch detection results.
2. The real-time road surface patch repair detection method based on the improved lightweight MANet according to claim 1, characterized in that, Step 1 specifically refers to: Step 1.1: Obtain grayscale images of strip repairs for road surface defects with uniform dimensions; Step 1.2: Use semantic segmentation and annotation software to draw polygonal labels for the strip repair of road surface defects, and obtain a mask file corresponding to the strip repair image; Step 1.3: Merge the striped patch images and the corresponding mask files into the same dataset, and divide them into training set, validation set and test set according to the preset ratio.
3. The real-time road surface patch repair detection method based on the improved lightweight MANet according to claim 1, characterized in that, The replacement of the main trunk specifically refers to: The original MANet backbone ResNet-50 was discarded and replaced with a ResNet-34 backbone. The ResNet-34 backbone adopts a basic residual block structure, which contains two consecutive 3×3 convolutions with two convolution kernels.
4. The real-time road surface patch repair detection method based on the improved lightweight MANet according to claim 1, characterized in that, The increased number of channels specifically refers to: In the Decoder Block, a constituent unit of the baseline MANet decoder, the two 1×1 convolutions at the beginning and end remain unchanged, while the number of channels of the 3×3 transposed convolution in the middle, which is responsible for upsampling the feature map, is increased by a preset factor.
5. The real-time road surface patch repair detection method based on the improved lightweight MANet according to claim 1, characterized in that, The addition of depthwise convolution specifically refers to: In the baseline MANet decoder, an additional 3×3 depthwise convolution is added after each feature fusion stage, with the number of groups equal to the number of channels. The expression for the feature fusion stage of the depthwise convolution is as follows: ; In the formula, output is the output feature map, DWConv3×3 is the additional 3×3 depthwise convolution, and FusionFeat is the fused feature map.
6. The real-time road surface patch repair detection method based on the improved lightweight MANet according to claim 1, characterized in that, Step 3 specifically refers to: Step 3.1: Set the hyperparameters for training the lightweight MANet, including the number of iterations, batch size, number of threads, floating-point precision selection, and number of segmentation categories; Step 3.2: Select the road surface patching image prediction mask of the benchmark MANet and the lightweight MANet for comparison, and train the lightweight MANet model according to the hyperparameters.