Road surface disease real-time detection method and system based on deep learning
By combining deep learning and improved image acquisition technology, dynamic adaptive deformation convolution and other methods are adopted to solve the challenges of asphalt pavement disease detection accuracy and efficiency in the prior art, and high-precision and efficient pavement disease detection are achieved.
Patent Information
- Application Number
- CN202510250251.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The prior art has the dual challenges of accuracy and efficiency in the automated detection of asphalt pavement crack details and large-area repairs. Especially when dealing with complex deformation and subtle disease details, the lightweight semantic segmentation algorithm is poorly robust, and the deep semantic segmentation algorithm is prone to losing feature information.
The real-time detection method of pavement diseases based on deep learning is adopted. By combining the 2D and 3D images collected by the PavementVision3D system on the digital highway data acquisition vehicle, the improved downsampling and upsampling algorithm is used, and dynamic adaptive deformation convolution is combined to improve detection accuracy and efficiency.
It significantly improves detection accuracy and efficiency, can accurately identify complex deformation or subtle disease characteristics, breaks through the bottleneck of insufficient feature extraction capabilities in complex scenarios, and achieves pixel-level recognition results.
Smart Images

Figure CN120088486A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of road operation and maintenance, and particularly to a real-time detection method and system for pavement diseases based on deep learning. Background Art
[0002] As a key component of road infrastructure, asphalt roads undertake important tasks of daily transportation. With the increase in traffic flow and transportation load, long-term high-load operation is likely to cause cracks on the road surface. This not only affects the comfort of drivers but may also gradually expand over time, ultimately seriously threatening the safety of the entire asphalt pavement structure. In addition to cracks, other common diseases on asphalt roads include potholes, closed cracks, repair patches, etc. These diseases accelerate the aging process of the road surface and even exacerbate the wear of the road surface, further affecting the service life and safety of the road. In addition to pavement diseases, certain surface design features (such as road markings, expansion joints, manhole covers, etc.) are also important factors for evaluating and optimizing road quality and performance. Whether road markings are clear, expansion joints are intact, and manhole covers are stable will directly affect driving safety and comfort. Therefore, how to efficiently and accurately detect pavement diseases and these surface design features on asphalt roads, and thus ensure driving safety and driving comfort, has become a major issue in current road engineering.
[0003] Although detection technologies based on deep learning have made remarkable progress in the field of computer vision, there are still many challenges in the automated detection of asphalt pavement crack details and large-area repairs, making it difficult for existing technologies to meet the dual requirements of accuracy and efficiency in practical applications. It should be particularly emphasized that the types and distribution characteristics of diseases on asphalt roads are highly diverse. Different road sections may present different forms of diseases, and at the same time, there are also significant differences in the manifestations of each disease. This means that the detection technology not only needs to have a strong information retention ability but also be able to adaptively process complex deformations and subtle disease details. Therefore, existing lightweight semantic segmentation algorithms have poor robustness in small target detection, and deep semantic segmentation algorithms usually easily lose important feature information, resulting in the inability to accurately identify complex deformations or subtle disease features in images. Summary of the Invention
[0004] To solve the problems existing in the above-mentioned prior art, the present invention provides a real-time pavement disease detection method and system based on deep learning, aiming to simultaneously identify and detect various pavement diseases such as cracks, sealant filling, potholes, repairs, markings, expansion joints, manhole covers, etc., and output pixel-level recognition results. By combining 2D and 3D images collected by the PavementVision3D system on a Digital Highway Data Collection Vehicle (DHDV) and inputting these images into the method designed based on the present invention for semantic segmentation, the detection accuracy and detection efficiency can be effectively improved, ensuring the accurate identification of various pavement diseases. It solves the technical problem that the existing semantic segmentation algorithm cannot accurately identify complex deformations or subtle disease features in images, which affects the accuracy of pavement automatic detection.
[0005] The real-time pavement disease detection method based on deep learning first preprocesses the image to be detected; then inputs the preprocessed image into the encoding layer for multiple downsamplings and convolutional feature extractions; then inputs the image output by the encoding layer into the decoding layer for multiple upsamplings and convolutional feature extractions; among which, the encoding and decoding layers of the first three resolution levels use dynamic adaptive convolution to dynamically extract features; the above encoding and decoding layers can be looped multiple times; finally, the image after multiple encodings and decodings is input into the dynamic adaptive deformable convolution to output the detection result map; both upsampling and downsampling adopt dynamic position encoding, lossless slicing, and dynamic branch mechanisms, and the lossless slicing order of the upsampling is the reverse order of the lossless slicing order of the downsampling; the dynamic adaptive deformable convolution can adaptively adjust the shape of the convolution kernel according to the shape features of the diseases in the input image.
[0006] Further, the preprocessing includes:
[0007] The image is downsampled using the bilinear interpolation method and adjusted to a fixed size.
[0008] The 3D and 2D images are processed through channel fusion to form a single-batch image with 2 channels.
[0009] Further, the downsampling includes:
[0010] (a) Input the feature image data with the size of W×H×C, and assign position information to each pixel through dynamic position encoding, where W is the width, H is the height, and C is the number of channels.
[0011] (b) Enhance the network's understanding of the position information through the Spatial Attention Mechanism (SAM).
[0012] (c) The image is sliced horizontally and vertically in the form of (W / 2, H / 2) to obtain four feature sub-images with the size of (W / 2, H / 2, C), and the feature sub-images are stacked along the number of channels to obtain a feature sub-image with the size of (W / 2, H / 2, 4C).
[0013] (d) The image processes the image features by dynamically selecting dilated convolution or depthwise separable convolution through a dynamic branching mechanism according to the learned size of the image.
[0014] Further, the upsampling includes:
[0015] (a) Input the feature image data with size W×H×C, and assign position information to each pixel through dynamic position encoding, where W is the width, H is the height, and C is the number of channels;
[0016] (b) Deepen the network's understanding of the position information through the spatial attention mechanism (SAM);
[0017] (c) According to the slicing order in the downsampling, combine and restore it in reverse order to form an image with size 2W×2H×C / 4;
[0018] (d) The restored image processes the image features by dynamically selecting dilated convolution or depthwise separable convolution through a dynamic branching mechanism according to the learned size of the image.
[0019] Further, the dynamic adaptive deformable convolution includes:
[0020] (1) Input the feature image data with size W×H×C, where W is the width, H is the height, and C is the number of channels;
[0021] (2) Generate the sampling point offsets for the feature image through depthwise separable convolution, focus on the disease itself, and output the offsets with shape (W×H×2K), where K is the size of the depthwise separable convolution kernel, and each position offset contains K offsets;
[0022] (3) The output offsets indicate that the convolution kernel dynamically adjusts the sampling positions to achieve adaptive dynamic deformation.
[0023] (4) Perform feature extraction based on the sampling positions after adaptive dynamic deformation and output the result map.
[0024] Further, both the upsampling and the downsampling include:
[0025] When the image size is greater than or equal to the size after downsampling n times, since the image size is relatively large, dilated convolution will be used to enhance the model's context awareness of the image.
[0026] When the image size is smaller than the size after downsampling n times, since the image resolution is too small, depth convolution and pointwise convolution will be performed in sequence. While reducing the computational amount, it can effectively extract spatial information and improve the detection accuracy.
[0027] Furthermore, it is implemented using a semantic segmentation algorithm model. Before using the semantic segmentation algorithm model, multiple 2D pavement images of different highways and 3D images showing pavement depth information are collected to construct a semantic segmentation database for pavement diseases. The data in the semantic segmentation database for pavement diseases are labeled to obtain a dataset, and the semantic segmentation algorithm model is trained and verified.
[0028] A real-time pavement disease detection system based on deep learning includes a data acquisition module and a trained semantic segmentation algorithm model. The data acquisition module obtains 2D images and 3D images of the pavement to be detected in real time. The trained semantic segmentation algorithm model processes the 2D images and 3D images of the pavement to be detected in sequence using a preprocessing unit, an encoding layer unit, a decoding layer unit, and a dynamic adaptive deformable convolution, and finally outputs a detection result map.
[0029] The beneficial effects of the present invention include:
[0030] The semantic segmentation algorithm ShuttleNetS3 proposed by the present invention shows significant adaptive advantages in the detection scenario of complex deformed diseases. This algorithm can accurately capture the characteristics of complex and changeable road traffic scenarios, and flexibly and efficiently extract the morphology and boundaries of diseases, breaking through the bottleneck of the insufficient feature extraction ability of traditional algorithms in complex scenarios.
[0031] In terms of the ability to comprehensively detect multiple diseases, ShuttleNetS3 has achieved significant performance improvements in all detection categories, comprehensively surpassing mainstream semantic segmentation networks and the original ShuttleNet, demonstrating excellent generalization ability and robustness, and providing stronger technical support for the intelligent detection of complex disease scenarios.
[0032] At the practical engineering application level, the real-time semantic segmentation detection algorithm proposed by the present invention can generate high-precision disease semantic segmentation maps in real time during the inspection process of road detection vehicles at normal driving speeds, ensuring the timeliness and accuracy of detection results. This technology not only provides intuitive and highly accurate decision-making support for the intelligent monitoring and diagnosis of road diseases, but also significantly reduces the complexity and workload of pavement maintenance, opening up a new efficient and intelligent way for road maintenance management work, and showing broad practical application value and promotion prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a flowchart of the semantic segmentation algorithm model related to the embodiment of the present application.
[0034] Figure 2 It is a legend of the flowchart of the semantic segmentation algorithm model related to the embodiment of the present application.
[0035] Figure 3It is a structural diagram for implementing downsampling involved in the embodiments of the present application.
[0036] Figure 4 It is a structural diagram for implementing upsampling involved in the embodiments of the present application.
[0037] Figure 5 It is the effect of dynamic adaptive deformable convolution and traditional convolution involved in the embodiments of the present application.
[0038] Figure 6 It is the implementation steps of dynamic adaptive deformable convolution involved in the embodiments of the present application. Detailed implementation manners
[0039] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part rather than all of the embodiments of the present application. Therefore, the detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but merely represents the selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0040] The types of diseases on asphalt pavements and their distribution characteristics are highly diverse. Different sections may present different forms of diseases. At the same time, there are also significant differences in the manifestation forms of each disease. This means that the detection technology not only needs to have a strong information retention ability, but also be able to adaptively process complex deformations and subtle disease details. Therefore, the existing lightweight semantic segmentation algorithms have poor robustness in small target detection, and the deep semantic segmentation algorithms usually easily lose important feature information, resulting in the inability to accurately identify complex deformation or subtle disease features in images.
[0041] To overcome this problem, the present invention realizes the "bidirectional mapping" mechanism of upsampling and downsampling through the improved downsampling SPsampling and upsampling SPupsampling algorithms. These improvements significantly enhance the information retention ability of image feature information. Especially when dealing with subtle diseases, it can effectively retain key texture information and avoid information loss.
[0042] Regarding the deformation problem of complex diseases, the dynamic adaptive deformation convolution method is innovatively used. By adaptively extracting the feature information of complex diseases, the flexibility and adaptability of the model are further improved. Compared with traditional convolution kernels, the dynamic adaptive convolution has a higher degree of freedom and can flexibly learn context information during the detection process, effectively avoiding the performance limitations brought by fixed receptive fields. This convolution method not only improves the recognition ability of complex deformed diseases but also demonstrates excellent adaptability when dealing with diseases of different scales and types.
[0043] Combined with the above technologies, the semantic segmentation algorithm proposed by the present invention can significantly improve the detection accuracy in practical engineering applications and is superior to existing similar technologies in terms of processing speed, showing a broader application prospect.
[0044] Through these innovations, the proposed technology can provide reliable technical support for the accurate positioning and automated repair of pavement diseases and provide a more efficient solution for future road maintenance and management.
[0045] A real-time detection method for pavement diseases based on deep learning, such as Figure 1-2 shown, first preprocesses the image to be detected; then inputs the preprocessed image into the encoding layer for multiple downsamplings and convolutional feature extractions. Each time of downsampling halves the image size, and the encoding layer can adjust the number of downsamplings according to the image size; then inputs the image output by the encoding layer into the decoding layer for multiple upsamplings and convolutional feature extractions. Each time of upsampling doubles the image size to the original size, and the number of upsamplings is the same as the number of downsamplings of the encoding layer; the above encoding and decoding layers can be cycled multiple times, and the image information of the previous encoding-decoding layer will be input into the next encoding-decoding layer through memory connections; finally, inputs the image after multiple encoding and decoding into the dynamic adaptive deformation convolution to output the detection result map; both upsampling and downsampling adopt dynamic position encoding, lossless slicing, and dynamic branch mechanisms, and the order of lossless slicing for upsampling is the reverse order of the lossless slicing order for downsampling; the dynamic adaptive deformation convolution can adaptively adjust the shape of the convolution kernel according to the shape characteristics of diseases in the input image.
[0046] Among them, the encoding and decoding layers of the first three resolution levels use dynamic adaptive deformation convolution to dynamically extract features, such as Figure 1 shown in the first to third layers from top to bottom. For the encoding layer image from top to bottom, each time of downsampling reduces the size to 1 / 2 , and for the decoding layer image, each time of upsampling doubles the size to the original size. Only the thick arrows in the first three layers are dynamic adaptive deformation convolutions, and the following are all ordinary convolutions.
[0047] In another embodiment, the preprocessing includes:
[0048] The image is downsampled using the bilinear interpolation method and adjusted to a fixed size;
[0049] The 3D and 2D images are processed through channel fusion to form a single-batch image with 2 channels.
[0050] In another embodiment, the downsampling includes:
[0051] (a) Input the feature image data with size W×H×C, and assign position information pixel by pixel through dynamic position encoding, where W is the width, H is the height, and C is the number of channels;
[0052] (b) Enhance the network's understanding of the position information through the spatial attention mechanism (SAM);
[0053] (c) The image is sliced horizontally and vertically in the form of (W / 2,H / 2) to obtain four feature sub-images with size (W / 2,H / 2,C), and the feature sub-images are stacked along the number of channels to obtain a feature sub-image with size (W / 2,H / 2,4C);
[0054] (d) The image passes through the dynamic branch mechanism and dynamically selects dilated convolution or depthwise separable convolution to process the image features by learning the size of the image.
[0055] In another embodiment, the upsampling includes:
[0056] (a) Input the feature image data with size W×H×C, and assign position information pixel by pixel through dynamic position encoding, where W is the width, H is the height, and C is the number of channels;
[0057] (b) Deepen the network's understanding of the position information through the spatial attention mechanism (SAM);
[0058] (c) According to the slicing order in the downsampling, combine and restore it in reverse order to an image with size 2W×2H×C / 4;
[0059] (d) The restored image passes through the dynamic branch mechanism and dynamically selects dilated convolution or depthwise separable convolution to process the image features by learning the size of the image.
[0060] In another embodiment, the dynamic adaptive deformable convolution includes:
[0061] (1) Input the feature image data with size W×H×C, where W is the width, H is the height, and C is the number of channels;
[0062] (2) Generate sampling point offsets for the feature image through depthwise separable convolution, focus on the disease itself, and output an offset with shape (W×H×2K), where K is the size of the depthwise separable convolution kernel, and each position offset contains K offsets;
[0063] (3) The output offset indicates that the convolutional kernel dynamically adjusts the sampling position to achieve adaptive dynamic deformation.
[0064] (4) Feature extraction is performed based on the sampling position after adaptive dynamic deformation, and the result map is output.
[0065] In another embodiment, both the upsampling and downsampling include:
[0066] When the image size is greater than or equal to the size after downsampling n times, the image size is large, and atrous convolution will be used to enhance the model's context awareness of the image.
[0067] When the image size is less than the size after downsampling n times, the image resolution is too small, and depth convolution and pointwise convolution will be performed in sequence. While reducing the computational amount, spatial information can be effectively extracted and the detection accuracy can be improved.
[0068] In another embodiment, before using the semantic segmentation algorithm model, 2D road surface images of multiple different highways and 3D images showing road surface depth information are collected to construct a semantic segmentation database for road surface diseases, and the data in the semantic segmentation database for road surface diseases are labeled to obtain a data set for training and validating the semantic segmentation algorithm model.
[0069] In another embodiment, the implementation process is as follows:
[0070] 1.1. Data acquisition
[0071] The road foreground images are collected by the PavementVision3D system on the Digital Highway Data Vehicle (DHDV).
[0072] The DHDV achieved full coverage of a 4-meter-wide lane at a scanning frequency of 30KHz, ensuring a high-precision resolution of 1 millimeter.
[0073] The collected image data includes 2D road surface images and 3D images showing road surface depth information, and these data are used to construct a semantic segmentation database for road surface diseases.
[0074] This database not only contains road surface images but also records the corresponding geographical location information and time stamps, providing important basic information for subsequent data processing and analysis.
[0075] In view of the differences in road surface conditions of different highways, it is necessary to collect road surface data from multiple different highways to ensure data diversity, so as to construct a sample database suitable for training the algorithm model.
[0076] The road surface training sample database covers various types of road sections, including but not limited to national highways, provincial highways, ring roads, urban expressways, and road sections with various other road conditions, fully reflecting the extensiveness and comprehensiveness of data collection.
[0077] 1.2 Professional pixel-level data annotation
[0078] To ensure the high quality of model training, all the collected data are manually annotated by experts to ensure pixel-level accuracy.
[0079] The original size of each image is 4096×2048 pixels, representing a road surface section that is 4 meters long and 2 meters wide.
[0080] 2. Road surface disease semantic segmentation
[0081] A new semantic segmentation algorithm is adopted to simultaneously identify multiple road surface diseases and surface design features, and its model structure is as Figure 1 shown.
[0082] 2.1. Semantic segmentation image preprocessing
[0083] In the process of road surface disease semantic segmentation, the image preprocessing link is crucial.
[0084] The bilinear interpolation method is used to downsample the image and adjust it to a fixed size of 512×256 pixels, which not only ensures the processing efficiency but also maximally retains the key feature information.
[0085] Next, 3D and 2D images are processed through channel fusion to form a single batch of images with 2 channels, and are input into the semantic segmentation algorithm model.
[0086] In this way, the semantic segmentation algorithm model can simultaneously combine pixel information and depth data, thereby achieving a more accurate disease detection effect.
[0087] 2.2. Downsampling (Spsampling) algorithm for retaining height information
[0088] Max pooling is a fixed non-linear method that ignores important potential representations, thus easily causing the loss of detailed information in semantic segmentation tasks.
[0089] In this part, the combination of lossless slice downsampling and dynamic position encoding is used, and a dynamic branch mechanism is used before output to customize the convolution to enhance the learning ability of downsampling.
[0090] Lossless slice downsampling ensures that all image feature information is retained when the image size is halved, while dynamic position encoding makes up for the loss of the model's understanding of sub-image position information during lossless slice downsampling.
[0091] The final dynamic branch provides the best convolutional solution for images of all sizes to better extract features. The structure diagram of SPsampling is as shown in Figure 3 and the specific implementation steps are as follows:
[0092] (1) Input the feature image data with the size of W×H×C (width×height×channels), and assign position information pixel by pixel through dynamic position encoding;
[0093] (2) Enhance the network's understanding of position information through the spatial attention mechanism (SAM);
[0094] (3) The image is sliced horizontally and vertically in the form of (W / 2, H / 2) to obtain four feature sub-images with the size of (W / 2, H / 2, C), and they are stacked along the number of channels to obtain a feature sub-image with the size of (W / 2, H / 2, 4C).
[0095] (4) The image passes through the dynamic branch mechanism, and dynamically selects dilated convolution or depthwise separable convolution to process the image features by learning the size of the image.
[0096] When the number of image downsampling times is less than 4, the image size is relatively large, and dilated convolution with a kernel size of 3×3 and a stride of 2 will be used to enhance the model's context awareness of the image.
[0097] When the number of image downsampling times is greater than 4, the image resolution is too small, and it will sequentially pass through a depth convolution with a kernel size of 3×3 and a pointwise convolution with padding of 0 and a stride of 1, effectively extracting spatial information and improving detection accuracy while reducing the computational cost.
[0098] 2.3. Upsampling (Spupsampling) algorithm for accurately restoring information
[0099] When detecting various road surface diseases and surface detail features, deconvolution and bilinear interpolation upsampling are prone to missing fine cracks and pixel artifacts in large-area road repair areas.
[0100] To achieve more accurate downsampling and upsampling, thus highly restoring image details, SPupsampling is proposed based on the principle of SPsampling. Through dynamic position encoding and inverse slicing mechanism, the "bidirectional mapping" mechanism of up and downsampling is realized while ensuring the high restoration of information details.
[0101] This improvement not only significantly enhances the overall performance of the model, but also greatly improves the recognition ability of fine cracks and large-area road repair areas, breaking through the limitations of traditional methods in detail restoration. The structure diagram of SPupsampling is as shown in Figure 4 and the specific implementation steps are as follows:
[0102] (1) Input the feature image data with dimensions of W×H×C (width × height × channels), and assign position information pixel by pixel through dynamic position encoding;
[0103] (2) Deepen the network's understanding of position information through the spatial attention mechanism (SAM);
[0104] (3) The image is combined and restored in reverse order according to the slicing order in SPsampling into an image with dimensions of (2W×2H×C / 4). This step is similar to the lossless slicing step in SPsampling played in reverse, so this step is also lossless.
[0105] (4) The restored image passes through the dynamic branch mechanism and dynamically selects dilated convolution or depthwise separable convolution to process the image features by learning the size of the image.
[0106] When the image size is greater than or equal to the size after downsampling 3 times, the image size is relatively large, and dilated convolution with a kernel size of 3×3 and a stride of 2 will be used to enhance the model's context awareness of the image.
[0107] When the image size is smaller than the size after downsampling 3 times, the image resolution is too small, and it will sequentially pass through a depthwise convolution with a kernel size of 3×3 and a pointwise convolution with padding of 0 and a stride of 1. This step is also similar to the mechanism of SPsampling, thus forming a "bidirectional mapping" mechanism with SPsampling as a whole.
[0108] 2.4. Dynamic Adaptive Deformable Convolution
[0109] In deep learning, a traditional 3×3 convolution is usually used as the basic unit for feature extraction. However, in pavement detection tasks, the targets often present extremely deformed complex shapes, such as cracks, potholes, and sealant joints, etc.
[0110] Traditional convolution uses a filter with a fixed size and a preset sliding stride. The characteristics of this fixed mode limit its flexibility in the information extraction process, thus affecting the network's ability to capture diverse disease shapes.
[0111] To solve this problem, a dynamic adaptive deformable convolution mechanism is proposed. Aimed at maintaining the lightweight of the network while being able to adaptively adjust the shape of the convolution kernel according to the shape characteristics of diseases in the input image. In this way, the network can more flexibly extract the potential target representations and context information in large-size images, thus enhancing the detection ability for complex deformed diseases. As Figure 5 shown, this method effectively improves the adaptability and recognition accuracy of the model to various disease forms.
[0112] As Figure 6 shown, the implementation steps of the dynamic adaptive deformable convolution are as follows:
[0113] (1) Input the feature image data with dimensions W×H×C (width×height×channels).
[0114] (2) Generate sampling point offsets for the feature image through depthwise separable convolution, focusing on the disease itself, and output offsets with a shape of (W×H×2K) (k is the size of the depthwise separable convolution kernel, and each position offset contains k offsets).
[0115] (3) The output offsets indicate that the convolution kernel dynamically adjusts the sampling positions to achieve adaptive dynamic deformation.
[0116] (4) Perform feature extraction based on the sampling positions after adaptive dynamic deformation and output the result map.
[0117] In another embodiment, it relates to a real-time road disease detection system based on deep learning, including a data acquisition module and a trained semantic segmentation algorithm model. The data acquisition module obtains 2D images and 3D images of the road surface to be detected in real time. The trained semantic segmentation algorithm model sequentially uses a preprocessing unit, a downsampling unit, an upsampling unit, and a dynamic adaptive deformation convolution to process the 2D images and 3D images of the road surface to be detected and finally outputs a detection result map.
[0118] In another embodiment, for the proposed semantic segmentation algorithm model, tests were conducted on 2000 road images with other current advanced semantic segmentation algorithms, namely Unet, DeepLabv3+, PSPnet, Segformer-B5, and the original ShuttleNet. The test environment was python3.8, the deep learning platform was Pytorch1.11, and the GPU hardware configuration was NVIDIA RTX3070. The results of the test metrics are shown in Table 1 below.
[0119] Table 1 Performance of Different Semantic Segmentation Algorithms
[0120]
[0121] Among them, the evaluation metrics adopted two of the most representative metrics in the current field of target detection intelligent algorithms, namely F1-Score and mIOU. The larger the values of these two metrics, the better the performance of the algorithm model. Compared with the current mainstream semantic segmentation algorithms: Unet, DeepLabv3+, PSPnet, Segformer-B5, and the original ShuttleNet, the algorithm proposed in the present invention has obvious advantages in pixel-level recognition of various road diseases and surface design features.
[0122] In another embodiment, the semantic segmentation algorithm proposed by the present invention was tested on the Crack500 dataset together with the currently advanced semantic segmentation algorithms Unet, DeepLabv3+, PSPnet, Segformer-B5, and the original ShuttleNet. The performance metrics of the test are shown in Table 2 below.
[0123] Table 2 Performance of Different Semantic Segmentation Algorithms
[0124]
[0125]
[0126] Compared with the currently mainstream Unet, DeepLabv3+, PSPnet, Segformer-B5, and the original ShuttleNet, the semantic segmentation algorithm proposed by the present invention is more robust in the comprehensive strength of pixel-level recognition of various pavement diseases and surface design features.
[0127] The above-described embodiments merely represent the specific implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the protection scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the technical solution of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application.
Claims
1. A real-time detection method for pavement defects based on deep learning, characterized in that: First, preprocess the image to be detected; The preprocessed image is then input into the encoding layer for multiple downsampling and convolution feature extraction; the image output from the encoding layer is then input into the decoding layer for multiple upsampling and convolution feature extraction; The encoding and decoding layers of the first three resolution levels use dynamic adaptive convolution to dynamically extract features; finally, the encoded and decoded images are input into the dynamic adaptive deformation convolution to output the detection result image; Both upsampling and downsampling adopt dynamic position encoding, lossless slicing and dynamic branching mechanism. The order of lossless slicing for upsampling is the reverse order of the order of lossless slicing for downsampling. The dynamic adaptive deformation convolution can adaptively adjust the shape of the convolution kernel according to the shape characteristics of the defect in the input image.
2. The real-time pavement disease detection method based on deep learning according to claim 1 is characterized in that: The encoding layer and the decoding layer are cycled multiple times, and finally the image after multiple encoding and decoding is input into the dynamic adaptive deformation convolution to output the detection result image.
3. The real-time detection method for pavement defects based on deep learning according to claim 1 is characterized in that: The pre-processing comprises: The image is downsampled using bilinear interpolation to adjust it to a fixed size; The 3D and 2D images are processed by channel fusion to form a single batch of images with 2 channels.
4. The real-time pavement disease detection method based on deep learning according to claim 1 is characterized in that: The down sampling includes: (a) Input feature image data of size W×H×C, and assign position information pixel by pixel through dynamic position encoding, where W is width, H is height, and C is the number of channels; (b) Enhance the network’s understanding of location information through the spatial attention mechanism; (c) The image is split horizontally and vertically in the form of (W / 2, H / 2) to obtain four feature sub-graphs of size (W / 2, H / 2, C). The feature sub-graphs are stacked along the number of channels to obtain feature sub-graphs of size (W / 2, H / 2, 4C). (d) The image is processed through a dynamic branching mechanism, which dynamically selects dilated convolution or depthwise separable convolution by learning the image size.
5. The real-time pavement disease detection method based on deep learning according to claim 1 is characterized in that: The upsampling comprises: (a) Input feature image data of size W×H×C, and assign position information pixel by pixel through dynamic position encoding, where W is width, H is height, and C is the number of channels; (b) Deepen the network’s understanding of location information through the spatial attention mechanism; (c) According to the slice order in downsampling, the slices are combined in reverse order to restore the image with a size of 2W×2H×C / 4; (d) The restored image is processed through a dynamic branching mechanism, which dynamically selects dilated convolution or depthwise separable convolution by learning the image size to process the image features.
6. The real-time pavement disease detection method based on deep learning according to claim 1 is characterized in that: The dynamic adaptive deformation convolution includes: (a) Input feature image data of size W×H×C, where W is width, H is height, and C is the number of channels; (b) Generate sampling point offsets for the feature image through depthwise separable convolution, focus on the disease itself, and output offsets of shape (W×H×2K), where K is the size of the depthwise separable convolution kernel, and each position offset contains K offsets; (c) The output offset instructs the convolution kernel to dynamically adjust the sampling position to achieve adaptive dynamic deformation; (d) Feature extraction is performed based on the sampling position after adaptive dynamic deformation, and the result image is output.
7. The real-time detection method for pavement defects based on deep learning according to claim 1 is characterized in that: The upsampling and downsampling also include: When the image size is greater than or equal to the size of the downsampled n times, the image size is large, and a dilated convolution will be used to improve the model's contextual perception of the image; When the image size is smaller than the size of downsampling n times, the image resolution is too small, and depth convolution and point-by-point convolution will be performed in sequence, which effectively extracts spatial information and improves detection accuracy while reducing the amount of calculation.
8. The real-time pavement disease detection method based on deep learning according to claim 1 is characterized in that: The semantic segmentation algorithm model is used for implementation. Before using the semantic segmentation algorithm model, 2D road surface images of multiple different highways and 3D images showing road surface depth information are collected to build a semantic segmentation database of road surface defects. The data in the semantic segmentation database of road surface defects are annotated to obtain a data set to train and verify the semantic segmentation algorithm model.
9. A real-time road disease detection system based on deep learning, characterized in that: A real-time pavement disease detection method based on deep learning as described in any one of claims 1 to 8 is adopted, including a data acquisition module and a trained semantic segmentation algorithm model. The data acquisition module acquires 2D images and 3D images of the road surface to be detected in real time. The trained semantic segmentation algorithm model sequentially processes the 2D images and 3D images of the road surface to be detected using a preprocessing unit, a coding layer unit, a decoding layer unit, and a dynamic adaptive deformation convolution, and finally outputs a detection result map.
Citation Information
Patent Citations
Real-time road image semantic segmentation method and system based on deep learning
CN113688836A
Improved semantic segmentation network construction method and system for bridge road crack recognition
CN119251842A