Real-time Detection Method and System for Road Surface Diseases Based on Deep Learning
Through improved downsampling and upsampling algorithms and dynamic adaptive deformation convolution, the problem of identifying complex deformation and subtle disease characteristics in asphalt pavement disease detection is solved, and high-precision and efficient disease detection is achieved, supporting intelligent road monitoring and diagnosis.
Patent Information
- Application Number
- CN202510250251.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The prior art is difficult to accurately identify complex deformation or subtle disease characteristics in asphalt pavement disease detection, resulting in insufficient detection accuracy and efficiency. In particular, the lightweight semantic segmentation algorithm is poorly robust, and the deep semantic segmentation algorithm is prone to loss of important feature information.
The real-time detection method of pavement diseases based on deep learning is adopted, combined with the PavementVision3D system of the digital highway data acquisition vehicle, and through improved downsampling and upsampling algorithms (SPsampling and SPupsampling) and dynamic adaptive deformation convolution, the precise processing of images is achieved, and the feature information retention ability and model adaptability are improved.
It significantly improves the accuracy and efficiency of road disease detection, can flexibly extract disease characteristics in complex scenarios, breaks through the bottleneck of traditional algorithms, and provides efficient and intelligent road disease monitoring and diagnosis support.
Smart Images

Figure CN120088486B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of road operation and maintenance, and particularly to a real-time detection method and system for pavement diseases based on deep learning. Background Art
[0002] As a key component of road infrastructure, asphalt roads undertake important tasks of daily transportation. With the increase in traffic flow and transportation load, long-term high-load operation is likely to cause cracks on the road surface, which not only affects the comfort of drivers, but may also gradually expand over time, ultimately seriously threatening the safety of the entire asphalt pavement structure. In addition to cracks, other types of diseases are also common in asphalt pavements, such as potholes, closed cracks, repair patches, etc. These diseases accelerate the aging process of the road surface and even exacerbate the wear of the road surface, further affecting the service life and safety of the road. In addition to pavement diseases, certain surface design features (such as road markings, expansion joints, manhole covers, etc.) are also important factors for evaluating and optimizing road quality and performance. Whether the road markings are clear, whether the expansion joints are intact, and whether the manhole covers are stable will directly affect the safety and comfort of driving. Therefore, how to efficiently and accurately detect pavement diseases and these surface design features of asphalt roads, and then ensure driving safety and driving comfort, has become a major issue in current road engineering.
[0003] Although detection technologies based on deep learning have made remarkable progress in the field of computer vision, there are still many challenges in the automated detection of asphalt pavement crack details and large-area repairs, making it difficult for existing technologies to meet the dual requirements of accuracy and efficiency in practical applications. It should be particularly emphasized that the types of diseases on asphalt pavements and their distribution characteristics are highly diverse. Different road sections may present different forms of diseases, and at the same time, there are also significant differences in the manifestation forms of each disease. This means that the detection technology not only needs to have a strong information retention ability, but also be able to adaptively process complex deformations and subtle disease details. Therefore, existing lightweight semantic segmentation algorithms have poor robustness in small target detection, and deep semantic segmentation algorithms usually easily lose important feature information, resulting in the inability to accurately identify complex deformations or subtle disease features in images. Summary of the Invention
[0004] To solve the problems existing in the above-mentioned prior art, the present invention provides a real-time pavement disease detection method and system based on deep learning, aiming to simultaneously identify and detect various pavement diseases such as cracks, grouting, potholes, repairs, markings, expansion joints, manhole covers, etc., and output pixel-level recognition results. By combining 2D and 3D images collected by the PavementVision3D system on a Digital Highway Data Collection Vehicle (DHDV) and inputting these images into the method designed based on the present invention for semantic segmentation, the detection accuracy and detection efficiency can be effectively improved, ensuring the accurate identification of various pavement diseases. It solves the technical problem that the existing semantic segmentation algorithm cannot accurately identify complex deformations or subtle disease features in images, which affects the accuracy of pavement automatic detection.
[0005] For the real-time pavement disease detection method based on deep learning, first, preprocess the image to be detected; then input the preprocessed image into the encoding layer for multiple downsamplings and convolutional feature extractions; then input the image output by the encoding layer into the decoding layer for multiple upsamplings and convolutional feature extractions; among them, the encoding and decoding layers of the first three resolution levels use dynamic adaptive convolution to dynamically extract features; the above encoding and decoding layers can be looped multiple times; finally, input the image after multiple encoding and decoding into the dynamic adaptive deformable convolution to output the detection result map; both upsampling and downsampling adopt dynamic position encoding, lossless slicing, and dynamic branch mechanisms, and the lossless slicing order of the upsampling is the reverse order of the lossless slicing order of the downsampling; the dynamic adaptive deformable convolution can adaptively adjust the shape of the convolution kernel according to the shape features of the diseases in the input image.
[0006] Further, the preprocessing includes:
[0007] Use the bilinear interpolation method to downsample the image and adjust it to a fixed size;
[0008] The 3D and 2D images are processed by channel fusion to form a single-batch image with 2 channels.
[0009] Further, the downsampling includes:
[0010] (a) Input the feature image data with the size of W×H×C, and assign position information to each pixel through dynamic position encoding, where W is the width, H is the height, and C is the number of channels;
[0011] (b) Enhance the network's understanding of the position information through the Spatial Attention Mechanism (SAM);
[0012] (c) The image is sliced horizontally and vertically in the form of (W / 2,H / 2) to obtain four feature sub-images with the size of (W / 2,H / 2,C), and the feature sub-images are stacked along the number of channels to obtain a feature sub-image with the size of (W / 2,H / 2,4C);
[0013] (d) The image processes the image features by dynamically selecting dilated convolution or depthwise separable convolution through a dynamic branching mechanism based on learning the size of the image.
[0014] Further, the upsampling includes:
[0015] (a) Input the feature image data with size W×H×C, and assign position information to each pixel through dynamic position encoding, where W is the width, H is the height, and C is the number of channels;
[0016] (b) Deepen the network's understanding of the position information through the spatial attention mechanism (SAM);
[0017] (c) According to the slicing order in the downsampling, combine and restore it in reverse order to an image with size 2W×2H×C / 4;
[0018] (d) The restored image processes the image features by dynamically selecting dilated convolution or depthwise separable convolution through a dynamic branching mechanism based on learning the size of the image.
[0019] Further, the dynamic adaptive deformable convolution includes:
[0020] (1) Input the feature image data with size W×H×C, where W is the width, H is the height, and C is the number of channels;
[0021] (2) Generate sampling point offsets for the feature image through depthwise separable convolution, focus on the disease itself, and output an offset with shape (W×H×2K), where K is the size of the depthwise separable convolution kernel, and each position offset contains K offsets;
[0022] (3) The output offsets indicate that the convolution kernel dynamically adjusts the sampling positions to achieve adaptive dynamic deformation.
[0023] (4) Perform feature extraction based on the sampling positions after adaptive dynamic deformation and output the result map.
[0024] Further, both the upsampling and the downsampling include:
[0025] When the image size is greater than or equal to the size after downsampling n times, the image size is relatively large, and dilated convolution will be used to enhance the model's context awareness of the image.
[0026] When the image size is smaller than the size after downsampling n times, the image resolution is too small, and depth convolution and pointwise convolution will be performed in sequence. While reducing the computational amount, it can effectively extract spatial information and improve the detection accuracy.
[0027] Furthermore, it is implemented using a semantic segmentation algorithm model. Before using the semantic segmentation algorithm model, multiple 2D road surface images of different highways and 3D images showing road surface depth information are collected to construct a semantic segmentation database for road surface diseases. The data in the semantic segmentation database for road surface diseases are labeled to obtain a dataset, and the semantic segmentation algorithm model is trained and verified.
[0028] A real-time road surface disease detection system based on deep learning includes a data acquisition module and a trained semantic segmentation algorithm model. The data acquisition module obtains 2D images and 3D images of the road surface to be detected in real time. The trained semantic segmentation algorithm model sequentially uses a preprocessing unit, an encoding layer unit, a decoding layer unit, and a dynamic adaptive deformable convolution to process the 2D images and 3D images of the road surface to be detected, and finally outputs a detection result map.
[0029] The beneficial effects of the present invention include:
[0030] The semantic segmentation algorithm proposed by the present invention, namely ShuttleNetS3, shows significant adaptive advantages in the detection scenario of complex deformed diseases. This algorithm can accurately capture the characteristics of complex and changeable road traffic scenarios, and flexibly and efficiently extract the morphology and boundaries of diseases, breaking through the bottleneck of the insufficient feature extraction ability of traditional algorithms in complex scenarios.
[0031] In terms of the ability to comprehensively detect multiple diseases, ShuttleNetS3 has achieved significant performance improvements in all detection categories, comprehensively surpassing mainstream semantic segmentation networks and the original ShuttleNet, demonstrating excellent generalization ability and robustness, and providing stronger technical support for the intelligent detection of complex disease scenarios.
[0032] At the level of actual engineering applications, the real-time semantic segmentation detection algorithm proposed by the present invention can generate high-precision disease semantic segmentation maps in real time during the inspection process of road detection vehicles at normal driving speeds, ensuring the timeliness and accuracy of detection results. This technology not only provides intuitive and highly accurate decision-making support for the intelligent monitoring and diagnosis of road diseases, but also significantly reduces the complexity and workload of road maintenance, opening up a new and efficient and intelligent way for road maintenance management work, showing broad practical application value and promotion prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a flowchart of the semantic segmentation algorithm model involved in the embodiment of the present application.
[0034] Figure 2 It is a legend of the flowchart of the semantic segmentation algorithm model involved in the embodiment of the present application.
[0035] Figure 3This is the structural diagram for implementing downsampling in the embodiments of this application.
[0036] Figure 4 This is the structural diagram for implementing upsampling in the embodiments of this application.
[0037] Figure 5 It shows the effects of dynamic adaptive deformable convolution and traditional convolution in the embodiments of this application.
[0038] Figure 6 These are the implementation steps of the dynamic adaptive deformable convolution in the embodiments of this application. Detailed implementation manners
[0039] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part rather than all of the embodiments of this application. Therefore, the detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative efforts fall within the scope of protection of this application.
[0040] The types and distribution characteristics of diseases on asphalt pavements are highly diverse. Different road sections may exhibit different forms of diseases. At the same time, there are also significant differences in the manifestation forms of each disease. This means that the detection technology not only needs to have a strong information retention ability but also be able to adaptively process complex deformations and fine disease details. Therefore, the existing lightweight semantic segmentation algorithms have poor robustness in small target detection, and deep semantic segmentation algorithms usually tend to lose important feature information, resulting in the inability to accurately identify complex deformation or fine disease features in images.
[0041] To overcome this problem, the present invention realizes the "bidirectional mapping" mechanism of upsampling and downsampling through the improved downsampling SPsampling and upsampling SPupsampling algorithms. These improvements significantly enhance the information retention ability of image feature information. Especially when dealing with fine diseases, they can effectively retain key texture information and avoid information loss.
[0042] Regarding the deformation problems of complex diseases, a dynamic adaptive deformation convolution method is innovatively used. By adaptively extracting the feature information of complex diseases, the flexibility and adaptability of the model are further improved. Compared with traditional convolution kernels, dynamic adaptive convolution has a higher degree of freedom and can flexibly learn context information during the detection process, effectively avoiding the performance limitations brought by fixed receptive fields. This convolution method not only improves the recognition ability of complex deformed diseases but also demonstrates excellent adaptability when dealing with diseases of different scales and types.
[0043] Combined with the above technologies, the semantic segmentation algorithm proposed by the present invention can significantly improve the detection accuracy in practical engineering applications and is superior to existing similar technologies in terms of processing speed, showing a broader application prospect.
[0044] Through these innovations, the proposed technology can provide reliable technical support for the precise positioning and automatic repair of pavement diseases and provide a more efficient solution for future road maintenance and management.
[0045] A real-time pavement disease detection method based on deep learning, as Figure 1-2 shown, first preprocesses the image to be detected; then inputs the preprocessed image into the encoding layer for multiple downsamplings and convolutional feature extractions. Each time of downsampling, the image size is halved, and the encoding layer can adjust the number of downsamplings according to the image size; then inputs the image output by the encoding layer into the decoding layer for multiple upsamplings and convolutional feature extractions. Each time of upsampling, the image size is expanded to twice the original size, and the number of upsamplings is equal to the number of downsamplings of the encoding layer; the above encoding and decoding layers can be cycled multiple times, and the image information of the previous encoding-decoding layer will be input into the next encoding-decoding layer through memory connections; finally, inputs the image after multiple encoding and decoding into the dynamic adaptive deformation convolution to output the detection result map; both upsampling and downsampling adopt dynamic position encoding, lossless slicing, and dynamic branch mechanisms, and the lossless slicing order of the upsampling is the reverse order of the lossless slicing order of the downsampling; the dynamic adaptive deformation convolution can adaptively adjust the shape of the convolution kernel according to the shape characteristics of the diseases in the input image.
[0046] Among them, the encoding and decoding layers of the first three resolution levels use dynamic adaptive deformation convolution to dynamically extract features, as Figure 1 shown in the first to third layers from top to bottom. Each time the image in the encoding layer is downsampled, the size will be reduced to 1 / 2, and each time the image in the decoding layer is upsampled, the size will be increased to 2 times the original. Only the thick arrows in the first three layers are dynamic adaptive deformation convolutions, and the following are all ordinary convolutions.
[0047] In another embodiment, the preprocessing includes:
[0048] The image is downsampled using the bilinear interpolation method and adjusted to a fixed size;
[0049] The 3D and 2D images are processed through channel fusion to form a single-batch image with 2 channels.
[0050] In another embodiment, the downsampling includes:
[0051] (a) Input the feature image data with the size of W×H×C, and assign position information to each pixel through dynamic position encoding, where W is the width, H is the height, and C is the number of channels;
[0052] (b) Enhance the network's understanding of the position information through the spatial attention mechanism (SAM);
[0053] (c) The image is sliced horizontally and vertically in the form of (W / 2, H / 2) to obtain four feature sub-images with the size of (W / 2, H / 2, C), and the feature sub-images are stacked along the number of channels to obtain a feature sub-image with the size of (W / 2, H / 2, 4C);
[0054] (d) The image passes through the dynamic branch mechanism and dynamically selects dilated convolution or depthwise separable convolution to process the image features by learning the size of the image.
[0055] In another embodiment, the upsampling includes:
[0056] (a) Input the feature image data with the size of W×H×C, and assign position information to each pixel through dynamic position encoding, where W is the width, H is the height, and C is the number of channels;
[0057] (b) Deepen the network's understanding of the position information through the spatial attention mechanism (SAM);
[0058] (c) According to the slicing order in the downsampling, the image is combined in reverse order to be restored to an image with the size of 2W×2H×C / 4;
[0059] (d) The restored image passes through the dynamic branch mechanism and dynamically selects dilated convolution or depthwise separable convolution to process the image features by learning the size of the image.
[0060] In another embodiment, the dynamic adaptive deformable convolution includes:
[0061] (1) Input the feature image data with the size of W×H×C, where W is the width, H is the height, and C is the number of channels;
[0062] (2) Generate the sampling point offset for the feature image through depthwise separable convolution, focus on the disease itself, and output an offset with the shape of (W×H×2K), where K is the size of the depthwise separable convolution kernel, and each position offset contains K offsets;
[0063] (3) The output offset indicates that the convolutional kernel dynamically adjusts the sampling position to achieve adaptive dynamic deformation.
[0064] (4) Feature extraction is performed based on the sampling position after adaptive dynamic deformation, and the result map is output.
[0065] In another embodiment, both the upsampling and downsampling include:
[0066] When the image size is greater than or equal to the size after downsampling n times, the image size is relatively large, and atrous convolution will be used to enhance the model's context awareness of the image.
[0067] When the image size is less than the size after downsampling n times, the image resolution is too small. Depth convolution and pointwise convolution will be performed in sequence to effectively extract spatial information and improve the detection accuracy while reducing the computational complexity.
[0068] In another embodiment, before using the semantic segmentation algorithm model, multiple 2D road surface images of different highways and 3D images showing road surface depth information are collected to construct a semantic segmentation database for road surface diseases. The data in the semantic segmentation database for road surface diseases is labeled to obtain a dataset for training and validating the semantic segmentation algorithm model.
[0069] In another embodiment, the implementation process is as follows:
[0070] 1.1. Data acquisition
[0071] The road foreground images are collected by the PavementVision3D system on the Digital Highway Data Vehicle (DHDV).
[0072] The DHDV achieves full coverage of a 4-meter-wide lane at a scanning frequency of 30KHz, ensuring a high-precision resolution of 1 millimeter.
[0073] The collected image data includes 2D road surface images and 3D images showing road surface depth information, and these data are used to construct a semantic segmentation database for road surface diseases.
[0074] This database not only contains road surface images but also records the corresponding geographical location information and timestamps, providing important basic information for subsequent data processing and analysis.
[0075] Given the differences in road surface conditions of different highways, it is necessary to collect road surface data from multiple different highways to ensure data diversity, thereby constructing a sample database suitable for training the algorithm model.
[0076] The road surface training sample database covers various types of road sections, including but not limited to national highways, provincial highways, ring roads, urban expressways, and road sections with various other road conditions, fully reflecting the extensiveness and comprehensiveness of data collection.
[0077] 1.2 Professional pixel-level data annotation
[0078] To ensure the high quality of model training, all the collected data are manually annotated by experts to ensure pixel-level accuracy.
[0079] The original size of each image is 4096×2048 pixels, representing a road surface section that is 4 meters long and 2 meters wide.
[0080] 2. Road surface disease semantic segmentation
[0081] A new semantic segmentation algorithm is adopted to simultaneously identify multiple road surface diseases and surface design features, and its model structure is as Figure 1 shown.
[0082] 2.1. Semantic segmentation image preprocessing
[0083] In the process of road surface disease semantic segmentation, the image preprocessing link is crucial.
[0084] The bilinear interpolation method is used to downsample the image and adjust it to a fixed size of 512×256 pixels, which not only ensures the processing efficiency but also maximally retains the key feature information.
[0085] Next, 3D and 2D images are processed through channel fusion to form a single batch of images with 2 channels and are input into the semantic segmentation algorithm model.
[0086] In this way, the semantic segmentation algorithm model can simultaneously combine pixel information and depth data, thus achieving a more accurate disease detection effect.
[0087] 2.2. Downsampling (Spsampling) algorithm for retaining height information
[0088] Max pooling is a fixed non-linear method that ignores important potential representations, thus easily causing the loss of detailed information in semantic segmentation tasks.
[0089] In this part, the combination of lossless slice downsampling and dynamic position encoding is used, and a dynamic branch mechanism is used before output to customize the convolution to enhance the learning ability of downsampling.
[0090] Lossless slice downsampling ensures that all image feature information is retained when the image size is halved, while dynamic position encoding makes up for the loss of the model's understanding of the sub-image position information during lossless slice downsampling.
[0091] The final dynamic branch provides the best convolution scheme for images of all sizes to better extract features. The SPsampling structure diagram is as Figure 3 shown, and its specific implementation steps are as follows:
[0092] (1) Input the feature image data with the size of W×H×C (width×height×channels), and assign position information to each pixel through dynamic position encoding;
[0093] (2) Enhance the network's understanding of position information through the spatial attention mechanism (SAM);
[0094] (3) The image is sliced horizontally and vertically in the form of (W / 2, H / 2) to obtain four feature sub-images with the size of (W / 2, H / 2, C), and they are stacked along the number of channels to obtain a feature sub-image with the size of (W / 2, H / 2, 4C).
[0095] (4) The image passes through the dynamic branch mechanism, and dynamically selects dilated convolution or depthwise separable convolution to process the image features by learning the size of the image.
[0096] When the number of image downsampling times is less than 4, the image size is relatively large, and dilated convolution with a kernel size of 3×3 and a stride of 2 will be used to enhance the model's context awareness of the image.
[0097] When the number of image downsampling times is greater than 4, the image resolution is too small, and it will sequentially pass through a depthwise convolution with a kernel size of 3×3 and a pointwise convolution with padding of 0 and a stride of 1, which can effectively extract spatial information and improve detection accuracy while reducing the computational cost.
[0098] 2.3. Upsampling (Spupsampling) algorithm for accurately restoring information
[0099] When detecting various road surface diseases and surface detail features, transposed convolution and bilinear interpolation upsampling are prone to missing detection of fine cracks and pixel artifacts in large-area road repair areas.
[0100] To achieve more accurate downsampling and upsampling, and thus highly restore image details, SPupsampling is proposed based on the principle of SPsampling. Through dynamic position encoding and inverse slicing mechanism, the "bidirectional mapping" mechanism of upsampling and downsampling is realized while ensuring the high restoration of information details.
[0101] This improvement not only significantly enhances the overall performance of the model, but also greatly improves the recognition ability of fine cracks and large-area road repair areas, breaking through the limitations of traditional methods in detail restoration. The structure diagram of SPupsampling is as Figure 4 shown, and its specific implementation steps are as follows:
[0102] (1) Input the feature image data with the size of W×H×C (width × height × channels), and assign position information pixel by pixel through dynamic position encoding;
[0103] (2) Deepen the network's understanding of position information through the spatial attention mechanism (SAM);
[0104] (3) The image is restored to an image with the size of (2W×2H×C / 4) in reverse order according to the slicing order in SPsampling. This step is similar to the lossless slicing step in SPsampling played in reverse, so this step is also lossless.
[0105] (4) The restored image passes through the dynamic branch mechanism, and dynamically selects dilated convolution or depthwise separable convolution to process the image features by learning the size of the image.
[0106] When the image size is greater than or equal to the size after downsampling 3 times, the image size is large, and dilated convolution with a kernel size of 3×3 and a stride of 2 will be used to enhance the model's context awareness of the image.
[0107] When the image size is smaller than the size after downsampling 3 times, the image resolution is too small, and it will successively pass through a depth convolution with a kernel size of 3×3 and a pointwise convolution with padding of 0 and a stride of 1. This step is also similar to the mechanism of SPsampling, thus forming a "bidirectional mapping" mechanism with SPsampling as a whole.
[0108] 2.4. Dynamic Adaptive Deformable Convolution
[0109] In deep learning, a traditional 3×3 convolution is usually used as the basic unit for feature extraction. However, in the road surface detection task, the targets often present extremely deformed complex shapes, such as cracks, potholes, and sealant joints.
[0110] Traditional convolution uses a filter with a fixed size and a preset sliding stride. The characteristics of this fixed mode limit its flexibility in the information extraction process, thus affecting the network's ability to capture diverse disease shapes.
[0111] To solve this problem, a dynamic adaptive deformable convolution mechanism is proposed, aiming to maintain the lightweight of the network while being able to adaptively adjust the shape of the convolution kernel according to the shape characteristics of the diseases in the input image. In this way, the network can more flexibly extract the potential target representations and context information in large-size images, thus enhancing the detection ability for complex deformed diseases. As Figure 5 shown, this method effectively improves the adaptability and recognition accuracy of the model to various disease morphologies.
[0112] As Figure 6 shown, the implementation steps of the dynamic adaptive deformable convolution are as follows:
[0113] (1) Input the feature image data with the size of W×H×C (width × height × channels).
[0114] (2) Generate the sampling point offsets for the feature image through depthwise separable convolution, focus on the disease itself, and output the offsets with the shape of (W×H×2K) (k is the size of the depthwise separable convolution kernel, and each position offset contains k offsets).
[0115] (3) The output offsets indicate that the convolution kernel dynamically adjusts the sampling positions to achieve adaptive dynamic deformation.
[0116] (4) Perform feature extraction based on the sampling positions after adaptive dynamic deformation and output the result map.
[0117] In another embodiment, it relates to a real-time road disease detection system based on deep learning, including a data acquisition module and a trained semantic segmentation algorithm model. The data acquisition module obtains the 2D image and 3D image of the road to be detected in real time. The trained semantic segmentation algorithm model sequentially uses a preprocessing unit, a downsampling unit, an upsampling unit, and a dynamic adaptive deformation convolution to process the 2D image and 3D image of the road to be detected and finally outputs the detection result map.
[0118] In another embodiment, for the proposed semantic segmentation algorithm model, tests were carried out on 2000 road images with other current advanced semantic segmentation algorithms, namely Unet, DeepLabv3+, PSPnet, Segformer-B5, and the original ShuttleNet. The test environment was python3.8, the deep learning platform was Pytorch1.11, and the GPU hardware configuration was NVIDIA RTX3070. The metric results of the test are shown in Table 1 below.
[0119] Table 1 Performance of Different Semantic Segmentation Algorithms
[0120]
[0121] Among them, the evaluation metrics adopted two of the most representative metrics in the current field of target detection intelligent algorithms, namely F1-Score and mIOU. The larger the values of these two metrics, the better the performance of the algorithm model. Compared with the current mainstream semantic segmentation algorithms: Unet, DeepLabv3+, PSPnet, Segformer-B5, and the original ShuttleNet, the algorithm proposed by the present invention has obvious advantages in the pixel-level recognition of various road diseases and surface design features.
[0122] In another embodiment, the semantic segmentation algorithm proposed by the present invention was tested on the Crack500 dataset together with the currently advanced semantic segmentation algorithms Unet, DeepLabv3+, PSPnet, Segformer-B5, and the original ShuttleNet. The performance metrics of the test are shown in Table 2 below.
[0123] Table 2 Performance of Different Semantic Segmentation Algorithms
[0124]
[0125]
[0126] Compared with the currently mainstream Unet, DeepLabv3+, PSPnet, Segformer-B5, and the original ShuttleNet, the semantic segmentation algorithm proposed by the present invention is more robust in the comprehensive strength of pixel-level recognition of various pavement diseases and surface design features.
[0127] The above-described embodiments merely represent the specific implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the protection scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the technical solution of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application.
Claims
1. A real-time detection method for pavement diseases based on deep learning, characterized in that, Preprocess the image to be detected first; Then input the preprocessed image into the encoding layer for multiple downsamplings and convolutional feature extractions; then input the image output by the encoding layer into the decoding layer for multiple upsamplings and convolutional feature extractions; Among them, the encoding layer and the decoding layer of the first three resolution levels use dynamic adaptive convolution to dynamically extract features; finally, input the encoded and decoded image into the dynamic adaptive deformable convolution to output the detection result map; Both upsampling and downsampling adopt dynamic position encoding, lossless slicing, and dynamic branching mechanisms, and the lossless slicing order of the upsampling is the reverse order of the lossless slicing order of the downsampling; The dynamic adaptive deformable convolution can adaptively adjust the shape of the convolution kernel according to the shape characteristics of the diseases in the input image; The dynamic adaptive deformable convolution includes: (a) Input the feature image data with size W×H×C, where W is the width, H is the height, and C is the number of channels; (b) Generate sampling point offsets for the feature image through depthwise separable convolution, focus on the diseases themselves, and output offsets with the shape of W×H×2K, where K is the size of the depthwise separable convolution kernel, and each position offset contains K offsets; (c) The output offsets indicate that the convolution kernel dynamically adjusts the sampling positions to achieve adaptive dynamic deformation; (d) Perform feature extraction based on the sampling positions after adaptive dynamic deformation and output the result map.
2. The real-time road surface disease detection method based on deep learning according to claim 1, characterized in that The encoding layer and the decoding layer loop multiple times, and finally input the image after multiple encodings and decodings into the dynamic adaptive deformable convolution to output the detection result map.
3. The real-time pavement disease detection method based on deep learning according to claim 1, wherein The preprocessing includes: Use the bilinear interpolation method to downsample the image and adjust it to a fixed size; Process 3D and 2D images through channel fusion to form a single-batch image with 2 channels.
4. The real-time pavement disease detection method based on deep learning according to claim 1, characterized in that The downsampling includes: (a) Input the feature image data with size W×H×C, and assign position information to each pixel through dynamic position encoding, where W is the width, H is the height, and C is the number of channels; (b) Enhance the network's understanding of the position information through the spatial attention mechanism; (c) Cut the image in the horizontal and vertical directions in the form of W / 2, H / 2 to obtain four feature sub-images with size W / 2, H / 2, C, and stack the feature sub-images along the number of channels to obtain a feature sub-image with size W / 2, H / 2, 4C; (d) The image passes through the dynamic branching mechanism and dynamically selects dilated convolution or depthwise separable convolution to process the image features by learning the size of the image.
5. The real-time road surface disease detection method based on deep learning according to claim 1, characterized in that, The upsampling includes: (a) Input the feature image data with size W×H×C, and assign position information to each pixel through dynamic position encoding, where W is the width, H is the height, and C is the number of channels; (b) Deepen the network's understanding of the position information through the spatial attention mechanism; (c) Combine and restore the image in the reverse order according to the slicing order in the downsampling to the image with size 2W×2H×C / 4; (d) The restored image passes through the dynamic branching mechanism and dynamically selects dilated convolution or depthwise separable convolution to process the image features by learning the size of the image.
6. The real-time pavement disease detection method based on deep learning according to claim 1, characterized in that, The upsampling and downsampling also include: When the image size is greater than or equal to the size after downsampling n times, use dilated convolution to enhance the model's context awareness of the image; When the image size is smaller than the size after downsampling n times, depth convolution and pointwise convolution are performed in sequence, which can effectively extract spatial information and improve detection accuracy while reducing the computational load.
7. The real-time road surface disease detection method based on deep learning according to claim 1, characterized in that It is implemented using a semantic segmentation algorithm model. Before using the semantic segmentation algorithm model, 2D road surface images of multiple different highways and 3D images showing road surface depth information are collected to construct a semantic segmentation database for road surface diseases. The data in the semantic segmentation database for road surface diseases are labeled to obtain a dataset, and the semantic segmentation algorithm model is trained and verified.
8. A real-time pavement disease detection system based on deep learning, characterized in that, The real-time road surface disease detection method based on deep learning described in any one of claims 1-7 is adopted, including a data acquisition module and a trained semantic segmentation algorithm model. The data acquisition module obtains 2D images and 3D images of the road surface to be detected in real time. The trained semantic segmentation algorithm model sequentially uses a preprocessing unit, an encoding layer unit, a decoding layer unit, and a dynamic adaptive deformable convolution to process the 2D images and 3D images of the road surface to be detected and finally outputs a detection result map.