Pavement disease detection method and system based on Mask R-CNN and transfer learning

The integration of Mask R-CNN and transfer learning with the Stable Diffusion 3.5 Large model and SCFN network addresses the limitations of existing road damage detection systems, enabling accurate and efficient detection of multiple road defects in complex backgrounds.

CN120318502AActive Publication Date: 2025-07-15GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510804461.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-15
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

The prior art has difficulty in multi-type and multi-object detection of road surface disease detection under complex backgrounds, and the existing methods are prone to repeated calculations and errors, which cannot meet the needs of fast and accurate detection.

Method used

Using a method of combining Mask R-CNN with transfer learning, the data set is expanded through the Stable Diffusion 3.5 Large model, and a variety of backbone feature extraction architectures and scale complementary feature fusion network are used to build an optimal detection model to realize multi-scale feature fusion and disease detection.

Benefits of technology

It significantly improves the accuracy and efficiency of multi-type and multi-target road disease detection in complex contexts, solves the multi-functional disease detection problem in complex contexts, and provides a highly reliable disease detection solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318502A_ABST
    Figure CN120318502A_ABST
Patent Text Reader

Abstract

The invention discloses a Mask R-CNN and transfer learning-based pavement disease detection method and system. The method comprises the steps of collecting pavement disease images to form a data set; carrying out disease marking on the data set by using a Label marking tool; carrying out expansion on the marked road disease data set by utilizing an optimized Stable Diffusion 3.5 Large model, and carrying out the expansion on the marked road disease data set by utilizing the optimized Stable Diffusion 3.5 Large model; performing multi-scale feature extraction from the expanded road disease data set through a plurality of backbone feature extraction architectures, fusing the extracted multi-scale features based on a scale complementary feature fusion network (SCFN), and constructing an optimal detection model based on a Mask R-CNN based on the fused multi-scale features; and quantifying a to-be-detected disease image by using the optimal detection model to complete disease detection under the complex road background. According to the method, the problems of multi-type and multi-target pavement disease detection and pixel quantification under a complex background are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of pavement disease detection, and particularly relates to a pavement disease detection method and system based on Mask R-CNN and transfer learning. Background Art

[0002] Pavement diseases such as pavement cracks (e.g., transverse cracks, longitudinal cracks, reticular cracks) and potholes are key factors affecting road safety and durability. These defects not only threaten the safety of vehicle driving, but also accelerate the degradation of the road structure, thus significantly shortening the service life of the highway. Therefore, achieving rapid and accurate detection of pavement diseases is of great significance for ensuring traffic safety and optimizing maintenance decisions. In recent years, artificial intelligence technologies (such as deep learning) have shown significant advantages in pavement disease detection, but there are still many challenges in practical applications.

[0003] Convolutional neural network is a deep learning algorithm that needs to be trained with a large amount of data to improve its performance. However, the current available amount of pavement disease data is limited, and there is no publicly available large-scale standard dataset for use. Therefore, correctly expanding pavement disease images is a prerequisite for CNN to obtain the best performance. With the continuous improvement of pavement maintenance requirements, the single functions of disease classification, disease location, and disease pixel segmentation can no longer meet the needs of precise maintenance, and the separate detection of multiple models takes a long time and is prone to problems such as multiple errors. Therefore, it is crucial to achieve multi-functional pavement disease detection. Currently, the road background has become increasingly complex, and most of the current research objects are still the research on cracks under a single background. Therefore, how to overcome the influence of complex backgrounds and achieve rapid and accurate disease detection is a prerequisite for road maintenance work. In the current road disease detection method, the image acquisition method is to continuously shoot or directly record road videos and process them into frames, and then directly input these frames into the model for detection. This method may cause a single disease to appear repeatedly in multiple pavement images or frames, resulting in repeated calculations in the model calculation, thus affecting the decision-making of the staff for maintenance work. Generally speaking, there is an urgent need to propose a road surface apparent disease detection method based on deep learning to achieve rapid and accurate detection of diseases. Summary of the Invention

[0004] To solve the problem of how to quickly and accurately detect pavement diseases under complex backgrounds, the present invention provides a pavement disease detection method and system based on Mask R-CNN and transfer learning, which effectively solves the problems of multi-type and multi-object pavement disease detection and pixel quantization under complex backgrounds.

[0005] To achieve the above object, the present invention provides the following solution:

[0006] A pavement disease detection method based on Mask R-CNN and transfer learning, the method comprising:

[0007] Step 1: Collect pavement distress images to form a dataset;

[0008] Step 2: Use the Labelme annotation tool to annotate the distress in the dataset;

[0009] Step 3: Use the optimized Stable Diffusion 3.5 Large model to augment the annotated road distress dataset;

[0010] Step 4: Extract multi-scale features from the augmented road distress dataset through multiple backbone feature extraction architectures, fuse the extracted multi-scale features based on the Scale Complementary Feature Fusion Network (SCFN), and construct an optimal detection model based on Mask R-CNN using the fused multi-scale features;

[0011] Step 5: Use the optimal detection model based on Mask R-CNN to quantify the distress images to be detected and complete distress detection in complex road backgrounds.

[0012] Preferably, in step 3, using the optimized Stable Diffusion 3.5 Large model to augment the annotated road distress dataset includes:

[0013] Use the VAE-encoder to extract the visual features of the input image;

[0014] According to the visual features, dynamically adjust the generation parameters and the prompt word embedding vector to generate synthetic images with structural diversity and style variations;

[0015] According to the synthetic images, introduce a domain prompt strategy based on the physical morphology of road distress to generate evolution features of cracks and potholes that are more in line with the actual situation;

[0016] Construct an automatic screening mechanism, use the trained convolutional neural network to detect the semantic consistency and feature fidelity of the synthetic images, and eliminate noise images and forged images.

[0017] Preferably, in step 4, the multiple backbone feature extraction architectures include: Select two deep convolutional networks, ResNet-50 and ResNet-101, as the main structures, and respectively combine the FPN (Pyramid Feature Network), C4 (the convolutional output of the fourth stage), and DC5 (the dilated convolution structure of the fifth stage) to construct feature extraction modules, and design six different feature paths;

[0018] Among them, the C4 structure extracts middle-level features from ResNet Stage 4; the DC5 architecture introduces dilated convolutions in Stage 5 to expand the receptive field and strengthen global semantic representation; the FPN structure fuses semantic information at different scales with the help of cross-layer lateral connections and upsampling mechanisms.

[0019] Preferably, in step 4, the multi-scale features extracted are fused based on a scale complementary feature fusion network, including:

[0020] Extract features with textures from the middle and shallow layers of the backbone network, where the resolution meets the preset requirements;

[0021] Obtain abstract semantic representations with global perception ability from the deep layers;

[0022] Meanwhile, fuse the multi-scale semantic features constructed in the FPN module to cover target information of different sizes.

[0023] Preferably, in step 4, based on the fused multi-scale features, an optimal detection model based on Mask R-CNN is constructed, including:

[0024] Combine ResNet50 / 101 as the backbone network, and respectively form eight types of feature combinations with FPN, C4, DC5, and SCFN, and substitute them into the Mask R-CNN framework to form a detection model based on Mask R-CNN;

[0025] Adjust the hyperparameters of the detection model to obtain the optimal detection model based on Mask R-CNN.

[0026] The present invention also provides a pavement disease detection system based on Mask R-CNN and transfer learning. The system is used to implement the aforementioned method, and the system includes: a collection module, a labeling module, an expansion module, a construction module, and a detection module;

[0027] The collection module is used to collect pavement disease images to form a dataset;

[0028] The labeling module is used to label the diseases in the dataset using the Labelme labeling tool;

[0029] The expansion module is used to expand the labeled road disease dataset using the optimized Stable Diffusion 3.5 Large model;

[0030] The building module is used to perform multi-scale feature extraction from the augmented road disease dataset through multiple backbone feature extraction architectures, fuse the extracted multi-scale features based on the scale complementary feature fusion network, namely SCFN, and construct an optimal detection model based on Mask R-CNN based on the fused multi-scale features.

[0031] The detection module is used to quantify the disease images to be detected by using the optimal detection model based on Mask R-CNN and complete the disease detection under the complex road background.

[0032] Preferably, the augmentation module includes: an extraction unit, an adjustment unit, an enhancement unit, and a screening unit.

[0033] The extraction unit is used to extract the visual features of the input image by using VAE-encoder.

[0034] The adjustment unit is used to dynamically adjust the generation parameters and the prompt word embedding vector according to the visual features to generate synthetic images with structural diversity and style variations.

[0035] The enhancement unit is used to introduce a domain prompt strategy based on the physical morphology of road diseases according to the synthetic images to generate evolutionary features of cracks and potholes that are more in line with the actual situation.

[0036] The screening unit is used to construct an automatic screening mechanism, and use the trained convolutional neural network to detect the semantic consistency and feature fidelity of the synthetic images, and eliminate the noise images and forged images.

[0037] Preferably, in the building module, the multiple backbone feature extraction architectures include: selecting two deep convolutional networks, ResNet-50 and ResNet-101, as the backbone structures, and respectively combining FPN (Pyramid Feature Network), C4 (the convolutional output of the fourth stage), and DC5 (the dilated convolution structure of the fifth stage) to construct a feature extraction module, and designing six different feature paths.

[0038] Among them, the C4 structure extracts the middle-level features from ResNet Stage 4; the DC5 architecture introduces dilated convolution in Stage 5 to expand the receptive field to enhance the global semantic expression; the FPN structure fuses the semantic information of different scales by means of cross-layer lateral connections and upsampling mechanisms.

[0039] Preferably, the building module includes: a resolution feature unit, a semantic representation unit, and a fusion unit.

[0040] The resolution feature unit is used to extract features with textures from the middle and shallow layers of the backbone network, and the resolution meets the preset requirements.

[0041] The semantic representation unit is used to obtain an abstract semantic representation with global perception ability from the deep layer;

[0042] The fusion unit is used to simultaneously fuse multi-scale semantic features constructed in the FPN module to cover target information of different sizes.

[0043] Preferably, in the construction module, an optimal detection model based on Mask R-CNN is constructed based on the fused multi-scale features, including:

[0044] Combining ResNet50 / 101 as the backbone network, respectively combining with FPN, C4, DC5 and SCFN to construct eight types of feature combinations, substituting them into the Mask R-CNN framework to form a detection model based on Mask R-CNN;

[0045] Adjust the hyperparameters of the detection model to obtain an optimal detection model based on Mask R-CNN.

[0046] Compared with the prior art, the beneficial effects of the present invention are:

[0047] 1. The present invention first introduces the improved Stable Diffusion 3.5 Large model into the road disease data enhancement process as the core sample generation engine. Different from traditional geometric transformation-based enhancement methods (such as rotation, mirroring, cropping, etc.), this method can generate crack and pothole images with significantly different structures and diverse visual styles through an explicitly guided Image-to-Image generation module, combined with prompt words (PromptEmbedding) and an image parameter control mechanism, significantly improving the disease diversity and scene complexity coverage ability of the dataset.

[0048] 2. The present invention uses the transfer learning method to compare the performance of two algorithms, Faster R-CNN and Mask R-CNN, in identifying road surface diseases under complex backgrounds. The results show that the performance of Mask R-CNN in road surface disease detection is better than that of Faster R-CNN. By adjusting the hyperparameters of the Mask R-CNN model, an optimal detection model is obtained, and this model is used to quantitatively verify the disease images to be detected, proving the effectiveness of this model in disease detection under complex road backgrounds. This method solves the problems of multi-type and multi-target road surface disease detection and pixel quantization under complex backgrounds, has high reliability, and will make a positive contribution to the field of intelligent detection of road surface diseases. Brief Description of the Drawings

[0049] To more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0050] Figure 1 Schematic diagram of some pavement disease images in the dataset of the embodiment of the present invention;

[0051] Figure 2 Schematic diagram showing some labeled transverse cracks, longitudinal cracks, reticular cracks and pothole images in the embodiment of the present invention;

[0052] Figure 3 Schematic diagram of the Stable Diffusion image augmentation process based on adaptive prompts and parameters in the embodiment of the present invention;

[0053] Figure 4 Schematic diagram of the scale complementary feature fusion network in the embodiment of the present invention;

[0054] Figure 5 Schematic diagram of the process of a pavement disease detection method based on Mask R-CNN and transfer learning in the embodiment of the present invention;

[0055] Figure 6 Schematic diagram comparing the performance of Faster R-CNN and Mask R-CNN under six backbone structures in the embodiment of the present invention, where (a) is the schematic diagram of AP and (b) is the schematic diagram of the test time;

[0056] Figure 7 Schematic diagram of the optimal detection model based on Mask R-CNN in the embodiment of the present invention. Detailed implementation manner

[0057] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0058] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the drawings and specific implementation manners.

[0059] Embodiment 1

[0060] The present invention discloses a pavement disease detection method based on Mask R-CNN and transfer learning, belonging to the field of pavement disease detection. Pavement diseases pose a serious threat to vehicle driving safety and also affect the service life of highways. Therefore, it is of great engineering significance to detect pavement diseases such as cracks and potholes in a timely and accurate manner. With the continuous improvement of pavement maintenance requirements, the single functions of disease classification, disease localization, and disease pixel segmentation can no longer meet the needs of precise maintenance. Moreover, the separate detection of multiple models takes a long time and is prone to problems such as multiple errors. Therefore, it is crucial to achieve multi-functional pavement disease detection. In addition, the current road background is becoming increasingly complex, and most of the current research objects are still cracks under a single background. It is still somewhat difficult to overcome the influence of complex backgrounds and achieve fast and accurate pavement disease detection. Thus, the present invention uses the Stable Diffusion 3.5 Large diffusion model to batch generate synthetic images containing multiple types of diseases (transverse cracks, longitudinal cracks, reticular cracks, and potholes) and complex backgrounds (shadows, low light, different weather conditions) through adaptive prompt guidance and input parameters (such as noise level, redrawing amplitude, noise type, etc.), significantly expanding the scale of the dataset and the diversity of scenarios. The transfer learning method is used to compare the performance of the Faster R-CNN and Mask R-CNN algorithms in identifying pavement diseases under complex backgrounds, proving that Mask R-CNN has better performance than Faster R-CNN in pavement disease detection. By adjusting the hyperparameters of the better Mask R-CNN model, the optimal detection model is obtained, and the optimal model is used to quantify the disease images to be detected, verifying the effect of the model in detecting diseases under complex road backgrounds. This method can effectively solve the problems of multi-type and multi-object pavement disease detection and pixel quantization under complex backgrounds. The specific implementation process is as follows:

[0061] As Figure 5 shown, this embodiment provides a pavement disease detection method based on Mask R-CNN and transfer learning, including the following implementation steps:

[0062] Step 1, collect pavement disease images to form a dataset.

[0063] The dataset used is composed of pavement disease images collected by on-vehicle cameras. This dataset contains 900 RGB pavement disease images with a resolution of 600×600, including different types of pavement diseases such as transverse cracks, longitudinal cracks, reticular cracks, and potholes (as Figure 1 ).

[0064] Step 2, use the Labelme annotation tool to annotate the diseases in the dataset.

[0065] For subsequent research needs, it is necessary to perform polygon annotation on road diseases and generate a file in JSON format. This is because object detection algorithms (such as Faster R-CNN) and instance segmentation algorithms (such as Mask R-CNN, BlendMask) are needed to identify and segment road diseases. In this process, we use the open-source image annotation tool Labelme to create and edit image annotation data. Labelme supports various annotation types such as points, lines, rectangles, and polygons, and can save image data and annotation information in JSON format. This makes data processing and analysis more convenient. During the annotation process, we first load the image folder, select the image to be annotated, and then manually annotate the boundaries of each disease. After the annotation is completed, Labelme will generate the corresponding JSON file, which contains information such as the name of the annotated image, the disease label, the polygon coordinates of the disease edge, and the size of the annotated image (such as Figure 2 ). In the present invention, the transfer models of the network are all pre-trained based on the COCO dataset, and all algorithms are implemented using the Python language.

[0066] Step 3, use the Stable Diffusion 3.5 Large model to expand the road disease dataset.

[0067] As Figure 4 shown, the generation enhancement process of road disease samples is based on the sample library constructed in the process of Step 1. Real image samples with clear annotations are preferably selected, covering various typical pavement damage forms such as transverse cracks, longitudinal cracks, reticular cracks, and potholes. On the locally deployed Stable Diffusion platform, this study introduced the StableDiffusion 3.5 Large model fine-tuned with the bridge disease dataset, and carried out sample expansion in combination with the Image-to-Image image generation module.

[0068] To improve the matching degree between the generated image and the target disease form, the system designs an adaptive generation mechanism based on image content: First, the VAE-encoder extracts visual features such as the texture and edge density of the input image, and the problem prompt words extract text vectors through the Bert model. Then, the image feature vectors and text vectors are aligned. The alignment strategy uses a cross-modal attention mechanism. Subsequently, the generation parameters and the prompt word embedding vectors are dynamically adjusted accordingly (such as Figure 3as shown), thus generating synthetic images with structural diversity and style variations. In addition, to further enhance the generation quality and relevance, a domain hint strategy based on the physical morphology of road diseases is introduced, making the generated content more in line with the evolution characteristics of actual cracks and potholes. Specifically, taking the crack image as an example, physical constraints are added through loss calculation in data augmentation, such as the coordination of the length and width of the crack, and the proportionality of the depth and shadow color.

[0069] In the data cleaning stage, a set of combined deep automated screening mechanisms is constructed. Specifically, the trained convolutional neural network is used to detect the semantic consistency and feature fidelity of the synthetic images, and the noise images and forged images are removed. Finally, combined with the manual review link, it is ensured that the images meet the training requirements in terms of visual integrity and disease expression, providing high-quality input for subsequent model training.

[0070] Step 4, compare the performance of Faster R-CNN and Mask R-CNN by migrating eight different backbone network structures.

[0071] To achieve high-precision detection of road diseases at different scales, this paper systematically evaluates the detection performance based on a variety of backbone feature extraction architectures. Specifically, two commonly used deep convolutional networks, ResNet-50 and ResNet-101, are selected as the backbone structures, and the feature extraction modules are constructed by combining FPN (Pyramid Feature Network), C4 (convolutional output of the fourth stage), and DC5 (atrous convolution structure of the fifth stage), respectively, as well as the SCFN network feature extraction module, thus designing eight different feature paths.

[0072] Among these structures, the C4 structure extracts the middle-level features from ResNet Stage 4 and is suitable for the detection of medium-sized targets; the DC5 architecture introduces atrous convolution in Stage 5 to expand the receptive field to enhance the global semantic expression; the FPN structure uses cross-layer lateral connections and upsampling mechanisms to fuse semantic information at different scales, significantly enhancing the model's robustness to small-sized targets and complex background interference.

[0073] To more effectively integrate the feature information from different levels, as Figure 4 shown, the present invention proposes a Scale-Complementary Feature Network (SCFN), aiming to improve the robustness and accuracy of the detection model when dealing with disease targets of different sizes. The network structure designs a dual-channel feature acquisition mechanism, focusing on the texture details of the shallow layer and the semantic structure information of the deep layer respectively, to achieve the joint modeling of small targets and large targets.

[0074] Specifically, the proposed SCFN network extracts high-resolution features with rich texture details from the middle and shallow layers of the backbone network (such as the C4 structure), and obtains more abstract semantic representations with stronger global perception ability from the deep layers (such as the DC5 structure). At the same time, it fuses the multi-scale semantic features constructed in the FPN module to comprehensively cover target information of different sizes.

[0075] To achieve efficient fusion between multi-scale features, first, the feature maps output by each branch are processed to unify the spatial dimensions (for example, resolution matching is achieved through bilinear interpolation). Immediately afterwards, a multi-channel attention mechanism (such as the improved EMA module) is introduced to dynamically model the importance of different-scale channels, guiding the network to flexibly adjust the fusion weights of features at each scale according to the context content of the current input image, thereby enhancing the perception ability of significant regions.

[0076] Among them, the improved EMA module is as follows. The module inputs a tensor, denoted as , where X is a multi-channel feature map containing multi-scale information generated by multiple feature extraction modules (including the C4 structure, DC5 structure, and FPN module). The process is as follows:

[0077] First, the input features are compressed in channels. A 1×1 convolution is used to reduce the number of channels from to (r is the dimensionality reduction ratio factor, and r usually takes values in [4, 8, 16]). Subsequently, LayerNorm is applied for normalization in the spatial dimension ( ) to standardize the feature distribution and improve stability. Then it enters the multi-scale feature extraction stage. The dimension-reduced features are respectively fed into three parallel branches: the first branch is processed by a 7×7 depthwise convolution. The second branch is processed by a 5×5 depthwise convolution. The third branch is processed by a 3×3 depthwise convolution. Each branch extracts features with different receptive fields to effectively perceive multi-scale spatial information. Subsequently, the outputs of the three branches are concatenated in the channel dimension to form comprehensive multi-scale features. The concatenated features are mixed in channels through a 1×1 convolution and LayerNorm is applied again for normalization. Finally, the number of channels of the features is restored to the initial , the output is used as the final result of the improved EMA module. In terms of the selection of the target detection framework, the present invention respectively introduces Faster R-CNN and Mask R-CNN. The former uses the Region Proposal Network (RPN) in conjunction with the regression and classification heads to construct the detection process; the latter adds a mask prediction branch on this basis to achieve fine-grained instance segmentation. For each group of structures (such as FPN / C4 / DC5 / SCFN), the RPN and RoI Align modules are uniformly connected to generate candidate boxes and align features.

[0078] Combined with ResNet50 / 101 as the backbone network, eight types of feature combinations are constructed with FPN, C4, DC5, and the SCFN proposed by the present invention respectively, and substituted into the Faster R-CNN and Mask R-CNN frameworks to form a total of sixteen model structures (as shown in Table 1). The model naming method is like "F-R50-FPN", which represents the Faster R-CNN + ResNet-50 + FPN structure. All models use transfer learning to load the COCO pre-trained weights and are fine-tuned on the disease image dataset amplified by the Stable Diffusion 3.5 Large guided by the adaptive prompt words proposed by the invention. Finally, the performance of each model is compared and analyzed through multi-dimensional indicators such as accuracy index (Accuracy), average precision (AP), and inference delay (Inference Time) to select the optimal solution.

[0079] Table 1 Backbone structures of Faster R-CNN and Mask R-CNN

[0080]

[0081] To evaluate the detection effect of the model, the study selected AP as the core evaluation index. AP calculates the area under the curve by calculating the precision and recall at different thresholds to measure the average performance of the model on all categories. Among them, precision represents the proportion of actual positive samples among the positive samples predicted by the model, while recall represents the proportion of actual positive samples correctly identified by the model. The calculation formulas for both are as follows:

[0082]

[0083]

[0084]

[0085] where T P(True Positive): It represents the number of positive samples predicted by the model and actually being positive samples, that is, the number of positive samples correctly predicted by the model as positive samples. F P (False Positive): It represents the number of positive samples predicted by the model but actually being negative samples, that is, the number of negative samples wrongly predicted by the model as positive samples. F N (False Negative): It represents the number of negative samples predicted by the model but actually being positive samples, that is, the number of positive samples wrongly predicted by the model as negative samples. In the present invention, the image average test time consumption and AP are used as two metrics to comprehensively evaluate the Faster R-CNN and Mask R-CNN algorithms, as Figure 6 shown. The experimental results show that Mask R-CNN performs better in comprehensive performance compared with Faster R-CNN.

[0086] Step 5: Adjust the hyperparameters of the better model to obtain the optimal detection model, use this model to quantify the disease images to be detected, and test its effect on disease detection under complex road backgrounds.

[0087] In the specific implementation process of the present invention, in order to realize the adaptability evaluation of different detection algorithms in the pavement disease recognition task, the average inference time consumption and average precision (Average Precision, AP) are selected as the evaluation criteria, and the performance tests of various typical detection algorithms are carried out, including representative methods such as Faster R-CNN and Mask R-CNN.

[0088] As Figure 6 shown, by uniformly evaluating the above models under the same dataset and hardware environment, and based on the comprehensive performance of the two metrics, in the subsequent optimization design, it is determined to use Mask R-CNN as the overall framework, and subsequently the ResNet-101 backbone network will be selected and the SCFN feature extraction module will be integrated as the basic detection model.

[0089] On this basis, in order to further improve the accuracy and adaptability of disease detection, the hyperparameters of the selected model are tuned, and the impact of the learning rate on the detection performance is mainly investigated. The present invention relies on the standard training process and sets multiple learning rates (including but not limited to 0.01, 0.001, 0.0001) respectively for model training to explore its specific impact on the model fitting ability and detection stability.

[0090] Finally, the Mask R-CNN (M-R101-SCFN structure) model with a learning rate of 0.001 is selected as the target detection model based on the comprehensive training performance, and using this model as the core framework, the disease area extraction and subsequent quantitative analysis processing are completed, asFigure 7 as shown

[0091] Embodiment 2

[0092] The present invention also provides a pavement disease detection system based on Mask R-CNN and transfer learning. The system is used to implement the foregoing method, and the system includes: a collection module, a labeling module, an expansion module, a construction module, and a detection module;

[0093] The collection module is used to collect pavement disease images to form a data set;

[0094] The labeling module is used to label diseases in the data set using the Labelme labeling tool;

[0095] The expansion module is used to expand the labeled road disease data set using the optimized Stable Diffusion 3.5 Large model;

[0096] The construction module is used to perform multi-scale feature extraction from the expanded road disease data set through multiple backbone feature extraction architectures, fuse the extracted multi-scale features based on the scale complementary feature fusion network, i.e., SCFN, and construct an optimal detection model based on Mask R-CNN based on the fused multi-scale features;

[0097] The detection module is used to quantify the disease images to be detected using the optimal detection model based on Mask R-CNN and complete disease detection in a complex road background.

[0098] In this embodiment, the expansion module includes: an extraction unit, an adjustment unit, an enhancement unit, and a screening unit;

[0099] The extraction unit is used to extract visual features of the input image using VAE-encoder;

[0100] The adjustment unit is used to dynamically adjust the generation parameters and the prompt word embedding vector according to the visual features to generate synthetic images with structural diversity and style variations;

[0101] The enhancement unit is used to introduce a domain prompt strategy based on the physical morphology of road diseases according to the synthetic images to generate evolution features of cracks and potholes that are more in line with the actual situation;

[0102] The screening unit is used to construct an automatic screening mechanism, detect the semantic consistency and feature fidelity of the synthetic images using a trained convolutional neural network, and eliminate noise images and forged images.

[0103] In this embodiment, in the construction module, multiple backbone feature extraction architectures include: selecting two deep convolutional networks, ResNet-50 and ResNet-101, as the backbone structures, and respectively combining FPN (i.e., Pyramid Feature Network), C4 (i.e., the convolutional output of the fourth stage), and DC5 (i.e., the dilated convolution structure of the fifth stage) to construct feature extraction modules, and designing six different feature paths;

[0104] Among them, the C4 structure extracts middle-level features from ResNet Stage 4; the DC5 architecture introduces dilated convolution in Stage 5 to expand the receptive field to enhance global semantic expression; the FPN structure uses cross-layer lateral connections and upsampling mechanisms to fuse semantic information of different scales.

[0105] In this embodiment, the construction module includes: a resolution feature unit, a semantic representation unit, and a fusion unit;

[0106] The resolution feature unit is used to extract features with textures whose resolution meets the preset requirements from the middle and shallow layers of the backbone network;

[0107] The semantic representation unit is used to obtain abstract semantic representations with global perception ability from the deep layers;

[0108] The fusion unit is used to simultaneously fuse multi-scale semantic features constructed in the FPN module to cover target information of different sizes.

[0109] In this embodiment, in the construction module, based on the fused multi-scale features, an optimal detection model based on Mask R-CNN is constructed, including:

[0110] Combining ResNet50 / 101 as the backbone network, respectively combining with FPN, C4, DC5, and SCFN to construct eight types of feature combinations, and substituting them into the Mask R-CNN framework to form a detection model based on Mask R-CNN;

[0111] Adjusting the hyperparameters of the detection model to obtain an optimal detection model based on Mask R-CNN.

[0112] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention should all fall within the protection scope determined by the claims of the present invention.

Claims

1. A pavement disease detection method based on Mask R-CNN and transfer learning, characterized in that, The method includes: Step 1: Collect pavement distress images to form a dataset; Step 2: Use the Labelme annotation tool to annotate the distress in the dataset; Step 3: Use the optimized Stable Diffusion 3.5 Large model to augment the annotated road distress dataset; Step 4: Perform multi-scale feature extraction from the augmented road distress dataset through multiple backbone feature extraction architectures. Based on the scale complementary feature fusion network, i.e., SCFN, fuse the extracted multi-scale features. Based on the fused multi-scale features, construct an optimal detection model based on Mask R-CNN; Step 5: Use the optimal detection model based on Mask R-CNN to quantify the distress images to be detected and complete the distress detection in complex road backgrounds.

2. The method according to claim 1, characterized in that In the said Step 3, using the optimized StableDiffusion 3.5 Large model to augment the annotated road distress dataset includes: Use VAE-encoder to extract the visual features of the input image; According to the visual features, dynamically adjust the generation parameters and the prompt word embedding vector to generate synthetic images with structural diversity and style variations; According to the synthetic images, introduce a domain prompt strategy based on the physical morphology of road distress to generate evolution features of actual cracks and potholes that are more fitting; Construct an automated screening mechanism, and use the trained convolutional neural network to detect the semantic consistency and feature fidelity of the synthetic images, and eliminate noise images and forged images.

3. The method according to claim 1, wherein In the said Step 4, the multiple backbone feature extraction architectures include: Select two deep convolutional networks, ResNet-50 and ResNet-101, as the backbone structures, and respectively combine the FPN, i.e., Pyramid Feature Network, C4, i.e., the convolutional output of the fourth stage, and DC5, i.e., the dilated convolution structure of the fifth stage, to construct feature extraction modules, and design six different feature paths; Among them, the C4 structure extracts the middle-level features from ResNet Stage 4; the DC5 architecture introduces dilated convolution in Stage 5 to expand the receptive field to enhance the global semantic expression; the FPN structure uses cross-layer lateral connections and upsampling mechanisms to fuse semantic information of different scales.

4. The method according to claim 3, wherein In the said Step 4, based on the scale complementary feature fusion network, fusing the extracted multi-scale features includes: Extract features with textures and resolutions meeting the preset requirements from the middle and shallow layers of the backbone network; Obtain abstract semantic representations with global perception ability from the deep layers; At the same time, fuse the multi-scale semantic features constructed in the FPN module to cover target information of different sizes.

5. The method according to claim 4, wherein In the said Step 4, based on the fused multi-scale features, constructing an optimal detection model based on Mask R-CNN includes: Combine ResNet50 / 101 as the backbone network, respectively combine with FPN, C4, DC5, and SCFN to construct eight types of feature combinations, and substitute them into the Mask R-CNN framework to form a detection model based on Mask R-CNN; Adjust the hyperparameters of the detection model to obtain the optimal detection model based on Mask R-CNN.

6. A road surface disease detection system based on Mask R-CNN and transfer learning, the system is used to implement the method described in any one of claims 1-5, characterized in that, The system includes: a collection module, an annotation module, an augmentation module, a construction module, and a detection module; The collection module is used to collect road disease images to form a dataset; The annotation module is used to annotate the diseases in the dataset using the Labelme annotation tool; The augmentation module is used to augment the annotated road disease dataset using the optimized Stable Diffusion 3.5 Large model; The construction module is used to perform multi-scale feature extraction from the augmented road disease dataset through multiple backbone feature extraction architectures, fuse the extracted multi-scale features based on the Scale Complementary Feature Fusion Network (SCFN), and construct the optimal detection model based on Mask R-CNN based on the fused multi-scale features; The detection module is used to quantify the disease images to be detected using the optimal detection model based on Mask R-CNN and complete the disease detection in the complex road background.

7. The system according to claim 6, wherein The augmentation module includes: an extraction unit, an adjustment unit, an enhancement unit, and a screening unit; The extraction unit is used to extract the visual features of the input image using VAE-encoder; The adjustment unit is used to dynamically adjust the generation parameters and the prompt word embedding vector according to the visual features to generate synthetic images with structural diversity and style variations; The enhancement unit is used to introduce a domain prompt strategy based on the physical morphology of road diseases according to the synthetic images to generate evolution features of cracks and potholes that are more in line with the actual situation; The screening unit is used to construct an automatic screening mechanism, detect the semantic consistency and feature fidelity of the synthetic images using a trained convolutional neural network, and eliminate noise images and forged images.

8. The system according to claim 6, characterized in that In the construction module, the multiple backbone feature extraction architectures include: selecting two deep convolutional networks, ResNet-50 and ResNet-101, as the backbone structures, and respectively combining the FPN (Pyramid Feature Network), C4 (the convolutional output of the fourth stage), and DC5 (the dilated convolution structure of the fifth stage) to construct feature extraction modules, and designing six different feature paths; Among them, the C4 structure extracts the middle-level features from ResNet Stage 4; the DC5 architecture introduces dilated convolution in Stage 5 to expand the receptive field to enhance the global semantic expression; the FPN structure fuses the semantic information of different scales through cross-layer lateral connections and upsampling mechanisms.

9. The system according to claim 8, wherein The construction module includes: a resolution feature unit, a semantic representation unit, and a fusion unit; The resolution feature unit is used to extract features with textures from the middle and shallow layers of the backbone network whose resolution meets the preset requirements; The semantic representation unit is used to obtain abstract semantic representations with global perception ability from the deep layers; The fusion unit is used to fuse the multi-scale semantic features constructed in the FPN module simultaneously to cover target information of different sizes.

10. The system according to claim 9, wherein In the described building block, an optimal detection model based on Mask R-CNN is constructed based on the fused multi-scale features, including: Combining ResNet50 / 101 as the backbone network, eight types of feature combinations are constructed with FPN, C4, DC5, and SCFN respectively, and substituted into the Mask R-CNN framework to form a detection model based on Mask R-CNN; Adjust the hyperparameters of the detection model to obtain the optimal detection model based on Mask R-CNN.

Citation Information

Patent Citations

  • Mask R-CNN-based underground drainage pipeline disease pixel level detection method

    CN112686217A

  • Domain alignment for object detection domain adaptation tasks

    US20210312232A1