Pavement defect detection method, device and equipment and storage medium
By using convolutional neural network and Grad-CAM technology in road defect detection, the feature map of pavement images is converted and evaluated, the target hidden layer is selected and the network is reconstructed, and the existing technology is too high in dependence on labeled data is solved, achieving efficient and accurate road defect detection.
Patent Information
- Application Number
- CN202510075092.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The existing pavement defect detection methods are too dependent on labeled data, resulting in insufficient interpretation and limited adaptability.
By inputting the pavement image into a convolutional neural network for feature extraction, the feature map is converted into a heat map using Grad-CAM, only pavement features that reach the preset defect level are displayed, and defect authenticity evaluation is performed on the pavement features displayed in the heat map, select a feature map set that meets the preset accuracy, determine the target hidden layer, and reconstruct the convolutional neural network based on the hidden layer to perform road defect detection.
It realizes that without large amounts of labeled data, locate the significant area of defects, reduces dependence on labeled data, improves detection effect, and reduces the consumption of computing resources.
Smart Images

Figure CN120163760A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to a road surface defect detection method, device, equipment, and storage medium. Background Art
[0002] With the rapid development of society, the connection between cities is becoming closer and closer, and the ground road traffic system is becoming more and more indispensable in daily life. In China, asphalt concrete roads with semi-rigid bases are generally adopted. Due to years of exposure to wind, sun, rain erosion, and repeated rolling by heavy vehicles, over time, structural damage will affect the durability of the road, and aging will inevitably lead to a decline in structural performance, thus affecting the safety of vehicle driving.
[0003] With the development of deep learning methods, the detection of continuous defects such as cracks is achieved through deep learning based on convolutional neural networks (CNNs) and semantic segmentation, which is also the most advanced application currently. However, in the current method, a large amount of data needs to be collected and trained in advance to ensure a certain degree of accuracy during defect detection. That is, there are generally problems such as high dependence on labeled data, insufficient interpretability, or limited adaptability. Summary of the Invention
[0004] This application provides a road surface defect detection method, device, equipment, and storage medium to solve the problem of excessive dependence on labeled data in existing detection methods.
[0005] In the first aspect of this application, a road surface defect detection method is provided. The method includes: inputting the acquired road surface image into a convolutional neural network for feature extraction, and extracting the feature map sets extracted by each hidden layer in the convolutional neural network. Among them, corresponding label information is set for each road surface feature in each feature map of the feature map set, and the label information is used to indicate the defect degree of the road surface feature; using Grad-CAM to convert each feature map set into a heat map, where only the road surface features reaching a preset defect degree are shown in the heat map; evaluating the defect authenticity of the road surface features shown in each heat map, and selecting the feature map set that meets the preset accuracy based on the evaluation result to determine the corresponding target hidden layer; reconstructing the convolutional neural network based on the target hidden layer to obtain a target convolutional neural network; using the target convolutional neural network to re-extract features from the road surface image, and detecting road surface defects based on the extracted features.
[0006] In a feasible implementation, converting each of the feature maps into a heat map using Grad-CAM includes: extracting the label information of the feature maps in each of the hidden layers using Grad-CAM, and extracting the concerned features in the feature maps based on the relationship between the label information and the defect degree to obtain the pavement defect concerned area; performing visualization processing on the pavement defect concerned area, and generating a heat map of the corresponding hidden layer based on the result of the visualization processing.
[0007] In a feasible implementation, performing visualization processing on the pavement defect concerned area includes: calculating the mean value of multiple feature maps corresponding to the pavement defect concerned area using a visualization processing formula to obtain the visualization result of the pavement defect concerned area.
[0008] In a feasible implementation, generating a heat map of the corresponding hidden layer based on the result of the visualization processing includes: squeezing multiple feature maps of each of the visualized hidden layers and calculating the average eigenvalue; generating a heat map of the corresponding hidden layer based on the average eigenvalue using a heat map generation rule.
[0009] In a feasible implementation, generating a heat map of the corresponding hidden layer based on the average eigenvalue using a heat map generation rule includes: using the ReLU activation function to remove the features in all the average eigenvalues that contribute nothing or have a negative impact on defect recognition, and performing normalization processing on the feature map after removal; determining the target color corresponding to the normalized feature map based on a preset color mapping rule to generate a heat map of the corresponding hidden layer.
[0010] In a feasible implementation, evaluating the defect authenticity of the pavement features shown in each of the heat maps, and selecting a feature map that meets a preset accuracy based on the evaluation result to determine the corresponding target hidden layer includes: calculating the distance between the pavement features shown in each of the heat maps and the features in the real defect area using a preset scoring formula to obtain a score; using a selective layer attention network to select a target feature map with a score reaching a preset similarity value, and determining the target hidden layer corresponding to the target feature map based on the corresponding relationship between the feature map and the hidden layer.
[0011] In a feasible implementation, reconstructing the convolutional neural network based on the target hidden layer to obtain a target convolutional neural network includes: retaining all the target hidden layers in the convolutional neural network, and loading the pre-trained weights of each of the target hidden layers for configuration to obtain a target convolutional neural network.
[0012] The second aspect of the present application provides a road surface defect detection device, including: an extraction module, configured to input the acquired road surface image into a convolutional neural network for feature extraction, and extract the feature map sets extracted by each hidden layer in the convolutional neural network, wherein corresponding label information is set for each road surface feature in each feature map of the feature map set, and the label information is used to indicate the defect degree of the road surface feature; a conversion module, configured to convert each of the feature map sets into a heat map by using Grad-CAM, wherein only the road surface features reaching a preset defect degree are displayed in the heat map; an evaluation module, configured to evaluate the authenticity of the road surface features displayed in each of the heat maps, and select the feature map set meeting the preset accuracy based on the evaluation result to determine the corresponding target hidden layer; a reconstruction module, configured to reconstruct the convolutional neural network based on the target hidden layer to obtain a target convolutional neural network; a detection module, configured to use the target convolutional neural network to re-extract features from the road surface image, and detect road surface defects based on the extracted features.
[0013] The third aspect of the present application provides an electronic device, including: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor calls the instructions in the memory to enable the electronic device to execute the above-mentioned road surface defect detection method.
[0014] The fourth aspect of the present application provides a computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium, and when the instructions run on a computer, the computer is enabled to execute the above-mentioned road surface defect detection method.
[0015] In the technical solution provided by this application, by extracting the feature maps extracted from the road surface image by each hidden layer in the convolutional neural network, where corresponding label information is set for each road surface feature in each feature map of the feature map set, and the label information is used to indicate the defect degree of the road surface feature; using Grad-CAM to convert each feature map set into a heat map that only shows the road surface features reaching the preset defect degree; evaluating the authenticity of the defects of the road surface features shown in each heat map, selecting the feature map set that meets the preset accuracy based on the evaluation result, and determining the corresponding target hidden layer; reconstructing the convolutional neural network based on the target hidden layer to obtain the target convolutional neural network; using the target convolutional neural network to re-extract features from the road surface image, and detecting road surface defects based on the extracted features. In this application, by using Grad-CAM to visualize the features extracted by each layer in the convolutional neural network, and then evaluating the authenticity of the defects of the features extracted by each layer, in order to select the target hidden layer with a high similarity to the real defect area, and finally reconstructing the convolutional neural network based on the target hidden layer, this method realizes the internal adjustment of the feature extraction layer in the convolutional neural network, enabling it to select the most informative and accurate feature layers, realizing the localization of the significant area of the defect even without a large amount of labeled data, reducing the dependence on labeled data, improving the detection effect and reducing the consumption of computing resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic diagram of an embodiment of the road surface defect detection method in an embodiment of this application;
[0017] Figure 2 It is a schematic diagram of another embodiment of the road surface defect detection method in an embodiment of this application;
[0018] Figure 3 It is the feature map set of each hidden layer represented by Grad-CAM in an embodiment of this application;
[0019] Figure 4 It is the calculation process of defect authenticity evaluation in an embodiment of this application;
[0020] Figure 5 It is a schematic diagram of the selective attention module in an embodiment of this application;
[0021] Figure 6 It is the maximum pooling process in an embodiment of this application;
[0022] Figure 7 It is a schematic diagram of an embodiment of the road surface defect detection device in an embodiment of this application;
[0023] Figure 8 It is a schematic diagram of another embodiment of the road surface defect detection device in an embodiment of this application;
[0024] Figure 9 This is a schematic diagram of an embodiment of the electronic device in the embodiments of the present application. Detailed implementation manners
[0025] The embodiments of the present application provide a road surface defect detection method, device, equipment and storage medium.
[0026] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "include" or "have" and any deformation thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or equipment that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.
[0027] It can be understood that the execution subject of the present application can be a road surface defect detection device, or a terminal or a server. Specifically, it is not limited here. The embodiments of the present application take the server as the execution subject as an example for illustration.
[0028] For ease of understanding, the specific process of the embodiments of the present application is described below. Please refer to Figure 1 , an embodiment of the road surface defect detection method in the embodiments of the present application includes:
[0029] 101. Input the obtained road surface image into a convolutional neural network for feature extraction, and extract the feature map sets extracted by each hidden layer in the convolutional neural network. Each feature map in the feature map set is provided with corresponding label information for each road surface feature, and the label information is used to indicate the defect degree of the road surface feature.
[0030] It should be noted that the road surface image refers to the road surface image after preprocessing, and the preprocessing can be image cleaning, image enhancement, etc.
[0031] The convolutional neural network is a model for analyzing road surface defect features or road surface information obtained by training with training data in advance. The model is composed of multiple hidden layers, and each hidden layer analyzes the road surface image respectively, extracts and outputs road surface features to obtain corresponding feature map sets.
[0032] That is, the obtained road surface image is input into the model, and each hidden layer of the model extracts features and outputs the corresponding feature map set. Then, the road surface defects are identified and detected for each feature map set, and a label information is generated based on the detection result. At the same time, the corresponding features are marked with the label information to be used as a part of the feature map set.
[0033] In practical applications, the convolutional neural network also includes a recognition network structure for road surface defects. After the features are extracted by each hidden layer, they are input into the recognition network structure to complete the recognition of whether there are defects and generate a label information. Among them, the label information can be color, mark, symbol, etc.
[0034] 102. Use Grad-CAM to convert each feature map set into a heat map, and only the road surface features reaching the preset defect degree are shown in the heat map.
[0035] In this embodiment, when using Grad-CAM for heat map conversion, specifically, Grad-CAM is used to extract the features in each feature map set that meet the defect determination for visualization processing to obtain the heat map of each feature map set. Specifically, Grad-CAM performs an unweighted average calculation on each feature map set, and compares the calculated result with the class-specific Grad-CAM heat map to screen out the features that meet the defect determination and perform visual display to generate a heat map. It should be noted that the visual display can be specifically implemented by using a color representation method, that is, matching the corresponding color according to the level of defect determination, and then replacing the color at the position of the feature in the feature map set with the corresponding color to achieve visualization.
[0036] 103. Evaluate the defect authenticity of the road surface features shown in each heat map, and select the feature map set that meets the preset accuracy based on the evaluation result to determine the corresponding target hidden layer.
[0037] Specifically, the matching degree between the road surface features shown in the feature map and the real road surface defects is calculated. The matching degree includes the distance between adjacent two road surface features and the similarity of the features themselves, and then a weighting coefficient is introduced to generate a score.
[0038] Based on the score, the better feature map sets are screened out, and then based on the output relationship between the screened feature map sets and the hidden layers, the target hidden layer corresponding to the better feature map sets is matched; among them, the screening method can be carried out by using a similarity comparison method.
[0039] 104. Reconstruct the convolutional neural network based on the target hidden layer to obtain the target convolutional neural network.
[0040] In this embodiment, the reconstruction here can be achieved by means of elimination and masking, that is, eliminating the hidden layers other than the target hidden layer in the original convolutional neural network to generate a new convolutional neural network, or setting masking identifiers for the hidden layers other than the target hidden layer in the original convolutional neural network. After reusing the convolutional neural network for feature extraction, the features output by each hidden layer are screened according to the masking identifiers, or the hidden layers with masking identifiers do not output features.
[0041] 105. Use the target convolutional neural network to extract features from the road surface image again, and detect road surface defects based on the extracted features.
[0042] It can be understood that when the target convolutional neural network extracts features from the road surface image, the features it outputs are only the output features of the target hidden layer, and then only the feature map set output by the target hidden layer is used to detect road surface defects.
[0043] In the embodiment of the present application, the features extracted from each layer in the convolutional neural network are visualized by using Grad-CAM, and then the authenticity of the defects of the features extracted from each layer is evaluated to select the target hidden layer with a high similarity to the real defect area. Finally, the convolutional neural network is reconstructed based on the target hidden layer. This method realizes the adjustment inside the feature extraction layer in the convolutional neural network, enabling it to select the most informative and accurate feature layers, and realizing the localization of the significant area of the defect even without a large amount of labeled data, reducing the dependence on labeled data.
[0044] Please refer to Figure 2 , another embodiment of the road surface defect detection method in the embodiment of the present application includes:
[0045] 201. Input the obtained road surface image into the convolutional neural network for feature extraction, and extract the feature map set extracted from each hidden layer in the convolutional neural network. Each feature map in the feature map set has corresponding label information for each road surface feature, and the label information is used to indicate the defect degree of the road surface feature.
[0046] In this embodiment, the road surface image can be acquired by devices such as drones or road detection robots. These devices can efficiently and comprehensively cover large detection areas, especially for areas that are difficult to reach (such as viaducts, tunnels or remote roads).
[0047] The collected road surface images are input into a convolutional neural network, which extracts road surface features from the road surface images and identifies the features with road surface defects. It should be noted that the convolutional neural network consists of multiple hidden layers, and each hidden layer extracts road surface features that may have defects from the road surface images to output a feature map set, which can be based on the road surface images and mark or frame the road surface features that may have defects in them.
[0048] Specifically, the feature map sets are extracted from all the hidden layers, and these feature map sets of different layers are generated by performing convolution operations through the corresponding hidden layers in the convolutional neural network. Each feature map set has multiple channels, and these channels represent the positions of different important features in the image. These feature map sets help to reveal the key areas of defects such as cracks and potholes.
[0049] 202. Use Grad-CAM to extract the label information of the feature maps in each hidden layer, and extract the concerned features in the feature maps based on the relationship between the label information and the defect degree to obtain the road surface defect concerned area.
[0050] It can be understood that the mean value of the multiple feature maps corresponding to the road surface defect concerned area is calculated using the visualization processing formula to obtain the visualization result of the road surface defect concerned area, where the visualization processing formula is: Among them, is the feature map set of the hidden layer, F k is obtained by compressing the feature map set into a two-dimensional feature map.
[0051] Specifically, by comparing the unweighted average feature map set with the class-specific Grad-CAM heat map, it can be seen whether the focus of the model is reasonable when detecting road surface defects. The average feature map set generated by the visualization processing formula can highlight these non-class-specific areas as potential defect areas to help identify subtle defects such as surface unevenness or slight cracks.
[0052] 203. Perform visualization processing on the road surface defect concerned area and generate a heat map corresponding to the hidden layer based on the result of the visualization processing.
[0053] In this embodiment, the multiple feature maps of each of the visualized hidden layers are squeezed and the average eigenvalue is calculated; using the heat map generation rule, a heat map corresponding to the hidden layer is generated based on the average eigenvalue. That is, Grad-CAM shows these significant areas in the form of a heat map to help us understand the attention distribution of the model when processing road surface structure images, thereby supporting the localization and identification of road surface defects. Grad-CAM is a method using heat maps to visualize feature maps.
[0054] Further, use the ReLU activation function to remove the features in all average eigenvalues that contribute nothing or have a negative impact on defect recognition, and perform normalization processing on the resulting feature map; based on the preset color mapping rules, determine the target color corresponding to the normalized feature map, and generate the heat map for the corresponding hidden layer.
[0055] In practical applications, the method for generating the heat map is to squeeze and average the eigenvalues of the N feature map sets in the k-th hidden layer. Then, through ReLU activation using the first formula, use the ReLU function to set all negative values to 0 and only retain positive values. Remove the regions that contribute nothing or have a negative impact on the target category, so that subsequent visualization only includes the regions that have a positive impact on the target category. Among them, the first formula is: x ij is the feature map value at the i-th row and j-th column in the k-th hidden layer.
[0056] Through normalization processing using the second formula, normalize the two-dimensional vector - heat map element (x i , x j ) to a value between 0 and 1, enhance the contrast of the image, ensure that the color mapping of the heat map is more evenly distributed over different intensity activation regions, and make the high-activation regions more prominent. After normalization, apply color mapping to map the numerical range to a color range, map the numerical range to a color range, so that different activation intensities correspond to different colors, and obtain the heat map. Among them, the second formula is: x′ ij is the normalized eigenvalue, representing the value after normalization processing; x ij is the original eigenvalue, representing the eigenvalue without normalization; max(x ij ) represents the maximum value in the original eigenvalue x ij and is used for the normalization operation.
[0057] Figure 3 The heat map is obtained through the above conversion method. There are differences in the size and heat map distribution of the feature maps of different hidden layers. Therefore, the feature representations obtained through heat map visualization are also different. Through quantitative analysis of the heat maps of each layer, we found that the feature maps generated by the intermediate layers are more complex than those of the last layer of the CNN model. This indicates that the earlier layers can capture subtle local defects on the road surface, such as cracks and small potholes, while the closer to the last layer, the network focuses more on the overall structural features to help identify larger-scale defect regions. Through this hierarchical analysis, we can more deeply understand how the network processes the detailed information in road images at each layer, thereby improving the accuracy and reliability of road structure defect detection and providing more accurate data support for maintenance and repair.
[0058] 204. Evaluate the authenticity of the road surface features shown in each heat map, select the feature maps that meet the preset accuracy based on the evaluation results, and determine the corresponding target hidden layer.
[0059] Select the feature layer that is most similar to the true defect area based on the evaluation results. This layer contains the most accurate attention area of the model for the defect area. Based on the evaluation results, the convolutional layer that can most accurately reflect the defect information can be screened out, and irrelevant or inaccurate layers can be excluded.
[0060] Select those layers with higher scores from all hidden layers through the Selective Layer Attention Network (SAN) method. Construct a new and simplified model that only retains these selected layers. The reconstructed model is more efficient, with a significant reduction in computational volume, avoiding unnecessary waste of computing resources.
[0061] In this embodiment, use the preset scoring formula to calculate the distance between the road surface features shown in each heat map and the features in the true defect area to obtain a score; use the Selective Layer Attention Network to select the target feature map whose score reaches the preset similarity value, and determine the target hidden layer corresponding to the target feature map based on the correspondence between the feature map and the hidden layer.
[0062] It should be noted that the evaluation result is actually an index quantitatively representing the similarity between the hidden layer feature map and the true label. Specifically, it is calculated through this scoring formula to estimate the perceptual accuracy of the extracted features, while considering the size of each layer of feature map. The scoring formula S(L k ) is expressed as follows:
[0063]
[0064] Among them, S(L k ) is the score of the feature map set of the kth hidden layer; E i,j is the absolute error of calculating the distance between the activation area extracted from the hidden layer and the labeled target defect area; E i,j is defined by the following formula.
[0065] E i,j =|x' x,j -t i,j |,
[0066] x'∈(0,1),
[0067] t∈{0,1},
[0068] Among them, x′ i,j is an element of the normalized heat map (V k ); t i,jis an element of the true label mask (Mk) with a binary conversion defect label. The true label mask consists of 0s and 1s, while the values of the heatmap are between 0 and 1. The accuracy of the heatmap is quantitatively estimated by calculating the difference between the heatmap vector value and the target value. As the feature vector (x i,j ) approaches the target vector (t i,j ), the consistency increases. Therefore, the sum of each consistency value is divided by the total number of feature vectors as an index for extracting the layer with the highest similarity to the expected localization, and the calculation process is as Figure 4 shown.
[0069] 205. Reconstruct the convolutional neural network based on the target hidden layer to obtain a target convolutional neural network.
[0070] 206. Use the target convolutional neural network to extract features from the road surface image again, and detect road surface defects based on the extracted features.
[0071] Specifically, when reconstructing the convolutional neural network, all target hidden layers in the convolutional neural network are retained, and the pre-trained weights of each target hidden layer are loaded for configuration to obtain the target convolutional neural network.
[0072] Loading pre-trained weights, loading the pre-trained weights of the selected layer in the new simplified model, without the need to retrain the entire network. This helps to maintain the performance of the existing model while significantly reducing the consumption of computing resources.
[0073] Furthermore, the reconstructing the convolutional neural network based on the target hidden layer to obtain a target convolutional neural network further includes: introducing an attention mechanism in each of the target hidden layers in the target convolutional neural network, where the attention mechanism is implemented by combining max unpooling and average pooling.
[0074] Max unpooling and average pooling: Apply the max unpooling operation in the selected convolutional layer to restore the spatial information lost during the max pooling process. To prevent the loss of spatial information, an average pooling operation is also added, which can further improve the integrity of the feature map.
[0075] Enhancing key features: Use the attention mechanism to strengthen the attention to key features. The attention module improves the accuracy of the final feature map by automatically emphasizing key features (such as defect locations and shapes) during the training process, thereby improving the accuracy of defect detection.
[0076] In practical applications, according to the selected layer, the convolutional network is reconstructed into a reduced convolutional network, where the model loads the pre-trained weights only from the input layer to the selected layer (L c ), while excluding the remaining layers.
[0077] Then, the proposed attention module is combined with the selected layer to provide enhanced prediction feature map information, more efficiently focus on the road surface defect area, enhance the ability to capture details such as cracks and potholes, and provide a more accurate and efficient solution for the road surface structure defect detection task. The attention module emphasizes the important features of the selected layer. The proposed attention mechanism extracts, enhances, and aligns feature information through progressive processing, thereby enhancing the model's attention to key areas, such as Figure 5 shown.
[0078] This selective layer attention module makes full use of the significant features extracted by the pooling layer and enhances the expression ability of key information by combining Max Unpooling and upsampling operations. The specific steps are as follows:
[0079] Max pooling records significant feature indices: First, extract the significant regions of the feature map from the max pooling layer and record the position indices (i.e., pooling indices) of the maximum values in each pooling window. These indices are used to indicate the original spatial positions of the activated features before pooling, ensuring that the positions of the significant features can be correctly restored during subsequent unpooling.
[0080] Max unpooling restores significant features: Using the saved pooling indices, restore the pooled activated features to their original positions in the feature map to generate a sparse feature map. This sparse feature map retains the key information (i.e., maximization) of the significant activation regions, while the non-significant regions are set to zero.
[0081] Upsampling projects significant features: Through upsampling operations, expand the sparse feature map to a higher resolution so that the significant features cover a larger spatial range. The upsampled feature map is as Figure 6 shown, where the largest number of activated features are projected and enhanced.
[0082] Second upsampling enhances unpooled features: Further upsample the feature map after max unpooling to refine the features and strengthen the information in the key regions. This process ensures that the unpooled features are fully expressed in the high-resolution feature map.
[0083] Complementary spatial information: Average pooling: To prevent over-concentration on activated features and neglect of other important spatial information during the pooling and upsampling processes, the module introduces an average pooling operation. Average pooling performs global spatial information smoothing aggregation on the feature map, supplementing the detailed information that may be missed in max unpooling and upsampling.
[0084] Generating attention weights from fused feature maps: Fuse the feature maps generated by max unpooling and average pooling to integrate the significantly activated regions and global spatial information, and generate an attention weight map. The attention weight map is used to perform a weighting operation on the original feature map to enhance the model's attention to key features.
[0085] In the embodiments of the present application, the features extracted from each layer in the convolutional neural network are visualized using Grad-CAM, and then the authenticity of the defects of the features extracted from each layer is evaluated to select a target hidden layer with a high similarity to the real defect region. Finally, the convolutional neural network is reconstructed based on the target hidden layer. This method realizes the adjustment inside the feature extraction layer in the convolutional neural network, enabling it to select the most informative and accurate feature layers, achieving the localization of the significant regions of defects even without a large amount of labeled data, reducing the dependence on labeled data, improving the detection effect, and reducing the consumption of computing resources.
[0086] The pavement defect detection method in the embodiments of the present application has been described above. Next, the pavement defect detection device in the embodiments of the present application will be described. Please refer to FIGS. 7 and 8. An embodiment of the pavement defect detection device in the embodiments of the present application includes:
[0087] An extraction module 710, configured to input the acquired pavement image into the convolutional neural network for feature extraction, and extract the feature map sets extracted from each hidden layer in the convolutional neural network. Among them, corresponding label information is set for each pavement feature in each feature map of the feature map set, and the label information is used to indicate the defect degree of the pavement feature;
[0088] A conversion module 720, configured to convert each of the feature map sets into a heat map using Grad-CAM, where only the pavement features reaching a preset defect degree are shown in the heat map;
[0089] An evaluation module 730, configured to evaluate the authenticity of the pavement features shown in each of the heat maps, and select a feature map set that meets the preset accuracy based on the evaluation result to determine the corresponding target hidden layer;
[0090] A reconstruction module 740, configured to reconstruct the convolutional neural network based on the target hidden layer to obtain a target convolutional neural network;
[0091] A detection module 750, configured to use the target convolutional neural network to re-extract features from the pavement image, and detect pavement defects based on the extracted features.
[0092] Optionally, the conversion module 720 is specifically configured to:
[0093] Extract the label information of the feature maps in each of the hidden layers using Grad-CAM, and extract the attention features in the feature maps based on the relationship between the label information and the defect degree to obtain the road surface defect attention area;
[0094] Perform visualization processing on the road surface defect attention area, and generate a heat map of the corresponding hidden layer based on the result of the visualization processing.
[0095] Optionally, the conversion module 720 is specifically configured to:
[0096] Use the visualization processing formula to calculate the mean value of the multiple feature maps corresponding to the road surface defect attention area to obtain the visualization result of the road surface defect attention area.
[0097] Optionally, the conversion module 720 is specifically configured to:
[0098] Squeeze the multiple feature maps of each of the visualized hidden layers, and calculate the average eigenvalue;
[0099] Use the heat map generation rule to generate a heat map of the corresponding hidden layer based on the average eigenvalue.
[0100] Optionally, the conversion module 720 is specifically configured to:
[0101] Use the ReLU activation function to remove the features in all the average eigenvalues that have no contribution or negative impact on defect recognition, and perform normalization processing on the feature maps after removal;
[0102] Based on the preset color mapping rule, determine the target color corresponding to the normalized feature map, and generate a heat map of the corresponding hidden layer.
[0103] Optionally, the evaluation module 730 includes:
[0104] A calculation unit 731, configured to calculate the distance between the road surface features shown in each of the heat maps and the features in the real defect area using a preset scoring formula to obtain a score;
[0105] An evaluation unit 732, configured to use the selective layer attention network to select the target feature maps with scores reaching a preset similarity value, and determine the target hidden layer corresponding to the target feature maps based on the corresponding relationship between the feature maps and the hidden layers.
[0106] Optionally, the reconstruction module 740 is specifically configured to:
[0107] Retain all the target hidden layers in the convolutional neural network, and load the pre-trained weights of each of the target hidden layers for configuration to obtain a target convolutional neural network.
[0108] In the embodiments of the present application, the features extracted by each layer in the convolutional neural network are visualized using Grad-CAM, and then the authenticity of the defects of the features extracted by each layer is evaluated to select a target hidden layer with a high similarity to the true defect area. Finally, the convolutional neural network is reconstructed based on the target hidden layer. This method realizes the adjustment inside the feature extraction layer in the convolutional neural network, enabling it to select the most informative and accurate feature layers, achieving the localization of the significant area of the defect even without a large amount of labeled data, reducing the dependence on labeled data, improving the detection effect, and reducing the consumption of computing resources.
[0109] Above Figure 7 and Figure 8 The pavement defect detection device in the embodiments of the present application is described in detail from the perspective of modular functional entities. Next, the electronic device in the embodiments of the present application is described in detail from the perspective of hardware processing.
[0110] See Figure 9 As shown, the electronic device includes a processor 900 and a memory 901. The memory 901 stores machine-executable instructions that can be executed by the processor 900, and the processor 900 executes the machine-executable instructions to implement the above pavement defect detection method.
[0111] Furthermore, Figure 9 The shown electronic device further includes a bus 902 and a communication interface 903. The processor 900, the communication interface 903, and the memory 901 are connected through the bus 902.
[0112] Among them, the memory 901 may include a high-speed random access memory (Random Access Memory, RAM), and may also include non-volatile memory, for example, at least one disk memory. Through at least one communication interface 903 (which can be wired or wireless), the communication connection between the system network element and at least one other network element is realized, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 902 can be an ISA bus, a PCI bus, or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 9 only a bidirectional arrow is used in
[0113] The processor 900 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 900 or instructions in the form of software. The above-mentioned processor 900 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 901, and the processor 900 reads the information in the memory 901 and combines its hardware to complete the method steps of the foregoing embodiments.
[0114] This application also provides an electronic device. The computer device includes a memory and a processor. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by the processor, the processor executes the steps of the road surface defect detection method in the above-mentioned various embodiments.
[0115] This application also provides a computer-readable storage medium. The computer-readable storage medium may be a non-volatile computer-readable storage medium, or the computer-readable storage medium may also be a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer executes the steps of the road surface defect detection method provided above.
[0116] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described in detail here.
[0117] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0118] As described above, the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of this application.
Claims
1. A road surface defect detection method, characterized in that: The method comprises: Inputting the acquired road surface image into a convolutional neural network for feature extraction, and extracting a feature atlas extracted by each hidden layer in the convolutional neural network, wherein each feature map in the feature atlas is provided with corresponding label information for each road surface feature, and the label information is used to indicate the degree of defect of the road surface feature; Using Grad-CAM to convert each of the feature atlases into a heat map, wherein only road surface features reaching a preset defect level are displayed in the heat map; Performing defect authenticity evaluation on the road surface features displayed in each of the heat maps, and selecting a feature atlas that meets a preset accuracy based on the evaluation results, and determining a corresponding target hidden layer; Reconstructing the convolutional neural network based on the target hidden layer to obtain a target convolutional neural network; The target convolutional neural network is used to re-extract features from the road surface image, and road surface defects are detected based on the extracted features.
2. The road surface defect detection method according to claim 1, characterized in that: The step of converting each of the feature maps into a heat map by using Grad-CAM comprises: Grad-CAM is used to extract label information of the feature graphs in each of the hidden layers, and based on the relationship between the label information and the degree of defects, the focus features in the feature graphs are extracted to obtain the road surface defect focus area; The road surface defect focus area is visualized, and a heat map of the corresponding hidden layer is generated based on the result of the visualization processing.
3. The road surface defect detection method according to claim 2, characterized in that: The visualizing of the road surface defect area of concern includes: The visualization processing formula is used to calculate the mean of multiple feature maps corresponding to the road surface defect focus area to obtain a visualization result of the road surface defect focus area.
4. The road surface defect detection method according to claim 2, characterized in that: The step of generating a heat map of the corresponding hidden layer based on the result of the visualization processing includes: Squeeze the visualized multiple feature maps of each hidden layer and calculate the average feature value; A heat map generation rule is used to generate a heat map of the corresponding hidden layer based on the average eigenvalue.
5. The road surface defect detection method according to claim 4, characterized in that: The step of using a heat map generation rule to generate a heat map of a corresponding hidden layer based on the average eigenvalue comprises: Use the ReLU activation function to remove features that have no contribution or negative impact on defect identification from all average eigenvalues, and normalize the feature maps after removal; Based on the preset color mapping rules, the target color corresponding to the normalized feature map is determined, and the heat map of the corresponding hidden layer is generated.
6. The road surface defect detection method according to claim 1, characterized in that: The defect authenticity evaluation of the road surface features displayed in each of the heat maps is performed, and based on the evaluation results, a feature map that meets a preset accuracy is selected to determine the corresponding target hidden layer, including: Using a preset scoring formula, the distance between the road surface features shown in each of the heat maps and the features in the actual defect area is calculated to obtain a score; A selective layer attention network is used to select a target feature map whose score reaches a preset similarity value, and a target hidden layer corresponding to the target feature map is determined based on the correspondence between the feature map and the hidden layer.
7. The road surface defect detection method according to claim 1, characterized in that: The reconstructing the convolutional neural network based on the target hidden layer to obtain a target convolutional neural network includes: All target hidden layers in the convolutional neural network are retained, and the pre-trained weights of each target hidden layer are loaded for configuration to obtain a target convolutional neural network.
8. A road surface defect detection device, characterized in that: The device comprises: An extraction module, used to input the acquired road surface image into a convolutional neural network for feature extraction, and extract a feature atlas extracted by each hidden layer in the convolutional neural network, wherein each feature map in the feature atlas is provided with corresponding label information for each road surface feature, and the label information is used to indicate the degree of defect of the road surface feature; A conversion module, configured to convert each of the feature atlases into a heat map using Grad-CAM, wherein only road surface features reaching a preset defect level are displayed in the heat map; An evaluation module, used to evaluate the authenticity of defects of the road surface features displayed in each of the heat maps, and select a feature atlas that meets a preset accuracy based on the evaluation results to determine the corresponding target hidden layer; A reconstruction module, used to reconstruct the convolutional neural network based on the target hidden layer to obtain a target convolutional neural network; A detection module is used to re-extract features of the road surface image using the target convolutional neural network, and detect road surface defects based on the extracted features.
9. An electronic device, characterized in that: The electronic device comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory so that the electronic device executes the road surface defect detection method as described in any one of claims 1-7.
10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the road surface defect detection method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Pruning method and device for neural network model
CN112749797A
Power transmission line small target equipment small sample defect detection method and device and storage medium
CN119107496A
Method for training artificial neural network
US20190370662A1
Attention weight calculation method and apparatus based on convolutional neural network, and device
WO2021068528A1
Cited By
Multi-modal target automatic identification and tracking method and system based on photoelectric pod
CN120913024A
A multi-modal target automatic identification and tracking method and system based on an optoelectronic pod
CN120913024B