A method and system for identifying cracks in a concrete road
By improving the DeepLabv3+ model, the hollow space pyramid pooling module and composite loss function are introduced, the low-precision problem caused by complex background interference is solved, the accuracy and adaptability of crack detection are improved, and the real-time requirements are met.
Patent Information
- Application Number
- CN202510252438.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-05
AI Technical Summary
In the prior art, due to significant interference in complex backgrounds, the detection accuracy is low and the ability to adapt to the diversified crack morphology is insufficient.
Using the improved DeepLabv3+ model, model parameters are optimized to improve the accuracy and adaptability of crack detection by introducing an improved hollow space pyramid pooling module and a composite loss function combining edge enhancement and geometric adaptability.
It significantly improves the accuracy of crack detection, especially in complex backgrounds and fine crack detection, enhances the adaptability to the diversified crack morphology, and meets the real-time requirements of road crack identification tasks.
Smart Images

Figure CN119762977B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of computer vision and road maintenance, and particularly relates to a method and system for identifying concrete road cracks. Background Art
[0002] Road cracks are one of the most common diseases in traffic infrastructure such as highways and urban roads. Their formation is mainly caused by the combined action of various factors such as traffic loads, environmental impacts, and construction quality. The appearance of cracks not only affects the service life and driving comfort of the road, but also increases the hidden dangers of driving safety, further leading to early damage of the road. Therefore, the timely detection and maintenance of road cracks are of great significance for ensuring traffic safety and extending the road life.
[0003] Traditional road crack detection mostly relies on manual observation and measurement. Staff identify cracks with the naked eye and use tools for measurement. This method has a low cost, but is inefficient, easily affected by human fatigue and subjective factors, and it is difficult to cover large areas of roads. Precision measurement of cracks is carried out using equipment such as crack width meters and laser scanners. This method has a higher accuracy, but the equipment is expensive, and it is impossible to efficiently complete large-scale detection tasks.
[0004] With the development of computer vision and artificial intelligence technologies, automated road crack detection technologies based on image processing have gradually become a research hotspot. These technologies mainly identify cracks automatically by taking road images and using algorithms, and have the advantages of high efficiency, precision, and low cost. In recent years, deep learning technologies such as convolutional neural networks (CNNs) have been widely used in the field of crack detection. For example, networks such as U-Net and DeepLab can complete pixel-level segmentation of cracks in an end-to-end manner, greatly improving the detection accuracy and robustness. Since the forms of cracks are diverse, they may appear as linear, reticular, blocky or other complex forms, the detection algorithm needs to have strong generalization ability for different crack types, and the road environment is complex, and crack detection may be interfered by shadows, oil stains, illumination changes, road surface textures, etc., resulting in a decrease in the detection accuracy. Summary of the Invention
[0005] The present invention provides a method and system for identifying concrete road cracks, which are used to solve the technical problems in the prior art that due to significant interference from complex backgrounds, the detection accuracy is low and the adaptability to diverse crack forms is insufficient.
[0006] In a first aspect, the present invention provides a method for identifying concrete road cracks, including:
[0007] Construct a concrete road crack recognition model based on the improved DeepLabv3+. The concrete road crack recognition model includes an encoder and a decoder directly connected to the encoder. The encoder includes a deep convolutional network layer and an improved atrous spatial pyramid pooling module. The decoder includes a bilinear upsampling module, a 1×1 convolutional module, and a 3×3 convolutional module.
[0008] Obtain concrete road crack images, annotate and preprocess the concrete road crack images to obtain a concrete road crack dataset.
[0009] Input the concrete road crack dataset into the concrete road crack recognition model, optimize the model parameters of the concrete road crack recognition model based on a composite loss function that combines edge enhancement and geometric adaptability, train to generate the weights of the concrete road crack recognition model, and obtain a target crack recognition model.
[0010] Input a real-time concrete road crack image into the target crack recognition model. The target crack recognition model outputs a preliminary recognition result, and optimizes the preliminary recognition result through a pseudo-detection elimination and missed detection repair algorithm to obtain a final recognition result.
[0011] In a second aspect, the present invention provides a concrete road crack recognition system, including:
[0012] A construction module configured to construct a concrete road crack recognition model based on the improved DeepLabv3+. The concrete road crack recognition model includes an encoder and a decoder directly connected to the encoder. The encoder includes a deep convolutional network layer and an improved atrous spatial pyramid pooling module. The decoder includes a bilinear upsampling module, a 1×1 convolutional module, and a 3×3 convolutional module.
[0013] An acquisition module configured to obtain concrete road crack images, annotate and preprocess the concrete road crack images to obtain a concrete road crack dataset.
[0014] An optimization module configured to input the concrete road crack dataset into the concrete road crack recognition model, optimize the model parameters of the concrete road crack recognition model based on a composite loss function that combines edge enhancement and geometric adaptability, train to generate the weights of the concrete road crack recognition model, and obtain a target crack recognition model.
[0015] A recognition module configured to input a real-time concrete road crack image into the target crack recognition model. The target crack recognition model outputs a preliminary recognition result, and optimizes the preliminary recognition result through a pseudo-detection elimination and missed detection repair algorithm to obtain a final recognition result.
[0016] In a third aspect, an electronic device is provided, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the concrete road crack recognition method according to any embodiment of the present invention.
[0017] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the program instructions are executed by a processor, the processor is enabled to execute the steps of the concrete road crack recognition method according to any embodiment of the present invention.
[0018] The concrete road crack recognition method and system of the present application, by improving the encoder and decoder structures of the DeepLabv3+ model and introducing an improved atrous spatial pyramid pooling module, significantly enhance the model's ability to extract multi-scale features, improve the accuracy of crack detection, and perform particularly well in detecting complex backgrounds and subtle cracks; adopting a composite loss function that combines edge enhancement and geometric adaptability can effectively improve the model's perception ability of crack edges, while considering the constraints of crack geometric characteristics, enhancing the adaptability to diverse crack morphologies, and enabling the model to work stably in different scenarios and conditions; improving the feature fusion and channel adjustment strategies of the decoder simplifies the calculation process, and combined with a lightweight processing strategy, greatly improves the model inference speed, meeting the real-time requirements of the road crack recognition task; through the pseudo-detection elimination and missed detection repair algorithms, the preliminary detection results are further optimized, significantly reducing the false detection rate and missed detection rate, and ensuring the accuracy and integrity of the final recognition results. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 It is a flowchart of a concrete road crack recognition method provided by an embodiment of the present invention;
[0021] Figure 2 It is a structural block diagram of a concrete road crack recognition system provided by an embodiment of the present invention;
[0022] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0024] Please refer to Figure 1 , which shows a flowchart of a method for identifying concrete road cracks in the present application.
[0025] As Figure 1 shown, the method for identifying concrete road cracks specifically includes the following steps:
[0026] Step S101, construct a concrete road crack identification model based on an improved DeepLabv3+ (a semantic segmentation model based on deep learning), where the concrete road crack identification model includes an encoder and a decoder directly connected to the encoder; the encoder includes a depth convolutional network layer and an improved atrous spatial pyramid pooling module; the decoder includes a bilinear upsampling module, a 1×1 convolutional module, and a 3×3 convolutional module.
[0027] In this step, the improved atrous spatial pyramid pooling module includes:
[0028] A dynamic multi-branch feature extraction module, which includes a 1×1 convolutional module, 3×3 dilated convolutional modules with dilation rates of 6, 12, and 18 respectively, a global average pooling module, and a dynamic asymmetric atrous convolution module. Among them, the dynamic asymmetric atrous convolution module adjusts the dilation rate in the horizontal direction adaptively according to the directional characteristics of crack distribution in the input feature map by introducing a dynamic receptive field adjustment factor and the dilation rate in the vertical direction , and the expression is:
[0029] ,
[0030] ,
[0031] In the formula, is the base dilation rate in the horizontal direction, , are both adjustment coefficients, is the input feature map the variance in the horizontal direction, is the base dilation rate in the vertical direction, is the input feature map Variance in the vertical direction;
[0032] The input features of the 3×3 convolution module with a dilation rate of 6 are composed of deep features and the output features of the 1×1 convolution module;
[0033] The input features of the 3×3 convolution module with a dilation rate of 12 are composed of deep features and the output features of the convolution module with a dilation rate of 6;
[0034] The input features of the 3×3 convolution module with a dilation rate of 18 are composed of deep features and the output features of the convolution module with a dilation rate of 12;
[0035] The crack shape sensitive attention module, which combines directional gradient and edge information to weight and adjust the output features of the 3×3 dilated convolution module, global average pooling module and dynamic asymmetric atrous convolution module with dilation rates of 6, 12, and 18 respectively. Among them, the calculation formula for the attention weight of the crack shape sensitive attention module is:
[0036] ,
[0037] In the formula, is the channel weight, is the Sigmoid function, , are the gradients of the feature map F in the horizontal and vertical directions respectively, is the weight matrix for fusing gradient features and original features.
[0038] The output features of the 1×1 convolution module, dynamic asymmetric atrous convolution module and crack shape sensitive attention module are merged through feature fusion operation to obtain the fused features;
[0039] The fused features are adjusted in channels through another 1×1 convolution module to generate the output deep features of the improved atrous spatial pyramid pooling module.
[0040] In this embodiment, the crack shape-sensitive attention module enhances features by combining channel attention and spatial attention mechanisms, targeting the long and slender strip characteristics and directional distribution characteristics of cracks. First, in the channel attention part, the input feature map undergoes global average pooling and global max pooling to generate global descriptors, thereby capturing global statistical information in the channel dimension. Then, the descriptors are non-linearly mapped through two layers of 1×1 convolutions and ReLU activation functions, and the Sigmoid function is used to generate channel weights, which are finally applied to the input feature map to strengthen key channel features and suppress irrelevant information. Subsequently, the enhanced feature map further captures the directional and spatial multi-scale features of cracks through dynamic asymmetric dilated convolutions and multi-scale asymmetric convolutions, dynamically adjusting the receptive field to adapt to the shape and length of cracks.
[0041] In the spatial attention part, spatial attention weights are generated using the horizontal and vertical direction gradient information and edge features of the feature map. The gradient and edge information are fused through 1×1 convolutions, and the weight values are normalized using the Sigmoid function to generate a spatially enhanced feature representation with increased saliency. Finally, the feature maps enhanced in both channels and space undergo a pointwise weighting operation to output a crack feature representation that includes both directionality and edge characteristics while enhancing the saliency of key regions, significantly improving the ability to detect and classify cracks.
[0042] In the improved DeepLabv3+ concrete crack road recognition model, the deep convolutional network layer is set as the lightweight model MobileNetv2, and the output shallow features are mainly used to capture the edge information and local texture details of cracks. These features are generated through the first few Bottleneck modules of the network, with the resolution gradually decreasing and the number of channels gradually increasing, as follows:
[0043] First, the size of the first shallow feature is 1 / 4H × 1 / 4W, and the number of channels is 24. It is derived from the initial convolutional layer (Conv1) of MobileNetV2 and the first Bottleneck module. After the image passes through Conv1 (a 3×3 standard convolution with a stride of 2 for downsampling), the resolution of the feature map is reduced to 1 / 2 of the original image, and the number of channels is 16; subsequently, the first Bottleneck module (with an expansion factor of 1) is used to extract shallow features, with the resolution further reduced to 1 / 4 and the number of channels increased to 24. This feature mainly captures the basic edge information and local details of cracks, retaining a relatively high spatial resolution.
[0044] Secondly, the size of the second shallow feature is 1 / 8H × 1 / 8W, and the number of channels is 32. It comes from the second Bottleneck module of the MobileNetV2 backbone network. After passing through the first Bottleneck module, the feature map enters the second Bottleneck module (expansion factor is 6). After depthwise separable convolution and feature transformation, the resolution is further reduced to 1 / 8, and the number of channels increases to 32. This feature introduces more semantic information on the basis of edge and texture information, and can capture more details in the crack area.
[0045] Finally, the size of the third shallow feature is 1 / 16H × 1 / 16W, and the number of channels is 64. It comes from the third Bottleneck module of the MobileNetV2 backbone network. After passing through multiple Bottleneck modules, the resolution is reduced to 1 / 16, and the number of channels is further increased to 64. This feature gradually contains the mesoscale structural information of the crack, providing support for subsequent deep feature fusion.
[0046] After these features enter the decoder, the following operations need to be performed: First, adjust the number of channels through a 1×1 convolution module to ensure the channel consistency of features at different scales; then use Instance Normalization to standardize the features and eliminate the feature distribution differences caused by different batches of data; finally, adopt the ReLU activation function to enhance the non-linear representation ability of the features.
[0047] The decoder part receives the shallow features and deep features output by the encoder. The shallow feature information is bilinearly upsampled to a feature map of the same size and then merged. The channels are compressed through a 1×1 convolution module to generate the first feature map; after the deep feature information extracts features through an improved atrous spatial pyramid pooling module and is bilinearly upsampled, the second feature map is obtained; after the first feature map is fused with the second feature map, features are extracted through a 3×3 convolution module, and the channels are adjusted through a 1×1 convolution module. Finally, classification prediction is completed, and the prediction result is restored to the same size as the input image using the bilinear upsampling module to obtain the preliminary prediction result.
[0048] Step S102: Obtain a concrete road crack image, and annotate and preprocess the concrete road crack image to obtain a concrete road crack dataset.
[0049] In this step, the collected concrete road crack image is subjected to image enhancement processing. The contrast of the crack area is enhanced through histogram equalization, gamma correction, or contrast adjustment, and the corresponding segmentation label map is added; the enhanced image and label are integrated into an image dataset, and the image dataset is divided into a training set and a test set.
[0050] In step S103, input the concrete road crack dataset into the concrete road crack recognition model, optimize the model parameters of the concrete road crack recognition model based on a composite loss function that combines edge enhancement and geometric adaptability, train to generate the weights of the concrete road crack recognition model, and obtain the target crack recognition model.
[0051] In this step, input the preprocessed crack dataset into the crack recognition model, calculate the predicted output of the model through forward propagation, and calculate a composite loss function that combines edge enhancement and geometric adaptability based on the true label and the predicted output. This composite loss function is composed of a cross-entropy loss and an edge-aware loss. The cross-entropy loss is used to optimize the pixel-level classification accuracy, and the edge-aware loss is used to enhance the model's ability to capture the edge details of the cracks. At the same time, in order to adapt to the geometric features of the cracks, the composite loss function dynamically adjusts the weight allocation during the calculation process, making the model more accurate in identifying the main body and edge regions of the cracks. Finally, optimize the model parameters through backpropagation and gradient update, and iteratively train until convergence to generate the final weights of the target crack recognition model.
[0052] The composite loss function includes a basic segmentation loss sub-function, an edge enhancement loss sub-function, and a geometric adaptability loss sub-function;
[0053] The expression of the composite loss function is:
[0054] ,
[0055] In the formula, is the basic segmentation loss sub-function, , are both balance weight parameters, is the edge enhancement loss sub-function, is the geometric adaptability loss sub-function.
[0056] By adjusting the , and values of the composite loss function, it can achieve adaptive balance between different loss terms to improve the accuracy and robustness of crack detection.
[0057] It should be noted that the expression of the basic segmentation loss sub-function is:
[0058] ,
[0059] In the formula, is the true label of the th pixel point, is the The predicted value of a pixel point, is the total number of pixel points;
[0060] The expression of the edge enhancement loss sub - function is:
[0061] ,
[0062] In the formula, is the dynamic weight based on edge pixels;
[0063] The expression of the geometric adaptability loss sub - function is:
[0064] ,
[0065] ,
[0066] In the formula, is the geometric weight, is the parameter controlling the geometric feature weight, is the non - linear function of the crack width, is the pixel position of the
[0067] Step S104, input the real - time concrete road crack image into the target crack recognition model, the target crack recognition model outputs a preliminary recognition result, and optimizes the preliminary recognition result through the pseudo - detection elimination and missed - detection repair algorithm to obtain the final recognition result.
[0068] In this step, according to the deep convolutional network layer, three different - scale shallow - layer feature information and one deep - layer feature information of the real - time concrete road crack image are extracted. The deep convolutional network layer is a lightweight network structure; the three different - scale shallow - layer feature information is processed by the bilinear up - sampling module into feature maps of the same size and then merged, and the channels are compressed by the 1×1 convolutional module to generate the first feature map; the improved atrous spatial pyramid pooling module processes the deep - layer feature information through multi - scale dilated convolution and is up - sampled by the bilinear up - sampling module to obtain the second feature map; the first feature map and the second feature map are fused, and after fusion, features are extracted by the 3×3 convolutional module, and the channels are adjusted by the 1×1 convolutional module. The prediction result is restored to the same size as the real - time concrete road crack image by using the bilinear up - sampling module to obtain the preliminary prediction result.
[0069] Furthermore, extract the multi - dimensional geometric features in the preliminarily recognized crack region , where, 、 、 and , is the area, is the perimeter, is the pixel value of the pixel with coordinates in the binary image, where 1 represents a crack and 0 represents the background, is the boundary of the region ; , is a certain pixel point on the boundary of the crack region with the abscissa and ordinate, , is another pixel point on the boundary of the crack region adjacent to with the abscissa and ordinate, is the shape factor, is the eccentricity of the circumscribed ellipse, is the major axis of the circumscribed ellipse, is the minor axis of the circumscribed ellipse;
[0070] Calculate the normalized score of each feature and combine the weighted coefficients to obtain the comprehensive score. The expression is:
[0071] ,
[0072] In the formula, is the comprehensive score, is the number of features, is the value of the i-th feature of the region , is the weight of the feature, is the standard deviation of feature i, is the mean of feature i;
[0073] If , then eliminate the false detection region, is the threshold of the comprehensive score;
[0074] Based on the missed detection repair of optimized tensor decomposition, construct a sparse adjacency tensor . The expression is:
[0075] ,
[0076] In the formula, is the Euclidean distance between two points, is the angle between two points, represents the adjacency relationship between the i-th and j-th crack connected regions, , is the centroid coordinates of the crack regions and , is the distance threshold, is the angle threshold, is the first - dimension length of T, representing the number of rows of the tensor, is the second - dimension length of T, representing the number of columns of the tensor, represents the real number field;
[0077] The tensor is decomposed into a low - rank form by non - negative tensor decomposition, and the expression is:
[0078] ,
[0079] In the formula, is the rank of the decomposition, is the k - th factor vector used to describe the characteristics of the tensor in the first dimension, is the outer - product operation of the tensor, is the k - th factor vector used to describe the characteristics of the tensor in the second dimension, is the k - th factor vector used to describe the characteristics of the tensor in the third dimension;
[0080] Define the objective function, and restore the connected region by optimizing the sparsity and smoothness constraints. The expression of the objective function is:
[0081] ,
[0082] In the formula, and are the constraint weight coefficients of sparsity and smoothness respectively, is the L2 norm, used to constrain smoothness, is the Frobenius norm, is the L1 norm, used to constrain sparsity, is the gradient of, , , are the factor vectors in the tensor decomposition, corresponding to the three dimensions of the tensor T respectively.
[0083] The reconstructed tensor obtained by optimization represents the possible crack connection region. The expression of the reconstructed tensor is:
[0084] ,
[0085] In the formula, is the i - th element representing the vector , is the j - th element representing the vector ;
[0086] The Map back to the image space to generate the connected crack image, and the expression is:
[0087] ,
[0088] In the formula, is the line connection operation between two points;
[0089] Fuse the crack image after pseudo-detection rejection and the reconstruction area , and output the final detection result, and the expression is:
[0090] ,
[0091] In the formula, is the final fused crack image.
[0092] In summary, the method of this application improves the encoder and decoder structures of the DeepLabv3+ model, introduces an improved atrous spatial pyramid pooling module, significantly enhances the model's ability to extract multi-scale features, improves the accuracy of crack detection, and performs particularly well in detecting complex backgrounds and fine cracks; adopts a composite loss function that combines edge enhancement and geometric adaptability, which can effectively improve the model's perception ability of crack edges, and at the same time considers the constraints of crack geometric characteristics, enhances the adaptability to diverse crack morphologies, and enables the model to work stably in different scenarios and conditions; improves the feature fusion and channel adjustment strategies of the decoder, simplifies the calculation process, and combines a lightweight processing strategy to greatly improve the model inference speed, which can meet the real-time requirements of the road crack recognition task; through the pseudo-detection rejection and missed detection repair algorithms, the preliminary detection results are further optimized, significantly reducing the false detection rate and missed detection rate, and ensuring the accuracy and integrity of the final recognition result.
[0093] In a specific embodiment, the simulation experiment is implemented under PYTORCH, and an NVIDIA(R) Tesla(R) V100 GPU is used for training and testing in the experiment. 100 epochs are performed on the constructed concrete road crack dataset. The first epoch ≤ 50 is the freezing stage, and the batch_size is 8; the latter epoch > 50 is the thawing stage, and the batch_size (batch size) is 4. During training, the image size is uniformly scaled to 512×512, SGD (stochastic gradient descent) is used as the optimizer, the initial learning rate is 7e-3, the learning rate decay method is step (equal-interval adjustment), the momentum is 0.9, and the weight_decay (weight decay) is set to 1e-4.
[0094] By comparing with the original DeepLabv3+ model and other three representative semantic segmentation models, the superior performance of the method of the present invention is verified. The three algorithms are PSPNet, HRNet, and U-Net in sequence. To evaluate the performance of the model in road crack detection, the evaluation metrics used are: mean pixel accuracy (MPA), mean intersection over union (MIoU), accuracy, precision, and F1-score, as well as the inference time (Inference time) for comparing the model complexity. The experimental results on the dataset of this embodiment are shown in Table 1.
[0095] Table 1: Comparison Table of Experimental Results on Self-built Dataset
[0096] ,
[0097] As can be seen from the semantic segmentation index results given in Table 1, in the process of concrete road crack identification by the method of the present application, the inference time is the shortest, and the MPA, MIoU, Accuracy, Precision, and F1-score are all better than the other four models, indicating that the method of the present application not only has significant advantages in recognition accuracy and robustness, but also has high efficiency, and can meet the dual requirements of real-time and accuracy in actual road crack detection.
[0098] Please refer to Figure 2 , which shows the structural block diagram of a concrete road crack identification system of the present application.
[0099] As Figure 2 shown, the concrete road crack identification system 200 includes a construction module 210, an acquisition module 220, an optimization module 230, and an identification module 240.
[0100] Among them, the construction module 210 is configured to construct a concrete road crack recognition model based on the improved DeepLabv3+. The concrete road crack recognition model includes an encoder and a decoder directly connected to the encoder. The encoder includes a deep convolutional network layer and an improved atrous spatial pyramid pooling module. The decoder includes a bilinear upsampling module, a 1×1 convolutional module, and a 3×3 convolutional module. The acquisition module 220 is configured to acquire a concrete road crack image, annotate and preprocess the concrete road crack image to obtain a concrete road crack data set. The optimization module 230 is configured to input the concrete road crack data set into the concrete road crack recognition model, optimize the model parameters of the concrete road crack recognition model based on a composite loss function that combines edge enhancement and geometric adaptability, train to generate the weights of the concrete road crack recognition model, and obtain a target crack recognition model. The recognition module 240 is configured to input a real-time concrete road crack image into the target crack recognition model. The target crack recognition model outputs a preliminary recognition result, and optimizes the preliminary recognition result through a pseudo-detection elimination and missed detection repair algorithm to obtain a final recognition result.
[0101] It should be understood that Figure 2 the modules described in Figure 1 correspond to the respective steps in the method described in Figure 2 Therefore, the operations, features, and corresponding technical effects described above for the method also apply to
[0102] In some other embodiments, the embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the program instructions are executed by a processor, the processor executes the concrete road crack recognition method in any of the above method embodiments;
[0103] As an implementation manner, the computer-readable storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are set as:
[0104] Construct a concrete road crack recognition model based on the improved DeepLabv3+. The concrete road crack recognition model includes an encoder and a decoder directly connected to the encoder. The encoder includes a deep convolutional network layer and an improved atrous spatial pyramid pooling module. The decoder includes a bilinear upsampling module, a 1×1 convolutional module, and a 3×3 convolutional module;
[0105] Acquire a concrete road crack image, annotate and preprocess the concrete road crack image to obtain a concrete road crack data set;
[0106] Input the concrete road crack data set into the concrete road crack recognition model, optimize the model parameters of the concrete road crack recognition model based on a composite loss function that combines edge enhancement and geometric adaptability, train to generate the weights of the concrete road crack recognition model, and obtain the target crack recognition model;
[0107] Input the real-time concrete road crack image into the target crack recognition model. The target crack recognition model outputs a preliminary recognition result, and optimizes the preliminary recognition result through a pseudo-detection elimination and missed detection repair algorithm to obtain the final recognition result.
[0108] A computer-readable storage medium may include a storage program area and a storage data area. Among them, the storage program area can store an operating system and application programs required for at least one function; the storage data area can store data created according to the use of the concrete road crack recognition system, etc. In addition, the computer-readable storage medium may include high-speed random access memory, and may also include memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the computer-readable storage medium may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the concrete road crack recognition system through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0109] Figure 3 is a schematic structural diagram of the electronic device provided by an embodiment of the present invention, as Figure 3 shown, the device includes: a processor 310 and a memory 320. The electronic device may further include: an input device 330 and an output device 340. The processor 310, the memory 320, the input device 330, and the output device 340 may be connected through a bus or other means, Figure 3 taking the connection through the bus as an example. The memory 320 is the above-mentioned computer-readable storage medium. The processor 310 executes various functional applications and data processing of the server by running non-volatile software programs, instructions, and modules stored in the memory 320, that is, implements the concrete road crack recognition method in the above method embodiment. The input device 330 can receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the concrete road crack recognition system. The output device 340 may include a display device such as a display screen.
[0110] The above electronic device can execute the method provided by the embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment can be found in the method provided by the embodiment of the present invention.
[0111] As an implementation, the above electronic device is applied to a concrete road crack identification system and is used for a client, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to:
[0112] Construct a concrete road crack identification model based on the improved DeepLabv3+, where the concrete road crack identification model includes an encoder and a decoder directly connected to the encoder; the encoder includes a depth convolution network layer and an improved atrous spatial pyramid pooling module; the decoder includes a bilinear upsampling module, a 1×1 convolution module, and a 3×3 convolution module;
[0113] Obtain a concrete road crack image, annotate and preprocess the concrete road crack image to obtain a concrete road crack data set;
[0114] Input the concrete road crack data set into the concrete road crack identification model, optimize the model parameters of the concrete road crack identification model based on a composite loss function that combines edge enhancement and geometric adaptability, train to generate the weights of the concrete road crack identification model, and obtain a target crack identification model;
[0115] Input a real-time concrete road crack image into the target crack identification model, the target crack identification model outputs a preliminary identification result, and optimize the preliminary identification result through a pseudo-detection elimination and missed detection repair algorithm to obtain a final identification result.
[0116] Through the description of the above implementation manners, those skilled in the art can clearly understand that each implementation manner can be realized by means of software plus a necessary general hardware platform, and of course, it can also be realized by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for identifying cracks in a concrete road, characterized in that: include: A concrete road crack recognition model based on improved DeepLabv3+ is constructed, wherein the concrete road crack recognition model includes an encoder and a decoder directly connected to the encoder; the encoder includes a deep convolutional network layer and an improved dilated space pyramid pooling module; the decoder includes a bilinear upsampling module, a 1×1 convolution module, and a 3×3 convolution module, wherein the improved dilated space pyramid pooling module includes: A dynamic multi-branch feature extraction module, which includes a 1×1 convolution module, a 3×3 dilated convolution module with dilation rates of 6, 12, and 18, a global average pooling module, and a dynamic asymmetric dilated convolution module. The dynamic asymmetric dilated convolution module adaptively adjusts the dilation rate in the horizontal direction according to the directional characteristics of the crack distribution in the input feature map by introducing a dynamic receptive field adjustment factor. and the vertical expansion rate , the expression is: , , In the formula, is the basic expansion rate in the horizontal direction, , are adjustment coefficients, is the variance of the input feature map in the horizontal direction, is the basic expansion rate in the vertical direction, is the variance of the input feature map in the vertical direction; The input features of the 3×3 convolution module with a dilation rate of 6 consist of the deep features and the output features of the 1×1 convolution module; The input features of the 3×3 convolutional module with a dilation rate of 12 are composed of the deep features and the output features of the convolutional module with a dilation rate of 6; The input features of the 3×3 convolutional module with a dilation rate of 18 are composed of the deep features and the output features of the convolutional module with a dilation rate of 12; The crack shape sensitive attention module combines directional gradients and edge information to weight the output features of the 3×3 dilated convolution module with dilation rates of 6, 12, and 18, the global average pooling module, and the dynamic asymmetric dilated convolution module. The attention weight calculation formula of the crack shape sensitive attention module is: , In the formula, is the channel weight, is the Sigmoid function, , are the gradients of the feature map F in the horizontal and vertical directions, respectively. is the weight matrix used to fuse gradient features with original features; Acquire a concrete road crack image, and annotate and preprocess the concrete road crack image to obtain a concrete road crack data set; The concrete road crack data set is input into the concrete road crack recognition model, model parameters of the concrete road crack recognition model are optimized based on a composite loss function combining edge enhancement and geometric adaptability, weights of the concrete road crack recognition model are generated through training, and a target crack recognition model is obtained; The real-time concrete road crack image is input into the target crack recognition model, the target crack recognition model outputs a preliminary recognition result, and the preliminary recognition result is optimized by a false detection elimination and missed detection repair algorithm to obtain a final recognition result.
2. A method for identifying cracks in a concrete road according to claim 1, characterized in that: The real-time concrete road crack image is input into the target crack recognition model, and the target crack recognition model outputs a preliminary recognition result, which includes: Extracting shallow feature information of three different scales and one deep feature information of the real-time concrete road crack image according to the deep convolutional network layer, wherein the deep convolutional network layer is a lightweight network structure; The shallow feature information of three different scales is processed into feature maps of the same size by the bilinear upsampling module and then merged, and the channel is compressed by the 1×1 convolution module to generate a first feature map; The improved atrous spatial pyramid pooling module processes deep feature information through multi-scale dilated convolution, and upsamples through a bilinear upsampling module to obtain a second feature map; The first feature map is fused with the second feature map, and after fusion, the features are extracted by the 3×3 convolution module, and the channels are adjusted by the 1×1 convolution module. The prediction results are restored to the same size as the real-time concrete road crack image using a bilinear upsampling module to obtain a preliminary prediction result.
3. A method for identifying cracks in a concrete road according to claim 1, characterized in that: The composite loss function includes a basic segmentation loss sub-function, an edge enhancement loss sub-function and a geometric adaptability loss sub-function; The expression of the composite loss function is: , In the formula, is the basic segmentation loss sub-function, , are all balance weight parameters, is the edge enhancement loss sub-function, is the geometric adaptability loss sub-function.
4. A method for identifying cracks in a concrete road according to claim 3, characterized in that: The expression of the basic segmentation loss subfunction is: , In the formula, For the The true label of each pixel, For the The predicted value of each pixel, is the total number of pixels; The expression of the edge enhancement loss subfunction is: , In the formula, is the dynamic weight based on edge pixels; The expression of the geometric adaptability loss subfunction is: , , In the formula, is the geometric weight, is a parameter that controls the weight of geometric features. is a nonlinear function of crack width, For the The pixel location of the crack area.
5. A method for identifying cracks in a concrete road according to claim 1, characterized in that: The optimization of the preliminary recognition result by using the false detection elimination and missed detection repair algorithm to obtain the final recognition result includes: Extract the initially identified crack area Multidimensional geometric features ,in, , , as well as , is the area, is the circumference, The coordinates in the binary image are The pixel value of the pixel, 1 represents cracks, 0 represents background, For Region The borders of , The crack area The horizontal and vertical coordinates of a pixel point on the boundary, , The crack area on the border with The horizontal and vertical coordinates of another adjacent pixel point, is the shape factor, is the eccentricity of the circumscribed ellipse, is the major axis of the circumscribed ellipse, is the minor axis of the circumscribed ellipse; Calculate the normalized score of each feature and combine it with the weighted coefficient to get the comprehensive score, which is expressed as: , In the formula, is the comprehensive score, is the number of features, For Region The value of the i-th feature of is the weight of the feature, is the standard deviation of feature i, is the mean of feature i; like , then remove the false detection area, is the threshold of the comprehensive score; Missed detection repair based on optimized tensor decomposition, constructing sparse adjacency tensors , the expression is: , In the formula, is the Euclidean distance between two points, is the angle between two points, represents the adjacency relationship between the i-th and j-th crack connected regions, , The crack area and The centroid coordinates of is the distance threshold, is the angle threshold, is the length of the first dimension of T, indicating the number of rows of the tensor, is the second dimension length of T, indicating the number of columns of the tensor, represents the field of real numbers; The tensor Decomposed into a low-rank form through non-negative tensor decomposition, the expression is: , In the formula, is the rank of the decomposition, is the kth factor vector used to describe the characteristics of the tensor in the first dimension, is the outer product operation of the tensor, is the kth factor vector used to describe the characteristics of the tensor in the second dimension, is the kth factor vector used to describe the characteristics of the tensor in the third dimension; Define an objective function to restore the connected area by optimizing the sparsity and smoothness constraints. The expression of the objective function is: , In the formula, and are the constraint weight coefficients for sparsity and smoothness, respectively. is the L2 norm, used to constrain smoothness, is the Frobenius norm, is the L1 norm, used to constrain sparsity, for The gradient of , , is the factor vector in tensor decomposition, corresponding to the three dimensions of tensor T; The reconstructed tensor obtained by optimization Represents possible crack connection areas and reconstructs the tensor The expression is: , In the formula, To represent the vector The i-th element of To represent the vector The jth element of ; Will Mapping back to the image space, generating the connected crack image, the expression is: , In the formula, It is a line connection operation between two points; Crack image after fusion of false detections and removal and reconstruction areas , output the final detection result, the expression is: , In the formula, The final fused crack image.
6. A concrete road crack identification system, characterized in that: include: A construction module is configured to construct a concrete road crack recognition model based on improved DeepLabv3+, wherein the concrete road crack recognition model includes an encoder and a decoder directly connected to the encoder; the encoder includes a deep convolutional network layer and an improved dilated space pyramid pooling module; the decoder includes a bilinear upsampling module, a 1×1 convolution module, and a 3×3 convolution module, wherein the improved dilated space pyramid pooling module includes: A dynamic multi-branch feature extraction module, which includes a 1×1 convolution module, a 3×3 dilated convolution module with dilation rates of 6, 12, and 18, a global average pooling module, and a dynamic asymmetric dilated convolution module. The dynamic asymmetric dilated convolution module adaptively adjusts the dilation rate in the horizontal direction according to the directional characteristics of the crack distribution in the input feature map by introducing a dynamic receptive field adjustment factor. and the vertical expansion rate , the expression is: , , In the formula, is the basic expansion rate in the horizontal direction, , are adjustment coefficients, is the variance of the input feature map in the horizontal direction, is the basic expansion rate in the vertical direction, is the variance of the input feature map in the vertical direction; The input features of the 3×3 convolution module with a dilation rate of 6 consist of the deep features and the output features of the 1×1 convolution module; The input features of the 3×3 convolutional module with a dilation rate of 12 are composed of the deep features and the output features of the convolutional module with a dilation rate of 6; The input features of the 3×3 convolutional module with a dilation rate of 18 are composed of the deep features and the output features of the convolutional module with a dilation rate of 12; The crack shape sensitive attention module combines directional gradients and edge information to weight the output features of the 3×3 dilated convolution module with dilation rates of 6, 12, and 18, the global average pooling module, and the dynamic asymmetric dilated convolution module. The attention weight calculation formula of the crack shape sensitive attention module is: , In the formula, is the channel weight, is the Sigmoid function, , are the gradients of the feature map F in the horizontal and vertical directions, respectively. is the weight matrix used to fuse gradient features with original features; An acquisition module is configured to acquire a concrete road crack image, and annotate and preprocess the concrete road crack image to obtain a concrete road crack data set; an optimization module, configured to input the concrete road crack data set into the concrete road crack recognition model, optimize the model parameters of the concrete road crack recognition model based on a composite loss function combining edge enhancement and geometric adaptability, train and generate weights of the concrete road crack recognition model, and obtain a target crack recognition model; The recognition module is configured to input the real-time concrete road crack image into the target crack recognition model, the target crack recognition model outputs a preliminary recognition result, and optimizes the preliminary recognition result through a false detection elimination and missed detection repair algorithm to obtain a final recognition result.
7. An electronic device, characterized in that: include: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Asphalt road crack detection method based on improved DeepLabv3 + network
CN117058386A
Attitude estimation method based on improved YOLOV8 algorithm
CN118609205A