Road crack detection method and related device
By introducing high-frequency information filters and dynamic cross-type convolution C3 modules in the YOLOv5 network model, the problem of insufficient crack feature extraction by traditional methods under background low-frequency information interference is solved, and the accuracy and robustness of road crack detection is significantly improved.
Patent Information
- Application Number
- CN202510212525.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional road crack detection methods based on deep neural networks are easily disturbed by background low-frequency information, resulting in insufficient fissure feature extraction and reducing the accuracy of detection results.
Using the improved YOLOv5 network model, the C3 module in the basic YOLOv5 network model is replaced with a C3 module based on high-frequency cross-type convolution, including a high-frequency information filter and a dynamic cross-type convolution.
It significantly improves the accuracy and robustness of road crack detection results, effectively enhances the extraction of edge details and high-frequency texture information, and suppresses background noise.
Smart Images

Figure CN120071091A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a road crack detection method and related device. Background Art
[0002] Regarding the problem of highway pavement aging, timely maintenance and repair work is particularly important; through regular inspections and detections, timely discovery and treatment of pavement defects can extend the service life of highways, improve driving safety and comfort. Traditional road detection methods mainly rely on manual inspections or semi-automatic detection equipment, which have many limitations.
[0003] With the continuous innovation of neural network technology, deep learning technology has emerged, bringing new breakthroughs to the field of road defect detection; among them, deep learning technology can use methods such as deep neural networks to analyze road images to identify and locate potential defects, and avoid traffic accidents and infrastructure damage caused by defects through timely road maintenance.
[0004] However, in the task of crack feature detection on roads, crack features usually show high-frequency features that are slender, continuous, and have obvious edges. However, the background of on-site road images is complex and there is a lot of noise. Traditional detection methods based on deep neural networks are extremely vulnerable to interference from the low-frequency information of the background, resulting in insufficient extraction of crack features and greatly reducing the accuracy of road crack detection results. Summary of the Invention
[0005] In view of the technical problems existing in the prior art, the present invention provides a road crack detection method and related device to solve the technical problem that traditional detection methods based on deep neural networks in the task of crack feature detection on roads are extremely vulnerable to interference from the low-frequency information of the background, resulting in insufficient extraction of crack features and greatly reducing the accuracy of road crack detection results.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows: The present invention provides a road crack detection method, including: Obtaining an original image of a road to be detected; Inputting the original image of the road to be detected into a road crack detection model, and outputting a road crack detection result of the road to be detected; wherein, the road crack detection model is an improved YOLOv5 network model; The improved YOLOv5 network model is obtained by replacing the C3 module in the basic YOLOv5 network model with a C3 module based on high-frequency cross-shaped convolution; wherein, the C3 module based on high-frequency cross-shaped convolution is a C3 module including a high-frequency information filter and a dynamic cross-shaped convolution.
[0007] Further, the C3 module based on the high-frequency cross-shaped convolution includes a high-frequency information filter, a first fusion unit, a first CBS unit, a second CBS unit, a third CBS unit, a CrossBottleneck module, a second fusion unit, and a fourth CBS unit; The high-frequency information filter is used to perform high-frequency enhancement processing on the input feature map to obtain a feature map after high-frequency enhancement processing; The first fusion unit is used to fuse the input feature map and the feature map after high-frequency enhancement processing to obtain a first fusion feature map; The first CBS unit is used to generate a first CBS feature map based on the first fusion feature map; The second CBS unit is used to generate a second CBS feature map based on the first CBS feature map; The third CBS unit is used to generate a third CBS feature map based on the first CBS feature map; The CrossBottleneck module is used to generate a dynamic convolution feature map based on the second CBS feature map; wherein, a dynamic cross-shaped convolution is introduced into the rossBottleneck module; The second fusion unit is used to fuse the third CBS feature map and the dynamic convolution feature map to obtain a second fusion feature map; The fourth CBS unit is used to generate a fourth CBS feature map based on the second fusion feature map.
[0008] Further, the CrossBottleneck module includes a fifth CBS unit, a dynamic cross-shaped convolution unit, and a third fusion unit; The fifth CBS unit is used to generate a fifth CBS feature map based on the second CBS feature map; The dynamic cross-shaped convolution unit is used to perform a dynamic cross-shaped convolution operation on the fifth feature map to obtain a convolution operation feature map; The third fusion unit is used to fuse the second CBS feature map and the convolution operation feature map to obtain a dynamic convolution feature map.
[0009] Further, in the dynamic cross-shaped convolution unit, the process of performing a dynamic cross-shaped convolution operation on the fifth feature map to obtain a convolution operation feature map includes: Using an offset convolutional layer to learn dynamic offsets from the input fifth feature map; Performing batch normalization, tanh activation, range limitation, and smoothing processing on the dynamic offsets so that the offsets are within a preset range to obtain convolution offsets; Generate new dynamic sampling coordinates by adding the convolutional offset to the fixed central coordinates of the input fifth feature map; Sample the deformed feature map from the input fifth feature map by bilinear interpolation based on the new dynamic sampling coordinates; Use a cross-shaped convolutional kernel to perform convolution processing on the deformed feature map to obtain a convolution operation feature map.
[0010] Furthermore, the process of performing high-frequency enhancement processing on the input feature map to obtain the feature map after high-frequency enhancement processing includes: Generate high-frequency information features based on the input feature map using a content encoder; Use the Softmax normalization method and the Hamming normalization method to normalize the high-frequency information features to obtain the normalized high-frequency information features; Use the adaptive filtering operation carafe to filter the normalized high-frequency information features and output the feature map after high-frequency enhancement processing.
[0011] Furthermore, the process of replacing the C3 module in the basic YOLOv5 network model with a C3 module based on high-frequency cross-shaped convolution includes: In the backbone network of the basic YOLOv5 network model, replace both the seventh-layer C3 module and the ninth-layer C module of the backbone network with C3 modules based on high-frequency cross-shaped convolution; In the neck network of the basic YOLOv5 network model, replace both the first-layer CBS module and the fourth-layer CBS module of the neck network with C3 modules based on high-frequency cross-shaped convolution.
[0012] The present invention also provides a road crack detection system, including: An image acquisition module for acquiring the original image of the road to be detected; A crack detection module for inputting the original image of the road to be detected into the road crack detection model and outputting the road crack detection result of the road to be detected; wherein, the road crack detection model is an improved YOLOv5 network model; The improved YOLOv5 network model is obtained by replacing the C3 module in the basic YOLOv5 network model with a C3 module based on high-frequency cross-shaped convolution; wherein, the C3 module based on high-frequency cross-shaped convolution is a C3 module including a high-frequency information filter and a dynamic cross-shaped convolution.
[0013] The present invention also provides a road crack detection device, including: A processor suitable for executing a computer program; A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the road crack detection method described above is executed.
[0014] The present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the road crack detection method is implemented.
[0015] The present invention also provides a computer program product including a computer program, and when the computer program is executed by a processor, the road crack detection method is implemented.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: The road crack detection method provided by the present invention uses the basic YOLOv5 network model as the basic framework of the road crack detection model, and introduces a C3 module including a high-frequency information filter and a dynamic cross-shaped convolution into the basic YOLOv5 network model, so that the model can more effectively enhance edge details and high-frequency texture information in road crack feature extraction, while suppressing background noise, thereby significantly improving the accuracy and robustness of road crack detection results; specifically, by introducing a high-frequency information filter, details and edge information in the crack image can be effectively extracted; among them, the high-frequency information filter can focus on the parts with faster changes in the image, thereby suppressing the low-frequency information and redundant noise of the background, making the edges and textures of the cracks more prominent, and providing a clearer input for subsequent feature extraction; by introducing a dynamic cross-shaped convolution, which uses dynamic offset technology to expand the traditional convolution kernel into a deformable cross-shaped structure, enabling the convolution kernel to swing adaptively in the up, down, left, and right directions, and its dynamic adjustment ability can better capture the ductility and morphological changes of cracks in different directions, thereby enhancing the extraction effect of crack features, and being able to more accurately distinguish fine cracks and background noise, especially suitable for actual scenarios with uneven crack widths and diverse morphologies.
[0017] The road crack detection system, road crack detection device, computer-readable storage medium, and computer program product provided by the present invention have all the advantages of the above road crack detection method. Description of the Drawings
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] Figure 1 Flow chart of the road crack detection method provided in Embodiment 1; Figure 2 Schematic structural diagram of the C3 module based on high-frequency cross-shaped convolution in Embodiment 1; Figure 3 Schematic structural diagram of the CrossBottleneck module in Embodiment 1; Figure 4 Schematic structural diagram of the improved YOLOv5 network model in Embodiment 1; Figure 5 Schematic diagram of the detection effect of the road crack detection model in Embodiment 1; Figure 6 Block diagram of the structure of the road crack detection system provided in Embodiment 2; Figure 7 Block diagram of the structure of the road crack detection device provided in Embodiment 3. Detailed implementation manners
[0020] In order to make the technical problems, technical solutions and beneficial effects solved by this application clearer and more understandable, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application; obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0021] Embodiment 1 As shown in the attached Figure 1 figures, a road crack detection method provided in this Embodiment 1 includes the following steps: Step 1, obtain the original image of the road to be detected.
[0022] Step 2, input the original image of the road to be detected into the road crack detection model, and output the crack detection result of the road to be detected; wherein, the road crack detection model is an improved YOLOv5 network model.
[0023] In this Embodiment 1, the improved YOLOv5 network model replaces the C3 module in the basic YOLOv5 network model with a C3 module based on high-frequency cross-shaped convolution, abbreviated as the C3_Hff_Cross module; among them, the C3 module based on high-frequency cross-shaped convolution is a C3 module that includes a high-frequency information filter and a dynamic cross-shaped convolution; specifically, the C3 module based on high-frequency cross-shaped convolution includes a high-frequency information filter, a first fusion unit, a first CBS unit, a second CBS unit, a third CBS unit, a CrossBottleneck module, a second fusion unit, and a fourth CBS unit, as shown in the appendix Figure 2 as shown.
[0024] The high-frequency information filter is used to perform high-frequency enhancement processing on the input feature map to obtain a feature map after high-frequency enhancement processing; the first fusion unit is used to fuse the input feature map and the feature map after high-frequency enhancement processing to obtain a first fused feature map; the first CBS unit is used to generate a first CBS feature map based on the first fused feature map; the second CBS unit is used to generate a second CBS feature map based on the first CBS feature map; the third CBS unit is used to generate a third CBS feature map based on the first CBS feature map; the CrossBottleneck module is used to generate a dynamic convolution feature map based on the second CBS feature map; among them, the rossBottleneck module introduces a dynamic cross-shaped convolution; the second fusion unit is used to fuse the third CBS feature map and the dynamic convolution feature map to obtain a second fused feature map; the fourth CBS unit is used to generate a fourth CBS feature map based on the second fused feature map.
[0025] In the high-frequency information filter, the process of performing high-frequency enhancement processing on the input feature map to obtain a feature map after high-frequency enhancement processing includes: Using a content encoder to generate high-frequency information features based on the input feature map; using the Softmax normalization method and the Hamming normalization method to perform normalization processing on the high-frequency information features to obtain normalized high-frequency information features; using the adaptive filtering operation carafe to perform filtering processing on the normalized high-frequency information features, and outputting a feature map after high-frequency enhancement processing.
[0026] It should be noted that in the road crack detection task, the thin and long shape of road cracks occupies more high-frequency information. In this Embodiment 1, a high-frequency information filter is used to filter and enhance road cracks, so that the feature map after high-frequency enhancement processing contains richer high-frequency information, which helps to improve the performance of the road crack target detection task. Specifically, high-frequency information features are generated from the input feature map through a content encoder. Secondly, the Softmax normalization method or the Hamming normalization method is used for normalization processing to further enhance the high-frequency filtering effect and ensure that the sum of the weights of the filter kernel is 1. Finally, the adaptive filtering operation carafe is used to enhance the high-frequency information of the features, and finally the feature map after high-frequency enhancement processing is output.
[0027] As shown in the Figure 3 accompanying drawings, the CrossBottleneck module includes a fifth CBS unit, a dynamic cross-shaped convolution unit, and a third fusion unit. The fifth CBS unit is used to generate a fifth CBS feature map based on the second CBS feature map. The dynamic cross-shaped convolution unit is used to perform a dynamic cross-shaped convolution operation on the fifth feature map to obtain a convolution operation feature map. The third fusion unit is used to fuse the second CBS feature map and the convolution operation feature map to obtain a dynamic convolution feature map.
[0028] Specifically, in the dynamic cross-shaped convolution unit, the process of performing a dynamic cross-shaped convolution operation on the fifth feature map to obtain a convolution operation feature map includes: learning a dynamic offset from the input fifth feature map using an offset convolutional layer; performing batch normalization, tanh activation, range limitation, and smoothing processing on the dynamic offset so that the offset is within a preset range to obtain a convolutional offset; adding the convolutional offset to the fixed center coordinates of the input fifth feature map to generate new dynamic sampling coordinates. The new dynamic sampling coordinates are used to determine the spatial roaming characteristics of the sampling position of the convolution kernel. Based on the new dynamic sampling coordinates, a deformed feature map is sampled from the input fifth feature map through bilinear interpolation. A cross-shaped convolution kernel is used to perform convolution processing on the deformed feature map to obtain a convolution operation feature map, thereby realizing the dynamic adaptive extraction of the input fifth feature map.
[0029] It should be noted that the road crack structure has small and fragile local features and complex and variable global morphologies, posing a huge challenge that is difficult to recognize for traditional convolution modules. In Embodiment 1, by introducing a dynamic cross-shaped convolution unit, long and narrow features are extracted at the horizontal and vertical positions of the input feature map. Among them, the position of the cross-shaped convolution kernel is dynamically adjusted by introducing a convolution offset, so that the cross-shaped convolution kernel can move and adapt in a more flexible manner on the input feature map, thereby more accurately capturing and strengthening the slender and tortuous features of the road crack structure target. Among them, the dynamic cross-shaped convolution unit is implemented through the idea of deformable convolution, that is, the sampling position of the convolution kernel is adjusted by using the dynamically learned offset, so that the originally fixed sampling grid is transformed into a sampling mode that can be stretched into a cross shape.
[0030] In Embodiment 1, the process of determining the convolution offset includes: generating a dynamic offset using standard convolution; performing a normalization operation on the dynamic offset and restricting the range to [-1, 1] through a trainable hyperparameter to avoid offsets beyond the design range, obtaining the normalized convolution kernel offset; then, smoothing the normalized convolution kernel offset through two-dimensional average pooling to obtain the convolution offset.
[0031] It should also be noted that before each iterative perception, based on the convolution offset and the position of the cross-shaped convolution kernel during the previous iterative perception, the position of the cross-shaped convolution kernel during the current iterative perception is determined to ensure that during each iterative perception, the receptive field of the cross-shaped convolution kernel always focuses on the core area of the target, avoiding loss of attention to the target due to excessive offsets. Secondly, during iterative perception, since the convolution offset may be a small value while the position coordinates of the cross-shaped convolution kernel must be integer values. In Embodiment 1, the bilinear interpolation method is used to perform weighted averaging according to the integer coordinates adjacent to the small value coordinates of the cross-shaped convolution kernel to process the cross-shaped convolution kernel at the small value coordinate position, obtaining an estimate closer to the actual value to more accurately represent the features at the decimal coordinate position.
[0032] In Embodiment 1, in the improved YOLOv5 network model, the process of replacing the C3 module in the basic YOLOv5 network model with the C3 module based on the high-frequency cross-shaped convolution includes: in the backbone network of the basic YOLOv5 network model, replacing the seventh-layer C3 module and the ninth-layer C module of the backbone network with the C3 module based on the high-frequency cross-shaped convolution; in the neck network of the basic YOLOv5 network model, replacing the first-layer CBS module and the fourth-layer CBS module of the neck network with the C3 module based on the high-frequency cross-shaped convolution, as shown in the appendix Figure 4 shown; appendix Figure 4The overall structure of the improved YOLOv5 network model is given, including a backbone network, a neck network, and a head network.
[0033] The backbone network includes a first CBS module, a second CBS module, a first C3 module, a third CBS module, a second C3 module, a fourth CBS module, a first C3_Hff_Cross module, a fifth CBS module, a second C3_Hff_Cross module, and an SPPF module connected in series in sequence.
[0034] The neck network includes a third C3_Hff_Cross module, a first upsampling module, a first fusion module, a fifth C3 module, a fourth C3_Hff_Cross module, a second upsampling module, a second fusion module, a sixth C3 module, a sixth CBS module, a third fusion module, a seventh C3 module, a seventh CBS module, a fourth fusion module, and an eighth C3 module.
[0035] The head network includes a first detection head, a second detection head, and a third detection head.
[0036] In the backbone network, the first CBS module is used to give the original image of the road to be detected and generate a first CBS feature; the second CBS module is used to generate a second CBS feature based on the first CBS feature; the first C3 module is used to generate a first C3 feature based on the second CBS feature; the third CBS module is used to generate a third CBS feature based on the first C3 feature; the second C3 module is used to generate a second C3 feature based on the third CBS feature; the fourth CBS module is used to generate a fourth CBS feature based on the second C3 feature; the first C3_Hff_Cross module is used to generate a first C3_Hff_Cross feature based on the fourth CBS feature; the fifth CBS module is used to generate a fifth CBS feature based on the first C3_Hff_Cross feature; the second C3_Hff_Cross module is used to generate a second C3_Hff_Cross feature based on the fifth CBS feature; the SPPF module is used to generate a pooled feature based on the second C3_Hff_Cross.
[0037] In the neck network, the third C3_Hff_Cross module is used to generate third C3_Hff_Cross features based on the pooled features; the first upsampling module is used to generate first upsampling features based on the third C3_Hff_Cross features; the first fusion module is used to fuse and generate first fusion features based on the first upsampling features and the first C3_Hff_Cross features; the fifth C3 module is used to generate fifth C3 features based on the first fusion features; the fourth C3_Hff_Cross module is used to generate fourth C3_Hff_Cross features based on the fifth C3 features; the second upsampling module is used to generate second upsampling features based on the fourth C3_Hff_Cross features; the second fusion module is used to fuse and generate second fusion features based on the second C3 features and the second upsampling features; the sixth C3 module is used to generate sixth C3 features based on the second fusion features; the sixth CBS module is used to generate sixth CBS features based on the sixth C3 features; the third fusion module is used to fuse and generate third fusion features based on the fourth C3_Hff_Cross features and the sixth CBS features; the seventh C3 module is used to generate seventh C3 features based on the third fusion features; the seventh CBS module is used to generate seventh CBS features based on the seventh C3 features; the fourth fusion module is used to fuse and generate fourth fusion features based on the third C3_Hff_Cross features and the seventh CBS features; the eighth C3 module is used to generate eighth C3 features based on the fourth fusion features.
[0038] It should be noted that the input of the third C3_Hff_Cross module is the pooled features output by the SPPF module; the inputs of the first fusion module include the first C3_Hff_Cross output by the first C3_Hff_Cross module and the first upsampling features output by the first upsampling module; the inputs of the second fusion module include the second C3 features output by the second C3 module and the second upsampling features output by the second upsampling module; the inputs of the third fusion module include the fourth C3_Hff_Cross features output by the fourth C3_Hff_Cross module and the sixth CBS features output by the sixth CBS module; the inputs of the fourth fusion module include the third C3_Hff_Cross features output by the third C3_Hff_Cross module and the seventh CBS features output by the seventh CBS module.
[0039] In the head network, the first detection head is used to detect defects based on the sixth C3 features, the second detection head is used to detect defects based on the seventh C3 features, and the third detection head is used to detect defects based on the eighth C3 features.
[0040] In this Embodiment 1, by introducing a high-frequency information filter, the details and edge information in the crack image can be effectively extracted. Specifically, the high-frequency filter focuses on the parts with faster changes in the image, thereby suppressing the low-frequency information of the background and redundant noise, making the edges and textures of the cracks more prominent, and providing a clearer input for subsequent feature extraction. By introducing a dynamic cross-shaped convolution and using the dynamic offset technology, the traditional convolution kernel is extended into a cross-shaped structure in the horizontal and vertical directions, enabling the convolution kernel to swing adaptively in the up, down, left, and right directions, so that it can better capture the ductility and morphological changes of the cracks in different directions, and further enhance the extraction effect of crack features, which is particularly important for the actual scenarios with uneven crack widths and diverse morphologies because it can more accurately distinguish fine cracks and background noise. In summary, by introducing the C3 module containing a high-frequency information filter and a dynamic cross-shaped convolution, the model can more effectively enhance the edge details and high-frequency texture information in crack feature extraction, while suppressing background noise, thus significantly improving the detection accuracy and robustness, and effectively avoiding the problem of insufficient crack feature extraction caused by low-frequency interference.
[0041] Verification experiment description: During the verification experiment process, an experimental environment was created with two RTX4090 GPUs and one Intel i9-13900k CPU. The model was trained with a batch size of 64, and the resolution of the input images was fixed at 640×640. To fully train the model and achieve a good convergence effect, the number of training epochs was set to 100. In terms of the optimizer and data augmentation, SGD was selected as the optimizer, and the mosaic data augmentation method was adopted. In the test stage, the resolution of the input images was fixed at 640×640, and the confidence threshold of 0.25 and the IOU (Intersection over Union) threshold of 0.45 were set. Among them, the selection of the above parameters was based on the experimental experience during the training and verification processes, aiming to achieve the best detection performance. All the experiments were completed on the same hardware device, ensuring the fairness and comparability of the experiments.
[0042] Using input images with a resolution of 640×640, a comparative experiment was conducted between the basic YOLOv5 network model and the road crack detection model in this Embodiment 1. Among them, the basic YOLOv5 network model was taken as YOLOv5m. The comparative experiment results of YOLOv5m and the road crack detection model in this Embodiment 1 are shown in Table 1 below. The schematic diagram of the road crack detection effect of the road crack detection model in this Embodiment 1 is as shown in the appendix Figure 5 as follows.
[0043] Table 1 Comparative experiment results of YOLOv5m and the road crack detection model in this Embodiment 1
[0044] As can be seen from Table 1 above, the road crack detection model in Embodiment 1 is superior to YOLOv5m in both overall mAP50 and overall mAP95, with an improvement of 1.3% in both; this indicates that the road crack detection method described in Embodiment 1 has certain advantages in overall detection performance; in addition, from the comparison of mAP50 of specific crack categories, the road crack model has varying degrees of improvement in each category.
[0045] From the attached Figure 5 it can be seen that for different cracks in different road scenarios, the road crack detection model described in Embodiment 1 can give accurate detection results; among them, each rectangular box marks the detected crack area and is accompanied by a corresponding confidence score for evaluating the reliability of model recognition.
[0046] Embodiment 2 As shown in the attached Figure 6 In this embodiment, a road crack detection system is provided, including an image acquisition module and a crack detection module; the image acquisition module is used to acquire the original image of the road to be detected; the crack detection module is used to input the original image of the road to be detected into the road crack detection model and output the road crack detection result of the road to be detected; among them, the road crack detection model is an improved YOLOv5 network model.
[0047] In this embodiment, the improved YOLOv5 network model replaces the C3 module in the basic YOLOv5 network model with a C3 module based on high-frequency cross-shaped convolution; among them, the C3 module based on high-frequency cross-shaped convolution is a C3 module including a high-frequency information filter and a dynamic cross-shaped convolution.
[0048] Embodiment 3 As shown in the attached Figure 7 In this embodiment, a road crack detection device is provided, including: a memory for storing a computer program; a processor for implementing the steps of the road crack detection method when executing the computer program.
[0049] When the processor executes the computer program, it implements the steps of the above road crack detection method, for example: Acquire the original image of the road to be detected; input the original image of the road to be detected into the road crack detection model and output the road crack detection result of the road to be detected; among them, the road crack detection model is an improved YOLOv5 network model.
[0050] Or, when the processor executes the computer program, it implements the functions of each module in the above road crack detection system, for example: An image acquisition module for acquiring the original image of the road to be detected; a crack detection module for inputting the original image of the road to be detected into a road crack detection model and outputting the road crack detection result of the road to be detected; wherein, the road crack detection model is an improved YOLOv5 network model.
[0051] Exemplarily, the computer program can be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of completing preset functions, and the instruction segments are used to describe the execution process of the computer program in the road crack detection device.
[0052] For example, the computer program can be divided into an image acquisition module and a crack detection module, and the specific functions of each module are as follows: an image acquisition module for acquiring the original image of the road to be detected; a crack detection module for inputting the original image of the road to be detected into a road crack detection model and outputting the road crack detection result of the road to be detected; wherein, the road crack detection model is an improved YOLOv5 network model.
[0053] The road crack detection device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The road crack detection device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above are examples of the road crack detection device, which do not constitute a limitation on the road crack detection device, and may include more components than the above, or combine some components, or different components. For example, the road crack detection device may further include an input / output device, a network access device, a bus, etc.
[0054] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc. The processor is the control center of the road crack detection device, and connects various parts of the entire road crack detection device through various interfaces and lines.
[0055] The memory can be used to store the computer program and / or module. By running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory, the processor realizes various functions of the road crack detection device.
[0056] The memory may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0057] Embodiment 4 Embodiment 4 of the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the road crack detection method described above are realized, for example: Obtain the original image of the road to be detected; input the original image of the road to be detected into the road crack detection model, and output the road crack detection result of the road to be detected; wherein, the road crack detection model is an improved YOLOv5 network model.
[0058] If the modules / units integrated in the road crack detection system are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
[0059] Based on such an understanding, to implement all or part of the processes in the above road crack detection method of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above road crack detection method can be realized. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or preset intermediate form, etc.
[0060] The computer-readable storage medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0061] It should be noted that the content included in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.
[0062] Embodiment 5 Embodiment 5 of the present invention provides a computer product. The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium; a processor of the road crack detection device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the road crack detection device can execute the road crack detection method described in Embodiment 1, which will not be elaborated here.
[0063] It should be noted that those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the above embodiments of the methods.
[0064] The above embodiments are only one of the implementation manners capable of implementing the technical solution of the present invention. The scope of protection required by the present invention is not only limited by this embodiment, but also includes any changes, substitutions, and other implementation manners that are easily conceivable by those skilled in the art within the technical scope disclosed by the present invention.
Claims
1. A road crack detection method, characterized in that: include: Obtaining the original image of the road to be detected; Inputting the original image of the road to be detected into a road crack detection model, and outputting a road crack detection result of the road to be detected; wherein the road crack detection model is an improved YOLOv5 network model; The improved YOLOv5 network model replaces the C3 module in the basic YOLOv5 network model with a C3 module based on high-frequency cross convolution; wherein the C3 module based on high-frequency cross convolution is a C3 module including a high-frequency information filter and a dynamic cross convolution.
2. A road crack detection method according to claim 1, characterized in that: The C3 module based on high-frequency cross convolution includes a high-frequency information filter, a first fusion unit, a first CBS unit, a second CBS unit, a third CBS unit, a CrossBottleneck module, a second fusion unit and a fourth CBS unit; The high-frequency information filter is used to perform high-frequency enhancement processing on the input feature map to obtain the feature map after high-frequency enhancement processing; The first fusion unit is used to fuse the input feature map and the feature map after high-frequency enhancement processing to obtain a first fused feature map; The first CBS unit is used to generate a first CBS feature map based on the first fusion feature map; The second CBS unit is configured to generate a second CBS feature map based on the first CBS feature map; The third CBS unit is configured to generate a third CBS feature map based on the first CBS feature map; The CrossBottleneck module is used to generate a dynamic convolution feature map based on the second CBS feature map; wherein the CrossBottleneck module introduces a dynamic cross convolution; The second fusion unit is used to fuse the third CBS feature map and the dynamic convolution feature map to obtain a second fused feature map; The fourth CBS unit is used to generate a fourth CBS feature map based on the second fusion feature map.
3. A road crack detection method according to claim 2, characterized in that: The CrossBottleneck module includes a fifth CBS unit, a dynamic cross convolution unit and a third fusion unit; The fifth CBS unit is used to generate a fifth CBS feature map based on the second CBS feature map; The dynamic cross convolution unit is used to perform a dynamic cross convolution operation on the fifth feature map to obtain a convolution operation feature map; The third fusion unit is used to fuse the second CBS feature map and the convolution operation feature map to obtain a dynamic convolution feature map.
4. A road crack detection method according to claim 3, characterized in that: In the dynamic cross convolution unit, a process of performing a dynamic cross convolution operation on the fifth feature map to obtain a convolution operation feature map includes: Use the offset convolution layer to learn the dynamic offset from the fifth feature map of the input; The dynamic offset is batch normalized, tanh activated, range limited, and smoothed to keep the offset within the preset range, thus obtaining the convolution offset. Generate new dynamic sampling coordinates by adding the convolution offset and the fixed center coordinates of the input fifth feature map; Based on the new dynamic sampling coordinates, a deformed feature map is sampled from the fifth feature map of the input by a bilinear interpolation method; A cross-shaped convolution kernel is used to perform convolution processing on the deformed feature map to obtain the convolution operation feature map.
5. A road crack detection method according to claim 2, characterized in that: The process of performing high-frequency enhancement processing on the input feature map to obtain the feature map after high-frequency enhancement processing includes: Using a content encoder to generate high-frequency information features based on the input feature map; Using a Softmax normalization method and a Hamming normalization method, the high-frequency information features are normalized to obtain normalized high-frequency information features; The normalized high-frequency information features are filtered using an adaptive filtering operation carafe, and a feature map subjected to high-frequency enhancement processing is output.
6. A road crack detection method according to claim 1, characterized in that: The process of replacing the C3 module in the basic YOLOv5 network model with the C3 module based on high-frequency cross convolution includes: In the backbone network of the basic YOLOv5 network model, the seventh-layer C3 module and the ninth-layer C module of the backbone network are replaced with C3 modules based on high-frequency cross convolution; In the neck network of the basic YOLOv5 network model, the first-layer CBS module and the fourth-layer CBS module of the neck network are replaced by the C3 module based on high-frequency cross convolution.
7. A road crack detection system, characterized in that: include: An image acquisition module is used to acquire the original image of the road to be detected; A crack detection module, used for inputting the original image of the road to be detected into a road crack detection model, and outputting a road crack detection result of the road to be detected; wherein the road crack detection model is an improved YOLOv5 network model; The improved YOLOv5 network model replaces the C3 module in the basic YOLOv5 network model with a C3 module based on high-frequency cross convolution; wherein the C3 module based on high-frequency cross convolution is a C3 module including a high-frequency information filter and a dynamic cross convolution.
8. A road crack detection device, characterized in that: include: a processor suitable for executing a computer program; A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the road crack detection method according to any one of claims 1 to 6 is executed.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the road crack detection method according to any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the road crack detection method according to any one of claims 1 to 6 is implemented.