A model and method for road surface defect detection
Through the improved CSPDarkNet53 backbone network, PAFPN detection neck and WTC detection head, combined with the star computing module, multi-dimensional auxiliary fusion module and wavelet transform convolution module, the limitations of the existing road surface defect detection method in complex scenarios are solved, and high-precision small object detection and complex background processing are achieved.
Patent Information
- Application Number
- CN202510178205.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The existing road surface defect detection methods show limitations in complex scenarios, especially in the treatment of small target detection, complex background interference and irregular shape cracks, which are difficult to meet the requirements of high precision.
A model for road defect detection is proposed. The improved CSPDarkNet53 backbone network, path aggregation feature pyramid network (PAFPN) detection neck and wavelet transform convolution (WTC) detection head are enhanced by introducing star operation module (SOM), multi-dimensional auxiliary fusion module (MAF) and wavelet transform convolution module.
It significantly improves the detection accuracy of small targets in the defect detection process, enhances the robustness in complex pavement scenarios, and improves the overall performance and reliability of the model in complex backgrounds and various types of defect processing.
Smart Images

Figure CN119649017B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection, and specifically relates to a model and method for road surface defect detection. Background Art
[0002] Existing road surface defect detection methods usually adopt traditional target detection algorithms such as YOLO, SSD, etc. Although these methods have real-time performance and good detection capabilities, they show obvious limitations when dealing with small target detection, complex background interference, and irregularly shaped cracks in complex scenarios. Traditional methods have deficiencies in the accuracy of small target detection, the ability to suppress complex backgrounds, and the capture of irregular crack shapes, and it is difficult to meet the high-precision requirements of actual engineering applications.
[0003] Therefore, there is an urgent need for an improved road surface defect detection method that can achieve high-precision detection of road surface defects in complex scenarios, especially for the detection requirements of small targets, complex backgrounds, and irregular crack shapes. Summary of the Invention
[0004] The technical solution of the present invention is as follows: A model for road surface defect detection is proposed, which includes a backbone network, a detection neck, and a detection head. The backbone network is an improved CSPDarkNet53 (a name of a backbone network), which includes four stages and its main function is feature extraction. The detection neck is an improved Path Aggregation Feature Pyramid Network (PAFPN), which includes three feature fusion processes of top-down, bottom-up, and top-down. The detection head mainly completes classification and regression tasks. Among them, a Star-shaped Operation Module (SOM) is introduced into CSPDarkNet53, and SOM replaces the ConvModule in the third and fourth stages of CSPDarkNet53. A Multi-dimensional Auxiliary Fusion (MAF) module is introduced into the original PAFPN. Before the bottom-up feature fusion, the features are first weighted through the MAF module, and the output features of the four stages of the backbone network are preliminarily fused. A Wavelet Transform Convolution (WTC) module is introduced into the detection head, and WTC is used to replace the first two ConvModules in the detection head to reduce the complexity of the model, and then the features are classified and regressed.
[0005] A method for road surface defect detection. After receiving the input road surface defect image, the model first enters the backbone network, which is the upgraded CSPDarkNet53. Specifically: in the third and fourth stages of the backbone network, we replace the traditional ConvModule module with the Star Operation Module (SOM) to enhance the feature extraction ability of the backbone network. 8 consecutive SOMs are applied in the third stage, and 5 consecutive SOMs are applied in the fourth stage. In the SOM, the depth convolution layer and the normalization layer perform preliminary processing on the features. The depth convolution (DW-Conv) and batch normalization (BN) are used to adjust the features in the spatial and channel dimensions to capture more context information. After activation by Relu6 (a variant of the rectified linear unit (Relu)), the features enhance the feature selection ability through element-wise multiplication, providing high-quality input for subsequent feature aggregation. The key of the SOM is to map the features to a high-dimensional space, enhancing the feature extraction ability. Subsequently, the output features are input into the improved Path Aggregation Feature Pyramid Network (PAFPN) for feature fusion processing. This part mainly performs feature fusion on the output of the backbone network. The improved PAFPN introduces the Multi-dimensional Auxiliary Fusion (MAF) module compared with the original PAFPN. The MAF processes the output from the backbone network. Here, changes are made to the output of the backbone, adding the output features of the first stage of CSPDarkNet53, which contain more detailed information. The MAF enables better fusion of the detailed information contained in the high-dimensional features and the semantic information in the low-dimensional features, and performs weighted processing on the fused features, significantly improving the detection effect for small targets. Subsequently, the output features of the PAFPN are input into the detection head, which mainly performs classification and regression tasks. We replace the original two-layer ConvModule in the detection head with the Wavelet Transform Convolution (WTC) module. In this part, the input features are extracted into high-frequency and low-frequency components through wavelet transform, and dimensionality reduction operations are performed on them respectively. Small convolution kernels are used to operate on the low-dimensional features to obtain a large receptive field, significantly reducing the number of model parameters. After passing through the WTC, classification and regression tasks are performed to obtain the output results of the model.
[0006] Furthermore, the specific processing of the star operation module is as follows:
[0007] CSPDarkNet53 processes the output features of the second stage, whose dimensions are H×W×C, where H, W, and C represent the length, width, and number of channels of the features, respectively. The H×W×C features are downsampled to H / 2×W / 2×C dimensions.
[0008] Then the features enter the deep convolution layer, that is, each channel has an independent convolution kernel for preliminary feature extraction, and the output dimension remains unchanged, that is, H / 2×W / 2×C.
[0009] After the first deep convolutional layer, batch normalization is performed to keep the feature dimension unchanged.
[0010] The output feature x is then convolved point-wise (1×1 convolution) simultaneously to obtain x1 and x2, which have the same dimension and the number of channels becomes three times the original, which is H / 2×W / 2×3C.
[0011] Apply the Relu6 activation function on x1 to introduce nonlinearity, and the dimension remains the original H / 2×W / 2×3C.
[0012] Then, the activated x1 is multiplied element-wise with x2, so that the activation value of x1 controls the feature of x2 and obtains feature x.
[0013] Next, x is input into a layer of point-by-point convolution (1×1 convolution), the number of channels of x is restored to C, and the output feature dimension is H / 2×W / 2×C. The output feature continues to input a layer of deep convolution, and the dimension remains unchanged, still H / 2×W / 2×C. This layer of deep convolution compresses and refines the features after fusion to further enhance important spatial information, remove redundant features, and maintain low computational complexity.
[0014] Finally, a layer of 1×1 point-by-point convolution is used to change the number of channels of the output feature, and the dimension is changed to H / 2×W / 2×2C. At this point, a complete star operation module operation is completed.
[0015] Furthermore, the third stage applies eight consecutive star operation modules, and the fourth stage applies five consecutive star operation modules.
[0016] Next, the output features are input into the Path Aggregation Feature Pyramid Network (PAFPN) for feature fusion processing, in which a multi-dimensional auxiliary fusion module (MAF) is introduced.
[0017] Furthermore, the multi-dimensional auxiliary fusion module processes specifically as follows:
[0018] The output features of the first stage of the backbone network (160×160×64) are convolved through the MAF module to (80×80×64). The output features of the second stage (80×80×128) pass through the channel attention mechanism and are concatenated with the features after convolution in the first stage in the channel dimension to become (80×80×192). Horizontally, it passes through a 1×1 convolutional layer, and the number of channels is adjusted to 128. Vertically, it passes through a convolutional layer and a 1×1 convolutional layer, and the features become (40×40×64). The output features of the third stage (40×40×256) pass through the channel attention mechanism and are concatenated with the features after fusion in the second stage, becoming (40×40×320). Horizontally, it passes through a 1×1 convolutional layer, and the number of channels is adjusted to 256. Vertically, it passes through a convolutional layer and a 1×1 convolutional layer, and the features become (20×20×64). The output features of the fourth stage (20×20×512) maintain the original resolution. After passing through the channel attention mechanism, they are concatenated with the features after fusion in the third stage in the channel dimension to generate a comprehensive multi-scale feature. The dimension after concatenation is (20×20×576). Horizontally, it passes through a 1×1 convolutional layer, and the number of channels is adjusted to 512.
[0019] Furthermore, inside the detection head, the wavelet transform convolutional module (WTC) extracts high-frequency and low-frequency components from the features through wavelet transform and performs dimensionality reduction operations on them respectively.
[0020] The wavelet transform convolutional module is configured as follows:
[0021] The input feature X first undergoes wavelet transform (WT) and is decomposed into four components: the low-frequency component X LL (1) , the horizontal high-frequency component X LH (1) , the vertical high-frequency component X HL (1) and the diagonal high-frequency component X HH (1) . The feature size is reduced to half of the original, and the resolution of each component is H / 2×W / 2×C. The low-frequency component X LL (1) retains the main structure and global information of the image. The high-frequency components X LH (1) , X HL (1) , X HH (1) capture edge and texture detail information. A second wavelet decomposition is performed on the low-frequency component X LL (1) obtained from the first decomposition to get four new components: X LL (2) , X LH(2) , X HL (2) , X HH (2) , the size of the newly decomposed feature is (H / 4)×(W / 4)×C. Convolution operations are performed on each component feature obtained from the first and second wavelet decompositions for X LL (1) , X LH (1) , X HL (1) , X HH (1) , and X LL (2) , X LH (2) , X HL (2) , X HH (2) Convolutions are performed sequentially to obtain the convolution result Y LL (1) , Y LH (1) , Y HL (1) , Y HH (1) , and Y LL (2) , Y LH (2) , Y HL (2) , Y HH (2) , The convolved features are fused through the inverse wavelet transform (IWT). For the first-level convolution result Y LL (1) , Y LH (1) , Y HL (1) , Y HH (1) , the inverse wavelet transform is performed to obtain the feature Z (1) , For the second-level convolution result Y LL (2) , Y LH (2) , Y HL (2) , Y HH (2) the inverse wavelet transform is performed to obtain the feature Z (2) , By performing the inverse transform fusion on the low-frequency and high-frequency components, the generated feature contains both the global information of the low-frequency component and the detailed information of the high-frequency component. The size of the feature after the inverse wavelet transform is restored to be the same as the input. The obtained feature Z (1) and Z(2) Fusion is performed by adding pixel by pixel to form the final output feature Z (0) The size of the output feature is the same as the input, which is H×W×C
[0022] The beneficial effects of the present invention are as follows
[0023] The present invention overcomes the problem of inaccurate target detection in the prior art in complex scenarios. The present invention can effectively detect defects on the road surface, especially solves the problems of insufficient performance in small target detection (small cracks on the road surface), interference from complex backgrounds, and poor handling of irregularly shaped cracks in the prior art. The present invention not only significantly improves the detection accuracy of small targets in the defect detection process of the model, but also enhances its robustness in complex road surface scenarios, and significantly improves the overall performance and reliability of the model in dealing with complex backgrounds and various types of defects. This improvement can ensure accurate detection of defects under various complex road conditions and improve the efficiency and quality of road maintenance work Description of the Drawings
[0024] Figure 1 is the overall structure diagram of the model
[0025] Figure 2 is the structure diagram of the star operation (SOM) module
[0026] Figure 3 is the spatial attention structure diagram (SA) of the multi-dimensional auxiliary fusion module
[0027] Figure 4 is the flow chart of the wavelet transform convolution (WTC) module
[0028] In the figure
[0029] CSPlayer cross-stage partial stacking layer, ConvModule convolution module, SOM star operation module, SPPF spatial pyramid pooling layer, SA spatial attention mechanism, Concat splicing layer, Conv1*11×1 convolution layer, Upsample upsampling layer, WTC wavelet transform convolution, Backbone backbone network, Neck detection neck, Head detection head, Input input, Stage stage, offset offset, hard sigmoid activation function Detailed Embodiments
[0030] The object of the present invention is to overcome the problem of inaccurate object detection in complex scenarios in the prior art, and a method for road surface defect detection is proposed. This method can effectively detect defects on the road surface, especially solve the problems of insufficient performance in small object detection (small cracks on the road surface), complex background interference, and poor handling of irregularly shaped cracks in the prior art. By integrating a Star Operation Module (SOM), a Multidimensional Auxiliary Fusion (MAF) module, and a Wavelet Transform Convolution (WTC) module, the present invention not only significantly improves the detection accuracy of small objects in the defect detection process of the model, but also enhances its robustness in complex road surface scenarios, and significantly improves the overall performance and reliability of the model in dealing with complex backgrounds and various types of defects. This improvement can ensure accurate detection of defects under various complex road conditions, and improve the efficiency and quality of road maintenance work.
[0031] The technical solution adopted by the present invention to achieve the above object is as follows:
[0032] In the first aspect, a method for road surface defect detection is provided.
[0033] When the network receives a road surface defect image, it first enters the CSPDarkNet53 backbone network for feature extraction. And there is a Star Operation Module (SOM) configured in CSPDarkNet53. Specifically: it replaces the ConvModule in the third and fourth stages of the network to enhance the feature extraction ability of the backbone network. SOM is located in the deep feature extraction stage of the backbone network. 8 consecutive SOMs are applied in the third stage, and 5 consecutive SOMs are applied in the fourth stage. Through experiments, the detection accuracy of the model can reach the highest when the depths of SOM in the two stages are 8 and 5. SOM can map data to a high-dimensional, non-linear feature space, where the hidden patterns and complex relationships of the data become more obvious and easy to analyze. Especially when the task is complex, the deep network can extract and combine more abstract and higher-level features layer by layer, and the road surface defect images in complex scenarios can be better processed. Subsequently, the output features are input into the Path Aggregation Feature Pyramid Network (PAFPN) for feature fusion processing. This part mainly performs feature fusion on the output of the backbone network. A Multi-dimensional Auxiliary Fusion (MAF) module is introduced in PAFPN. MAF processes the output from the backbone network. Here, changes are made to the output of the backbone, and the output features of the first stage of CSPDarkNet53 are added, which contain more detailed information. MAF enables better fusion of the detailed information contained in the high-dimensional features and the semantic information in the low-dimensional features, and performs weighted processing on the fused features, significantly improving the detection effect for small targets. Subsequently, the output features of PAFPN are input into the detection head (YOLOv8-head). The detection head mainly performs classification and regression tasks. But before classification and regression, the two-layer ConvModule originally included in YOLOv8 is replaced with a Wavelet Transform Convolution (WTC) module. In this part, the input features are subjected to wavelet transform to extract high-frequency and low-frequency components, and dimensionality reduction operations are performed on them respectively. Small convolution kernels are used to operate on the low-dimensional features to obtain a large receptive field, significantly reducing the number of model parameters. After passing through WTC, classification and regression tasks are performed to obtain the output results of the model.
[0034] First, the output tensor of the second stage of CSPDarkNet53 is processed. The dimension of the output tensor of the second stage is H×W×C, where H, W, and C represent the length, width, and number of channels of the feature respectively. The tensor of H×W×C is downsampled, and the tensor dimension becomes H / 2×W / 2×C.
[0035] The tensor is input into a layer of depthwise convolution, i.e., each channel has an independent convolution kernel for preliminary feature extraction. The depthwise convolution operation reduces the computational amount and maintains the extraction of spatial features, and the output dimension remains unchanged, i.e., H / 2×W / 2×C.
[0036] Batch normalization is enabled after the first layer of depthwise convolution, which helps to stabilize the training process, accelerate the convergence speed, and prevent the problems of gradient vanishing or explosion. The tensor dimension remains unchanged.
[0037] Then, the output tensor x is simultaneously subjected to pointwise convolution (1×1 convolution) to obtain x1 and x2, which have the same dimension and the number of channels becomes three times the original, i.e., H / 2×W / 2×3C. The change in the number of channels here can be regarded as a linear combination of each input channel, used to fuse information from different channels. By increasing the feature dimension, the model can learn richer expressions, thereby capturing more complex patterns and features. This is the amplification and enhancement of features, expressing the features in a higher-dimensional space. The Relu6 activation function is applied to x1 to introduce non-linearity, and the dimension remains the original H / 2×W / 2×3C.
[0038] Then, the activated x1 and x2 are multiplied element-wise, allowing the activation values of x1 to control the features of x2, resulting in the tensor x. The dimension remains unchanged in this step. This element-wise multiplication generates more interactions of details and high-order features in the high-dimensional space, helping the model focus on more useful information while ignoring irrelevant information.
[0039] Next, x is input into a layer of pointwise convolution (1×1 convolution) to restore the number of channels of x to C, and the output tensor dimension is H / 2×W / 2×C. Because after the element-wise multiplication in the high dimension, the feature space contains more complex relationships and high-order information, this pointwise convolution layer integrates and extracts the most important parts of these high-order features, ensuring that the subsequent features can still retain important representation information while reducing the dimension.
[0040] The output tensor continues to be input into a layer of depthwise convolution, and the dimension remains unchanged, still H / 2×W / 2×C. This layer of depthwise convolution compresses and refines the features after fusion to further enhance important spatial information, remove redundant features, and at the same time maintain a low computational complexity.
[0041] Finally, a 1×1 convolution is used to change the number of channels of the output tensor, and the dimension changes to H / 2×W / 2×2C. So far, the operation of a complete star operation module is completed.
[0042] The original output of the Backbone of the Yolov8 model is three tensors, coming from the second, third, and fourth stages respectively. The backbone is modified so that the tensor of its first stage is also input into the neck. This adjustment provides more detailed information support for the subsequent multi-dimensional auxiliary fusion module (MAF), which helps to achieve better results in small object detection and complex background suppression. After the modification, the output tensors of the Backbone are expanded from the original three to four. These tensors correspond to the features extracted at different stages. The specific dimensional changes are as follows:
[0043] Output tensor of the first stage:
[0044] Dimensions: (160×160×64)
[0045] Feature content: The tensor of the first stage contains high-resolution features, retaining a large amount of detailed information in the image. Adding the output of the first stage to the Neck layer can combine more low-level detailed features during subsequent feature fusion, which helps to improve the effect of small object detection and is especially important when detecting subtle defects or small cracks.
[0046] Output tensor of the second stage:
[0047] Dimensions: (80×80×128)
[0048] Feature content: This tensor has undergone one downsampling and contains medium-resolution features, extracting partial detailed and shape information of the image. After inputting this feature into the Neck layer and combining it with the high-resolution features of the first stage, it can further enhance the performance of the model in detecting larger objects and effectively suppress the interference of complex backgrounds.
[0049] Output tensor of the third stage:
[0050] Dimensions: (40×40×256)
[0051] Feature content: The features of this stage contain strong semantic information and can identify structural features such as the direction and edges of cracks. In complex backgrounds, these semantic features provide strong support to ensure that the model can accurately identify irregularly shaped defects.
[0052] Output tensor of the fourth stage:
[0053] Dimensions: (20×20×512)
[0054] Feature content: This feature describes the global information of the image at the lowest resolution and has a higher-level abstract semantics. It provides a global feature representation for the model, which helps to ensure the consistency of the detection results
[0055] In the Neck layer, the Multi-dimensional Auxiliary Fusion Module (MAF) combines the features of the above four stages and completes feature fusion through a series of processing steps, specifically including the following dimensional change processes:
[0056] Feature concatenation and fusion:
[0057] The features of the first stage (160×160×64) are convolved through the MAF module to (80×80×64). The features of the second stage (80×80×128) pass through the channel attention mechanism and are concatenated with the features after the convolution operation of the first stage in the channel dimension to become (80×80×192). Horizontally, it passes through a 1×1 convolutional layer, and the number of channels is adjusted to 128. Vertically, it passes through a convolutional layer and a 1×1 convolutional layer, and the features become (40×40×64). The features of the third stage (40×40×256) pass through the channel attention mechanism and are concatenated with the features after the second-stage fusion, becoming (40×40×320). Horizontally, it passes through a 1×1 convolutional layer, and the number of channels is adjusted to 256. Vertically, it passes through a convolutional layer and a 1×1 convolutional layer, and the features become (20×20×64). The features of the fourth stage (20×20×512) maintain the original resolution, and all features are concatenated in the channel dimension to generate a comprehensive multi-scale feature. The dimension after concatenation is (20×20×576). Horizontally, it passes through a 1×1 convolutional layer, and the number of channels is adjusted to 512.
[0058] In stage 2, 3, and 4, after the output, a channel attention mechanism is added to dynamically allocate the weights of each channel using the attention mechanism. The weights of low-level features are higher to highlight small targets and detailed information; the weights of high-level features are relatively lower and are used to capture global information.
[0059] By introducing the output tensor of the first stage and combining the Multi-dimensional Auxiliary Fusion Module (MAF), the present invention achieves higher flexibility and robustness in feature fusion. The multi-scale feature processing, convolutional operation, weighted processing, and attention mechanism of the MAF module significantly improve the detection accuracy and reliability of the model, giving it obvious advantages in complex road surface defect detection tasks, especially in small target detection and the suppression of complex background interference.
[0060] The input feature map X first undergoes wavelet transform (WT) and is decomposed into four components: the low-frequency component X LL (1) 、the horizontal high-frequency component X LH (1) 、the vertical high-frequency component X HL (1) and the diagonal high-frequency component X HH (1), the size of the feature map is reduced to half of the original, and the resolution of each component is H / 2×W / 2×C. The low-frequency component X LL (1) retains the main structure and global information of the image, which helps to capture macroscopic features. The high-frequency component X LH (1) , X HL (1) , X HH (1) captures detailed information such as edges and textures, enabling the model to perform better in detecting small features such as edges and cracks.
[0061] Next, a wavelet decomposition is performed again on the low-frequency component X LL (1) obtained from the first decomposition, resulting in four new components: X LL (2) , X LH (2) , X HL (2) , X HH (2) After the new decomposition, the size of the feature map is (H / 4)×(W / 4)×C, further reducing the spatial resolution. The secondary decomposition provides a lower resolution and higher abstract information for the feature map, enabling the model to effectively handle global complex features and background interference. In the multi-scale irregular crack scenario, such decomposition helps the model to better identify targets of different scales.
[0062] Then, convolution operations are performed on the respective component feature maps obtained from the first and second wavelet decompositions. For X LL (1) , X LH (1) , X HL (1) , X HH (1) , and X LL (2) , X LH (2) , X HL (2) , X HH (2) Convolution is performed in sequence to obtain the convolution results Y LL (1) , Y LH (1) , Y HL (1) , Y HH (1) , and Y LL (2) , Y LH(2) ,Y HL (2) ,Y HH (2) This step convolves the low - frequency and high - frequency components separately, which can enhance the capture of detailed features while maintaining the global structural information. After reducing the spatial resolution, the convolution operation with a small convolution kernel can cover the necessary receptive field and reduce the computational cost, improving the efficiency of the model.
[0063] The convolved feature maps are fused through the inverse wavelet transform (IWT). First, perform the inverse wavelet transform on the first - level convolution result Y LL (1) ,Y LH (1) ,Y HL (1) ,Y HH (1) to obtain the feature map Z (1) . Similarly, perform the inverse wavelet transform on the second - level convolution result Y LL (2) ,Y LH (2) ,Y HL (2) ,Y HH (2) to obtain the feature map Z (2) . By inversely transforming and fusing the low - frequency and high - frequency components, the generated feature map contains both the global information of the low - frequency component and the detailed information of the high - frequency component. The size of the feature map after the inverse wavelet transform is restored to be the same as the input.
[0064] Fuse the obtained feature maps Z (1) and Z (2) by adding them pixel - by - pixel to form the final output feature map Z (0) . The size of the output feature map is the same as the input, i.e., (H×W×C). The pixel - by - pixel addition fusion method allows information at different scales to be combined at the same level, thus forming a comprehensive feature map containing multi - scales and multi - frequencies, which retains multi - scale features and has rich detailed information.
[0065] The above - mentioned is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modification, equivalent replacement, and improvement made within the spirit and principle of the present invention should be covered within the protection scope of the present invention. At the same time, the content not described in detail in this specification belongs to the prior art well - known to those skilled in the art.
Claims
1. A method for detecting road surface defects, characterized in that: include: The invention discloses a backbone network, a detection neck and a detection head in four stages, wherein the backbone network is an improved CSPDarkNet53, whose main function is feature extraction, and a star operation module is introduced in the third and fourth stages to replace the original convolution module; the detection neck is an improved path aggregation feature pyramid network, which includes three feature fusion processes from top to bottom, from bottom to top and from top to bottom, and a multi-dimensional auxiliary fusion module is introduced in the path aggregation feature pyramid network. Before the feature fusion from bottom to top, the features are weighted by the multi-dimensional auxiliary fusion module, and the output features of the four stages of the backbone network are preliminarily fused; the detection head mainly completes the classification and regression tasks, and a wavelet transform convolution module is introduced in the detection head. The wavelet transform convolution module is used to replace the first two convolution modules in the detection head to reduce the complexity of the model, and then the features are classified and regressed. Eight consecutive star operation modules are applied in the third stage, and five consecutive star operation modules are applied in the fourth stage; the multi-dimensional auxiliary fusion module is specifically processed as follows: the output feature 160×160×64 of the first stage of the backbone network is The convolution operation of the multi-dimensional auxiliary fusion module is increased to 80×80×64. The second stage output feature 80×80×128 is concatenated with the features after the convolution operation in the first stage on the channel dimension after the channel attention mechanism to become 80×80×192. It passes through a 1×1 convolution layer horizontally, and the number of channels is adjusted to 128. It passes through a convolution layer and a 1×1 convolution layer vertically, and the feature becomes 40×40×64. The third stage output feature 40×40×256 is concatenated with the features after the second stage fusion through the channel attention mechanism. The row splicing operation becomes 40×40×320, and it passes through a 1×1 convolution layer horizontally, and the number of channels is adjusted to 256. It passes through a convolution layer and a 1×1 convolution layer vertically, and the features become 20×20×64. The fourth stage outputs the feature 20×20×512 to maintain the original resolution. After the channel attention mechanism, it is spliced with the features fused in the three stages in the channel dimension to generate a comprehensive multi-scale feature. The spliced dimension is 20×20×576, and it passes through a 1×1 convolution layer horizontally to adjust the number of channels to 512.
2. A method for road surface defect detection according to claim 1, characterized in that: The following steps are involved: The S1 model receives the input road defect image and processes it in the star operation module of the backbone network. The deep convolution layer and normalization layer perform preliminary processing on the features, adjust the spatial and channel dimensions of the features through deep convolution and batch normalization, and use the fully connected layer to expand the dimension of the features to capture context information. S2 inputs the features processed in step S1 into the path aggregation feature pyramid network for feature fusion, multi-dimensional auxiliary fusion module processing, and performs weighted processing on the fused features; S3 inputs the features processed by step S2 into the detection head, processes them using the wavelet transform convolution module, performs wavelet transform on the features to extract high-frequency and low-frequency components, performs dimensionality reduction operations on them respectively, performs classification and regression tasks, and obtains the output results of the model.
3. A method for road surface defect detection according to claim 2, characterized in that: The star operation module is processed as follows: CSPDarkNet53 processes the output features of the second stage. The output feature dimension of the second stage is H×W×C, where H, W, and C represent the length, width, and number of channels of the feature, respectively. The H×W×C feature is downsampled and the dimension becomes H / 2×W / 2×C. Then the feature enters the deep convolution layer, that is, each channel has an independent convolution kernel for preliminary feature extraction, and the output dimension remains unchanged, that is, H / 2×W / 2×C. After the first deep convolution layer, batch normalization is performed, and the feature dimension remains unchanged. Then, the output feature x is convolved point by point at the same time to obtain x1 and x2, which have the same dimension, and the number of channels becomes three times the original, which is H / 2×W / 2×3C. ReLU6 activation is applied to x1 Function, introduce nonlinearity, the dimension is still the original H / 2×W / 2×3C, then multiply the activated x1 and x2 element by element, let the activation value of x1 control the feature of x2, and get feature x, then input x into a layer of point-by-point convolution, restore the number of channels of x to C, and the output feature dimension is H / 2×W / 2×C, the output feature continues to input a layer of deep convolution, the dimension remains unchanged, still H / 2×W / 2×C, this layer of deep convolution compresses and refines the features after fusion to further enhance important spatial information, remove redundant features, and maintain low computational complexity, finally use a layer of 1×1 point-by-point convolution to change the number of channels of the output feature, and the dimension changes to H / 2×W / 2×2C, so far, a complete star operation module operation is completed.
4. A method for road surface defect detection according to claim 2, characterized in that: The wavelet transform convolution module processes specifically as follows: the input feature X is first transformed by wavelet transform and decomposed into four components: the low-frequency component X LL (1) , horizontal high frequency component X LH (1) , vertical high frequency component X HL (1) and the diagonal high frequency component X HH (1) , the feature size is reduced to half of the original, the resolution of each component is H / 2×W / 2×C, and the low-frequency component X LL (1) The main structure and global information of the image are retained, and the high-frequency component X LH (1) , X HL (1) , X HH (1) Capture edge and texture detail information, and obtain the low-frequency component X in the first decomposition LL (1) Perform another wavelet decomposition on it to obtain four new components: X LL (2) , X LH (2) , X HL (2) , X HH (2) , the size of the newly decomposed feature is (H / 4)×(W / 4)×C, and the convolution operation is performed on each component feature obtained by the first and second wavelet decompositions, respectively. LL (1) , X LH (1) , X HL (1) , X HH (1) , and X LL (2) , X LH (2) , X HL (2) , X HH (2) Perform convolution in sequence to obtain the convolution result Y LL (1) , Y LH (1) , Y HL (1) , Y HH (1) , and Y LL (2) , Y LH (2) , Y HL (2) , Y HH (2) , the convolutional features are fused through inverse wavelet transform, and the first-level convolution result Y LL (1) , Y LH (1) , Y HL (1) , Y HH (1) , perform inverse wavelet transform and get feature Z (1) , for the secondary convolution result Y LL (2) , Y LH (2) , Y HL (2) , Y HH (2) Perform inverse wavelet transform to obtain feature Z (2) By inverse transforming and fusing the low-frequency and high-frequency components, the generated features contain both the global information of the low-frequency components and the detailed information of the high-frequency components. The feature size after inverse wavelet transform is restored to be consistent with the input. The obtained feature Z (1) and Z (2) The final output feature Z is formed by adding pixels one by one. (0) , the size of the output feature is consistent with the input, that is, H×W×C.
Citation Information
Patent Citations
Asphalt pavement disease detection method based on YOLOv8n model
CN118247636A