Road defect detection method based on heavy parameter multi-scale fusion

By adopting the method of re-parameterization technology and multi-scale feature fusion strategy in road defect detection, a lightweight neural network model is built, which solves the problem of insufficient efficiency and accuracy of road defect detection in the existing technology, and achieves efficient and accurate road defect detection.

CN120070417AActive Publication Date: 2025-05-30CHINA JILIANG UNIV +1

Patent Information

Application Number
CN202510525317.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-30
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The prior art is difficult to achieve efficient and accurate road defect detection in resource-constrained environments. The traditional high-precision object detector model is huge and has low execution efficiency, making it difficult to meet the needs of autonomous driving and real-time reasoning scenarios.

Method used

The road defect detection method based on reparameterization technology and multi-scale feature fusion strategy is adopted. By building a lightweight road defect detection neural network model, reparameterized partial convolution module and multi-scale efficient convolution module are used, combined with the IoU perception query mechanism and efficient upsampling module, to achieve fast and accurate detection of road defects.

Benefits of technology

It significantly improves the speed and accuracy of road defect detection, reduces model complexity and calculation amount, adapts to resource-constrained application scenarios, and ensures detection accuracy and real-timeness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070417A_ABST
    Figure CN120070417A_ABST
Patent Text Reader

Abstract

The invention discloses a road defect detection method based on heavy parameter multi-scale fusion, and belongs to the technical field of target detection in a deep learning neural network, and the method comprises the steps: selecting a data set of road defects from a public data source, and dividing the data set into a training set, a verification set and a test set; a lightweight road defect detection neural network model based on heavy parameter multi-scale fusion is constructed, the model adopts a heavy parameterization technology and a multi-scale feature fusion strategy, and various road defects in the image can be identified and positioned; the constructed neural network model is trained by using the divided training set and verification set, an optimal lightweight road defect detection model is obtained through index evaluation, then reasoning acceleration of the lightweight detection model is carried out, and the lightweight detection model is embedded into an edge computing terminal for deployment; a vehicle-mounted camera is used for collecting image data, collected road images are input into an edge computing terminal to be processed, and types and positions of road defects existing in the vehicle-mounted images are output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of object detection in deep learning neural networks, and specifically relates to a road defect detection method based on reparameterized multi-scale fusion. Background Art

[0002] In the field of computer vision, the development of lightweight object detection algorithms aims to meet the requirements of efficient and accurate object recognition in resource-constrained environments, and is particularly suitable for application scenarios with limited computing resources such as mobile devices and embedded systems. These algorithms optimize the network structure, reduce the number of parameters and floating-point operation counts (FLOPs), and introduce efficient feature extraction methods to significantly reduce the model complexity while ensuring the detection accuracy. Traditional high-precision object detectors are difficult to meet the requirements of real-time inference scenarios such as autonomous driving and security monitoring due to their large model size and low execution efficiency. Summary of the Invention

[0003] To solve the deficiencies of the prior art and achieve the purpose of improving the speed and accuracy of road defect detection, the present invention adopts the following technical solutions:

[0004] A road defect detection method based on reparameterized multi-scale fusion, comprising the following steps:

[0005] Step 1: Obtain road defect data;

[0006] Step 2: Construct a road defect detection neural network model, which adopts reparameterization technology and multi-scale feature fusion strategy to identify and locate various road defects in the image; the construction of the road defect detection neural network model includes the following steps:

[0007] Step 2.1: Sequentially add a group of reparameterizable partial convolution modules to the backbone network for hierarchical feature extraction, optimize the convolutional kernel weights through the dynamic parameter recombination mechanism, and simultaneously construct a hierarchical feature reuse channel for cross-layer feature interaction and information complementation;

[0008] Step 2.2: Perform in-scale feature interaction on the outputs of the first and last reparameterizable partial convolution modules to obtain the corresponding interacted features of the last layer;

[0009] Step 2.3: Construct a multi-path scale feature fusion network to perform multi-scale feature fusion on the corresponding interacted features of the last layer and the outputs of the previous reparameterizable partial convolution modules, and extract a set of features for decoding;

[0010] Step 2.4: Through the IoU-aware query mechanism, a certain number of image features are screened out from the encoded output sequence as the initial object queries for decoding. Subsequently, through the decoding process equipped with an auxiliary prediction head, the object queries are iteratively optimized, and finally the target boundaries and their corresponding categories are generated;

[0011] Step 3: Use the road defect data to divide the training set and the validation set, train the road defect detection neural network model, and use the trained model for road defect detection.

[0012] Furthermore, for the reparameterizable partial convolution module in Step 2.1, feature extraction of a multi-branch structure is performed on some of the input channels. While reducing the number of parameters, it effectively extracts the target features. Reparameterize the RepVGG Block in the detection head during training, convert the 1x1 convolution and the two branches without data processing into 3x3 convolutions, and then fuse them. To simplify the operations in the inference stage, the batch normalization (BN) layer needs to be fused into the corresponding convolution layer. The formula is as follows:

[0013]

[0014] where x represents the input feature, represents the output feature, μ, σ, γ, and β respectively represent the mean, variance, scaling factor, and offset factor of the normalization layer, and ε represents a small constant to avoid division by zero errors;

[0015] When fusing the BN and Conv layers, the new convolution kernel W’ and bias b’ can be calculated by the following formula:

[0016]

[0017]

[0018] where W represents the weight matrix of the original convolution layer, and W’ and b’ respectively represent the fused convolution kernel and bias term;

[0019] Finally, fuse the reparameterizable partial convolution module PRConv with the BasicBlock in the backbone network ResNet to obtain a lightweight backbone network based on reparameterizable partial convolution.

[0020] The reparameterizable partial convolution module combines the advantages of partial convolution (PConv) and reparameterizable convolution (RepConv), aiming to reduce the model parameters while improving the feature extraction ability and detection accuracy. It divides the input channels and only performs feature extraction on a part of them, while the rest remains unchanged.

[0021] Further, in step 2.2, a scale-internal feature interaction module is adopted. After the output of the first layer of the reparameterizable partial convolution module passes through the variability attention module and the first dropout layer, it is connected with the output of the last layer of the reparameterizable partial convolution module in a residual connection and then normalized. The normalized feature map is connected with the first convolution layer, activation layer, second dropout layer, and second convolution layer in a residual connection, and then the output feature map obtained after normalization is used as the feature after interaction.

[0022] Further, in step 2.3, based on four layers of reparameterizable partial convolution modules, the outputs S 2 , S 3 , S 4 and the corresponding feature X 5 after interaction of the last layer are input into the efficient multi-path scale feature fusion network (EMBS-FFN) together. The input feature map first undergoes a series of convolution operations, including multiple 3x3 convolution layers, for extracting the basic features of the image. X 5 passes through a 1×1 convolution and then is fused with S 4 through the multi-scale feature weighted fusion module (BiFusion), and then is input into the multi-scale efficient convolution module (CSP-MSDC) to obtain the F 1 feature map. F 1 passes through the efficient upsampling module (E-Upsample) to obtain , and then is fused with S 3 , S 4 through the multi-scale feature weighted fusion module. The fused feature passes through the multi-scale efficient convolution module to obtain the fused feature map. F 2 passes through the efficient upsampling module (E-Upsample) to obtain , and then is fused with S 2 , S 3 through the multi-scale feature weighted fusion module. After passing through the multi-scale efficient convolution module, it obtains the fused feature map. The feature map is further fused with through the multi-scale feature weighted fusion module to obtain . Finally, passes through the multi-scale efficient convolution module to obtain the first output , which is input into the decoder; is further fused with and through the multi-scale feature weighted fusion module and then passes through the multi-scale efficient convolution module to obtain the second output , it is input into the decoder; Then, it is combined with and F 1 After multi-scale feature weighted fusion, the third output is obtained through a multi-scale efficient convolution module , and it is input into the decoder.

[0023] The efficient multi-path scale feature fusion network combines a multi-scale efficient convolution module and a global heterogeneous kernel selection mechanism, which can intelligently select the most suitable convolution kernel for feature layers of different scales during the intra-scale feature fusion stage, so as to obtain the best multi-scale perception field information. At the same time, a multi-scale feature weighted fusion architecture is adopted, and the Add operation is used instead of the traditional Concat method to reduce the number of parameters and the amount of computation. Without affecting the performance, the model size is further compressed, and at the same time, it can adaptively select weighted fusion according to the importance of features at each scale to improve the quality of feature representation. Finally, an efficient upsampling module is also adopted in the efficient multi-path scale feature fusion network, which can maintain a relatively high operation efficiency while maintaining a certain effect, which is crucial for real-time processing of a large number of road images.

[0024] The multi-scale feature weighted fusion module introduces a two-way information flow, enabling the feature maps of each layer to exchange information from top to bottom and from bottom to top, ensuring that the feature maps of each layer can fully capture information at different scales. This two-way information flow mechanism makes feature fusion more comprehensive and refined.

[0025] Furthermore, in the multi-scale feature weighted fusion module in step 2.3, during the fusion process, the weighting coefficient of the feature map is obtained through neural network training and can be optimized through backpropagation. The weighted feature maps are fused by summation, and the formula is as follows:

[0026]

[0027] Among them, represents the feature map of the i-th layer, represents the weighting coefficient related to the feature map of the i-th layer, which is used to represent the contribution degree of this feature map in the fusion, and O represents the final fused feature map;

[0028] To ensure the effectiveness of the weights, the multi-scale feature weighted fusion normalizes the weights so that the sum of all weights is 1. The formula for the normalized weights is as follows:

[0029]

[0030] Among them, represents the normalized weighting coefficient, Indicates the original weights w i Perform ReLU activation function processing on the weights to ensure that the weights are positive; ε represents a small constant to avoid division by zero errors. w j Represents the weighted coefficient related to the feature map of the j-th layer.

[0031] The multi-scale feature weighted fusion module is not just a simple feature fusion. It adopts a feature fusion process with multiple iterations, enabling the features of each layer to be fully optimized. After each fusion, the feature map will be propagated through a two-way information flow to further optimize the feature representation of each layer. Through this multi-level feature fusion, the output of each layer not only depends on the current features but is also affected by the features of the upper and lower layers, thereby enhancing the expressive ability of the features; the output feature map of each layer Can be represented by the following formula:

[0032]

[0033] Wherein, Represents the i-th feature map of the k-th layer. Represents the weighted coefficient of the feature map of the i-th layer. Represents the output feature map of the k-th layer after weighted fusion.

[0034] The multi-scale feature weighted fusion module enables better fusion of low-level features (with rich detail information) and high-level features (with strong semantic information) through a two-way information flow, thereby providing a more accurate feature representation when detecting small objects; through a learnable weighting mechanism, it can dynamically adjust the weights of features at different scales according to the feedback during the training process, enabling the network to flexibly optimize the feature fusion strategy according to the requirements of the actual task; although the two-way information flow and the weighted fusion mechanism are added, its design still focuses on computational efficiency, avoiding excessive computation and being able to reduce the computational overhead while ensuring performance.

[0035] Furthermore, the multi-scale efficient convolution module in step 2.3 combines multi-scale depth convolution (MSDC) and cross-stage partial connection (CSP) designs. By replacing the Bottleneck block in the traditional C2f module with a multi-scale depth convolution module, it aims to improve the multi-scale feature extraction ability of the network and enhance the perception ability of targets at different scales while maintaining a low computational complexity; in the multi-scale efficient convolution module, first, the input feature map is divided into x 1 and x 2 two parts through cross-stage partial connection. The x 1 part is directly passed to the subsequent layer through a skip connection, while the x 2A part enters the multi-scale depth convolution module for processing. The main advantage of this segmentation design is to reduce the computational load, and rich information in the input feature map is retained through partial connections, avoiding excessive information loss. Next, the multi-scale depth convolution module performs multi-scale feature extraction on x 2 ; After being processed by the multi-scale depth convolution module, the number of channels of the feature map is restored to the original dimension through a new 1×1 pointwise convolution (PWC 2 ), followed by a batch normalization (BN) layer and a ReLU6 activation function, which helps the network maintain a stable training process and introduces non-linear features; In this way, the features processed by the multi-scale depth convolution module can not only contain information of more scales, but also enhance the interaction between channels, improving the network's ability to express complex features; Finally, in the fusion stage, the features of x 1 and x 2 are fused through concatenation, and the fused features will be processed by a further convolutional layer to output the final feature map. Through this structural design, the multi-scale efficient convolution module can effectively combine low computational cost and powerful feature representation ability to meet the detection requirements of targets of different scales.

[0036] Furthermore, inside the multi-scale depth convolution module, first, a 1×1 pointwise convolution (PWC 1 ) is used to expand the channels of x 2 , increasing the number of channels to enhance the network's representation ability; Then, multi-scale depth convolutions (DWConv) are used to perform convolution operations on different convolutional kernel sizes ks (such as 3×3, 5×5, 7×7, etc.), so as to extract features from different scales; The design of multi-scale convolutions makes the network more sensitive when processing targets of different sizes, thus enhancing the detection ability for small and large targets; To further improve the feature expression ability, a channel rearrangement operation is introduced into the multi-scale depth convolution module. Channel rearrangement can effectively break the independence between channels, enhance the information interaction between channels, and thus improve the diversity and representation ability of the feature map. The formula is as follows:

[0037]

[0038]

[0039]

[0040] Among them, PWC 1 represents the 1×1 pointwise convolution operation, BN represents the batch normalization operation, and R6 represents the ReLU6 activation operation. Denotes a depth convolution block, which uses depth convolution operations for each convolution kernel size ks and combines batch normalization (BN) and ReLU6 activation. CS represents the channel shuffle operation.

[0041] Furthermore, for the efficient upsampling module in step 2.3, its design goal is to improve the spatial resolution of the feature map while maintaining computational efficiency. First, the size of the input feature map is scaled up by a factor of 2 through an upsampling operation. Then, depthwise separable convolution (DWConv) is applied. Different from traditional convolution methods, depthwise separable convolution performs convolution operations independently on each input channel, thus significantly reducing the computational amount and the number of parameters, and thereby improving computational efficiency. Next, the convolution result is processed through a batch normalization (BN) layer to stabilize the training process. Subsequently, the ReLU activation function is applied to introduce non-linearity. Finally, a 1×1 convolution is used to adjust the number of channels of the output feature map to match the input requirements of the next decoding stage.

[0042] Furthermore, in step 2.4, the traditional method selects the top K high-scoring samples from the encoded layer features based on classification confidence to generate initial query samples. However, due to the evaluation bias between classification confidence and localization accuracy, the spatial overlap between some prediction boxes with high classification scores and the ground truth boxes is low, resulting in the screening mechanism being prone to missing high-quality candidates with low classification scores but high IoU. Therefore, the present invention adopts an IoU-aware candidate screening mechanism. By imposing optimization constraints during the model training stage, the model is forced to learn to map feature representations with high spatial overlap IoU to high classification scores, while suppressing the classification confidence of features with low spatial overlap IoU. After this joint optimization, the system finally extracts the optimal K candidate samples from the encoded layer features according to the optimized classification scores. The optimization objectives are as follows:

[0043]

[0044]

[0045] Among them, and represent the predicted value and the ground truth value respectively, and , where c and b represent the class confidence and the target bounding box respectively, L box represents the target bounding box prediction loss function, L cls represents the class prediction loss function.

[0046] Based on the above-mentioned road defect detection method based on reparameterized multi-scale fusion, in the model training of step 3, through index evaluation, the best lightweight road defect detection neural network model is obtained. Subsequently, the inference acceleration of the lightweight model is carried out and embedded into the edge computing terminal for deployment; it also includes step 4, using an in-vehicle camera to collect road image data and inputting it into the deployed edge computing terminal to obtain the types and positions of road defects existing in the in-vehicle images.

[0047] The advantages and beneficial effects of the present invention are as follows:

[0048] The present invention optimizes road defect detection by integrating the reparameterized multi-scale technology. Among them, by introducing partial convolution and reparameterizable convolution (PRConv), the multi-scale feature capture ability is enhanced; by constructing an efficient multi-path scale feature fusion network (EMBS-FFN), combining the CSP-MSDC module and the global heterogeneous kernel selection mechanism, the in-scale feature adaptive fusion is realized, and the Add operation is used to reduce the complexity; the detection effect and computational efficiency are balanced through the efficient upsampling module (E-Upsample). While inheriting the advantages of lightweight design, the present invention significantly improves the accuracy and efficiency of road defect detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is the flowchart of the method of the embodiment of the present invention.

[0050] Figure 2 is the schematic diagram of the network structure of the road defect detection model in the embodiment of the present invention.

[0051] Figure 3 is the schematic diagram of the structure of the reparameterizable partial convolution module in the embodiment of the present invention.

[0052] Figure 4 is the schematic diagram of the structure of the multi-scale efficient convolution module in the embodiment of the present invention.

[0053] Figure 5 is the schematic diagram of the structure of the efficient upsampling module in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The following further describes in detail the specific embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not intended to limit the present invention.

[0055] As Figure 1 shown, a road defect detection method based on reparameterized multi-scale fusion specifically includes the following steps:

[0056] Step 1: Select a dataset containing typical road defects such as longitudinal cracks, transverse cracks, reticular cracks, and potholes from public data sources, and divide it into a training set, a validation set, and a test set;

[0057] Step 2: Construct a lightweight road defect detection neural network model based on reparameterized multi-scale fusion. This model adopts reparameterization technology and multi-scale feature fusion strategy, and can identify and locate various road defects in images;

[0058] Among them, the specific structure of the lightweight road defect detection neural network model based on reparameterized multi-scale fusion is implemented as follows: Figure 2 as shown below:

[0059] Step 2.1: The lightweight road defect detection neural network model based on reparameterized multi-scale fusion is based on the ResNet18 network architecture. In the backbone network, a reparameterizable partial convolution module (PRConv) is innovatively proposed: adopting a phased feature extraction strategy, the convolution kernel weights are optimized through a dynamic parameter recombination mechanism, and at the same time, a hierarchical feature reuse channel is constructed to achieve cross-layer feature interaction and information complementation. Reparameterizable partial convolution modules are added to the last four layers of the backbone network. Finally, after the image is input into the backbone network, four feature maps of different scales will be output, which are respectively marked as S 2 ,S 3 、S 4 、S 5 ,where S 2 comes from the fifth layer of the backbone network, S 3 comes from the 6th layer of the backbone network, S 4 comes from the seventh layer of the backbone network, S 5 comes from the eighth layer of the backbone network.

[0060] The specific structure of the reparameterizable partial convolution module is as shown below: Figure 3 The input of the module is divided into a backbone and a branch. The input feature passes through a convolution module and a partial convolution PRConv module in sequence on the backbone. The input feature passes through an identity mapping module Identity on the branch and is concatenated with the output of the partial convolution PRConv module, and then outputs through a Relu activation module. The implementation process is as follows:

[0061] The reparameterizable partial convolution module combines the advantages of partial convolution (PConv) and reparameterizable convolution (RepConv), aiming to reduce the model parameter quantity while improving the feature extraction ability and detection accuracy. It divides the input channels and only extracts features from a part of them, and the rest remains unchanged. Specifically, the original convolution operation calculates using different convolution kernels for each input channel, and its parameters are A 2×C, while the PRConv module only takes a part of it (usually 1 / 4) for calculation, and its number of parameters is reduced to A 2 ×C / 4. At the same time, to alleviate the problem of accuracy decline caused by the reduction of the number of parameters, it effectively extracts target features through a multi-branch structure, which plays a role in improving the model accuracy in the PRConv module. Specifically, PRConv reparameterizes the RepVGG Block used in the detection head during training, converts the 1x1 convolution and two branches without data processing into 3x3 convolutions, and then fuses them. To simplify the operations in the inference stage, the batch normalization (BN) layer needs to be fused into the corresponding convolutional layer. Given the parameters μ (mean), σ (variance), γ (scaling factor), and β (offset factor) of a BN layer, and the input , its formula can be expressed as:

[0062]

[0063] where ε is a very small constant to avoid division by zero errors.

[0064] When fusing the BN and Conv layers, the new convolutional kernel W’ and bias b’ can be calculated by the following formula:

[0065]

[0066]

[0067] where W represents the weight matrix of the original convolutional layer, and W’ and b’ represent the fused convolutional kernel and bias term respectively.

[0068] Finally, PRConv is fused with the BasicBlock in the backbone network ResNet to obtain a lightweight backbone network based on reparameterized partial convolution.

[0069] Step 2.2: Input the feature map S extracted in Step 2.1 5 into the intra-scale feature interaction module AIFI (Attention-based Intra-scale Feature Interaction). In the intra-scale feature interaction module, the feature map S 2 passes through the deformable attention module, dropout, and then is connected with S 5 through a residual connection and then normalized. The feature map after the first normalization process is connected with the convolutional layer fc1, the activation function layer GELU, the dropout layer, and the convolutional layer fc2 through a residual connection and then passes through the normalization layer to finally output the feature map X 5 .

[0070] Step 2.3: S 2, S 3 , S 4 , X 5 are input into the Efficient Multi-Path Scale Feature Fusion Network (EMBS-FFN) together. The input feature maps first go through a series of convolution operations, including multiple 3×3 convolution layers, which are used to extract the basic features of the image. X 5 After passing through a 1×1 convolution, it is combined with S 4 and undergoes feature fusion through the Multi-Scale Feature Weighted Fusion Module (BiFusion), and then is input into the Multi-Scale Efficient Convolution Module (CSP-MSDC) to obtain F 1 feature map. F 1 After passing through the Efficient Upsampling Module (E-Upsample), it gets , Then, it undergoes feature fusion with S 3 , S 4 again through the Multi-Scale Feature Weighted Fusion Module. The fused features pass through the Multi-Scale Efficient Convolution Module to obtain F 2 fused feature map. F 2 After passing through the Efficient Upsampling Module (E-Upsample), it gets , Then, it undergoes feature fusion with S 2 , S 3 again through the Multi-Scale Feature Weighted Fusion Module. After passing through the Multi-Scale Efficient Convolution Module, it obtains F 3 fused feature map. F 3 The feature map is then combined with through multi-scale feature weighted fusion to obtain . Finally it passes through the Multi-Scale Efficient Convolution Module to obtain the first output P 1 , which is input into the decoder; P 1 Then, it is combined with , F 2 and through multi-scale feature weighted fusion and then passes through the Multi-Scale Efficient Convolution Module to obtain the second output P 2 , which is input into the decoder; P 2 Then, it is combined with F 2 and F 1 through multi-scale feature weighted fusion and then passes through the Multi-Scale Efficient Convolution Module to obtain the third output P 3 , which is input into the decoder.

[0071] The structural diagram of the Efficient Multi-Path Scale Feature Fusion Network is implemented as shown Figure 2 .

[0072] The efficient multi-path scale feature fusion network combines a multi-scale efficient convolution module and a global heterogeneous kernel selection mechanism, which can intelligently select the most suitable convolution kernel for feature layers of different scales during the in-scale feature fusion stage, so as to obtain the best multi-scale perception field information. At the same time, a multi-scale feature weighted fusion architecture is adopted, and the Add operation is used instead of the traditional Concat method to reduce the number of parameters and the amount of computation. Without affecting the performance, the model size is further compressed, and at the same time, the weighted fusion can be adaptively selected according to the importance of each scale feature, improving the quality of feature representation. Finally, an efficient upsampling module is also adopted in the efficient multi-path scale feature fusion network, which can maintain a relatively high operation efficiency while maintaining a certain effect, which is crucial for real-time processing of a large number of road images.

[0073] The multi-scale feature weighted fusion module enables information exchange of each layer of feature maps from top to bottom and from bottom to top by introducing a two-way information flow, ensuring that each layer of feature maps can fully capture information of different scales. This two-way information flow mechanism makes feature fusion more comprehensive and refined.

[0074] At the same time, a weighted fusion strategy is introduced, so that when each layer of feature maps is fused, different weights are assigned according to their importance. The idea of weighted fusion is to dynamically adjust the contributions of feature maps of different scales through learnable weights, thereby improving the fusion effect. During the fusion process, the weighted coefficient of the feature map is obtained through neural network training and can be optimized through backpropagation. The weighted feature maps are fused by summation, and the formula is as follows:

[0075]

[0076] where, represents the feature map of the i-th layer, represents the weighted coefficient related to the feature map of the i-th layer, indicating the contribution degree of this feature map in the fusion, and O represents the final fused feature map.

[0077] To ensure the effectiveness of the weights, the multi-scale feature weighted fusion normalizes the weights so that the sum of all weights is 1. The formula for the normalized weights is as follows:

[0078]

[0079] where, represents the normalized weighted coefficient, represents applying the ReLU activation function to the original weight w i to ensure that the weight is positive; ε is a small constant to avoid division by zero errors, w jRepresents the weighting coefficient related to the feature map of the j-th layer.

[0080] Multi-scale feature weighted fusion is not just a simple feature fusion. It adopts a feature fusion process with multiple iterations, enabling the features of each layer to be fully optimized. After each fusion, the feature map will be propagated through a two-way information flow to further optimize the feature representation of each layer. Through this multi-level feature fusion, the output of each layer not only depends on the current features but is also affected by the features of the upper and lower layers, thereby enhancing the expressive power of the features.

[0081] The output feature map of each layer Can be represented by the following formula:

[0082]

[0083] Where, Represents the i-th feature map of the k-th layer, Represents the weighting coefficient of the feature map of the i-th layer, Represents the output feature map of the k-th layer after weighted fusion.

[0084] The multi-scale feature weighted fusion module enables better fusion of low-level features (with rich detailed information) and high-level features (with strong semantic information) through a two-way information flow, thereby providing a more accurate feature representation when detecting small objects; through a learnable weighting mechanism, it can dynamically adjust the weights of features at different scales according to the feedback during the training process, enabling the network to flexibly optimize the feature fusion strategy according to the actual task requirements; although the two-way information flow and weighted fusion mechanism are added, its design still focuses on computational efficiency, avoiding excessive computation and being able to reduce the computational overhead while ensuring performance.

[0085] Among them, the multi-scale efficient convolution module is implemented as Figure 4 Shown as follows:

[0086] The multi-scale efficient convolution module is an innovative combination of multi-scale depth convolution (MSDC) and cross-stage partial connection (CSP) designs based on the C2f module in YOLO. It aims to improve the multi-scale feature extraction ability of the network and enhance the perception ability of targets at different scales by replacing the Bottleneck block in the traditional C2f module with a multi-scale depth convolution module, while maintaining a low computational complexity.

[0087] In the multi-scale efficient convolution module, first, the input feature map is divided into two parts through cross-stage partial connection. Specifically, the input feature map x is split into x 1 And x 2 Two parts, x 1Part is directly passed to the subsequent layer through skip connections, while x 2 Part enters the multi-scale depth convolution module for processing. The main advantage of this segmentation design is to reduce the computational load, and rich information in the input feature map is retained through partial connections, avoiding excessive information loss.

[0088] Next, the multi-scale depth convolution module performs multi-scale feature extraction on x 2 Inside the multi-scale depth convolution module, first, a 1×1 pointwise convolution (PWC 1 ) is used to expand the channels of x 2 to increase the number of channels and enhance the representation ability of the network. Then, multi-scale depth convolutions (DWConv) are used to perform convolution operations on different convolutional kernel sizes ks (such as 3×3, 5×5, 7×7, etc.) to extract features from different scales. The design of multi-scale convolutions makes the network more sensitive when processing targets of different sizes, thus enhancing the detection ability for small and large targets. To further improve the feature expression ability, a channel shuffle operation is introduced into the multi-scale depth convolution module. Channel shuffle can effectively break the independence between channels, enhance the information interaction between channels, and thus improve the diversity and representation ability of the feature map. Its formula can be expressed as:

[0089]

[0090]

[0091]

[0092] where PWC 1 represents the 1×1 pointwise convolution operation, BN represents the batch normalization operation, R6 represents the activation operation, represents the depth convolution block, which uses depth convolution operations for each convolutional kernel size ks and combines batch normalization (BN) and ReLU6 activation, and CS represents the channel shuffle operation.

[0093] After being processed by the multi-scale depth convolution module, the number of channels of the feature map is restored to the original dimension, which is achieved through a new 1×1 pointwise convolution (PWC 2 ), followed by a batch normalization (BN) layer and a ReLU6 activation function, which helps the network maintain a stable training process and introduce non-linear features. In this way, the features processed by the multi-scale depth convolution module can not only contain more scale information but also enhance the interaction between channels, improving the network's ability to express complex features.

[0094] At the end of the multi-scale efficient convolution module, in the fusion stage, x 1 and x 2The features of the two parts are merged, and this fusion operation is achieved through splicing. The fused features are further processed by the convolution layer to output the final feature map. Through this structural design, the multi-scale efficient convolution module can effectively combine low computational cost and powerful feature representation capabilities to meet the detection needs of targets of different scales.

[0095] The efficient upsampling module is as follows Figure 5 As shown in the implementation:

[0096] The efficient upsampling module is designed to efficiently upsample feature maps. The design goal of this module is to increase the spatial resolution of feature maps while maintaining computational efficiency. Its structure is divided into several key steps. First, the efficient upsampling module scales up the size of the input feature map by a factor of 2 through an upsampling operation. After the upsampling operation, a depthwise separable convolution (DWConv) is applied. Unlike traditional convolution methods, depthwise separable convolution performs convolution operations independently on each input channel, so it can significantly reduce the amount of computation and the number of parameters, thereby improving computational efficiency. Next, the convolution result is processed through a batch normalization (BN) layer to stabilize the training process, and then the ReLU activation function is applied to introduce nonlinear characteristics. Finally, a 1×1 convolution is performed to adjust the number of channels of the output feature map to match the input requirements of the next decoding stage.

[0097] Step 2.4: Through the IoU-aware query mechanism, a certain number of image features are selected from the encoder's output sequence as the decoder's initial object query. Subsequently, the decoder equipped with an auxiliary prediction head iteratively optimizes these object queries and finally generates the target bounding box and its corresponding confidence score.

[0098] The optimization principle of the IoU-aware candidate screening mechanism is as follows: the traditional method screens the top K high-scoring samples from the coding layer features based on classification confidence to generate the initial query sample, but due to the evaluation deviation between classification confidence and positioning accuracy, some prediction boxes with high classification scores have low spatial overlap with the annotation boxes, which makes it easy for the screening mechanism to miss high-quality candidates with low classification scores but high IoU. To this end, the IoU-aware strategy imposes optimization constraints in the model training phase, forcing the model to learn to map feature representations with high spatial overlap (IoU) to high classification scores, while suppressing the classification confidence of low IoU features. After this joint optimization, the system finally extracts the best K candidate samples from the coding layer features based on the optimized classification scores. The optimization objectives of the detector are as follows:

[0099]

[0100]

[0101] in, and represent the predicted value and the true value respectively, and , where c and b represent the category and the bounding box respectively, L box represents the target bounding box prediction loss function, L cls represents the category prediction loss function.

[0102] Step 3: Use the training set and the validation set divided in Step 1 to train the constructed neural network model. After metric evaluation, obtain the best lightweight road defect detection model. Subsequently, perform inference acceleration on the lightweight detection model and embed it for deployment on the edge computing terminal;

[0103] To evaluate the performance of the object detection model, the metric mean average precision (mAP) is used. It comprehensively considers the accuracy and coverage rate of the model in different categories and averages the detection effects of various categories. Among them, accuracy (Precision) refers to the proportion of samples that are actually positive among all samples predicted as positive by the model; while coverage rate (Recall) measures the proportion of actual positive samples that the model can identify, that is, the ratio of the number of samples successfully predicted as positive to the total number of all actual positive samples. mAP50 specifically refers to the standard of using an IoU (Intersection over Union) threshold of 50% when calculating mAP to determine whether the predicted bounding box is correct. Specifically, only when the overlap degree (IoU) between the predicted bounding box and the true annotation bounding box reaches or exceeds 50%, the prediction is considered valid. This setting ensures that the evaluation not only considers the classification accuracy of the model but also examines its localization precision. The calculation formulas for recall rate (R) and precision rate (P) are as follows:

[0104]

[0105]

[0106] where R represents the recall rate, P represents the precision rate, F represents the number of object detection errors, T represents the number of correctly detected objects, and G represents the total number of objects in the test set.

[0107] A comprehensive evaluation index system is also adopted to evaluate the lightweight performance of the model, including three key indicators: the number of parameters (Params), the amount of computation (GFLOPs), and the number of frames detected per second (FPS). Params refers to the total number of parameters of the model, reflecting the spatial complexity of the model. The larger the number of parameters, usually the more storage space and memory resources the model occupies; GFLOPs represents the floating-point operation amount of the model, which is an important indicator to measure the computing performance of the model. A lower GFLOPs value means that the model requires less computing resources during operation, which helps to improve the energy efficiency ratio; FPS (Frames Per Second) represents the detection speed of the model, that is, the number of image frames that can be processed per second. A higher FPS indicates that the model can complete the detection task faster, which is particularly important for application scenarios with high real-time requirements.

[0108] This evaluation system not only focuses on the lightweight characteristics of the model (measured by Params and GFLOPs) to ensure its efficient operation in resource-constrained environments, but also attaches importance to the speed of the model (measured by FPS) and accuracy (measured by mAP) to guarantee the performance in practical applications.

[0109] To illustrate the effectiveness of this method, the complete network was trained on the classic road defect data GRDDC2020, and the indicators in four aspects of the number of parameters (Params), the amount of computation (GFLOPs), the number of frames detected per second (FPS), and the mean average precision (mAP) were compared with 8 real-time object detection methods including YOLOv5-s, YOLOv6-s, YOLOv8-s, DINO-Deformable-DETR, and RT-DETR through the test set. The experimental environment uses the Linux operating system and is implemented in the Pytorch deep learning development framework. The development language is Python. The training strategy and hyperparameters of the decoder basically follow the DINO setting. 4 NVIDIA GeForce GTX 1080 Ti graphics cards with 11 GB video memory are used for training and testing. The batch size is set to 40 for the GRDDC2020 dataset, and 250 epochs are iterated; the AdamW optimizer is adopted, the base learning rate is 0.0001, and the weight decay coefficient is also 0.0001. The global gradient clipping norm of 0.1 and the warm start of the linear learning rate in the first 2,000 steps are applied; the learning rate adjustment strategy of the backbone network refers to the experiments in the previous chapter, and the exponential moving average (EMA) technology is used with a decay rate of 0.9999 to improve the model performance; during the training process, data augmentation operations such as color distortion, dilation, cropping, flipping, and resizing are randomly applied to increase the robustness and generalization ability of the model. The comparison results on the GRDDC2020 dataset are shown in Table 1:

[0110] Table 1 Comparative experiment table of the present invention and other 8 methods on the classic road defect data GRDDC2020

[0111]

[0112] Step 4: Use an in-vehicle camera to collect image data, and input the collected road images into the edge computing terminal in Step 3 for processing, and output the road defect categories and positions existing in the in-vehicle images.

[0113] Specifically, input the road images captured by the in-vehicle camera into the road defect detection model with the highest comprehensive evaluation index in the test set. After the operation and processing of this model, finally output the road defect categories existing in the in-vehicle images and the positions where these defects are located.

[0114] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A road defect detection method based on heavy parameter multi-scale fusion, characterized in that The steps include: Step 1: Obtain road defect data; Step 2: Construct a neural network model for road defect detection. The model uses reparameterization technology and multi-scale feature fusion strategy to identify and locate various road defects in the image. The construction of the neural network model for road defect detection includes the following steps: Step 2.1: Add a set of reparameterizable partial convolution modules to the backbone network in turn to extract hierarchical features, optimize the convolution kernel weights through a dynamic parameter reorganization mechanism, and construct a hierarchical feature reuse channel to perform cross-layer feature interaction and information complementation; Step 2.2: Perform intra-scale feature interaction on the outputs of the first and last layers of reparameterizable partial convolution modules to obtain the corresponding interactive features of the last layer; Step 2.3: Construct a multi-path scale feature fusion network, fuse the interactive features corresponding to the last layer with the output of the previous reparameterizable partial convolution module, and extract a set of features for decoding; Step 2.4: Through the perceptual query mechanism, features are selected from the encoding output as the initial object query for decoding. Subsequently, the object query is iteratively optimized through the decoding process equipped with the auxiliary prediction head, and finally the object boundary and its corresponding category are generated; Step 3: Using the road defect data, train the road defect detection neural network model, and use the trained model for road defect detection.

2. The road defect detection method based on heavy parameter multi-scale fusion according to claim 1 is characterized by: The reparameterizable partial convolution module in step 2.1 performs feature extraction of a multi-branch structure for some input channels, reparameterizes the reparameterization block used in the detection head during training, and fuses the batch normalization layer into the corresponding convolution layer. The formula is as follows: , Among them, x represents the input feature, represents the output features, μ, σ, γ, β represent the mean, variance, scaling factor and offset factor of the normalization layer respectively, and ε represents a small constant; When the BN and Conv layers are fused, the new convolution kernel W' and bias b' can be calculated by the following formula: , , Among them, W represents the weight matrix of the original convolution layer, W' and b' represent the fused convolution kernel and bias term respectively; Finally, the reparameterizable partial convolution module is integrated with the basic blocks in the backbone network to obtain a lightweight backbone network based on reparameterized partial convolution.

3. The road defect detection method based on heavy parameter multi-scale fusion according to claim 1 is characterized by: In the step 2.2, an intra-scale feature interaction module is used to connect the output of the first layer of reparameterizable partial convolutional module to the residual of the output of the last layer of reparameterizable partial convolutional module after passing through the variable attention module and the first random drop layer, and then normalize the output feature map after the normalization processing is residually connected with the first convolutional layer, the activation layer, the second random drop layer, and the second convolutional layer, and then the output feature map obtained by normalization is used as the feature after interaction.

4. The road defect detection method based on heavy parameter multi-scale fusion according to claim 1 is characterized in that: In the step 2.3, based on the four-layer reparameterizable partial convolution module, the outputs S2, S3, S4 of the first three layers and the interactive features X5 corresponding to the last layer are input into the multi-path scale feature fusion network. The input feature map is first subjected to a series of convolution operations to extract the basic features of the image. X5 and S4 are subjected to feature fusion by the multi-scale feature weighted fusion module and then input into the multi-scale convolution module to obtain the F1 feature map. F1 is obtained after passing through the upsampling module. , Then, the multi-scale feature weighted fusion module is used to fuse the features with S3 and S4. The fused features are obtained after passing through the multi-scale efficient convolution module. The fused feature map, F2, is obtained after passing through the upsampling module , Then, the multi-scale feature weighted fusion module is used to fuse the features with S2 and S3, and the multi-scale convolution module is used to obtain Fusion feature map, The feature map is then Perform multi-scale feature weighted fusion to obtain ,at last After the multi-scale convolution module, the first output is obtained , input to the decoder; Again with , and After multi-scale feature weighted fusion, the second output is obtained through the multi-scale convolution module. , input to the decoder; Again with After multi-scale feature weighted fusion with F1, the third output is obtained through the multi-scale convolution module. , input to the decoder.

5. The road defect detection method based on heavy parameter multi-scale fusion according to claim 1 is characterized in that: In the multi-scale feature weighted fusion module in step 2.3, during the fusion process, the weighted coefficient of the feature map It is obtained through neural network training and can be optimized through back propagation. The weighted feature maps are fused by summing. The formula is as follows: , in, represents the feature map of the i-th layer, represents the weight coefficient associated with the feature map of the i-th layer, and O represents the final fused feature map; Multi-scale feature weighted fusion will normalize the weights, the formula is as follows: , in, represents the normalized weighting coefficient, Represents the original weight w i Perform ReLU activation function processing; ε represents a small constant, w j Represents the weighted coefficient related to the j-th layer feature map; The multi-scale feature weighted fusion module adopts a multi-iteration feature fusion process. After each fusion, the feature map will be propagated through a bidirectional information flow. The output feature map of each layer It can be expressed by the following formula: , in, represents the i-th feature map of the k-th layer, represents the weight coefficient of the feature map of the i-th layer, Represents the k-th layer output feature map after weighted fusion.

6. The road defect detection method based on heavy parameter multi-scale fusion according to claim 1 is characterized by: The multi-scale efficient convolution module in step 2.3 combines multi-scale deep convolution and cross-stage partial connection. First, the input feature map is divided into two parts, x1 and x2, through the cross-stage partial connection. The x1 part is directly passed to the subsequent layer through the jump connection, and the x2 part enters the multi-scale deep convolution module for processing; next, the multi-scale deep convolution module performs multi-scale feature extraction on x2; after processing by the multi-scale deep convolution module, the number of channels of the feature map will be restored to the original dimension through point convolution, followed by a batch normalization layer and an activation function; finally, in the fusion stage, the features of the two parts x1 and x2 are fused by splicing, and the fused features are processed by further convolution layers to output the final feature map.

7. The road defect detection method based on heavy parameter multi-scale fusion according to claim 1 is characterized by: Inside the multi-scale deep convolution module, firstly, the channel of x2 is expanded by point convolution to increase the number of channels; then, the multi-scale deep convolution is used to perform convolution operations on different convolution kernel sizes ks; the design of multi-scale convolution enables the network to be more sensitive when processing targets of different sizes; the channel rearrangement operation is introduced into the multi-scale deep convolution module, and its formula is as follows: , , , Among them, PWC1 represents a 1×1 point convolution operation, BN represents a batch normalization operation, and R6 represents an activation operation. represents a depthwise convolution block, which uses a depthwise convolution operation for each kernel size ks, combined with batch normalization and activation, and CS represents a channel shuffling operation.

8. The road defect detection method based on heavy parameter multi-scale fusion according to claim 1 is characterized by: The upsampling module in step 2.3 firstly scales up the size of the input feature map by upsampling operation; then, applies depthwise separable convolution to perform convolution operation independently on each input channel; next, the convolution result is processed by batch normalization layer; then, an activation function is applied to introduce nonlinear characteristics; finally, the number of channels of the output feature map is adjusted by convolution to match the input requirements of the next decoding stage.

9. The road defect detection method based on heavy parameter multi-scale fusion according to claim 1 is characterized by: In step 2.4, an overlap-aware candidate screening mechanism is adopted. By imposing optimization constraints in the model training stage, the model is forced to learn to map feature representations with high spatial overlap into high classification scores, while suppressing the classification confidence of low spatial overlap features. The optimal K candidate samples are extracted from the coding layer features based on the optimized classification scores. The optimization objectives are as follows: , , in, and denote the predicted value and the true value respectively. and , c and b represent the category confidence and target bounding box respectively, L box represents the target bounding box prediction loss function, L cls Represents the category prediction loss function.

10. The road defect detection method based on heavy parameter multi-scale fusion according to claim 1, characterized in that: In the model training of step 3, the optimal lightweight road defect detection neural network model is obtained through indicator evaluation. Subsequently, the reasoning of the lightweight model is accelerated and embedded in the edge computing terminal for deployment. It also includes step 4, using the on-board camera to collect road image data and input it into the deployed edge computing terminal to obtain the road defect category and location in the on-board image.

Citation Information

Patent Citations

  • Plate defect detection method based on 2D and 3D images and related equipment

    CN118134927A

  • Road defect detection method combining YOLOv8 and RTDETR

    CN118196103A

  • Lightweight road defect detection method based on dynamic deformable attention mechanism

    CN118279562A

  • Deep learning system for cuboid detection

    US20200202554A1

Cited By

  • VHPC-DETR-based violent target detection method

    CN120411736A

  • Multi-scale road vehicle detection method based on re-parameterized visual converter

    CN120564164A

  • Unmanned aerial vehicle lightweight real-time small target recognition device and method based on SOD-DETR

    CN120747791A

  • Lightweight target detection model capable of adapting to multi-scene remote sensing

    CN120931912A

  • A lightweight target detection model adaptable to various remote sensing scenarios

    CN120931912B