Automatic identification method, storage medium and device for hidden road defects integrating multiple attention mechanisms

By adopting the YOLOv8 model with multiple attention mechanisms in ground penetrating radar image recognition, the problem of artificial reliance on recognition and low recognition accuracy in ground penetrating radar image recognition is solved, and more efficient feature extraction and disease detection are achieved.

CN118941841BActive Publication Date: 2025-05-06HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410940285.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-15
Publication Date
2025-05-06
Estimated Expiration
2044-07-15

AI Technical Summary

Technical Problem

Ground penetrating radar image identification is extremely dependent on artificiality, and the accuracy and robustness of the road hidden disease recognition model in complex scenarios are poor.

Method used

The YOLOv8 model that uses a fusion of multiple attention mechanisms is used to replace the standard convolutional layer into a depth-separable convolutional layer in the backbone network, and add a multi-attention mechanism module to the neck network to perform feature fusion and response fusion.

Benefits of technology

It improves the accuracy and robustness of ground penetrating radar image recognition, can better extract tiny target feature information in hidden roads, and enhances detection accuracy and real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118941841B_ABST
    Figure CN118941841B_ABST
Patent Text Reader

Abstract

This invention relates to an automatic road hidden defect identification method, storage medium, and device that integrates a multi-attention mechanism, belonging to the field of ground-penetrating radar (GPR) road non-destructive testing technology. It addresses the problem of GPR image identification being highly dependent on manual intervention, and the poor accuracy and robustness of road hidden defect identification models in complex scenarios. The invention preprocesses the raw images acquired by GPR and feeds them into an improved YOLOv8-based network, specifically a YOLOv8 model integrating a multi-attention mechanism, for identification. The standard convolutional layers in the backbone network are replaced with depth-separable convolutional layers. During feature fusion in the neck network, a multi-attention mechanism (MSE) module, formed by parallel connections of a hybrid attention mechanism module and an SE attention module, is inserted after each upsampled C2f module in the neck network. Road hidden defects are then identified based on the YOLOv8 model integrating the multi-attention mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of non-destructive detection of roads by ground penetrating radar, and in particular relates to a method, storage medium and equipment for automatically identifying hidden road defects. Background Art

[0002] During daily use, roads will be exposed to various hidden diseases due to temperature and humidity changes, seasonal freezing and thawing, traffic loads and other factors. Hidden diseases are invisible on the road surface and are highly hidden, making them difficult to detect, posing a huge threat to road safety. Ground penetrating radar technology can detect hidden road diseases of varying sizes and types, effectively helping road maintenance personnel to detect road damage as early as possible and repair it in a timely manner, thereby extending the service life of the road and saving maintenance costs.

[0003] Ground penetrating radar technology uses the penetrating characteristics of high-frequency electromagnetic waves to perform non-destructive testing on the internal structure of roads. When electromagnetic waves encounter underground targets or underground media with different electromagnetic characteristics during propagation, they are reflected and refracted. The transmitting antenna radiates high-frequency electromagnetic waves into the ground, and the receiving antenna receives the reflected echo signals of the electromagnetic characteristics (dielectric constant, conductivity, etc.) of various road materials in different states. Through the differences in the amplitude, phase and other characteristics of the reflected signals of different types of defects, their physical properties, geometric shapes and geographical locations can be inferred, thereby further determining the type and scale of the defect.

[0004] The amplitude and phase characteristics of different types of hidden defects are relatively small, but their sizes are very different, making them difficult to analyze and evaluate. At present, the identification of hidden defects by ground-penetrating radar mainly relies on the analysis and judgment of experienced technicians, or the identification of hidden defects of a single detection scene and type through the current mainstream artificial intelligence methods. This will result in a low degree of digitization of ground-penetrating radar technology, time-consuming and labor-intensive, greatly reducing the efficiency of ground-penetrating radar recognition, resulting in missed detection and false detection, and affecting the detection effect of ground-penetrating radar on hidden road defects.

[0005] The network structure of the artificial intelligence model contains many parameters that need to be learned, and updating the parameter values ​​of each layer of the network requires a large amount of correctly labeled sample data. The model can only have strong capabilities after the network is fully trained. Otherwise, problems such as underfitting and overfitting are prone to occur during the model training process, which is not conducive to the improvement of model capabilities. Since it is difficult to determine the category of damage by analyzing the features presented by ground penetrating radar images, and road radar image data is seriously lacking in labeled samples, the current artificial intelligence method cannot take into account the detection accuracy of targets of different sizes and types when identifying hidden damages of ground penetrating radar, which restricts the development of automatic detection of hidden damages of ground penetrating radar. Summary of the invention

[0006] The purpose of the present invention is to solve the problem that ground penetrating radar image identification is extremely dependent on manual work, and the problem that the accuracy and robustness of the road hidden disease recognition model in complex scenes are poor.

[0007] An automatic road hidden disease recognition method integrating multiple attention mechanisms is proposed. The original images collected by the ground penetrating radar are preprocessed and then sent to the YOLOv8 model integrating multiple attention mechanisms for recognition to obtain hidden road diseases.

[0008] The YOLOv8 integrating multiple attention mechanisms is an improved network of YOLOv8. The YOLOv8 integrating multiple attention mechanisms includes a backbone network for feature extraction, a neck network for feature fusion, and a head network for detection. The improvements to YOLOv8 include:

[0009] Replace the standard convolutional layers in the backbone network of YOLOv8 with depthwise separable convolutional layers, that is, all convolutional layers in the backbone network are DWConv layers;

[0010] During the feature fusion process in the neck network, a multi-attention mechanism MSE module formed by connecting a hybrid attention mechanism module and a SE attention module in parallel is inserted after each C2f module after upsampling in the neck network.

[0011] Furthermore, the improvement of YOLOv8 also includes multiple attention mechanism response fusion method, the specific processing process is as follows:

[0012] The features of the second DWConv+C2f module in the backbone network are defined as the original feature branch; the upsampled features of the first C2f+MSE module in the neck network are defined as the feature branch after multiple attention mechanisms;

[0013] The original feature branch and the feature branch after the multiple attention mechanism are subjected to convolution cross-correlation operations respectively, and then the two convolution cross-correlation operation results are linearly fused, and the fusion result is sent to the second C2f+MSE module in the neck network for processing.

[0014] Furthermore, the fusion formula for linearly fusion of the two convolution cross-correlation operation results is as follows:

[0015] F=c1F1+c2F2

[0016] Where: F1 is the response of the original feature branch after convolution cross-correlation operation; F2 is the response of the feature branch after the multiple attention mechanism after convolution cross-correlation operation; c1 is the weight coefficient of the original feature branch; c2 is the weight coefficient of the feature branch after the multiple attention mechanism.

[0017] Furthermore, the hybrid attention mechanism module mixes the channel attention and the spatial attention mechanism;

[0018] The channel attention mechanism processing process in the hybrid attention mechanism is as follows:

[0019] Perform global average pooling on the input single feature layer along the spatial direction, then perform MLP processing on the pooling result and send it to the sigmoid function to obtain the weight of each channel of the input feature layer, and finally multiply the weight by the original input feature layer as the output;

[0020] The spatial attention mechanism processing process in the hybrid attention mechanism is as follows:

[0021] The input is the output of the channel attention mechanism. The maximum value and the average value are taken on the channel of each feature point, and then the two results are stacked. The channel is reduced in dimension using a convolution with a channel of 1, and the reduced dimension result is passed through the sigmoid function to obtain the weight of each feature point in the input feature layer. Finally, the weight is multiplied by the channel attention output result.

[0022] Furthermore, preprocessing is performed on the original image collected by the ground penetrating radar, including background removal and filtering.

[0023] Furthermore, in the preprocessing process, the background removal is achieved by using a principal component analysis method; the filtering process is performed by filtering through a finite impulse response filter and a FR filter, and deconvolution is performed using a Wiener filter.

[0024] Furthermore, the YOLOv8 model integrating multiple attention mechanisms is pre-trained, and the training set construction process for training the YOLOv8 model integrating multiple attention mechanisms includes the following steps:

[0025] Step 1: pre-process the raw image data collected by the ground penetrating radar: perform static correction removal and interference suppression, gain, background removal and filtering, and deconvolution processing in sequence;

[0026] Step 2: Based on the image processed in step 1, determine the hidden road damage area, and use Labelimg to annotate the damage features in the image to create a data set. The damage features include shafts, voids, and cracks in the image.

[0027] Furthermore, based on the image processed in step 1, the hidden road damage area is determined, and the process of labeling the damage features in the image using Labelimg includes:

[0028] The original data set is sliced, that is, the image processed in step 1 is sliced ​​into many small images of size 512×512, the small images are screened based on the diseased area, and then the diseased area in the image is annotated using Labelimg software to form a corresponding txt file, which contains the type, size and location information of each detected target.

[0029] A computer storage medium stores a computer program, which is loaded and executed by a processor to implement a method for automatically identifying hidden road defects that integrates multiple attention mechanisms.

[0030] A device for automatically identifying hidden road defects that integrates multiple attention mechanisms. The device includes a processor and a memory. The memory stores a computer program. The computer program is loaded and executed by the processor to implement the method for automatically identifying hidden road defects that integrates multiple attention mechanisms.

[0031] Beneficial effects:

[0032] The present invention can realize the recognition of ground penetrating radar images and can effectively solve the problem that ground penetrating radar image identification is extremely dependent on manual work. At the same time, the invented model improves YOLOv8, uses DWConv to replace the standard convolution layer in the original backbone network, and adds the MSE attention mechanism to the neck network part for multi-feature perception fusion so that the model can obtain more accurate and useful information of the focused features. While better mapping more target features to the feature map, the feature information of some tiny targets in the disease can also be well extracted to enrich the semantic information, thereby improving the detection accuracy; the present invention also designs a multiple attention mechanism response fusion calculation method, which can enhance the output effect of the algorithm. The model of the present invention as a whole can well improve the accuracy and robustness of the road hidden disease recognition model in complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a flow chart of the method for automatically identifying hidden road defects incorporating multiple attention mechanisms of the present invention;

[0034] Figure 2 This is a schematic diagram of the network structure of the improved YOLOv8 in the present invention;

[0035] Figure 3 The network architecture of DWConv;

[0036] Figure 4 It is the SE attention mechanism structure;

[0037] Figure 5 It is the MSE attention mechanism structure;

[0038] Figure 6 Visual comparison of the improved YOLOv8 hidden disease recognition network of the present invention and other methods. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention is described below by the specific embodiments shown in the accompanying drawings. However, it should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the description of well-known structures and technologies is omitted to avoid unnecessary confusion of the concept of the present invention.

[0040] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, only structures and / or processing steps closely related to the solutions according to the present invention are shown in the accompanying drawings, while other details that are not closely related to the present invention are omitted.

[0041] Specific implementation method 1: Combination Figure 1-Figure 4 To explain this embodiment,

[0042] This embodiment proposes a method for automatically identifying hidden road defects by integrating multiple attention mechanisms, including the following steps:

[0043] Step 1: pre-process the raw image data collected by the ground penetrating radar, and perform static correction removal and interference suppression, gain, background removal and filtering, and deconvolution processing in sequence; the whole process includes:

[0044] Step 101: Perform static correction and cut-off and interference suppression on the original image data, cut off the propagation path of the radar incident wave in the air, and eliminate the interference of ground undulations on the image; perform spectrum analysis on the original data, identify and suppress the electromagnetic interference generated by various components in the system, and replace it with linear interpolation data.

[0045] Step 102, performing a gain on the image data processed in step 101, amplifying the electromagnetic wave energy signal, and increasing the amplitude of the deep signal;

[0046] During the gain process, according to the specific conditions such as detection purpose, detection depth, data acquisition center frequency, etc., consider performing overall image gain (enhancing image contrast, suitable for when the overall image imaging effect is not obvious), exponential gain (suitable for identifying deeper targets) or piecewise linear gain (suitable for paying special attention to targets at a specific depth).

[0047] Step 103: performing background removal and filtering processing on the image data after gain in step 102.

[0048] The principal component analysis method is used to remove the background and weaken the echo of the ground layer and the structural layer.

[0049] When filtering, the finite impulse response filter and the FR filter are used for filtering, and the Wiener filter deconvolution technique is used to filter the signals with large differences in the center frequency of the ground penetrating radar pulses.

[0050] Step 104: Deconvolve the image data after filtering in step 103 to remove a series of repeated medium interfaces in the image data.

[0051] When the ground penetrating radar is working, it will perform mixing processing and sampling operations before receiving the reflected echo signal. During this process, the output of various components in the ground penetrating radar system will produce DC offset, which will cause strong interference to the collected electromagnetic wave signal. Therefore, the DC offset must be removed before target detection. Background removal can highlight the location of hidden diseases in underground space to a certain extent. In some special cases, if the radar wave encounters a clear separation between two layers of medium, the electromagnetic wave signal will be reflected multiple times on this interface. The original signal actually undergoes repeated convolution operations, which appear as a series of repeated medium interfaces on the radar image. Therefore, deconvolution processing is used to remove these repetitions and reproduce the original signal.

[0052] Step 2: Based on the image processed in step 1, determine the hidden road damage area, and use Labelimg to annotate the damage features in the image to create a data set. The damage features include shafts, voids, and cracks in the image.

[0053] In the process of making the data set, when searching for damaged areas (shafts, cavities, cracks, etc.) in the ground penetrating radar images through core sampling, manual image analysis and expert interpretation methods, the original radar image contains a wide detection range, the target signal response occupies a small area, and the hyperbolic feature is not obvious. Therefore, it is necessary to slice the original data set, that is, to cut the image processed based on step 1 into many small images of size 512×512, and then use Labelimg software to annotate the damaged areas (shafts, cavities, cracks, etc.) in the image to form a corresponding txt file, which contains the type, size and location information of each detected target.

[0054] In this embodiment, 986 pieces of valid image data are acquired.

[0055] Step 3: Perform data augmentation on the data set obtained in step 2 to expand the image data set, and divide the expanded image data set into a training set, a validation set, and a test set.

[0056] In this embodiment, the data enhancement methods to expand the image data set include but are not limited to random noise, horizontal mirroring, color conversion, Mixup, and can also include inversion, scaling grayscale data enhancement and other methods to expand the data set. In this embodiment, the number of samples in the expanded data set reaches 3225.

[0057] In the process of deep learning, sufficient data sets are extremely important for model training. Insufficient data sets may lead to difficulties in convergence or overfitting, which makes it difficult to reflect the advantages of deep learning. For this reason, the sample data sets need to be augmented. After data enhancement, not only can the number of samples be effectively increased, the background information of the image can be enriched, and the network can learn more robust and deeper target features, thereby avoiding problems such as non-convergence and overfitting during model training, and ultimately improving the detection precision and accuracy of the model. Before training the network, methods such as adding Gaussian noise, mirroring, and color transformation are used. This method can better amplify the data set and detect the performance of the model without affecting the accuracy of road hazard detection.

[0058] In this implementation, the training set, test set, and validation set are divided in a ratio of 7:2:1.

[0059] Step 4: Use the training set divided in step 3 to train and optimize the YOLOv8 model that integrates multiple attention mechanisms to obtain the trained weight file.

[0060] In the backbone network of traditional YOLOv8, it mainly includes Conv module, C2f module and SPPF module. Among them, Darknet53 is used as the feature extraction network model. The Darknet53 network structure is a pure convolutional neural network, which contains 53 convolutional layers. Each convolutional layer is usually followed by batch normalization BN (Batch Normalization) and LeakyReLU activation function to improve the training speed and stability, so that the model can learn rich feature representations. At the same time, residual connections are used to allow the network to learn identity mapping technology, and to a certain extent, the gradient disappearance problem in deeper networks is solved, so that the network can be effectively trained. Although VGG and Resnet models have been very maturely applied in target recognition and image classification, YOLOv8 supports end-to-end training, which means that from the input image to the final detection result, the whole process can be completed at one time by gradient descent, which helps to improve training efficiency and model performance. The present invention is based on the traditional YOLOv8 basic network model, and proposes a YOLOv8 model based on the improved YOLOv8 algorithm, that is, the YOLOv8 model of the fusion multiple attention mechanism proposed by the present invention. The YOLOv8 model that integrates multiple attention mechanisms and is improved on the basis of the traditional YOLOv8 also includes a backbone network (Backbone) for feature extraction, a neck network (Neck) for feature fusion, and a head network (Head) for detection. The structure of the YOLOv8 model that integrates multiple attention mechanisms is as follows: Figure 2 As shown, it should be noted that Figure 2 It is just that the improved YOLOv8 schematic diagram of the present invention based on the traditional YOLOv8 does not display all modules. For example, for the convenience of representing the neck network (Neck), the convolutional layer in the neck network Neck is not displayed: the first two modules of the traditional neck network are C2f modules, and the last two modules are Conv+C2f modules. When the present invention adds the MSE module, the first two modules of the neck network of the present invention are C2f+MSE modules, and the last two modules are Conv+C2f+MSE modules. Figure 2 And the Conv in the Conv+C2f+MSE module is demonstrated.

[0061] The present invention improves the traditional YOLOv8 as the basic network model, and the improvements to YOLOv8 mainly include:

[0062] A. Since YOLOv8 continues the CSP structure idea of ​​YOLOv5 and SPPF for global feature fusion, the Bottleneck CSP module with a large number of parameters is used in the path aggregation part, which also causes the pyramid structure to consume a lot of time in the feature extraction process, making it difficult to ensure the real-time detection of hidden road diseases. The present invention considers improving YOLOv8 by replacing the original standard convolution layer ( Figure 3 ), reducing the amount of calculation and parameters, in order to improve the efficiency and speed of computer recognition. The existing convolution algorithms include conventional convolution, grouped convolution, transposed convolution, dilated convolution, deformable convolution, depth-separable convolution and Ghost convolution. To improve the real-time performance of YOLOv8, it is necessary to consider the characteristics of different convolution algorithms, compare their performance with each other, and finally adopt the best replacement scheme. Conventional convolution performs convolution on each convolution kernel and each channel of the input feature map during feature extraction, which leads to relatively large parameters and calculations; grouped convolution divides the input and output channels into the same number of groups to reduce the number of parameters and calculations, but it groups channels to a certain extent, which limits the information exchange between channels; transposed convolution is mainly used for the expansion of feature maps and does not have the ability to improve computational efficiency; dilated convolution captures more information by expanding the receptive field of the convolution kernel, and also does not reduce the number of parameters or calculations of the model; deformable convolution introduces an offset to adapt to the geometric transformation of the input data, enhancing the geometric adaptability of the model. However, this flexibility comes at the cost of adding additional offset parameters and computational complexity. Ghost convolution reduces the amount of computation and parameters by generating some "ghost" feature maps, which are obtained from the original feature maps through simple linear transformations. Although this method reduces the complexity of the model, it sacrifices the richness of feature extraction to a certain extent. In contrast, deep separable convolution (DWConv) can achieve effective fusion between channels through point-by-point convolution, significantly reducing the complexity of the model while maintaining spatial perception capabilities, and effectively processing input features through efficient convolution without adding additional parameters.

[0063] Therefore, the present invention considers improving the backbone network in YOLOv8, replacing the original standard convolution layer with the depthwise separable convolution DWConv, reducing the amount of calculation and the amount of parameters, so as to improve the efficiency and speed of computer recognition operation. In the improved convolution layer, one convolution channel of DWConv corresponds to one filter, and the original convolution layer in YOLOv8 performs convolution calculations on one channel and all filters. Therefore, the use of DWConv can significantly reduce the number of calculation parameters and reduce the dimension of the output image, especially under the condition of limited computing resources. The network architecture of DWConv is as follows: Figure 3As shown, DWConv (Depthwise Separable Convolution) is divided into two parts: 1) Depthwise Convolution: In the convolution process, each hidden layer performs convolution calculations with only a single filter in the convolution calculation, and the convolution layer is mainly used to generate and capture the spatial state information of the input data; 2) Pointwise Convolution: Pointwise convolution is a 1×1 convolution operation, that is, all data on the entire spatial channel are convolved one by one, which is used to perform linear combination calculations on the spatial data information generated by the depth convolution layer. The present invention uses this method to reduce parameters and calculations to provide an efficient feature extraction method.

[0064] In the backbone network of this embodiment, the input passes through the DWConv module, the DWConv module+C2f module, the DWConv module+C2f module, the DWConv module+C2f module, the DWConv module+C2f module and the SPPF module in sequence.

[0065] B. After the backbone network performs the corresponding convolution calculation, the obtained spatial data is subjected to feature fusion calculation in the neck network, and a multi-attention mechanism module MSE is inserted after each C2f module after upsampling in the neck network, as shown in Figure 5 As shown in Figure 2, the MSE attention module can enhance the channel features of the input feature data without changing the size of the input feature map.

[0066] Considering that the model pays less attention to the features of the region of interest, the present invention designs an improved hybrid attention mechanism module ( Figure 5 The Mixture Attention in the CNN is connected in parallel with the SE (Squeeze-and-Excitation) attention module to form a multi-attention mechanism MSE module (such as Figure 5 As shown in Figure 2, different feature branches are tracked, searched and located respectively, so that the algorithm can focus on the channel, space and position feature information of the target object, so as to help the algorithm locate and identify the features of the target more accurately. Figure 4 shown.

[0067] Mixture Attention mixes channel attention with spatial attention mechanism. For the channel attention mechanism in the hybrid attention mechanism, a global average pooling operation is performed on the input single feature layer along the spatial direction, and then the pooling result is processed by MLP and sent to the sigmod function to obtain the weight of each channel of the input feature layer, and finally the weight is multiplied by the original input feature layer as the output. For the input of the spatial attention mechanism in the hybrid attention mechanism, the output of the channel attention mechanism is taken as the maximum value and average value on the channel of each feature point, and then the two results are stacked, and then the channel is reduced in dimension using convolution with a channel of 1, and the dimension reduction result is passed through the sigmod function to obtain the weight of each feature point in the input feature layer, and finally the weight is multiplied by the channel attention output result. After the above two steps, the improvement of the hybrid attention mechanism is completed, and the algorithm can learn to pay attention to the channel and spatial features of the target.

[0068] After experimental analysis, the algorithm is not very sensitive to spatial features. The present invention connects the above hybrid attention mechanism with the SE attention module in parallel. It first performs channel attention operation on the input feature map, that is, compresses the spatial dimension of the input feature map, learns the importance of different channels, and then performs spatial attention operation on the feature map, compresses the channel dimension of the feature map, and learns the importance of different spatial parts. Therefore, it can effectively solve the problem of feature loss caused by the different proportions of different channels in the convolution pooling process of the model, and can make the network model pay more attention to the channel, space and position information of the target more accurately, thereby greatly improving the recognition speed and accuracy of the yolov8 model for different hidden road diseases, and providing a fast and accurate analysis tool for ground penetrating radar in actual road hidden disease detection.

[0069] C. In order to enhance the output effect of the algorithm, the present invention also designs a multi-attention mechanism to respond to fusion calculation. As mentioned above, the improved hybrid attention mechanism Mixture Attention and SE are used for feature branch tracking, search and positioning respectively, and the features output by the second DWConv+C2f module in the backbone network are defined as the original feature branch ( Figure 2 The light blue branch in the middle of the neck network is defined as the upsampled features of the first C2f+MSE module in the neck network as the feature branch after multiple attention mechanisms. Then, the original feature branch and the feature branch after multiple attention mechanisms are convolutionally correlated, and then the two results are linearly fused. The fusion result is sent to the second C2f+MSE module in the neck network for processing.

[0070] The linear fusion formula is as follows:

[0071] F=c1F1+c2F2

[0072] Where: F1∈R C×H×W is the original feature branch after convolution cross-correlation operation response; F2∈R C×H×W is the convolution cross-correlation operation response of the feature branch after the multiple attention mechanism; c1 is the weight coefficient of the original feature branch; c2 is the weight coefficient of the feature branch after the multiple attention mechanism.

[0073] In the improved YOLOv8, the input first passes through the backbone network to extract the feature map of the input image, and then enters the SPPF module after repeated convolution. This process can effectively capture features of different scales and realize the fusion of local features and global features. The extracted feature map enters the neck network (Neck). By constructing a multi-scale feature pyramid and realizing the beneficial transmission of cross-layer information, the network can better perceive the features of targets of different scales. The model uses anchor boxes on feature maps of multiple scales to predict the bounding box and predict the position offset, confidence and category probability of each anchor box. Finally, it enters the detection network. The region will extract and optimize the feature map to obtain a feature map with region proposals. The fully connected operation is used to locate the target, so as to perform bounding box regression and classification regression of image defects. The prediction results are post-processed, including threshold filtering and non-maximum suppression, to remove low-confidence predictions and merge overlapping boxes. Finally, the target category, bounding box coordinates and confidence are output, and the precise information of the detected underground target space body is finally obtained.

[0074] Step 5: Use the improved YOLOv8 trained in step 4 to accurately identify hidden road defects collected by the ground penetrating radar.

[0075] In the recognition process, the newly collected ground-penetrating radar original image is used as the input of the trained neural network model, so that the model automatically recognizes various hidden defects in the ground-penetrating radar image, and finally marks the ground-penetrating radar scanned image with hidden road defects. In this embodiment, the improved YOLOv8 obtained by training is tested using the test set and verified using the validation set.

[0076] Example

[0077] A high-precision three-dimensional ground-penetrating radar is used to detect underground hidden diseases on urban roads of about 50 km. According to the above steps, the collected raw image data of the ground-penetrating radar are subjected to DC removal, background removal, filtering, and deconvolution processing. Then the model structure is designed. At the same time, data enhancement and image annotation are required to produce a data set. Finally, model tests and comparative analysis are carried out.

[0078] The model test results are shown in Table 1:

[0079] Table 1 Comparison of iterative evaluation indicators of each model

[0080]

[0081] The recognition effect comparison of each model on the test set is as follows: Figure 6 shown.

[0082] Based on the rich practical experience and professional knowledge of such products, the present invention designs a method for automatically identifying hidden road defects that integrates multiple attention mechanisms, which can solve the problems of difficult data interpretation, large workload of target identification and classification, low detection accuracy, and poor real-time performance in the current ground penetrating radar in target detection to a certain extent. Compared with the prior art, the model of this application replaces the original standard convolution layer with DWConv, and provides an efficient feature extraction method by reducing parameters and calculation amount, which is particularly suitable for fast and accurate deep learning model deployment in resource-constrained environments; and the introduction of SE attention mechanism can obtain channel-level global features and the weights of each channel, and then learn the relationship between each channel for multi-feature perception fusion, which solves the problem of feature loss caused by different proportions of different channels in the convolution pooling process of the model to a certain extent. Through the above optimization and improvement measures, the imbalance problem between difficult samples and easy samples can be effectively reduced, so that the model can be better trained, and the real-time and accuracy of the ground penetrating radar road internal cavity detection can be improved. Specific implementation method 2:

[0084] This embodiment is a computer storage medium, in which a computer program is stored. The computer program is loaded and executed by a processor to implement an automatic identification method for hidden road defects that integrates multiple attention mechanisms.

[0085] It should be noted that the computer program is loaded and executed by the processor to realize a method for automatically identifying hidden road defects by integrating multiple attention mechanisms, which is mainly a process of pre-processing the original image collected by the ground penetrating radar and then sending it to the YOLOv8 model integrating multiple attention mechanisms for identification to obtain hidden road defects. It can also include the training process of the YOLOv8 model integrating multiple attention mechanisms.

[0086] In addition, it should be understood that the storage medium described in this embodiment includes but is not limited to magnetic storage media and optical storage media; the magnetic storage media includes but is not limited to RAM, ROM, and other storage media such as hard disks and USB flash drives. Specific implementation method three:

[0088] This embodiment is a device for automatically identifying hidden road defects that integrates multiple attention mechanisms. The device includes a processor and a memory. The memory stores a computer program. The computer program is loaded and executed by the processor to implement a method for automatically identifying hidden road defects that integrates multiple attention mechanisms.

[0089] It should be noted that the computer program is loaded and executed by the processor to realize a method for automatically identifying hidden road defects by integrating multiple attention mechanisms, which is mainly a process of pre-processing the original image collected by the ground penetrating radar and then sending it to the YOLOv8 model integrating multiple attention mechanisms for identification to obtain hidden road defects. It can also include the training process of the YOLOv8 model integrating multiple attention mechanisms.

[0090] In addition, it should be understood that the device described in this embodiment includes but is not limited to a device including a processor and a memory, and may also include other devices corresponding to units or modules with information collection, information interaction, and control functions, for example, the device may also include a signal collection device, etc. The device includes but is not limited to a PC, a workstation, a mobile device, etc.

[0091] The above shows and describes the basic principles and main features of the present invention and the advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A method for automatically identifying hidden road defects by integrating multiple attention mechanisms, characterized in that: The original images collected by the ground penetrating radar are preprocessed and sent to the YOLOv8 model with multiple attention mechanisms for identification to obtain hidden road defects. The YOLOv8 integrating multiple attention mechanisms is an improved network of YOLOv8, and the YOLOv8 integrating multiple attention mechanisms includes a backbone network for feature extraction, a neck network for feature fusion, and a head network for detection; Improvements to YOLOv8 include: Replace the standard convolutional layers in the backbone network of YOLOv8 with depthwise separable convolutional layers, that is, all convolutional layers in the backbone network are DWConv layers; During the feature fusion process in the neck network, a multi-attention mechanism MSE module formed by connecting a hybrid attention mechanism module and a SE attention module in parallel is inserted after each C2f module after upsampling in the neck network.

2. According to claim 1, a method for automatically identifying hidden road defects by integrating multiple attention mechanisms is characterized in that: The improvement of YOLOv8 also includes the multi-attention mechanism response fusion method. The specific processing process is as follows: The features of the second DWConv+C2f module in the backbone network are defined as the original feature branch; the upsampled features of the first C2f+MSE module in the neck network are defined as the feature branch after multiple attention mechanisms; The original feature branch and the feature branch after the multiple attention mechanism are subjected to convolution cross-correlation operations respectively, and then the two convolution cross-correlation operation results are linearly fused, and the fusion result is sent to the second C2f+MSE module in the neck network for processing.

3. The method for automatically identifying hidden road defects by integrating multiple attention mechanisms according to claim 2 is characterized in that: The fusion formula for linearly fusing the results of two convolution cross-correlation operations is as follows: F=c1F1+c2F2 Where: F1 is the response of the original feature branch after convolution cross-correlation operation; F2 is the response of the feature branch after the multiple attention mechanism after convolution cross-correlation operation; c1 is the weight coefficient of the original feature branch; c2 is the weight coefficient of the feature branch after the multiple attention mechanism.

4. According to claim 1, the method for automatically identifying hidden road defects by integrating multiple attention mechanisms is characterized in that: The hybrid attention mechanism module mixes channel attention and spatial attention mechanisms; The channel attention mechanism processing process in the hybrid attention mechanism is as follows: Perform global average pooling on the input single feature layer along the spatial direction, then perform MLP processing on the pooling result and send it to the sigmoid function to obtain the weight of each channel of the input feature layer, and finally multiply the weight by the original input feature layer as the output; The spatial attention mechanism processing process in the hybrid attention mechanism is as follows: The input is the output of the channel attention mechanism. The maximum value and the average value are taken on the channel of each feature point, and then the two results are stacked. The channel is reduced in dimension using a convolution with a channel of 1, and the reduced dimension result is passed through the sigmoid function to obtain the weight of each feature point in the input feature layer. Finally, the weight is multiplied by the channel attention output result.

5. According to claim 1, the method for automatically identifying hidden road defects by integrating multiple attention mechanisms is characterized in that: Preprocessing of the original images collected by the ground penetrating radar includes background removal and filtering.

6. The method for automatically identifying hidden road defects by integrating multiple attention mechanisms according to claim 5 is characterized in that: In the preprocessing process, the background removal is achieved by principal component analysis; the filtering process is performed by filtering through a finite impulse response filter and a FR filter, and deconvolution is performed using a Wiener filter.

7. The method for automatically identifying hidden road defects by integrating multiple attention mechanisms according to any one of claims 1 to 6, characterized in that: The YOLOv8 model integrating multiple attention mechanisms is pre-trained, and the training set construction process for training the YOLOv8 model integrating multiple attention mechanisms includes the following steps: Step 1: pre-process the raw image data collected by the ground penetrating radar: perform static correction removal and interference suppression, gain, background removal and filtering, and deconvolution processing in sequence; Step 2: Based on the image processed in step 1, determine the hidden road damage area, and use Labelimg to annotate the damage features in the image to create a data set. The damage features include shafts, voids, and cracks in the image.

8. The method for automatically identifying hidden road defects by integrating multiple attention mechanisms according to claim 7 is characterized in that: Based on the image processed in step 1, the process of determining the hidden road damage area and labeling the damage features in the image using Labelimg includes: The original data set is sliced, that is, the image processed in step 1 is sliced ​​into many small images of size 512×512, the small images are screened based on the diseased area, and then the diseased area in the image is annotated using Labelimg software to form a corresponding txt file, which contains the type, size and location information of each detected target.

9. A computer storage medium, characterized in that: The storage medium stores a computer program, which is loaded and executed by a processor to implement the method for automatically identifying hidden road defects integrating multiple attention mechanisms as described in any one of claims 1 to 6.

10. An automatic road hidden disease recognition device integrating multiple attention mechanisms, characterized in that: The device includes a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the method for automatically identifying hidden road defects integrating multiple attention mechanisms as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Steel surface defect detection method based on improved YOLOv5s

    CN115829991A

  • GIS infrared feature recognition system and method based on improved YOLOv5

    CN116342894A