Ground penetrating radar underground pipeline identification method based on improved YOLOv8
By adding the CBAM attention mechanism and replacement loss function in YOLOv8, the problems of low accuracy and poor small target detection performance in ground penetrating radar GPR data detection are solved, and higher detection accuracy and small target detection capabilities are achieved.
Patent Information
- Application Number
- CN202411781732.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-05-06
AI Technical Summary
YOLOv8 has problems with low accuracy and poor small target detection performance in GPR data detection of ground penetrating radar.
Added the CBAM attention mechanism to the feature extraction network of YOLOv8, and replaced the CIoU loss function with the WIoU border loss function to improve the detection accuracy and small object detection capability of the model.
By improving the network structure and loss function of YOLOv8, the detection accuracy of underground pipelines in the ground penetrating radar image is significantly improved, and the error detection rate and miss detection rate are reduced.
Smart Images

Figure CN119942146A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of radar image processing, and in particular to a ground penetrating radar underground pipeline recognition method based on improved YOLOv8. Background Art
[0002] Ground Penetrating Radar (GPR) is an efficient non-destructive detection method with advantages such as high detection accuracy, strong penetration and mature technology. It is also easy to combine with data processing technologies such as artificial intelligence. It has shown excellent application value in the fields of underground target exploration, road and bridge quality inspection and archaeological research. Ground Penetrating Radar (GPR) locates underground objects by utilizing the natural laws of electromagnetic wave propagation and reflection between various materials. This technology involves transmitting high-frequency electromagnetic waves into the ground and collecting the signals reflected by these waves when they encounter boundaries with different electromagnetic characteristics. However, the complexity of the underground environment often causes a large amount of noise and interference signals in the image. In addition, large-scale detection will generate massive amounts of data. Relying solely on manual interpretation is not only inefficient, but may also lead to misidentification of certain targets. Therefore, it is urgent to explore more efficient and accurate ground penetrating radar (GPR) data interpretation methods to improve the effect and efficiency of ground penetrating radar (GPR) image interpretation.
[0003] In order to give full play to the advantages of GPR in underground space detection, researchers began to combine data processing technology with GPR. Machine learning, as one of the key technologies in the field of artificial intelligence, is known for its excellent data processing and analysis capabilities, as well as its self-learning ability. Traditional GPR image machine learning detection methods mainly include generalized Hough transform (GHT), support vector machine (SVM), clustering algorithm and BP neural network. These methods require artificially designed features, are sensitive to training data and have limited generalization capabilities, making it difficult to adapt to complex and changeable actual detection scenarios.
[0004] With the continuous advancement of deep learning technology, it has shown significant superiority in applications such as image segmentation and target recognition. This method does not require manual feature design, and can automatically learn and extract complex and diverse features, which is particularly suitable for GPR image detection in complex environments. At present, the mainstream deep learning models used for GPR image recognition include Faster R-CNN, SSD, and YOLO. As a two-stage detection algorithm, Faster R-CNN treats target recognition and classification as two independent steps, resulting in a relatively slow detection process. YOLO and SSD are single-stage detection algorithms. They do not rely on the generation of candidate regions, but directly predict the category and bounding box coordinates of the target in the network. In addition, YOLO's recognition accuracy is better than SSD, and it is more suitable for GPR image detection task scenarios.
[0005] With the continuous iteration of the YOLO series of algorithms, it has significantly improved the accuracy of recognition while maintaining fast detection, and has been successfully applied to target recognition in open-pit mines, dangerous goods detection, forest fire monitoring and other fields. In the field of ground penetrating radar image recognition, Yang Bisheng et al. used YOLOv3 to detect road structure targets on GPR B-Scan images in "Real-time Detection Method of Underground Targets by Vehicle-mounted Ground Penetrating Radar". The initial data set was small, and conventional enhancement methods such as cutting, cropping, and flipping were used to expand it by 6.2 times. Excessive enhancement of data may lead to potential overfitting of the model.
[0006] Hu et al.
[0007] "A study of automatic recognition and localization of pipeline for ground penetration rating based on deep learning" uses Faster R-CNN, YOLOv5 and SSD to perform recognition on a self-made dataset, and uses the center position of the upper boundary of the predicted bounding box as the buried position of the pipeline, thereby realizing the positioning function. This method is relatively accurate in positioning on standard hyperbolic images, but has a large error in non-standard hyperbolic positioning under complex conditions.
[0008] The bounding box loss function plays an important role in the convergence speed and accuracy of the network. Choosing a suitable loss function can significantly optimize the model performance. The YOLOv8 network model uses the CIoU loss function as its bounding box loss function by default. The CIoU loss function comprehensively measures the overlapping area, center point distance, and aspect ratio between the predicted box and the true box, but there is a certain ambiguity in the aspect ratio processing. The problem of balancing difficult and easy samples is not considered, and too much emphasis is placed on the regression of low-quality bounding boxes, which restricts the improvement of network detection performance. Summary of the invention
[0009] In order to effectively solve the problems of low accuracy and poor small target detection performance of YOLOv8 in GPR data detection, the present invention provides a ground penetrating radar underground pipeline recognition method based on improved YOLOv8, which adds a CBAM attention mechanism to the YOLOv8 backbone neural network, replaces the original CIoU loss function of YOLOv8 with the WIoU border loss function, and inputs the data set to be detected into the trained target detection network to realize the detection of common underground pipeline targets.
[0010] The technical solution adopted by the present invention is:
[0011] The underground pipeline recognition method of ground penetrating radar based on improved YOLOv8 includes the following steps:
[0012] Step 1: Obtain ground penetrating radar target image data and build a data set;
[0013] Step 2: Perform data enhancement on the target image data obtained in step 1;
[0014] Step 3: Improve the YOLOv8 network model, add the CBAM attention mechanism to the feature extraction network of the YOLOv8 network model; replace the CIoU loss function of the YOLOv8 network model with the WIoU border loss function;
[0015] Step 4: Input the data set to be detected into the trained improved YOLOv8 network model to realize underground pipeline location and category detection.
[0016] In step 1, the target image data is acquired based on the ground penetrating radar nondestructive testing technology. The ground penetrating radar mainly uses the high-frequency pulse electromagnetic waves emitted by the antenna to detect the target. The detection tool uses the Swedish MALA series ground penetrating radar. The ground penetrating radar is used to scan the measured scene multiple times at different speeds and positions. During the scanning, the time window, sampling frequency, sampling spacing and other parameters can be adjusted to finally complete the acquisition of the target image data.
[0017] In step 1, the constructed data set consists of two parts: one part is generated by simulating ground penetrating radar images using the GprMax3.0 toolbox as a training set, and the other part is used to collect targets in the constructed measured scene, and the collected target image data is used as a test set; the same parameters are set for the two data sets;
[0018] In the step 1, the ground penetrating radar detects underground pipelines by transmitting signals using a shielded antenna with a center frequency of 800 MHz, with a maximum penetration of 2.5 m. Ground penetrating radar data is obtained to form a B-Scan image, and the obtained B-Scan image is annotated with hyperbolic features using annotation software; a total of 1,440 simulated ground penetrating radar images are generated, and 360 ground penetrating radar images in actual scenarios are collected.
[0019] In the step 2, the obtained target image data is enhanced based on the use of Mixup software. Mixup is a data enhancement algorithm that minimizes local risk. Compared with similar enhancement methods, this algorithm improves the generalization ability of the model by reducing the risk of overfitting, increases the diversity of the data set, and makes its application scope more extensive. Therefore, in order to enrich the ground penetrating radar image samples and improve the accuracy of recognition, the Mixup algorithm is used to expand the ground penetrating radar image samples.
[0020] The process can be expressed as
[0021] batch' x =λ×batch xi +(1-λ)batch xj (1);
[0022] batch' y =λ×batch yi +(1-λ)batch yj (2);
[0023] Among them, × represents multiplication operation, batch xi 、batch xj and batch yi 、batch yj Represents ground penetrating radar images and corresponding labels respectively. This is randomly selected training data. x ' and batch y 'represent the mixed GPR image samples and their corresponding labels respectively, λ represents the proportional factor between two randomly selected training data, which is usually the mixing coefficient calculated by Beta distribution and is a constant.
[0024] In ground penetrating radar image enhancement, Mixup can mix ground penetrating radar images from different pipelines to expand the training data set.
[0025] In step 3, the improved YOLOv8 network model includes three parts: feature extraction network (Backbone), feature fusion network (Neck) and detection head (Head);
[0026] The construction of the improved YOLOv8 network model includes the following steps:
[0027] Step 3.1: Build the feature extraction network (Backbone) of the improved YOLOv8 network model. The feature extraction network (Backbone) uses the CSPDarknet network structure, which reduces the number of model parameters and improves the feature extraction capability, thereby achieving higher detection accuracy and speed. The present invention proposes to add the CBAM attention mechanism to the feature extraction network (Backbone), which can effectively reduce the false detection rate and missed detection rate of underground pipelines in the ground penetrating radar B-SCAN image, and improve the detection accuracy of the model;
[0028] The CBAM attention mechanism is a lightweight convolutional attention mechanism module, which consists of two sub-modules, CAM and SAM, which are responsible for the calculation of channel attention and spatial attention respectively. This module can be easily integrated into the existing network structure to achieve plug-and-play functionality. When performing image classification tasks, the channel attention sub-module focuses on identifying image information that has a positive impact on classification, while the spatial attention sub-module focuses on determining the specific location of this information in the image.
[0029] The CBAM attention mechanism structure includes input features, channel attention module, spatial attention module and output features;
[0030] The input feature is represented as F∈R C*H*W , where C represents the number of channels, H and W represent the height and width respectively;
[0031] The channel attention module is implemented through a one-dimensional convolution M C ∈R C*1*1 The features are processed and then multiplied by the original features. Next, the output of the channel attention module is used as the input of the spatial attention module, which uses a two-dimensional convolution M S ∈R 1*H*W Further refine the features. Finally, the output of the spatial attention module is multiplied with the original features to obtain the final feature representation.
[0032] The details are as follows:
[0033] The channel attention module maintains the consistency of the channel dimension while reducing the spatial dimension. The module first processes the input feature map in parallel through maximum pooling (MaxPool) and average pooling (AvgPool), converting the size of the feature map from C*H*W to C*1*1; then, these feature maps pass through a multi-layer shared perceptron module, which first reduces the number of channels to the original 1 / r, and then expands back to the original number of channels. Through the ReLU activation function, two activated feature maps can be obtained. After superimposing the two feature maps, the channel attention map is generated by applying the sigmoid activation function. This map is then used to weight the original feature map to enhance important features and suppress unimportant parts, and finally restore it to the size of C*H*W. Its formula can be expressed as:
[0034]
[0035] In formula (3), M c (F) represents the generated channel attention, MaxPool(F) represents the input features through the MaxPool layer, AvgPool(F) represents the input features through the AvgPool layer, MLP represents multi-layer perceptron, Represents the features of the maximum pooling layer on the channel axis, represents the features of the average pooling layer on the channel axis, W1 and W0 represent the weight matrices of MLP, and σ represents the Sigmoid activation function.
[0036] The spatial attention module maintains the integrity of the spatial dimension while compressing the channel dimension. This module receives the output of the channel attention module and obtains two 1*H*W feature maps through maximum pooling and average pooling operations respectively. These two feature maps are then concatenated together and converted into a single channel feature map through a 7*7 convolution layer. Subsequently, the spatial attention map is generated using the sigmoid function. This spatial attention map is used to weight the original feature map to strengthen the key information and suppress the non-critical parts, and finally expand it back to the size of C*H*W. Its formula can be expressed as:
[0037]
[0038] In formula (4), M s (F) represents the generated spatial attention, f 7*7 represents a 7×7 convolution operation, represents the maximum pooling feature map, Represents the average pooling feature map
[0039] CBAM is a lightweight attention mechanism module that integrates two core components: channel attention (CAM) and spatial attention (SAM). This technology aims to capture important features in both channel and spatial dimensions, and can be easily integrated into existing network structures as a flexible plug-in. When performing image classification tasks, the channel attention mechanism focuses on extracting key information in the image that contributes to classification, while the spatial attention mechanism focuses on identifying the specific areas where this information is located. Adding the CBAM attention mechanism to the feature extraction network of the YOLOv8 network model can effectively reduce the false detection rate and missed detection rate of underground pipelines in the ground penetrating radar B-SCAN images, and improve the detection accuracy of the model.
[0040] Step 3.2: Build the feature fusion network (Neck) part of the improved YOLOv8 network model. The feature fusion network (Neck) part is located between the feature extraction network (Backbone) and the detection head (Head). The improved YOLOv8 network is composed of FPN (Feature Pyramind Network) and PAN (Path Aggregation Network), which mainly plays the role of feature fusion. It uses multi-scale feature fusion technology to capture the features of targets of different scales, thereby improving the accuracy and reliability of target recognition.
[0041] Step 3.3: Build the detection head part of the improved YOLOv8 network model. The detection head is responsible for multi-scale target recognition based on the feature map extracted by Backbone.
[0042] The improved YOLOv8 network model adopts a decoupled head structure, in which two parallel paths are responsible for extracting category features and spatial location features respectively. Each path finally passes through a 1×1 convolution layer to achieve classification and positioning functions. By efficiently integrating feature maps at all levels, comprehensive information capture is achieved, further enhancing the accuracy of detection.
[0043] In step 3, the CIoU loss function of the YOLOv8 network model is replaced with the WIoU border loss function; including the following:
[0044] The WIoU border loss function is an improved bounding box loss function that uses a dynamic non-monotonic focusing mechanism and measures the accuracy of the bounding box through a parameter β, while introducing an efficient gradient allocation strategy. This loss function reduces the competition between high-quality anchor boxes while also reducing the negative impact of low-quality borders on the gradient during model training. Therefore, the improved YOLOv8 described in the present invention uses the WIoU border loss function to replace the original CIoU loss function.
[0045] The WIoU border loss function is defined as follows:
[0046] L WIoU =rR WIoU L IoU (5);
[0047] Where, L IoU represents the intersection-over-union loss function, R WIoU represents the normalized distance between the center point of the target box and the predicted box, and r represents the gradient gain.
[0048] R WIoU The definition is as follows:
[0049]
[0050] Among them, W g Indicates the width of the minimum bounding box; H g Indicates the width and height of the minimum bounding box; (x gt ,y gt ) represents the coordinates of the center point of the real box; (x, y) represents the coordinates of the center point of the predicted box; ()* represents separation from the calculation graph, and will not be calculated during back propagation. g ×H g The gradient associated with the tensor of dimension .
[0051] The gradient gain r is defined as follows:
[0052]
[0053] Where: β is the outlier degree, α and δ are two hyperparameters.
[0054] The outlier degree β is defined as follows:
[0055]
[0056] in, Indicates L IoU Separate from the computational graph, represents the sliding average with momentum m. The WIoU loss function weighs the learning of low-quality samples and high-quality samples, suppresses the gradient interference caused by low-quality borders, enhances the robustness and adaptability of the detection model, and improves the convergence rate of model training.
[0057] In step 4, the detection head is used to predict the location and category of underground pipelines:
[0058] YOLOv8 adopts a decoupled head structure. Two parallel branches extract category features and position features respectively, and then use a layer of 1×1 convolution to complete the classification and positioning tasks.
[0059] Since the hyperbola of the pipeline is indirectly obtained by ground penetrating radar, mainstream detectors generally use rectangular prediction boxes to mark the detected area of interest, and the exact position of the pipeline cannot be directly determined in the prediction box. The present invention uses the vertex of the hyperbola as a positioning point to determine the location of the underground pipeline. YOLOv8 has two detection heads, one for classification and the other for predicting the bounding box. The above-mentioned underground pipeline position prediction task is added to the detection network as the third regression head, which consists of two parameters x and y, where x and y are pixel values along the horizontal axis and the vertical axis. The horizontal axis x is proportional to the actual distance along the measurement direction, and the vertical axis y is proportional to the two-way propagation time of the wave, so that the pipeline position can be obtained through the positioning point coordinates. The detection branch head uses the Wing-Loss loss function, which is expressed as:
[0060]
[0061] loss L (s) = ∑wing(ss′) (10);
[0062] Among them, w represents the range of the nonlinear part of the function (-w, w), ε represents the curvature of the nonlinear region, C is a constant, s represents the positioning point vector, and Loss L (s) represents the loss function of the anchor point vector s, and s' represents the predicted coordinate vector.
[0063] The enhanced data set is put into the improved YOLOv8 network for training, so that the improved YOLOv8 network can recognize the images of each pipeline and distinguish the image characteristics of each pipeline. When conducting field detection, it can automatically distinguish the category of the target pipeline.
[0064] The enhanced features obtained by the feature extraction network (Backbone) and FPN are regarded as a collection of countless feature points for image classification, so as to obtain the position and category of the hyperbolic target in the ground penetrating radar image.
[0065] The present invention provides a ground penetrating radar underground pipeline identification method based on improved YOLOv8, and the technical effects are as follows:
[0066] 1) Advantages of step 1 of the present invention: The radar data of underground pipelines is collected based on the ground penetrating radar non-destructive testing technology, which can detect underground pipelines without destroying the structure, detect underground pipeline distribution problems and provide high-resolution images.
[0067] 2) Advantages of step 2 of the present invention: Data enhancement is performed on the obtained target image data based on Mixup. Mixup is a data enhancement algorithm that minimizes local risk. Compared with similar enhancement methods, this algorithm improves the generalization ability of the model by reducing the risk of overfitting. At the same time, it increases the diversity of the data set, making it more widely applicable.
[0068] 3) Advantages of step 3 of the present invention: Based on the improved YOLOv8 ground penetrating radar image target detection method, the improved YOLOv8 adds a CBAM attention mechanism to its feature extraction network. The present invention proposes to add a CBAM attention mechanism to Backbone, which can effectively reduce the false detection rate and missed detection rate of underground pipelines in ground penetrating radar B-SCAN images, and improve the detection accuracy of the model.
[0069] 4) Advantages of step 3 of the present invention: Based on the improved YOLOv8 ground penetrating radar image target detection method, the improved YOLOv8 replaces the original CIoU loss function with the WIoU border loss function. YOLOv8 uses CIoU as its bounding box loss function by default. The CIoU loss function comprehensively measures factors such as the overlapping area, center point distance, and aspect ratio between the predicted box and the true box. However, there is a certain ambiguity in the processing of the aspect ratio. The problem of balancing difficult and easy samples is not considered, and too much emphasis is placed on the regression of low-quality bounding boxes, thereby restricting the improvement of network detection performance. WIoU adopts an innovative method. By dynamically calculating the IoU loss in category prediction, it not only reduces the competitive pressure between high-quality anchor boxes, but also reduces the gradient noise caused by low-quality borders to model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The present invention will be further described below in conjunction with the accompanying drawings and embodiments:
[0071] Figure 1 This is a flow chart of the underground pipeline identification method based on the improved YOLOv8 ground penetrating radar.
[0072] Figure 2 This is the GprMax forward simulation flow chart.
[0073] Figure 3 To perform hyperbolic feature annotation on B-scan image data;
[0074] Figure 3 The objects detected from left to right are: cable, PVC pipe, plastic bottle containing water.
[0075] Among them, Cable represents cable, PVC represents PVC pipe, and Water bottle represents plastic bottle containing water.
[0076] Figure 4(a) shows the image captured for the target at the actual measurement site;
[0077] Figure 4(b) is a simulation image;
[0078] Figure 4(c) shows the image after being processed by Mixup.
[0079] Figure 5 This is the structural diagram of the CBAM attention mechanism.
[0080] Figure 6 Schematic diagram of the improved YOLOv8 network model structure.
[0081] Figure 7 This is the prediction map of underground pipeline location and category by the improved YOLOv8 network detection head. DETAILED DESCRIPTION
[0082] The common underground pipeline recognition method based on the improved YOLOv8 ground penetrating radar includes the following steps:
[0083] Step 1: Acquire image data based on ground penetrating radar nondestructive testing technology;
[0084] Step 2: Perform data enhancement on the obtained target image data using Mixup software;
[0085] Step 3: A ground penetrating radar image target detection method based on an improved YOLOv8 network model, wherein the improved YOLOv8 adds a CBAM attention mechanism to its feature extraction network;
[0086] Step 4: A ground penetrating radar image target detection method based on an improved YOLOv8 network model, wherein the improved YOLOv8 network model replaces the original CIoU loss function with a WIoU border loss function;
[0087] Step 5: Input the data set to be detected into the trained target detection network to detect common underground pipeline targets.
[0088] In the step 1, image data is acquired based on the ground penetrating radar nondestructive testing technology. The steps include:
[0089] The dataset consists of two parts: one part uses the GprMax3.0 toolbox to simulate the generation of ground-penetrating radar images as a training set, and the other part uses the actual measurement scene to collect targets, and the collected target image data is used as a test set; the two parts of the dataset set have the same parameters; the ground-penetrating radar uses a shielded antenna with a center frequency of 800MHz to transmit signals to detect underground pipelines with a maximum penetration of 2.5m. The ground-penetrating radar data is obtained to form a B-Scan image, and the hyperbolic features of the acquired B-Scan image are annotated using annotation software; a total of 1,376 simulated ground-penetrating radar images are generated and 344 images in actual scenes are collected.
[0090] In the step 2, data enhancement is performed on the obtained target image data using Mixup software, which includes the following steps:
[0091] Mixup is a data enhancement algorithm that minimizes local risk. Compared with similar enhancement methods, this algorithm improves the generalization ability of the model by reducing the risk of overfitting, increases the diversity of the data set, and makes it more widely applicable. Therefore, in order to enrich the ground penetrating radar image samples and improve the recognition accuracy, the Mixup algorithm is used to expand the ground penetrating radar image samples.
[0092] In step 3, the ground penetrating radar image target detection method is based on the improved YOLOv8 network model, and the improved YOLOv8 network model adds a CBAM attention mechanism to its feature extraction network. The steps include:
[0093] The above improved YOLOv8 network model consists of three parts: feature extraction network (Backbone), feature fusion network (Neck) and detection head (Head).
[0094] S3.1: Build the Backbone part of the improved YOLOv8 network model. Backbone uses the CSPDarknet network structure, which reduces the number of model parameters and improves the feature extraction capability, thereby achieving higher detection accuracy and speed. The present invention proposes to add the CBAM attention mechanism to Backbone, which can effectively reduce the false detection rate and missed detection rate of underground pipelines in the ground penetrating radar B-SCAN image, and improve the detection accuracy of the model. CBAM is a lightweight convolutional attention mechanism module, which consists of two sub-modules, CAM and SAM, which are responsible for the calculation of channel attention and spatial attention respectively. This module can be easily integrated into the existing network structure to achieve plug-and-play functionality. When performing image classification tasks, the channel attention sub-module focuses on identifying image information that has a positive impact on classification, while the spatial attention sub-module focuses on determining the specific location of this information in the image.
[0095] S3.2: Build the Neck part of the improved YOLOv8 network model. The Neck part is located between the Backbone and the Head. It is still composed of FPN and PAN. It mainly plays the role of feature fusion. It uses multi-scale feature fusion technology to capture the features of targets of different scales, thereby improving the accuracy and reliability of target recognition.
[0096] S3.3: Build the head part of the improved YOLOv8 network model, which is responsible for multi-scale target recognition on the feature maps extracted by the backbone network. The improved YOLOv8 network model adopts a decoupled head structure, in which two parallel paths are responsible for extracting category features and spatial location features respectively, and each path finally realizes the classification and positioning functions through a 1×1 convolution layer. By efficiently integrating the feature maps of each level, comprehensive information capture is achieved, further enhancing the accuracy of detection.
[0097] In the step 4, based on the improved YOLOv8 network model ground penetrating radar image target detection method, the improved YOLOv8 network model replaces the original CIoU loss function with the WIoU border loss function. The steps include:
[0098] The WIoU loss function is an improved bounding box loss function that uses a dynamic non-monotonic focusing mechanism and measures the accuracy of the bounding box through the parameter β, while introducing an efficient gradient allocation strategy. This loss function reduces the competition between high-quality anchor boxes while also reducing the negative impact of low-quality borders on the gradient during model training. Therefore, the improved YOLOv8 network model described in the present invention uses the WIoU border loss function instead of the original CIoU loss function, which is defined as follows:
[0099] L WIoU =rR WIoU L IoU
[0100]
[0101] Among them, R WIOU The range of is set to [1, e), which will significantly increase the L of the normal quality anchor box. IoU Value. IoU The value is in the range [0, 1], which will significantly reduce the R of the high-quality anchor box. WIOU In addition, when the overlap between the anchor box and the target box is large, the loss function reduces the emphasis on the deviation of the center point. WIOU The gradient problem that makes the model difficult to converge is solved by g and height H gBy separating it from the calculation process, the gradient problem that hinders the convergence of the model has been successfully removed, thereby enhancing the stability and generalization ability of the detection model and accelerating the convergence of the model training.
[0102] In step 5, the data set to be detected is input into the trained target detection network to detect common underground pipeline targets. The following steps are included:
[0103] The detection head is used to predict the location and category of underground pipelines. The essence is to regard the enhanced features obtained by Backbone and FPN as a collection of countless feature points for image classification, so as to obtain the category of hyperbolic targets in the ground penetrating radar image.
[0104] YOLOv8 adopts a decoupled head structure. Two parallel branches extract category features and position features respectively, and then use a layer of 1×1 convolution to complete the classification and positioning tasks.
[0105] Since the hyperbola of the pipeline is indirectly obtained by ground penetrating radar, mainstream detectors generally use rectangular prediction boxes to mark the detected area of interest, and the exact position of the pipeline cannot be directly determined in the prediction box. The present invention uses the vertex of the hyperbola as a positioning point to determine the location of the underground pipeline. The original YOLOv8 has two detection heads, one for classification and the other for predicting the bounding box. The above-mentioned underground pipeline position prediction task is added to the detection network as the third regression head, consisting of two parameters x and y, where x and y are pixel values along the horizontal axis and the vertical axis. The horizontal axis x is proportional to the actual distance along the measurement direction, and the vertical axis y is proportional to the two-way propagation time of the wave, so that the pipeline position can be obtained by the positioning point coordinates. The structure of the positioning task is consistent with the prediction box and the category discrimination branch. The detection branch head uses the Wing-Loss loss function, and the formula can be expressed as:
[0106]
[0107] loss L (s) = ∑wing(ss′)
[0108] w indicates that the range of the nonlinear part of the function is (-w, w), ε indicates the curvature of the nonlinear region, C is a constant, s indicates the positioning point vector, and Loss L (s) represents the loss function of the anchor point vector s, Σ represents the summation, and s' represents the predicted coordinate vector.
[0109] For the prediction of the above-mentioned underground pipeline categories, the improved YOLOv8 network can be trained based on the above-mentioned enhanced data set, so that the improved YOLOv8 network can recognize the images of each pipeline and distinguish the image characteristics of each pipeline. During field detection, it can automatically distinguish the category of the target pipeline.
[0110] Figure 4(a) is a B-scan image collected at the actual measurement site. It can be seen that there are three different underground pipelines with the same spacing. From left to right, the burial depths of the first and second pipelines are similar, and the burial depth of the third pipeline is slightly shallower.
[0111] FIG4(b) is an image obtained by simulation. It can be seen that two identical pipelines are located in the same simulation model, but their buried depths are different. The buried depth of the pipeline on the left is shallower than that of the pipeline on the right. The background medium of the B-scan image obtained by the simulation model is evenly distributed.
[0112] Figure 4(c) is the image processed by Mixup, which can mix the GPR images of different pipelines. It can be seen that Mixup extracts the hyperbolic features in Figure 4(a) and Figure 4(b) and mixes them together to generate a new B-scan image, thereby achieving the purpose of expanding the data set.
[0113] Figure 7 The following is a prediction of the location and category of underground pipelines by the improved YOLOv8 network detector. It can be seen that the predicted underground pipeline category is cable, and the predicted depths are 0.72 meters and 0.79 meters respectively.
Claims
1. The underground pipeline identification method based on ground penetrating radar based on improved YOLOv8 is characterized by The following steps are involved: Step 1: Obtain ground penetrating radar target image data and build a data set; Step 2: Perform data enhancement on the target image data obtained in step 1; Step 3: Improve the YOLOv8 network model, add the CBAM attention mechanism to the feature extraction network of the YOLOv8 network model; replace the CIoU loss function of the YOLOv8 network model with the WIoU border loss function; Step 4: Input the data set to be detected into the trained improved YOLOv8 network model to realize underground pipeline location and category detection.
2. According to the method for identifying underground pipelines by ground penetrating radar based on improved YOLOv8 according to claim 1, it is characterized in that In step 1, the constructed data set consists of two parts: one part uses the GprMax3.0 toolbox to simulate the generation of ground penetrating radar images as a training set, and the other part performs target acquisition in the constructed measured scene, and the collected target image data is used as a test set; the same parameters are set for the two data sets.
3. According to the method for identifying underground pipelines by ground penetrating radar based on improved YOLOv8 according to claim 2, it is characterized in that : In the step 1, the ground penetrating radar detects underground pipelines by using a shielded antenna with a center frequency of 800 MHz to transmit signals, with a maximum penetration of 2.5 m. The ground penetrating radar data is obtained to form a B-Scan image and the obtained B-Scan image is annotated with hyperbolic features using annotation software; a total of 1440 simulated ground penetrating radar images are generated, and 360 ground penetrating radar images in actual scenes are collected.
4. According to the method for identifying underground pipelines by ground penetrating radar based on improved YOLOv8 according to claim 1, it is characterized in that : In the step 2, the obtained target image data is enhanced by using Mixup software to expand the ground penetrating radar image samples; The process can be expressed as batch' x =λ×batch xi +(1-λ)batch xj (1); batch' y =λ×batch yi +(1-λ)batch yj (2); Among them, × represents multiplication operation, batch xi 、batch xj and batch yi 、batch yj Represents ground penetrating radar images and corresponding labels, which are randomly selected training data; batch x ' and batch y 'represent the mixed GPR image samples and their corresponding labels, and λ represents the scaling factor between two randomly selected training data.
5. According to the method for identifying underground pipelines by ground penetrating radar based on improved YOLOv8 according to claim 1, it is characterized in that : In step 3, the improved YOLOv8 network model includes three parts: feature extraction network Backbone, feature fusion network Neck and detection head Head.
6. According to the method for identifying underground pipelines by ground penetrating radar based on improved YOLOv8 according to claim 5, it is characterized in that :The construction of the improved YOLOv8 network model includes the following steps: Step 3.1: Build the feature extraction network Backbone part of the improved YOLOv8 network model. The feature extraction network Backbone uses the CSPDarknet network structure and adds the CBAM attention mechanism to the feature extraction network Backbone. The CBAM attention mechanism structure includes input features, channel attention module, spatial attention module and output features; the input features are represented as F∈R C*H*W , where C represents the number of channels, H and W represent the height and width respectively; the channel attention module is implemented by a one-dimensional convolution M C ∈R C*1*1 The features are processed and then multiplied by the original features. Next, the output of the channel attention module is used as the input of the spatial attention module, which uses a two-dimensional convolution M S ∈R 1*H*W Further refine the features; finally, multiply the output of the spatial attention module with the original features to obtain the final feature representation; Step 3.2: Build the feature fusion network Neck part of the improved YOLOv8 network model. The feature fusion network Neck part is located between the feature extraction network Backbone and the detection head Head; Step 3.3: Build the detection head of the improved YOLOv8 network model. The detection head is responsible for multi-scale target recognition based on the feature maps extracted by the feature extraction network Backbone.
7. The underground pipeline identification method based on ground penetrating radar based on improved YOLOv8 according to claim 6 is characterized in that The improved YOLOv8 network model adopts a decoupled head structure, in which two parallel paths are responsible for extracting category features and spatial position features respectively, and each path finally realizes the classification and positioning functions through a 1×1 convolutional layer.
8. The underground pipeline identification method based on ground penetrating radar based on improved YOLOv8 according to claim 6, characterized in that: In step 3.1, the channel attention module first processes the input feature map in parallel through the maximum pooling (MaxPool) and the average pooling (AvgPool), and converts the size of the feature map from C*H*W to C*1*1; then, these feature maps pass through a multi-layer shared perceptron module, which first reduces the number of channels to the original 1 / r, and then expands back to the original number of channels. Through the ReLU activation function, two activated feature maps can be obtained; after superimposing the two feature maps, the channel attention map is generated by applying the sigmoid activation function; the map is then used to weight the original feature map to enhance important features and suppress unimportant parts, and finally restore it to the size of C*H*W; specifically expressed as: In formula (3), M c (F) represents the generated channel attention, MaxPool(F) represents the input features through the MaxPool layer, AvgPool(F) represents the input features through the AvgPool layer, MLP represents multi-layer perceptron, Represents the features of the maximum pooling layer on the channel axis, represents the features of the average pooling layer on the channel axis, W1 and W0 represent the weight matrices of MLP, and σ represents the Sigmoid activation function; The spatial attention module receives the output of the channel attention module and obtains two 1*H*W feature maps through maximum pooling and average pooling operations respectively; these two feature maps are then concatenated together and converted into a single channel feature map through a 7*7 convolution layer; then, the spatial attention map is generated using the sigmoid function; This spatial attention map is used to weight the original feature map to strengthen the key information and suppress the non-critical parts, and finally expand it back to the size of C*H*W; its formula can be expressed as: In formula (4), M s (F) represents the generated spatial attention, f 7*7 represents a 7×7 convolution operation, represents the maximum pooling feature map, Represents the average pooled feature map.
9. The underground pipeline identification method based on ground penetrating radar based on improved YOLOv8 according to claim 1, characterized in that: In step 3, the CIoU loss function of the YOLOv8 network model is replaced with the WIoU border loss function; the WIoU border loss function is defined as follows: L WIoU =rR WIoU L IoU (5); Where, L IoU represents the intersection-over-union loss function, R WIoU represents the normalized distance between the center point of the target box and the predicted box, and r represents the gradient gain; R WIoU The definition is as follows: Among them, W g Indicates the width of the minimum bounding box; H g Indicates the width and height of the minimum bounding box; (x gt ,y gt ) represents the coordinates of the center point of the real box; (x, y) represents the coordinates of the center point of the predicted box; ()* represents separation from the calculation graph, and will not be calculated during back propagation. g ×H g The gradient associated with a tensor of dimension; The gradient gain r is defined as follows: Where: β is the outlier degree, α and δ are two hyperparameters; The outlier degree β is defined as follows: in, Indicates L IoU Separate from the computational graph, represents a sliding average with momentum m.
10. The underground pipeline identification method based on ground penetrating radar based on improved YOLOv8 according to claim 1, characterized in that: In step 4, the detection head is used to predict the location and category of underground pipelines: The vertices of the hyperbola are used as positioning points to determine the location of underground pipelines. YOLOv8 has two detection heads, one for classification and the other for predicting bounding boxes. The underground pipeline location prediction task is added to the detection network as the third regression head, which consists of two parameters x and y, where x and y are the pixel values along the horizontal and vertical axes. The horizontal axis x is proportional to the actual distance along the measurement direction, and the vertical axis y is proportional to the two-way propagation time of the wave. In this way, the pipeline location can be obtained through the coordinates of the positioning points. The detection branch head uses the Wing-Loss loss function, which is expressed as: loss L (s)=∑wing(s-s′) (10); Among them, w represents the range of the nonlinear part of the function (-w, w), ε represents the curvature of the nonlinear region, C is a constant, s represents the positioning point vector, and Loss L (s) represents the loss function of the anchor point vector s, and s' represents the predicted coordinate vector.
Citation Information
Cited By
Red tide outbreak early warning method and system based on weak target detection
CN120472352A
Three-view underground pipeline intelligent detection method based on three-dimensional ground penetrating radar
CN120634986A
Concrete reinforcement parameter extraction method based on mixed loss AE and YOLOv8
CN122199553A