Method, system, electronic equipment and medium for detecting small underground targets by hyperbolic echo

By designing the SSL-YOLOv8 network model, integrating the downsampling convolution module and attention mechanism module, and using the dual-channel detection module, the problems of low detection accuracy and slow detection speed in small and medium-sized ground penetrating radar are solved, and higher detection accuracy and speed are achieved.

CN119131330BActive Publication Date: 2025-05-06CHANGSHA THUNDER TEST TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411351776.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-26
Publication Date
2025-05-06
Estimated Expiration
2044-09-26

AI Technical Summary

Technical Problem

In the field of ground penetrating radar technology, the existing target detection network structure is complex and has a deep layer, which leads to the loss of scattered hyperbolic echo characteristics of small targets, affecting the accuracy rate. In addition, traditional algorithms ignore the hyperbolic shape of the target in the GPR B-Scan image, and cannot effectively extract feature information, resulting in inaccurate target recognition and unsatisfactory detection speed.

Method used

A SSL-YOLOv8 network model was designed, and the feature fusion module was integrated into the feature fusion module and the attention mechanism module, and the dual-channel detection module was replaced by replacing the original detection module, which improved the feature extraction and fusion process and enhanced the extraction and capture capabilities of small-objective scattered hyperbolic echo features.

Benefits of technology

The detection accuracy and detection speed of small objects are improved, especially when processing complex GPR B-Scan images, the accuracy and detection speed are significantly improved, and the problems of low detection accuracy and slow detection speed of small objects in the prior art are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119131330B_ABST
    Figure CN119131330B_ABST
Patent Text Reader

Abstract

The present invention relates to a method, system, electronic device and medium for detecting underground small target hyperbolic echoes. The method comprises: obtaining a GPR B-Scan image of an underground small target area, constructing a data set and dividing it into a training set, a validation set and a test set; designing a downsampling convolution module SSDConv, an attention mechanism module SMRF and a dual-channel detection module LDCD and applying them on a YOLOv8 architecture to construct a small target detection model; training the model based on the training set to obtain a trained model; inputting the test set into the trained model to output a small target detection result. The present invention improves the detection accuracy and detection speed of small target scattered hyperbolic echoes in GPR B-Scan images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ground penetrating radar, and in particular to a method, system, electronic equipment and medium for detecting a small underground target hyperbolic echo. Background Art

[0002] With the continuous development of deep learning theory and the significant improvement of computing power, target detection technology has been widely used in many industrial fields, including industrial digital design, intelligent warehousing and autonomous driving. Target detection technology shows its unique advantages especially when processing highly complex image data. Among them, the YOLO algorithm, as a typical single-stage target detection algorithm, has performed well in industrial applications with its simple structure, high efficiency, multi-scale detection capabilities, excellent global context information utilization and multi-task learning performance.

[0003] In the field of object detection, common dataset image sizes are usually 640*640 or 320*320 pixels. Based on the absolute size of the target, a small target is usually defined as a target with an absolute size less than 32*32 pixels. Another definition based on relative scale is that when the side length of the target box is less than one tenth of the side length of the image, the target is considered a small target.

[0004] However, in the field of Ground Penetrating Radar (GPR) technology, especially when processing B-Scan images, target detection faces many complex challenges. First, the existing target detection network structure is complex and deep, resulting in the loss of scattered hyperbolic echo features of small targets during the convolution process. This is particularly evident in the detection of small targets whose target box side length is less than one-tenth of the image side length, thus affecting the precision rate. Secondly, traditional target detection algorithms usually ignore the hyperbolic shape commonly presented by targets in GPR B-Scan images, and cannot effectively extract feature information specific to GPR images, resulting in inaccurate target recognition. In addition, the detection speed of existing target detection networks is often not ideal when processing complex GPR B-Scan images. Summary of the invention

[0005] The present invention provides a method, system, electronic device and medium for detecting small underground targets by hyperbolic echoes, the purpose of which is to improve the detection accuracy of small targets, especially to effectively identify targets with hyperbolic shapes in GPR B-Scan images and to increase the processing speed in complex environments.

[0006] To achieve the above object, the first aspect of the present invention provides a method for detecting a small underground target by hyperbolic echo. The present invention includes:

[0007] Scan the underground area where the small target is buried along a one-dimensional survey line to obtain a B-Scan image;

[0008] Use annotation tools to annotate the small target scattering hyperbolic echo area in the B-Scan image, and divide the annotated B-Scan image into a training set, a validation set, and a test set in proportion;

[0009] An SSL-YOLOv8 network model is constructed, wherein the SSL-YOLOv8 network model is formed by sequentially connecting a feature extraction module, an improved feature fusion module, and an improved positioning recognition module; wherein a downsampling convolution module and an attention mechanism module are integrated into the feature fusion module of the original YOLOv8 network model to obtain an improved feature fusion module; a dual-channel detection module is used to replace the detection module in the original YOLOv8 network model to obtain an improved positioning recognition module; the downsampling convolution module includes two oblique edge extraction Sobel operators; the attention mechanism module includes a global branch, a large branch, and a local branch;

[0010] The training set is input into the SSL-YOLOv8 network model for iterative training. The validation set is used to evaluate the performance during the training process. When the SSL-YOLOv8 network model converges iteratively and the performance on the validation set reaches stability, a trained SSL-YOLOv8 network model is obtained. The test set is input into the trained SSL-YOLOv8 network model to output the small target detection result.

[0011] Furthermore, the method of inputting the test set into the trained SSL-YOLOv8 network model and outputting the small target detection result includes:

[0012] Use the feature extraction module to extract downsampled feature maps of different multiples from the B-Scan image;

[0013] The obtained down-sampled feature maps are sent to an improved feature fusion module, in which the down-sampled feature maps are merged and fused;

[0014] The fused down-sampled feature map is sent to the improved positioning and recognition module, and the dual-channel detection module is used to separate the fused down-sampled feature map by channel to determine the location and category of the target.

[0015] Furthermore, the feature extraction module includes a 10-layer structure, which is composed of a convolution layer, a feature extraction layer and a feature aggregation layer in sequence; wherein:

[0016] In the convolutional layers, the number of output channels, the convolutional kernel size and the step size of the convolutional kernels of the first convolutional layer to the fifth convolutional layer are set;

[0017] In the feature extraction layer, the number of repetitions and the number of output channels of the first feature extraction layer to the fourth feature extraction layer are set, and each feature extraction layer includes a bottleneck structure;

[0018] In the feature aggregation layer, the number of output channels and the pooling kernel size of the feature aggregation layer are set.

[0019] Furthermore, the improved feature fusion module includes a 14-layer structure, which is composed of an upsampling layer, a splicing layer, a feature extraction layer, a convolution layer, a downsampling convolution module and an attention mechanism module in sequence, wherein:

[0020] In the convolutional layer, the number of output channels, the convolutional kernel size and the step size of the convolutional kernel in the sixth convolutional layer to the seventh convolutional layer are set;

[0021] In the feature extraction layer, the number of repetitions and the number of output channels of the fifth feature extraction layer to the eighth feature extraction layer are set, and each feature extraction layer does not include a bottleneck structure.

[0022] Furthermore, two Sobel operators with fixed convolution kernel parameters are set in the downsampling convolution module, and the two Sobel operators have the same size but different directions.

[0023] Furthermore, the global branch includes branches using a dual-domain attention module and a spatial attention module; the large branch includes a first depth-separable convolution for obtaining a receptive field and a second depth-separable convolution for capturing long strip features; the local branch includes a third depth-separable convolution for extracting local information.

[0024] Furthermore, the dual-channel detection module adopts deep separable convolution to replace the original conventional convolution, and processes the classification and positioning tasks of small targets respectively through dual-channel separation.

[0025] Furthermore, the SSL-YOLOv8 network is composed of a feature extraction module Backbone, an improved feature fusion module Neck, and an improved positioning and recognition module Head connected in sequence. The feature extraction module Backbone is responsible for effectively extracting downsampled feature maps of different multiples from the B-Scan image. The improved feature fusion module Neck enhances the recognizability of the small target scattered hyperbolic echo features through a designed fusion strategy. The improved positioning and recognition module Head locates and recognizes small targets based on the enhanced small target scattered hyperbolic echo features.

[0026] To achieve the above object, the second aspect of the present invention further provides a system based on the underground small target hyperbolic echo detection method as described above, comprising the following modules:

[0027] A data acquisition module is used to scan the underground area where the small target is buried along a one-dimensional survey line to obtain a B-Scan image;

[0028] The data annotation and division module uses annotation tools to annotate the small target scattering hyperbolic echo area in the B-Scan image, generates an annotation file containing the coordinates and category information of each circumscribed rectangular box vertex, and divides the B-Scan image into training set, validation set and test set in proportion;

[0029] A model building module is used to build an SSL-YOLOv8 network model, wherein the SSL-YOLOv8 network model is formed by sequentially connecting a feature extraction module, an improved feature fusion module, and an improved positioning recognition module; wherein the downsampling convolution module and the attention mechanism module are integrated into the feature fusion module of the original YOLOv8 network model to obtain an improved feature fusion module; the detection module in the original YOLOv8 network model is replaced by a dual-channel detection module to obtain an improved positioning recognition module; the downsampling convolution module includes two oblique edge extraction Sobel operators; the attention mechanism module includes a global branch, a large branch, and a local branch;

[0030] A model training module is used to input the training set into the SSL-YOLOv8 network model for iterative training, and use the validation set to perform performance evaluation during the training process. When the SSL-YOLOv8 network model is iteratively converged and the performance on the validation set is stable, a trained SSL-YOLOv8 network model is obtained;

[0031] The model testing module inputs the test set into the trained SSL-YOLOv8 network model and outputs the small target detection result.

[0032] To achieve the above object, the third aspect of the present invention further provides an electronic device, comprising: at least one memory and at least one processor;

[0033] The at least one memory is used to store a readable program;

[0034] The at least one processor is used to call the readable program to execute the detection method.

[0035] To achieve the above-mentioned purpose, the fourth aspect of the present invention further provides a computer-readable medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the processor executes the detection method.

[0036] The beneficial effect of the present invention is that the present invention discloses a method for detecting the hyperbolic echo of underground small targets, which improves the detection effect of small targets such as those whose side length of the target frame is less than one tenth of the side length of the image to be detected. The original YOLOv8 network includes a feature extraction module, a feature fusion module and a positioning and recognition module. In the designed SSL-YOLOv8 network, the feature extraction module remains unchanged. In the improved feature fusion module, a downsampling convolution module and an attention mechanism module are designed, and the two modules are integrated into the improved feature fusion module. In the downsampling convolution module, two oblique edge extraction Sobel operators are designed to improve the extraction effect of the scattered hyperbolic echo features of small targets in the B-Scan image. In the attention mechanism module, a dual-domain attention module and a spatial attention module are designed to improve the capture ability of the scattered hyperbolic echo features of small targets. In the improved positioning and recognition module, a dual-channel detection module is designed to replace the original detection module, and a dual-channel separation mode is designed, and each channel is used to complete the positioning and recognition tasks of small targets respectively. Through the three modules designed above, after network training, the trained network achieves higher detection accuracy and detection speed for small targets in B-Scan. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0038] Figure 1 A flowchart of a deep learning network model designed for small target detection in B-Scan images by the present invention;

[0039] Figure 2 The present invention relates to a B-Scan image data set construction process;

[0040] Figure 3 The SSL-YOLOv8 network model structure diagram designed for the present invention;

[0041] Figure 4 is the parameter quantity of each layer of the SSL-YOLOv8 network model;

[0042] Figure 5 is the computational effort of each layer of the SSL-YOLOv8 network model;

[0043] Figure 6 It is a schematic diagram of the oblique Sobel operator parameters;

[0044] Figure 7The structure diagram of the attention mechanism module SMRF designed for the present invention;

[0045] Figure 8 (a) is the module structure diagram of the dual-domain attention module DA designed by the present invention, Figure 8 (b) is the structural diagram of the spatial attention module SA module;

[0046] Fig. 9 The structure diagram of the dual-channel detection module LDCD designed for the present invention;

[0047] Fig.10 This is a comparison chart of the actual detection differences between the original YOLOv8 network model and the SSL-YOLOv8 network model of the present invention for small target detection in B-Scan images. DETAILED DESCRIPTION

[0048] like Figure 1 As shown, the present invention provides a method for detecting underground small targets by hyperbolic echo, the method comprising the following steps:

[0049] Step S100, scanning the underground area where the small target is buried along a one-dimensional survey line to obtain a B-Scan image;

[0050] One-dimensional survey line refers to scanning along a fixed straight line on the ground during the detection process of ground penetrating radar (GPR). Usually, the GPR antenna moves along this survey line, sending and receiving radar signals point by point. By recording the echo data at each position, the corresponding underground image can be generated. The scanning result of the one-dimensional survey line can display the underground targets and their characteristic information along the line, and it only reflects the underground structure in a single dimension.

[0051] B-Scan image is a commonly used image representation in geological radar detection, also known as profile or scan image. It is a two-dimensional image generated by multiple sets of radar echo data collected along a one-dimensional survey line. The horizontal axis represents the horizontal displacement of the GPR antenna along the survey line (i.e. the scanning position), and the vertical axis represents the underground depth or the time delay of the signal. The echo data at each location point is arranged vertically to form a profile view reflecting the underground target and stratum. Through the B-Scan image, the echo signal of the underground target can be observed, especially some features such as the echo trajectory scattered by the target.

[0052] like Figure 2 As shown, measured data is collected by radar equipment to form a two-dimensional array, and the two-dimensional array is normalized and then drawn into a B-Scan grayscale image; in some implementations, the B-Scan grayscale image can be cropped to a desired size, for example, into a square image of 640*640 or 320*320 pixels, and finally constructed into a data set.

[0053] Step S200: annotate the small target scattering hyperbolic echo area in the B-Scan image using an annotation tool, and divide the annotated B-Scan image into a training set, a validation set, and a test set in proportion;

[0054] When the electromagnetic waves emitted by GPR encounter small targets underground (such as metal, rock, etc.), the target will reflect the echo signal. Due to the relative relationship between the movement of the antenna and the position of the target, the echo of the small target will appear as a hyperbolic trajectory on the B-Scan image. The reason for this is that the distance between the different positions of the GPR antenna and the target changes, causing the time delay of the echo signal to change, thus forming a hyperbola.

[0055] The data set consists of multiple B-Scan images containing small target scattered hyperbolic echoes, and the small targets in each B-Scan image are annotated. The annotation is to mark the coordinates of each vertex of each circumscribed rectangular box where the small target scattered hyperbolic echo area is located and the category it represents. First, use a specific annotation tool to accurately annotate the small target area in the B-Scan image. These small targets are usually shown as hyperbolic scattered echoes on the image. During the annotation process, it is necessary to draw the circumscribed rectangular box of each small target scattering area, and record the four vertex coordinates of the box and the corresponding target category. After the annotation is completed, the annotation information is generated into a annotation file, which contains the location information and classification information of all targets. Next, the B-Scan image data set is divided into a training set, a validation set, and a test set according to a specific ratio, such as a ratio of 3:1:1, for subsequent model training and evaluation.

[0056] Step S300, constructing an SSL-YOLOv8 network model, wherein the SSL-YOLOv8 network model is formed by sequentially connecting a feature extraction module, an improved feature fusion module, and an improved positioning recognition module; wherein the downsampling convolution module and the attention mechanism module are integrated into the feature fusion module of the original YOLOv8 network model to obtain an improved feature fusion module; the detection module in the original YOLOv8 network model is replaced by a dual-channel detection module to obtain an improved positioning recognition module; the downsampling convolution module includes two oblique edge extraction Sobel operators; the attention mechanism module includes a global branch, a large branch, and a local branch;

[0057] It can be understood that the original YOLOv8 network model includes a feature extraction module Backbone, a feature fusion module Neck, and a positioning and recognition module Head. In the designed SSL-YOLOv8 network model, the feature extraction module Backbone remains unchanged. In the improved feature fusion module Neck, the downsampling convolution module SSDConv (Sobel Spatial-to-Depth Convolution) and the attention mechanism module SMRF (Split Multi-scalar Receptive Field) are designed, and these two modules are integrated into the feature fusion module Neck. In the downsampling convolution module SSDConv, two oblique edge extraction Sobel operators are designed to improve the extraction effect of the scattered hyperbolic echo features of small targets in the B-Scan image. In the attention mechanism module SMRF, the dual-domain attention module DA (Dual-domain Attention) and the spatial attention module SA (Spatial Attention) are designed to improve the ability to capture the scattered hyperbolic echo features of small targets. In the improved positioning and recognition module Head, a dual-channel detection module LDCD (Layered Dual Channel Detection) is designed to replace the original detection module, and a dual-channel separation mode is designed, using one channel each to complete the positioning and recognition tasks of small targets.

[0058] Step S400, inputting the training set into the SSL-YOLOv8 network model for iterative training, using the validation set to perform performance evaluation during the training process, and obtaining a trained SSL-YOLOv8 network model when the SSL-YOLOv8 network model iteratively converges and the performance on the validation set reaches stability;

[0059] First, the prepared training set is input into the SSL-YOLOv8 network model, which learns the data in the training set through iteration. At each iteration, the SSL-YOLOv8 network model adjusts its internal parameters to reduce the prediction error, thereby better detecting small targets in B-Scan images. During the training process, the network weights are updated using optimization algorithms (such as gradient descent), and the performance of the model is measured by the loss function. In order to prevent overfitting and evaluate the generalization ability of the model, a validation set is used in each iteration to monitor the changes in model performance in real time. The training process is repeated until the loss function of the model tends to be stable, and indicators such as precision and recall are no longer significantly improved, that is, the model reaches a convergence state. Finally, an optimized and trained SSL-YOLOv8 network model that can accurately detect small targets is obtained.

[0060] Step S500: input the test set into the trained SSL-YOLOv8 network model and output the small target detection result.

[0061] In this step, the previously retained test set is input into the trained SSL-YOLOv8 network model. The model will make predictions for each test image and output the detection results, including the category and location of each small target (i.e., the coordinates of the circumscribed rectangular box), as well as the detection accuracy and confidence of the model. By comparing with the real annotated data of the test set, the performance of the model in small target detection is evaluated, and the actual detection effect of the model is measured by indicators such as the precision and recall of the detection results, thereby verifying the generalization ability of the model.

[0062] In step S500, the test set is input into the trained SSL-YOLOv8 network model, and the method of outputting the small target detection result includes:

[0063] Step S501, using a feature extraction module to extract down-sampled feature maps of different multiples from the B-Scan image;

[0064] Step S502: sending the obtained downsampled feature map to an improved feature fusion module, in which the downsampled feature map is merged and fused;

[0065] Step S503: Send the fused down-sampled feature map to the improved positioning and recognition module, and use the dual-channel detection module to separate the fused down-sampled feature map by channel to determine the location and category of the target.

[0066] Specifically, the method is to pass the P2-4 times downsampled feature map, that is, the downsampled feature map whose size becomes 2 to the power of 2 of the original image after 2 times downsampling, that is, 4 times, into the downsampling convolution module SSDConv, extract the information of the small target scattered hyperbolic echo feature, and then splice the above P2-4 times downsampled feature map with the P3-8 times downsampled feature map, P4-16 times downsampled feature map and P5-32 times downsampled feature map extracted by ordinary convolution, and obtain the composite P3-8 times downsampled feature map while keeping the number of detection branches unchanged. Then, this composite P3-8 times downsampled feature map is passed to the attention mechanism module SMRF, and two attention mechanisms are used on the global scale to capture valuable information in the image, and then the features of the three scales of global-large-small are integrated, emphasizing the fusion of the small target scattered hyperbolic echo features. Finally, the feature maps of three sizes, namely P3-8 times downsampling feature map, P4-16 times downsampling feature map and P5-32 times downsampling feature map after feature fusion, are passed to the dual-channel detection module LDCD, and the depthwise separable convolution DWConv (Depthwise SeparableConvolution) is used on two different channels for positioning and recognition respectively.

[0067] like Figure 3-Figure 5 As shown, the following will explain in detail the construction of the SSL-YOLOv8 network model for small target detection in GPR B-Scan images in three steps:

[0068] The first step is to add a downsampling convolution module SSDConv to enhance the extraction of small target scattered hyperbolic echo features. The working principle of the downsampling convolution module SSDConv is mainly the following three points:

[0069] (1) Sobel convolution with two fixed parameters

[0070] Since it is known that the target presents a downward-opening hyperbolic shape, the present invention designs two Sobel operators suitable for extracting oblique information. After their preliminary feature extraction, the feature map containing the hyperbolic feature information is passed to subsequent operations.

[0071] (2) Space to depth layer

[0072] The function of the spatial-to-depth layer is to compress the spatial dimensions (height and width) of the input feature map to the channel dimension. The specific method is to block the input feature map according to the given block size, and then stack the elements of each block along the channel dimension. This can retain all the information in the original feature map, but transfer it from the spatial dimension to the channel dimension.

[0073] (3) Non-strided convolutional layer

[0074] Non-stride convolution introduces a dilation rate between convolution kernel elements to increase the receptive field, which can preserve the spatial resolution of the feature map without losing the fine-grained information in the original feature map.

[0075] The second step is to fuse the information of the P2-4 times down-sampling feature map, the P3-8 times down-sampling feature map, the P4-16 times down-sampling feature map, and the P5-32 times down-sampling feature map to form a composite P3-8 times down-sampling feature map, and send it to the attention mechanism module SMRF to extract the small target scattered hyperbolic echo features. The attention mechanism module SMRF consists of three major branches:

[0076] (1) Large branch: This branch uses depthwise separable convolution DWConv to obtain a large receptive field. In order to further expand the receptive field, 1*K and K*1 depthwise separable convolution DWConv are also used to obtain strip-shaped feature information.

[0077] (2) Global branch: This branch uses the dual-domain attention module DA and the spatial attention module SA.

[0078] The dual-domain attention module DA first uses the fast Fourier transform FFT to convert the input feature map into the frequency domain, and then performs channel attention weighting on the frequency domain features through a 1*1 convolution layer, thereby realizing the global small target scattered hyperbolic echo feature extraction based on spectrum information. Next, the dual-domain attention module DA converts the adjusted frequency domain features back to the spatial domain through the inverse fast Fourier transform IFFT, and performs spatial attention weighting on the channel features through another 1*1 convolution layer. This process can further extract the spatial position information of the small target scattered hyperbolic echo features.

[0079] The purpose of the spatial attention module SA is to further perform frequency domain attention weighting on the feature map output by the dual-domain attention module DA. The spatial attention module SA first performs channel attention weighting on the features output by the dual-domain attention module DA through 1*1 convolution, and then converts it to the frequency domain using the fast Fourier transform FFT for spectral attention weighting. Finally, the spatial attention module converts the weighted frequency domain features back to the spatial domain through the inverse fast Fourier transform IFFT. Under the action of the dual attention mechanism in the frequency domain and the spatial domain, the network's ability to locate the scattered hyperbolic echo features of small targets in the GPR B-Scan image is further improved.

[0080] (3) Local branch: This branch uses 1*1 depthwise separable convolution to supplement local information.

[0081] The third step is to send the enhanced and fused P3-8 times downsampled feature map, P4-16 times downsampled feature map, and P5-32 times downsampled feature map to the dual-channel detection module LDCD. The dual-channel detection module LDCD uses a dual-channel separation method to process classification and regression tasks respectively, and replaces the conventional 3*3 convolution in the original network with a 3*3 deep separable convolution DWConv, which greatly reduces the number of parameters.

[0082] Preferably, the feature extraction module includes a 10-layer structure, which is sequentially: a first Conv layer, a second Conv layer, a first C2f layer, a third Conv layer, a second C2f layer, a fourth Conv layer, a third C2f layer, a fifth Conv layer, a fourth C2f layer, and an SPPF layer. The Conv layer is a convolutional layer, the C2f layer is a feature extraction layer, and the SPPF layer is a feature aggregation layer.

[0083] Among them, in the convolution layer (Conv layer), the output channel numbers of the convolution kernels from the first convolution layer to the fifth convolution layer are set to 64, 128, 256, 512 and 1024 respectively, the convolution kernel sizes are all 3*3, and the step sizes are all 2.

[0084] In the feature extraction layer (C2f layer), the number of repetitions from the first feature extraction layer to the fourth feature extraction layer is set to 3, 6, 6, and 3 respectively; the number of output channels is set to 128, 256, 512, and 1024 respectively; and each feature extraction layer contains a bottleneck structure to reduce parameters, increase network depth, and alleviate the gradient vanishing problem;

[0085] In the feature aggregation layer (i.e., SPPF layer), the number of output channels of the feature aggregation layer is set to 1024 and the pooling kernel size is set to 5.

[0086] It should be understood that bottleneck is a classic structure in convolutional neural networks, which aims to improve the efficiency of the model and avoid overfitting by reducing the number of parameters and calculations.

[0087] Preferably, the feature fusion module includes a 14-layer structure, which is sequentially: the first Upsample layer, the first Concat layer, the fifth C2f layer, the second Upsample layer, the downsampling convolution module SSDConv, the second Concat layer, the attention mechanism module SMRF, the sixth C2f layer, the sixth Conv layer, the third Concat layer, the seventh C2f layer, the seventh Conv layer, the fourth Concat layer, and the eighth C2f layer. The Upsample layer is an upsampling layer, the Concat layer is a splicing layer, the C2f layer is a feature extraction layer, and the Conv layer is a convolution layer, wherein:

[0088] The upsample layer increases the size of the input feature map by 2 times.

[0089] The concatenation layer (Concat layer) concatenates feature maps of the same size but different numbers of channels.

[0090] In the convolutional layer (Conv layer), the number of output channels of the convolution kernels in the sixth Conv layer to the seventh Conv layer are set to 256 and 512 respectively, the convolution kernel size is 3*3, and the step size is 2.

[0091] In the feature extraction layer (C2f layer), the number of repetitions of the fifth C2f layer to the eighth C2f layer are set to 3, 3, 3, and 3 respectively; the number of output channels are 512, 256, 512, and 1024 respectively; the difference is that the C2f layers here do not have bottlenecks.

[0092] The above Conv layer, C2f layer, SPPF layer, Upsample layer and Concat layer are consistent with the original YOLOv8 network model.

[0093] like Figure 6 As shown, the scattering curve of the small target in the GPR B-Scan image presents a hyperbolic shape with an opening downward. In this embodiment, an improved Sobel operator for extracting oblique edge features is designed to extract the scattering features of the small target. Different from the traditional Sobel operator for extracting edge information in the horizontal and vertical directions, this embodiment designs two oblique Sobel operators with fixed convolution kernel parameters of size 3*3, namely, the 45° direction Sobel operator and the 135° direction Sobel operator. A downsampling convolution module SSDConv is designed, which first uses the above two Sobel operators to preliminarily extract hyperbolic features, and then divides the obtained feature map into multiple sub-feature maps, which are spliced ​​along the channel dimension, thereby reducing the spatial dimension and increasing the channel dimension, which is used to replace the stride convolution layer and pooling layer in the traditional convolutional neural network.

[0094] like Figure 7 As shown in the figure, the attention mechanism module SMRF consists of three branches, namely the global branch, the large branch and the local branch, which learns the feature information from the local to the global, thereby improving the detection performance of the scattered hyperbolic echo of the small target. The large branch uses the depth-separable convolution DWConv with a convolution kernel parameter of 31*31 to obtain a large square receptive field, and combines the depth-separable convolution DWConv of 1*K and K*1 to capture strip information and obtain the strip receptive field without significantly increasing the computational overhead. The global branch uses the dual-domain attention module DA and the spatial attention module SA to enhance the global modeling capability. Figure 8As shown in the figure, the dual-domain attention module DA performs attention weighting on channel features in the frequency domain and spatial domain, while the spatial attention module SA performs frequency domain attention weighting on spatial features to capture the scattered hyperbolic echo features of small targets in the GPR B-Scan image. The local branch uses a 1*1 depthwise separable convolution DWConv to supplement local information. The attention mechanism module SMRF separates 25% of the channels from the composite P3-8 times down-sampling feature map after merging the P2-4 times down-sampling feature map and the P3-8 times down-sampling feature map, and performs multi-scale receptive field extraction operations.

[0095] like Fig. 9 As shown in the figure, the positioning and recognition module Head, that is, the dual-channel detection module LDCD, has three inputs, namely the composite P3-8 times down-sampled feature map, P4-16 times down-sampled feature map and P5-32 times down-sampled feature map. It uses a dual-channel separation method to process the positioning and recognition tasks respectively, and replaces the conventional 3*3 convolution in the original YOLOv8 with a 3*3 depth-separable convolution DWConv, which greatly reduces the number of parameters.

[0096] In this embodiment, after obtaining the trained SSL-YOLOv8 network model, the training parameter configuration table of this embodiment is shown in Table 1:

[0097] Table 1 Training parameter configuration table

[0098]

[0099] Input the test set into the trained small target detection model and output the small target detection results.

[0100] In order to verify the effectiveness of the three modules used in this experiment, this embodiment sets up an ablation experiment to analyze the improvement of the designed network on the small target detection performance.

[0101] In the YOLO series of models, the main indicators for evaluating its network performance are as follows: precision (Precision, P), recall (Recall, R), F1, meanAveragePrecision (mAP). This experiment uses mAP50 and mAP50:95 as performance reference indicators. mAP@0.5 and mAP@0.5:0.95 represent the mAP value when the IoU threshold is 0.5 and the average mAP value when the IoU starts from 50% and increases to 95% with a step size of 5%. The larger the average precision mAP, the higher the overall accuracy of the model. The calculation formulas for each indicator are as follows:

[0102] (1)

[0103] (2)

[0104] (3)

[0105] (4)

[0106] (5)

[0107] Among them, TP is the number of correctly predicted positive samples, FN is the number of incorrectly predicted negative samples, and FP is the number of incorrectly predicted positive samples. In addition to the above indicators, it is also necessary to introduce parameters (Params), computational complexity (Comps, unit: GFLOPs) and network detection image speed Fps (unit: pictures / second).

[0108] As shown in Table 2, this embodiment adopts ablation experiments to verify the effectiveness of the three modules designed by the present invention. The original YOLOv8 network and the SSL-YOLOv8 network with the three modules designed by the present invention added are trained on the same training set and tested on the same test set. The parameter quantity and calculation amount of the above two networks in the training process, as well as the precision rate, recall rate and other indicators in the test process are shown in Table 2 below, where the first row of Table 2 is the performance of the original YOLOv8 network, and the last row is the performance of this embodiment of the SSL-YOLOv8 designed by the present invention.

[0109] Table 2 Ablation experiment results

[0110]

[0111] As can be seen from Table 2, this embodiment comprehensively considers the low detection accuracy and slow detection speed of small target scattered hyperbolic echo in GPR B-Scan target detection, and designs the SSL-YOLOv8 network. Compared with the baseline YOLOv8 network model, the precision rate of this embodiment on the test set is improved by 1.4%, and the detection speed of GPR B-Scan images is increased by 15.3 images / second; the number of parameters on the training set is reduced by 15.8%, and the amount of calculation in the training process is increased by 12.3%.

[0112] Compared with the original YOLOv8 network, this network adds the downsampling convolution module SSDConv and the attention mechanism module SMRF, which improves the feature extraction effect of the small target scattered hyperbolic echo in the B-scan image. The dual-channel detection module LDCD designed in this network assigns positioning and recognition to two channels for separate processing, and uses the deep separable convolution DWConv to significantly reduce the network parameters. Compared with the original YOLOv8 network model, the SSL-YOLOv8 network model with three modules has fewer overall parameters. The attention mechanism SMRF is introduced in the second module feature fusion module, which makes the network take longer to train. After the network training is completed, the generated weight file is more concise than the original YOLOv8 network, and the detection speed of the small target scattered hyperbolic echo in the B-scan image is improved. This design improves the detection accuracy and speed of the network. Fig.10 This is a comparison chart of the actual detection difference. Fig.10 In the original YOLOv8 method, false detection and missed detection occurred.

[0113] This embodiment also provides a system based on the underground small target hyperbolic echo detection method, including:

[0114] A data acquisition module is used to scan the underground area where the small target is buried along a one-dimensional survey line to obtain a B-Scan image;

[0115] The data annotation and division module uses annotation tools to annotate the small target scattering hyperbolic echo area in the B-Scan image, generates an annotation file containing the coordinates and category information of each circumscribed rectangular box vertex, and divides the B-Scan image into training set, validation set and test set in proportion;

[0116] A model building module is used to build an SSL-YOLOv8 network model, wherein the SSL-YOLOv8 network model includes a feature extraction module, an improved feature fusion module and an improved positioning recognition module; wherein the downsampling convolution module and the attention mechanism module are integrated into the feature fusion module of the original YOLOv8 network model to obtain an improved feature fusion module; the detection module in the original YOLOv8 network model is replaced by a dual-channel detection module to obtain an improved positioning recognition module; the downsampling convolution module includes two oblique edge extraction Sobel operators; the attention mechanism module includes a global branch, a large branch and a local branch;

[0117] A model training module is used to input the training set into the SSL-YOLOv8 network model for iterative training, and use the validation set to perform performance evaluation during the training process. When the SSL-YOLOv8 network model is iteratively converged and the performance on the validation set is stable, a trained SSL-YOLOv8 network model is obtained;

[0118] The model testing module inputs the test set into the trained SSL-YOLOv8 network model and outputs the small target detection result.

[0119] Furthermore, this embodiment also provides an electronic device, including: at least one memory and at least one processor;

[0120] The at least one memory is used to store a readable program;

[0121] The at least one processor is used to call the readable program to execute the detection method.

[0122] This embodiment further provides a computer-readable storage medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the processor executes the detection method.

[0123] A person skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of protection of the present application is limited to these examples. In line with the concept of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of different aspects of one or more embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.

[0124] One or more embodiments of the present application are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the present application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present application should be included in the protection scope of the present application.

Claims

1. A method for detecting small underground targets by hyperbolic echo, characterized in that: The steps include: Scan the underground area where the small target is buried along a one-dimensional survey line to obtain a B-Scan image; Use annotation tools to annotate the small target scattering hyperbolic echo area in the B-Scan image, and divide the annotated B-Scan image into a training set, a validation set, and a test set in proportion; An SSL-YOLOv8 network model is constructed, wherein the SSL-YOLOv8 network model is formed by sequentially connecting a feature extraction module, an improved feature fusion module, and an improved positioning recognition module; wherein a downsampling convolution module and an attention mechanism module are integrated into the feature fusion module of the original YOLOv8 network model to obtain an improved feature fusion module; a dual-channel detection module is used to replace the detection module in the original YOLOv8 network model to obtain an improved positioning recognition module; the downsampling convolution module includes two oblique edge extraction Sobel operators; the attention mechanism module includes a global branch, a large branch, and a local branch; the global branch includes a dual-domain attention module and a spatial attention module; the large branch includes a first depth-separable convolution for obtaining a receptive field and a second depth-separable convolution for capturing long strip features; the local branch includes a third depth-separable convolution for extracting local information; The training set is input into the SSL-YOLOv8 network model for iterative training, and the validation set is used to perform performance evaluation during the training process. When the SSL-YOLOv8 network model is iteratively converged and the performance on the validation set is stable, a trained SSL-YOLOv8 network model is obtained; The test set is input into the trained SSL-YOLOv8 network model, and the small target detection result is output.

2. The method for detecting underground small targets by hyperbolic echo according to claim 1, characterized in that: The method of inputting the test set into the trained SSL-YOLOv8 network model and outputting the small target detection result includes: Use the feature extraction module to extract downsampled feature maps of different multiples from the B-Scan image; The obtained down-sampled feature maps are sent to an improved feature fusion module, in which the down-sampled feature maps are merged and fused; The fused down-sampled feature map is sent to the improved positioning and recognition module, and the dual-channel detection module is used to separate the fused down-sampled feature map by channel to determine the location and category of the target.

3. The method for detecting underground small targets by hyperbolic echo according to claim 1, characterized in that: The feature extraction module includes a 10-layer structure, which is composed of a convolution layer, a feature extraction layer and a feature aggregation layer in sequence; wherein: In the convolutional layers, the number of output channels, the convolutional kernel size and the step size of the convolutional kernels of the first convolutional layer to the fifth convolutional layer are set; In the feature extraction layer, the number of repetitions and the number of output channels of the first feature extraction layer to the fourth feature extraction layer are set, and each feature extraction layer includes a bottleneck structure; In the feature aggregation layer, the number of output channels and the pooling kernel size of the feature aggregation layer are set.

4. The method for detecting underground small targets by hyperbolic echo according to claim 1, characterized in that: The improved feature fusion module includes a 14-layer structure, which is composed of an upsampling layer, a splicing layer, a feature extraction layer, a convolution layer, a downsampling convolution module and an attention mechanism module, wherein: In the convolutional layer, the number of output channels, the convolutional kernel size and the step size of the convolutional kernel in the sixth convolutional layer to the seventh convolutional layer are set; In the feature extraction layer, the number of repetitions and the number of output channels of the fifth feature extraction layer to the eighth feature extraction layer are set, and each feature extraction layer does not include a bottleneck structure.

5. The underground small target hyperbolic echo detection method according to claim 1, characterized in that: Two Sobel operators with fixed convolution kernel parameters are set in the downsampling convolution module, and the two Sobel operators have the same size but different directions.

6. The method for detecting underground small targets by hyperbolic echo according to claim 1, characterized in that: The dual-channel detection module adopts deep separable convolution to replace the original conventional convolution, and processes the classification and positioning tasks of small targets respectively through dual-channel separation.

7. A system based on the underground small target hyperbolic echo detection method according to claim 1, characterized in that: Includes the following modules: A data acquisition module is used to scan the underground area where the small target is buried along a one-dimensional survey line to obtain a B-Scan image; The data annotation and division module uses annotation tools to annotate the small target scattering hyperbolic echo area in the B-Scan image, generates an annotation file containing the coordinates and category information of each circumscribed rectangular box vertex, and divides the B-Scan image into training set, validation set and test set in proportion; A model building module is used to build an SSL-YOLOv8 network model, wherein the SSL-YOLOv8 network model is formed by sequentially connecting a feature extraction module, an improved feature fusion module and an improved positioning recognition module; wherein the downsampling convolution module and the attention mechanism module are integrated into the feature fusion module of the original YOLOv8 network model to obtain an improved feature fusion module; the detection module in the original YOLOv8 network model is replaced by a dual-channel detection module to obtain an improved positioning recognition module; the downsampling convolution module includes two oblique edge extraction Sobel operators; the attention mechanism module includes a global branch, a large branch and a local branch; the global branch includes a dual-domain attention module and a spatial attention module; the large branch includes a first depth-separable convolution for obtaining a receptive field and a second depth-separable convolution for capturing long strip features; the local branch includes a third depth-separable convolution for extracting local information; A model training module is used to input the training set into the SSL-YOLOv8 network model for iterative training, and use the validation set to perform performance evaluation during the training process. When the SSL-YOLOv8 network model is iteratively converged and the performance on the validation set is stable, a trained SSL-YOLOv8 network model is obtained; The model testing module inputs the test set into the trained SSL-YOLOv8 network model and outputs the small target detection result.

8. An electronic device, characterized in that: include: at least one memory and at least one processor; The at least one memory is used to store a readable program; The at least one processor is used to call the readable program to execute the detection method according to any one of claims 1-6.

9. A computer-readable medium, characterized in that The computer readable medium stores computer instructions, and when the computer instructions are executed by a processor, the processor executes the detection method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Small target detection method, system and equipment based on improved YOLOv8

    CN118552935A

  • YOLOv8s traffic sign detection method based on edge and shape feature fusion

    CN118609092A