Anchor net detection and center hole coordinate identification method based on improved YOLOv8

By improving YOLOv8's anchor network detection method, using ADown downsampling module, C2f_Dual module, dynamic upsampler and SEAM attention mechanism, the anchor network detection model is optimized, and the problem of inaccurate anchor network positioning in coal mine tunnels is solved, high-precision anchor network center hole recognition is achieved, and construction efficiency is improved.

CN120279236APending Publication Date: 2025-07-08TAIYUAN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510364084.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Traditional object detection algorithms are prone to missed or missed detection when positioning anchor nets and central holes in coal mine tunnels, and it is difficult to achieve high-precision image recognition in complex environments.

Method used

Using the improved YOLOv8 anchor network detection method, the feature extraction and bounding box regression of the anchor network detection model are optimized by introducing ADown downsampling module, C2f_Dual module, dynamic upsampler and SEAM attention mechanism, and using the WIoU loss function.

Benefits of technology

It improves the accuracy and robustness of anchor net detection, can achieve high-precision anchor net center hole recognition under complex backgrounds and occlusions, and improves the construction efficiency of the anchor rod support process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279236A_ABST
    Figure CN120279236A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of target detection, and provides an anchor net detection and center hole coordinate identification method based on improved YOLOv8 in order to solve the problem that the positioning effect of an existing target detection algorithm on an underground anchor net is poor, an anchor net detection model is constructed on the basis of YOLOv8, a standard convolutional layer in a backbone network is replaced with an ADown down-sampling module, and a target detection model is constructed on the basis of the ADown down-sampling module. A C2f module is replaced by a C2fDual module, a dynamic up-sampler is adopted in a neck network, an attention mechanism is added in the neck network for feature extraction, an original loss function is replaced by a WI oU loss function in a head network, a historical coal mine tunnel anchor net sample data set is used for training an anchor net detection model, an optimal anchor net detection model is obtained, and the optimal anchor net detection model is obtained. And finally, inputting a to-be-detected anchor net image into the optimal anchor net detection model for detection to obtain an anchor net detection result and an anchor net center hole detection result, thereby realizing high-precision identification of the anchor net center hole.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target detection, and specifically relates to a method for detecting anchor nets and identifying the coordinates of central holes based on improved YOLOv8. Background Art

[0002] In China, coal mining mainly relies on underground operations, which leads to a rather heavy workload for the excavation and support of coal mine roadways. With the significant increase in the intensity, scale, and output of coal mining, the demand for roadway support has become increasingly strict. In this context, the bolt support technology has become an ideal choice for coal mine roadway support due to its high efficiency and economy. However, in the actual application process, key steps such as drilling, charging, and sending bolts still rely on manual operations, resulting in problems such as difficult anchor net positioning and low efficiency.

[0003] With the development of science and technology and the progress of society, target detection technology has developed rapidly and has been favored by more and more researchers. Currently, target detection technologies based on deep learning are mainly divided into two categories. One is the target detection method based on candidate boxes, such as R-CNN, Fast R-CNN, Faster R-CNN, Mask R-CNN; the other is the target detection method based on regression, such as SSD, YOLO series, Anchor Free, etc. Because the target detection algorithms based on deep learning have characteristics such as high detection accuracy and fast response, they have been more and more widely used in the coal mine field, such as the detection of belt deviation and tearing of coal mine belt conveyors, and the detection and positioning of the central holes of anchor nets in coal mine underground.

[0004] Currently, when traditional target detection algorithms are actually applied to coal mine roadways, problems such as missed detection or false detection will occur, and the positioning effect of underground anchor nets is not good. Therefore, it is necessary to design a method that can perform high-precision image recognition in complex environments. Summary of the Invention

[0005] In order to solve at least one of the above technical problems existing in the prior art, the present invention provides a method for detecting anchor nets and identifying the coordinates of central holes based on improved YOLOv8.

[0006] The present invention is implemented by adopting the following technical solutions: A method for detecting anchor nets and identifying the coordinates of central holes based on improved YOLOv8 includes the following steps:

[0007] Obtain a historical coal mine roadway anchor net sample data set, label the anchor nets and their central hole coordinates in the historical coal mine roadway anchor net sample data set, and divide the historical coal mine roadway anchor net sample data set into a training set, a validation set, and a test set according to a preset ratio;

[0008] Build an anchor net detection model based on YOLOv8. The anchor net detection model is based on YOLOv8. Replace the standard convolutional layer in the backbone network of YOLOv8 with an ADown downsampling module, replace the C2f module with a C2f_Dual module, adopt a dynamic upsampler in the neck network, and add an attention mechanism in the neck network for feature extraction. Replace the original loss function with a WIoU loss function in the head network;

[0009] Use the training set to train the anchor net detection model, and use the validation set for validation to obtain the optimal anchor net detection model. Finally, use the test set to test the detection performance of the optimal anchor net detection model;

[0010] Input the anchor net image to be detected into the optimal anchor net detection model for detection to obtain the anchor net detection result and the anchor net center hole detection result.

[0011] Preferably, the image feature extraction method of the ADown downsampling module is as follows:

[0012] First, perform average pooling on the first input feature map input to the ADown downsampling module, and then divide the first input feature map after average pooling into a first convolutional feature map and a first max-pooling feature map in the channel dimension. The first convolutional feature map performs a convolutional operation, and the first max-pooling feature map first performs a max-pooling operation and then a convolutional operation. Finally, the processed first convolutional feature map and the first max-pooling feature map are concatenated to obtain the ADown downsampling feature map.

[0013] Preferably, the image feature extraction method of the C2f_Dual module is as follows:

[0014] First, perform a convolutional operation on the second input feature map input to the C2f_Dual module, and then divide the second input feature map after the convolutional operation into two parts in the channel dimension. One part passes through an improved Bottleneck module, and the other part performs feature fusion with the feature map processed by the improved Bottleneck module. Finally, perform another convolutional operation to obtain the C2f_Dual module feature map. The improved Bottleneck module replaces the Conv in the original Bottleneck module with a DualConv module.

[0015] Preferably, the image feature extraction method of the dynamic upsampler is as follows:

[0016] First, process the third input feature map input to the dynamic upsampler through a sampling point generator to obtain a sampling set, and then process the sampling set through a grid sampling function to obtain the dynamic upsampler feature map.

[0017] Preferably, the method for obtaining the detection result of the center hole of the anchor net is as follows:

[0018] Using the anchor box and the bounding box prediction algorithm, generate candidate bounding boxes;

[0019] The optimal anchor net detection model adjusts the coordinates and sizes of the candidate bounding boxes according to the anchor net image to be detected, and obtains a center hole screening box;

[0020] Determine the center hole screening box with the highest confidence according to a preset threshold, and obtain the detection result of the center hole of the anchor net based on the center hole screening box with the highest confidence.

[0021] Compared with the prior art, the beneficial effects of the present invention are:

[0022] (1) The backbone network introduces an ADown downsampling module. As a convolutional block for downsampling operations, the ADown downsampling module can efficiently reduce the spatial dimensions (i.e., width and height) of the feature map, and at the same time may increase the number of channels to capture higher-level feature representations, which helps the anchor net detection model constructed by the present invention to capture the features of the image at a higher level while reducing the computational amount;

[0023] (2) Adding a C2f_Dual module replaces the C2f module. It helps to utilize the spatial feature extraction ability of large-size convolutional kernels and the computational efficiency of small-size convolutional kernels simultaneously, thereby reducing the number of model parameters and computational costs while maintaining accuracy;

[0024] (3) Adopting a dynamic upsampler in the neck network. The design of the dynamic upsampler is very lightweight, with fewer parameters and computational amounts, and by adopting a dynamic sampling mechanism, it can generate an upsampled feature map according to the local information of the input feature map, and can better retain feature details;

[0025] (4) Introducing the SEAM attention mechanism. The SEAM attention mechanism aims to improve the key point detection problem caused by occlusion in anchor net recognition. It compensates for the information loss of occluded key points by enhancing the response of unoccluded key points, making the anchor net detection model more focused on the detection of the anchor net;

[0026] (5) Replacing the original loss function with the WIoU loss function. As a bounding box regression loss, WIoU contains a dynamic non-monotonic mechanism and designs a gradient gain allocation strategy, which reduces the large gradients or harmful gradients that appear in extreme samples, enabling the anchor net detection model to pay more attention to the processing of ordinary-quality samples in the calculation, thereby improving the generalization ability of the anchor net detection model;

[0027] (6) Through the anchor net detection model, by fusing features from different network levels and global information, high-precision recognition of the center hole of the anchor net can be achieved, which has strong robustness for anchor net detection in complex backgrounds and occlusion situations, can realize the automatic recognition of the center hole of the anchor net during the bolt support process, and improve the construction efficiency of the bolt support process. Description of the Drawings

[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0029] Figure 1 It is a schematic flowchart of the anchor net detection and center hole coordinate recognition method based on the improved YOLOv8 in the present invention;

[0030] Figure 2 It is a structural diagram of the improved YOLOv8 neural network in the present invention;

[0031] Figure 3 It is a structural diagram of the Adown downsampling module in the present invention;

[0032] Figure 4 It is a structural diagram of the C2f_Dual module in the present invention;

[0033] Figure 5 It is a structural diagram of the DySample upsampling module in the present invention;

[0034] Figure 6 It is a structural diagram of the SEAM attention mechanism in the present invention;

[0035] Figure 7 It is a comparison diagram of the anchor net detection effects of the improved model at the bolt support site in the present invention;

[0036] Figure 8 It is an effect diagram of the anchor net center hole coordinate recognition of the improved model at the bolt support site in the present invention. Detailed Embodiment

[0037] Combined with the drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other implementation manners obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present invention.

[0038] It should be noted that the structures, proportions, sizes, etc. shown in the attached drawings of this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the qualified conditions for the implementation of the present invention. Therefore, they do not have substantial technical significance. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should fall within the scope covered by the technical content disclosed in the present invention. It should be noted that in this specification, relational terms such as first and second are only used to distinguish one entity from several other entities, and do not necessarily require or imply any actual relationship or order between these entities.

[0039] Aiming at the problems of inaccurate detection of anchor nets and inaccurate identification of the center hole coordinates of anchor nets in the operation of bolt support in coal mines, traditional object detection algorithms have poor effects when dealing with the positioning of anchor nets and the positioning of their center holes, and are prone to missed detection or false detection. Due to the complex working environment in coal mine roadways, traditional object detection algorithms are difficult to meet the coal mine construction standards in terms of real-time performance and accuracy. Therefore, the present invention provides an anchor net detection and center hole coordinate identification method based on improved YOLOv8 that can perform anchor net detection and center hole coordinate identification in real time and accurately.

[0040] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0041] The configuration environment used in each embodiment of the present invention is as follows: The operating system is Windows 10, the CPU is Intel(R) Core(TM) i7-9750H CPU@2.60GHz, the memory is 16G, the GPU is NVIDIA GeForce GTX 1650, and the video memory is 4G; the programming language is Python 3.8, the deep learning framework is Pytorch 1.12.1, and CUDA 10.2 is used to accelerate the GPU. In actual applications, the present invention can be applied to downhole drilling and bolting robots, and the present invention is not limited thereto.

[0042] As Figure 1 shown, the embodiment of the present invention provides an anchor net detection and center hole coordinate identification method based on improved YOLOv8, including the following steps:

[0043] S1: Obtain a historical coal mine roadway anchor net sample data set, label the anchor nets and their center hole coordinates in the historical coal mine roadway anchor net sample data set, and divide the historical coal mine roadway anchor net sample data set into a training set, a validation set, and a test set according to a preset ratio.

[0044] In this embodiment, the environment of a coal mine roadway and the placement of the anchor net are simulated in the laboratory. An anchor net image is captured by an industrial camera MV-CS050-10UC with 5 million pixels from Hikvision, equipped with an industrial lens MVL-KF0618M-12MPE with 12 million pixels. Meanwhile, to better simulate the complex environment underground in a coal mine roadway, gneiss, sandstone, limestone, granite, basalt, soil blocks, and coal blocks are prepared as the background of the anchor net. The captured images are preprocessed, and the preprocessing methods include, but are not limited to, adding Gaussian noise and salt-and-pepper noise to the captured images, and performing data augmentation operations such as flipping, rotating, scaling, and mirroring on the images, to obtain a historical coal mine roadway anchor net sample dataset. Then, the anchor nets and their center hole coordinates in the historical coal mine roadway anchor net sample dataset are labeled, and the historical coal mine roadway anchor net sample dataset is divided into a training set, a validation set, and a test set according to a preset ratio. During the actual experiment process, due to site limitations, there are 306 anchor net pictures in the captured images, and there are 2,745 image samples in the historical coal mine roadway anchor net sample dataset obtained after preprocessing. Then, the Imagelabels software is used for labeling, and it is divided into a training set, a validation set, and a test set according to the ratio of 8:1:1.

[0045] S2: Construct an anchor net detection model based on YOLOv8. The anchor net detection model is based on YOLOv8. Replace the standard convolutional layer in the backbone network of YOLOv8 with an ADown downsampling module, replace the C2f module with a C2f_Dual module, use a dynamic upsampler in the neck network, and add an attention mechanism in the neck network for feature extraction. Replace the original loss function with a WIoU loss function in the head network.

[0046] In this embodiment, as Figure 2 shown, the backbone network consists of layers 0 to 9. The input picture pixel size is 640×640×3. It is processed by a convolutional layer configured with a 3×3 convolutional kernel, a stride of 2, and a padding of 1, and the output size becomes 320×320×64. Subsequently, the backbone network alternately uses 4 ADown downsampling modules that process the input corresponding feature maps through multi-path processing and a C2f_Dual module that combines 3×3 and 1×1 convolutional kernels and group convolution technology. This alternating processing process finally obtains a feature map of 20×20×1024 pixels. Finally, through an SPPF module, which converts the input feature map into a fixed-size feature vector through multiple downsamplings and pooling operations at different scales, ensuring that the backbone network can process input images of any size, and finally outputs a feature map of 20×20×1024.

[0047] The neck network part consists of 10 to 24 layers, establishing a feature pyramid structure of FPN+PAN. The traditional upsampling layer is replaced by a dynamic upsampler. The output result of the dynamic upsampler is concatenated with the output of the 6th layer C2f_Dual module in the backbone network, and then an output with a size of 40×40×1536 is generated. Then, through the processing of the C2f_Dual module and the DySample dynamic upsampler, a feature map of 80×80×512 is generated, which is then concatenated with the output of the 4th layer C2f_Dual module in the backbone network to obtain an output size of 80×80×768. The processing of this size is divided into two paths: one is the feature extraction through the C2f_Dual module and the attention mechanism, and the recognized target pixel size of 80×80×256 is output and sent to the detection head part; the other path obtains a feature map of 40×40×256 through the Conv operation, and then performs a Concat operation with the output of the 12th layer C2f_Dual module of the neck network to obtain a feature map of 40×40×256, which is input into the C2f_Dual module and undergoes feature extraction through the attention mechanism, and finally a picture of 40×40×512 is output to the detection head; at the same time, the output feature map of 40×40×512 also undergoes a Conv operation and is concatenated with the output of the ninth layer in the backbone network to generate a feature map of 20×20×1536, and finally through the processing of the C2f_Dual module and the SEAM attention mechanism, a feature map of 20×20×1024 is obtained, and the features of each size of the fused and strengthened anchor network are input into the Decoupled Head detection head part.

[0048] The Decoupled Head detection head receives three fused feature maps of different sizes respectively. Each fused feature map will be operated in two branches, and information is extracted through 2 3x3 convolutions and a 1x1 two-dimensional convolution respectively, and finally Bbox.loss and Cls.loss are calculated separately; the two branches include a regression branch and a classification branch. The regression branch is responsible for calculating the position difference between the predicted box and the true box, and inputting these differences into the regression head for loss evaluation, and finally outputting a four-dimensional vector, which represents the upper left corner (x, y) and the lower right corner (x, y) coordinates of the target box, and the loss function CIoU of the regression branch is replaced with the loss function WIoU; the classification branch, for each candidate box extracted by the Anchor Free method, generates a classifier output tensor through region of interest pooling and convolution processing, where the value at each position reflects the probability that the corresponding candidate box belongs to each category, and finally the final detection result is selected from all candidate boxes through the non-maximum suppression technique.

[0049] Optionally, the image feature extraction method of the ADown downsampling module is as follows: First, perform average pooling on the first input feature map input to the ADown downsampling module, and then divide the first input feature map after average pooling into a first convolutional feature map and a first max-pooling feature map in the channel dimension. The first convolutional feature map performs a convolutional operation, and the first max-pooling feature map first performs a max-pooling operation and then a convolutional operation. Finally, the processed first convolutional feature map and the first max-pooling feature map are concatenated to obtain the ADown downsampling feature map.

[0050] In this embodiment, there are 5 convolutional layers in the backbone network of the original YOLOv8 network. The latter 4 are replaced with the ADown downsampling module. By using the ADown downsampling module to improve the feature extraction ability of the anchor net detection model, it can better capture the complex features of the input image and improve the recognition ability of the anchor net detection model for the anchor net.

[0051] In this embodiment, as Figure 3 shown in the network structure of the ADown downsampling module, first perform average pooling on the first input feature map, and then divide it into two parts, namely the first convolutional feature map and the first max-pooling feature map, in the channel dimension. The first convolutional feature map performs a convolutional operation, and the first max-pooling feature map first performs a max-pooling operation and then a convolutional operation. Finally, a concatenation operation is performed on the processing results of the two parts to output the final ADown downsampling feature map. This module reduces the spatial dimension of the feature map and adopts a multi-path processing strategy to efficiently extract image features while reducing the computational amount of the network.

[0052] Optionally, the image feature extraction method of the C2f_Dual module is as follows: First, perform a convolutional operation on the second input feature map input to the C2f_Dual module, and then divide the second input feature map after the convolutional operation into two parts in the channel dimension. One part passes through the Bottleneck module, and the other part performs feature fusion with the feature map processed by the Bottleneck module. Finally, perform another convolutional operation to obtain the C2f_Dual module feature map.

[0053] In this embodiment, the C2f_Dual module includes at least two convolutional operations and the Bottleneck module; the DualConv module is used to replace the Conv in the original Bottleneck module.

[0054] In this embodiment, there are 8 C2f modules in the original YOLOv8 network, all of which are replaced by C2f_Dual modules. The C2f_Dual module uses DualConv to replace the Conv in the Bottleneck module. The C2f_Dual module includes two convolutional operations and multiple Bottleneck modules. Through splicing and convolutional operations, the aggregation and compression of the second input feature map are achieved. Each Bottleneck module contains multiple Convs and DualConvs, and the use of residual connections can be configured. The design of the C2f_Dual module helps the anchor network detection model to utilize the spatial feature extraction ability of large-size convolutional kernels and the computational efficiency of small-size convolutional kernels simultaneously, thus maintaining the detection accuracy while reducing the number of parameters and computational cost of the anchor network detection model.

[0055] In this embodiment, as Figure 4 shown in the schematic diagram of the C2f_Dual module structure, the DualConv design is introduced into the C2f_Dual module. The DualConv combines the 3×3 and 1×1 convolutional kernels. The 3×3 convolutional kernel is used to extract the spatial features of the second input feature map, while the 1×1 convolutional kernel is used to integrate these features and reduce the parameters of the model. At the same time, the C2f_Dual module incorporates the group convolution technique. The corresponding input and output feature maps are divided into multiple groups, and the convolutional filters in each group only process a part of the corresponding input feature map, reducing the complexity of the model. DualConv uses this technique to further reduce the computational cost. It allows different convolutional kernels within a group to process the same group of input channels in parallel, optimizing the information flow and feature extraction efficiency while maintaining the representational ability of the network.

[0056] In this embodiment, the second input feature map transmitted to the C2f_Dual module first undergoes channel expansion through a 1×1 standard convolutional layer. The convolutional kernel, in conjunction with the batch normalization layer (BN) and the SiLU activation function, forms a feature enhancement unit to double the number of input channels. The expanded feature map is divided into two parallel data streams along the channel dimension through a channel splitting operation: the main path retains the first part of the channels and directly reaches the final fusion node, while the branch path inputs the other part of the channels into n improved Bottleneck modules for in-depth feature learning. The Bottleneck module combines a standard convolutional unit and a dual-modal DualConv unit. The standard convolutional unit uses a 3×3 convolutional kernel to extract spatial features, and the dual-modal DualConv unit adopts a collaborative convolution mechanism to achieve feature fusion through parallel processing of grouped convolution and point convolution: the grouped convolutional layer divides the input features into g independent subspaces, and each group of convolutional filters only processes the feature information of the corresponding group. Its output is element-wise added to the global features of the 1×1 point convolution. The Bottleneck module provides a dynamically configurable residual connection mechanism, which controls two feature processing modes through the boolean parameter shortcut: when shortcut = True and the input / output channels are equal, the residual enhancement mode is executed, and the input of the Bottleneck module is element-wise added to the output of the dual-modal DualConv; otherwise, the direct processing mode is executed, and the feature stream sequentially passes through the standard convolution and the dual-modal DualConv units. This dual-mode design not only ensures the feature integrity of the shallow network but also enhances the gradient propagation ability of the deep network. The feature maps processed by n levels of Bottlenecks are concatenated with the main path features in the channel dimension to form fusion features with cross-level perception ability. Finally, feature recalibration is achieved through a 1×1 convolutional kernel, and the dimensional compression of the feature space is completed in conjunction with the batch normalization layer and the SiLU activation function.

[0057] Optionally, the image feature extraction method of the dynamic upsampler is as follows: First, the third input feature map input to the dynamic upsampler is processed by a sampling point generator to obtain a sampling set, and then the sampling set is processed by a grid sampling function to obtain the dynamic upsampler feature map.

[0058] In this embodiment, the DySample dynamic upsampler is used to replace the Upsample module in the original YOLOv8 network. The network structure of the DySample dynamic upsampler is as Figure 5 shown.

[0059] In this embodiment, the size of the third input feature map X is set to C×H1×W1. After being processed by a Sampling Point Generator, a Sampleset with a size of 2g×sH×sW is output. Finally, a dynamic upsampling feature map with a size of C×sH×sW is output through the grid sample function, where s is the upsampling scale factor, and 2g represents the x and y coordinates of the third input feature map X.

[0060] In this embodiment, the DySample dynamic upsampler realizes the improvement of the spatial resolution of the corresponding feature map through a dynamic sampling mechanism. Its working process can be described as follows: The input feature map X (size C×H1×W1) first enters the sampling point generator. This generator learns the spatial offset through a 1×1 convolutional network, combines the initialized basic grid coordinates (uniformly distributed in the interval [-0.25, 0.25]), and generates a coordinate set S containing 2g groups of dynamic sampling points (size 2g×sH×sW, where s is the upsampling scale factor, and 2g corresponds to the x / y bidirectional offset of each group of features); Subsequently, the grid_sample function is used to perform deformable interpolation on the original features, mapping the learned offset coordinates to the high-resolution space enlarged by s times, and finally outputting the feature map X' (size C×sH×sW). The DySample dynamic upsampler supports two processing paths through the "style" parameter (LP samples first and then reorganizes, PL reorganizes the channels first and then samples), which can flexibly adapt to the requirements of different tasks for local fine-grained or global consistency while maintaining the channel-space balance.

[0061] Optionally, the attention mechanism is the SEAM attention mechanism, and the SEAM attention mechanism is added to the 15th, 19th, and 23rd layers of the neck network respectively.

[0062] In this embodiment, the network structure of the SEAM attention mechanism is as Figure 6 shown. The SEAM attention mechanism combines depthwise separable convolution and residual connection, uses patches of different sizes to extract multi-scale features, and learns the correlation between the spatial dimension and channels through depthwise separable convolution, enabling SEAM to more effectively handle occlusion problems and improving the accuracy of anchor network recognition.

[0063] In this embodiment, the input feature map passing through the SEAM attention mechanism enters the CSMM modules with three different receptive fields (patch = 6 / 7 / 8) in parallel. Each module uses depthwise separable convolution kernels of corresponding sizes, and performs multi-scale spatial feature extraction through depthwise convolution combined with GELU activation and batch normalization, while keeping the channel dimension unchanged. The features of the three branches are fused through channel concatenation in the embedding layer, and then compressed to the original number of channels through pointwise convolution to complete cross-scale feature interaction. The fused features enter a depth convolution sequence composed of multiple residual blocks: in the first layer of each residual block, spatial features are refined through depthwise convolution, and after GELU activation and batch normalization, the original input is superimposed to form a residual connection; in the second layer, the channel expression ability is enhanced through pointwise convolution, and local feature optimization is completed again through GELU and batch normalization. This process is gradually strengthened through n times of stacking. The global context information is obtained by compressing through adaptive average pooling, and after passing through the flatten layer, a channel attention vector is generated through a bottleneck structure composed of a double-layer fully connected layer. Finally, the significant channel weights are amplified through exponential operation and multiplied with the original input feature map at the channel level, enabling the network to dynamically focus on the key feature channels. The entire process retains the original information through residual connections, captures spatial details through multi-scale convolution, and realizes feature selection through channel attention, forming an end-to-end adaptive feature enhancement mechanism.

[0064] In this embodiment, the calculation formula of the WIoU loss function is:

[0065]

[0066] L IoU = 1 - IoU

[0067]

[0068] In the formula, L WIoU represents the WIoU loss, L IoU represents the bounding box loss function IOU, R WIoU represents the distance attention, β represents the outlier degree, α and δ are both hyperparameters, exp is the base of the natural logarithm, x and y respectively represent the horizontal and vertical coordinates of the center point of the predicted box, x gt and y gt are respectively the horizontal and vertical coordinates of the center point of the ground truth box, W g and H g are respectively the width and height of the minimum bounding rectangle of the predicted box and the ground truth box, * represents the separate operation, represents the monotonic focusing coefficient, represents the mean value of L IoU .

[0069] In this embodiment, the WIoU loss function introduces the concept of "outlier degree" and designs a gradient gain allocation strategy. Specifically, the larger the outlier degree, the lower the quality of the anchor box. Therefore, a larger gradient gain is assigned to the anchor box with a smaller "outlier degree", and a smaller gradient gain is assigned to the anchor box with a larger "outlier degree". This strategy not only suppresses the excessive proportion of high-quality anchor boxes but also reduces the harmful gradients generated by low-quality samples, thereby reducing the negative impact of these low-quality samples on the anchor box regression process.

[0070] S3: Train the anchor network detection model using the training set, and verify it using the validation set to obtain the optimal anchor network detection model. Finally, test the detection performance of the optimal anchor network detection model using the test set.

[0071] In this embodiment, ablation experiments are conducted using the validation set. A represents adding the ADown module, B represents adding the C2f_Dual module, C represents adding the DySample upsampler, D represents adding the SEAM attention mechanism, and E represents adding the WIoU loss function. The comparison results are shown in Table 1 below.

[0072] Table 1

[0073]

[0074] In Table 1, mAP is the mean average precision, and FPS is the frames per second.

[0075] From the above experimental data, it can be seen that after introducing the ADown module into the original YOLOv8 model, both the number of parameters and the computational complexity decrease. At the same time, the recognition accuracy improves by 1.4 percentage points compared to the baseline model. However, only introducing the ADown module causes the frames per second (FPS) to drop by 12, and its real-time detection speed cannot meet industrial requirements; on this basis, after introducing the C2f_Dual module, the number of parameters and the computational complexity of the model continue to show a decreasing trend. Although this improvement slightly reduces the recognition accuracy, the FPS increases by 7, and the real-time detection performance is improved; after continuing to add the DySample upsampler, it is found that its impact on the number of parameters and the computational complexity is minimal, almost remaining unchanged. At the same time, the recognition accuracy slightly decreases, but the FPS increases by 3 again, further enhancing the real-time detection performance; then adding the SEAM attention mechanism, although the number of parameters and the computational complexity increase, the recognition accuracy rises by 0.2 percentage points, and the FPS increases by 2; finally, replacing the loss function with WIoU, the number of parameters and the computational complexity remain unchanged, the recognition accuracy rises by 0.2 percentage points again, and the FPS increases by 2. Through this ablation experiment, it shows that the anchor network detection model not only maintains a high recognition accuracy but also enhances the real-time detection performance.

[0076] S4: Input the anchor net image to be detected into the optimal anchor net detection model for detection to obtain the anchor net detection result and the anchor net center hole detection result.

[0077] Optionally, the method for obtaining the anchor net center hole detection result is as follows: Use the anchor point box and the bounding box prediction algorithm to generate candidate bounding boxes; the optimal anchor net detection model adjusts the coordinates and sizes of the candidate bounding boxes according to the anchor net image to be detected to obtain the center hole screening boxes; determine the center hole screening box with the highest confidence according to the preset threshold, and obtain the anchor net center hole detection result based on the center hole screening box with the highest confidence.

[0078] In this embodiment, a series of candidate bounding boxes are generated by using the anchor point box and the bounding box prediction algorithm. The model will adjust the coordinates and sizes of these bounding boxes according to the information extracted from the feature map to make them more accurately enclose the target object. On this basis, the anchor net detection model will further calculate the objectness score of each prediction box and screen out the center hole screening boxes with high confidence according to the preset threshold. Finally, the screened center hole screening boxes are output as the final detection results, including the coordinates of the bounding boxes, the class labels, and the confidence scores.

[0079] To further verify the beneficial effects of the anchor net detection model constructed by the present invention, a test set is used for testing, and the test results are as Figure 7 shown. Figure 7 The left image in it is the anchor net detection result of the original YOLOv8 network. Figure 7 The right image in it is the anchor net detection result of the anchor net detection model constructed by the present invention. It can be clearly seen from this figure that the original YOLOv8 network has a missed detection phenomenon in the process of anchor net recognition, indicating that in complex coal mine roadways, the detection performance of the original YOLOv8 network for anchor nets is limited. In contrast, the detection accuracy of the anchor net detection model constructed by the present invention for anchor nets has been improved, effectively avoiding the occurrence of missed detection situations. It not only enhances the robustness of the algorithm but also improves the safety and efficiency of the bolt support operation in coal mine roadways.

[0080] Figure 8 As shown, during the bolt support process, after accurately positioning the anchor net by using the anchor net detection model constructed by the present invention, the system will generate a center hole screening box, and each corresponding center hole screening box gives the coordinates of the anchor net center hole, that is, the position of the center hole of the anchor net in the pixel coordinate system.

[0081] The present invention provides a method for anchor net detection and center hole coordinate recognition based on improved YOLOv8. Compared with the original model, the number of parameters has decreased by 15.3%, the computational volume has been reduced by 15.9%, the mAP has increased by 1.1%, and the frame rate has increased by 2. The standard convolutional layer is replaced with the ADown downsampling module to capture higher-level anchor net feature representations by reducing the spatial dimension of the feature map while increasing the number of channels. The C2f_Dual module is designed by introducing DualConv to replace the C2f module, integrating the techniques of large and small convolutional kernels and group convolution to improve the network's representation ability. The DySample dynamic upsampler is used in the neck network to better reconstruct the high-resolution anchor net feature map. At the same time, the added SEAM attention mechanism combines the advantages of depthwise separable convolution and residual connection, and can more efficiently handle occlusion problems, thereby improving the accuracy of anchor net recognition. The original loss function is optimized and replaced at the detection head to provide the generalization performance of the model. By introducing and optimizing each module, the anchor net detection model shows more excellent performance in extracting contour details and anchor net features, reducing the risk of complex environment interference and feature loss. The anchor net detection model of the present invention realizes the dual improvement of recognition accuracy and real-time detection performance while maintaining a low number of parameters and computational volume.

[0082] As mentioned above, the above is only the preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. An anchor net detection and center hole coordinate recognition method based on improved YOLOv8, characterized in that, It includes the following steps: Obtain a historical coal mine roadway anchor net sample dataset, label the anchor nets and their central hole coordinates in the historical coal mine roadway anchor net sample dataset, and divide the historical coal mine roadway anchor net sample dataset into a training set, a validation set, and a test set according to a preset ratio; Construct an anchor net detection model based on YOLOv8. The anchor net detection model is based on YOLOv8. Replace the standard convolutional layer in the backbone network of YOLOv8 with an ADown downsampling module, replace the C2f module with a C2f_Dual module, adopt a dynamic upsampler in the neck network, and add an attention mechanism in the neck network for feature extraction. Replace the original loss function with a WIoU loss function in the head network; Use the training set to train the anchor net detection model, and use the validation set for validation to obtain an optimal anchor net detection model. Finally, use the test set to test the detection performance of the optimal anchor net detection model; Input the anchor net image to be detected into the optimal anchor net detection model for detection to obtain the anchor net detection result and the anchor net central hole detection result.

2. The method for anchor net detection and center hole coordinate recognition based on improved YOLOv8 according to claim 1, characterized in that, The image feature extraction method of the ADown downsampling module is as follows: First, perform average pooling on the first input feature map input to the ADown downsampling module. Then, divide the first input feature map after average pooling into a first convolutional feature map and a first max-pooling feature map in the channel dimension. The first convolutional feature map performs a convolutional operation. The first max-pooling feature map first performs a max-pooling operation and then a convolutional operation. Finally, splice the processed first convolutional feature map and the first max-pooling feature map to obtain an ADown downsampling feature map.

3. The method for detecting anchor nets and identifying the coordinates of the central hole based on the improved YOLOv8 according to claim 1, wherein, The image feature extraction method of the C2f_Dual module is as follows: First, perform a convolutional operation on the second input feature map input to the C2f_Dual module. Then, divide the second input feature map after the convolutional operation into two parts in the channel dimension. One part passes through an improved Bottleneck module, and the other part performs feature fusion with the feature map processed by the improved Bottleneck module. Finally, perform another convolutional operation to obtain the C2f_Dual module feature map. The improved Bottleneck module replaces the Conv in the original Bottleneck module with a DualConv module.

4. The method for anchor net detection and center hole coordinate recognition based on improved YOLOv8 according to claim 1, wherein, The image feature extraction method of the dynamic upsampler is as follows: First, process the third input feature map input to the dynamic upsampler through a sampling point generator to obtain a sampling set. Then, process the sampling set through a grid sampling function to obtain a dynamic upsampler feature map.

5. The method for detecting anchor nets and identifying the coordinates of the central hole based on the improved YOLOv8 according to claim 1, wherein The method for obtaining the anchor net central hole detection result is as follows: Use an anchor box and a bounding box prediction algorithm to generate candidate bounding boxes; The optimal anchor net detection model adjusts the coordinates and sizes of the candidate bounding boxes according to the anchor net image to be detected to obtain a central hole screening box; Determine the central hole screening box with the highest confidence according to a preset threshold, and obtain the anchor net central hole detection result based on the central hole screening box with the highest confidence.

Citation Information

Cited By

  • Online monitoring system for automatic operation of distribution network

    CN121253982A