Multi-scale Detection Method and System for Coal Gangue Based on YOLOv8 Network
By improving the YOLOv8 network model, the CBAM attention mechanism and feature fusion module were introduced, the problem of insufficient detection of small and medium-sized targets was solved, and the accurate identification of coal gangue was achieved, which was suitable for coal gangue sorting equipment.
Patent Information
- Application Number
- CN202411814781.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-12-11
AI Technical Summary
The existing target detection methods are difficult to effectively identify coal gangue with large and large particle-grade differences in multi-scale detection, especially small target detection capabilities, which leads to the inability to accurately detect coal gangue under high overlap rates.
Improve the YOLOv8 network model, enhance feature information interaction by introducing CBAM attention mechanism, fast spatial pyramid pooling module and bidirectional feature fusion module, and adopting EIoU_Loss regression loss to optimize small object detection capabilities.
Accurate detection of 6mm~50mm particle-grade coal gangue has been achieved, and multi-scale detection capabilities have been improved, and it is suitable for coal gangue sorting equipment.
Smart Images

Figure CN119295953B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of coal gangue detection, and particularly relates to a multi-scale detection method and system for coal gangue based on the YOLOv8 network. Background Art
[0002] The coal mining process is very complex and is generally in the rock layer under the sedimentary rock, where a large amount of gangue is often mixed. Mixing coal gangue will reduce the combustion efficiency of coal and seriously pollute the environment. Therefore, separating coal gangue from the mined ore is of great significance for environmental protection and the combustion efficiency of coal. The common coal gangue separation methods mainly include manual gangue picking method, wet separation method and dry separation method. Among them, the manual gangue picking method has low efficiency and is unstable, and the wet separation method has problems such as wasting water resources and polluting the environment. Therefore, the current mainstream method is the dry separation method. In the dry separation method, the hardness identification method has a large limitation in the use scenario, the ray identification method is highly dangerous, the optoelectronic identification process is complex and has high requirements, while the target detection method has high identification efficiency and low cost and has broad application prospects.
[0003] The target detection method includes two steps: detection and classification. Among them, the traditional target detection method is the threshold method, which sets a threshold by the pixel value difference between the detection part and the picture background part to segment the target to be detected. This method can relatively accurately detect whether the target to be detected in the picture is the target or the background in coal gangue detection, but for overlapping targets, it can only detect them as one and cannot correctly detect the number. In actual application scenarios, such as in carbon gold equipment, when the coal processing volume is less than 200T / h, the coal overlap degree is not high and this method can be used, but when it is higher than 200T / h, due to the high overlap rate, many targets cannot be detected and this method cannot be used. The current mainstream target detection method is to use deep learning methods, which are mainly divided into two categories: two-stage target detection algorithms; one-stage target detection algorithms.
[0004] The two-stage target detection algorithms include R-CNN, Fast R-CNN, Faster R-CNN, Mask R-CNN, etc.; the one-stage target detection algorithms are mainly SSD and YOLO series. Although using deep learning methods for detection can solve the problem of overlapping target detection, in the case of multi-scale detection, if the sizes of the detection targets vary too much, small target features may be lost in a certain network layer, ultimately resulting in the inability to identify small targets; through splicing pictures of various large-grained coals and small-grained coals and using various current mainstream deep learning methods to train the model and then actually detect, it is found that when the size difference of the detection targets is more than a certain multiple, small targets below a certain grain size may not be recognized. Summary of the Invention
[0005] Based on this, in the embodiments of the present invention, a multi-scale detection method and system for coal gangue based on the YOLOv8 network are provided, aiming to improve the YOLOv8 network model by referring to the small target detection method, so that the improved model can optimize the detection problem of small targets in multi-scale detection while maintaining the recognition speed, and achieve accurate detection of coal gangue.
[0006] In the first aspect of the embodiments of the present invention, a multi-scale detection method for coal gangue based on the YOLOv8 network is provided, which is used to detect coal gangue with a particle size of 6 mm to 50 mm. The method includes:
[0007] Obtain an image of a mixture of coal and coal gangue, and label the image of the mixture of coal and coal gangue to obtain a data set;
[0008] Input the data set into the improved YOLOv8 network model for training to obtain a target network model;
[0009] Input the image to be detected into the target network model to determine the coal and coal gangue in the image to be detected;
[0010] The improved YOLOv8 network model includes a backbone network composed of 5 CBS modules, 4 CBC2f modules and 1 fast spatial pyramid pooling module, a bidirectional feature fusion module composed of 2 CBS modules and 4 CBC2f modules in the neck, and a detection structure finally sent to the head and composed of 3 detection heads with different sizes;
[0011] Among them, after the input stream in the CBC2f module undergoes a splitting operation, the original input stream is split into two equal parts. One part is sent into the remaining CBS modules in the CBC2f module, and the other part focuses on important features through a convolutional block attention module. Finally, the feature maps extracted by three convolutional blocks are concatenated with the feature map generated by the convolutional block attention module, and finally output. The convolutional block attention module adopts an integrated channel and spatial attention mechanism;
[0012] After the fast spatial pyramid pooling module performs convolutional processing on the feature map output by the previous layer, the obtained feature map is subjected to max pooling through three pooling layers of 5×5, 9×9, and 13×13, and is concatenated with the feature maps output by the three pooling layers in the channel. At the same time, an additional convolutional layer is set to extract features and perform a skip connection between the finally obtained features and the input features;
[0013] The bidirectional feature fusion module increases the detection ability by setting skip connections between the feature pyramids;
[0014] The detection head includes two CBS modules connected in sequence, and a convolution block respectively connected to the two CBS modules.
[0015] Further, the convolutional block attention module is used to perform average pooling and max pooling on the input feature map in the spatial dimension to obtain two first feature maps, which respectively represent the average value and the maximum value of this channel;
[0016] The two first feature maps are respectively input into two fully-connected networks to generate two channel attention vectors and add them together, and then normalized through the sigmoid function to generate the final channel attention weight;
[0017] The original feature map is multiplied element-wise by the channel attention weight to obtain a weighted feature map;
[0018] The weighted feature map is input into the spatial attention unit. Similarly, average pooling and max pooling are first performed in the channel dimension to obtain two second feature maps;
[0019] The two second feature maps are concatenated in the channel dimension to obtain a two-channel feature map, and then convolution operation is performed through a 7×7 convolution kernel to obtain a spatial attention map;
[0020] The original feature map is multiplied element-wise by the spatial attention map to obtain a weighted target feature map.
[0021] Further, the convolution blocks for feature fusion that perform regression loss and classification loss in the detection head are merged together, while the convolution for conversion to the output remains unchanged, so as to achieve the purpose of lightweight.
[0022] Further, in the step of performing average pooling and max pooling on the input feature map in the spatial dimension to obtain two first feature maps, the calculation formula is:
[0023]
[0024] Among them, F avg and F max are respectively the pooling results after channel average pooling and max pooling, is the feature description at the position of in the c-th channel, is the scale range to which the original feature map belongs, C is the number of channels, H is the height, W is the width, is the size of the pooled feature map with the number of channels being C and the height and width both being 1.
[0025] Further, in the step of multiplying the original feature map element-wise with the channel attention weight to obtain the weighted feature map, the calculation formula is:
[0026]
[0027] where M is the channel attention weight, σ is the sigmoid activation function, and MLP is a shared multi-layer perceptron composed of a learnable weight matrix and a ReLU activation function.
[0028] Further, the classification loss uses BCE_Loss for multi-label classification tasks.
[0029] Further, the regression loss consists of EIoU_Loss and Distribution Focal Loss. Among them, the calculation formula of EIoU_Loss is:
[0030]
[0031] where IOU is the intersection over union, b and b gt are the center points of bounding box A and bounding box B respectively, c is the diagonal length of the smallest bounding rectangle that can contain both the predicted box and the ground truth box, A is the predicted box, B is the ground truth box, w gt and h gt are the width and height of the predicted box, w and h are the width and height of the target box, is the Euclidean distance between two points, C w and C h are the width and height of the smallest bounding rectangle that can contain both the predicted box and the ground truth box respectively.
[0032] The second aspect of the embodiments of the present invention provides a coal gangue multi-scale detection system based on the YOLOv8 network, which is used to implement the coal gangue multi-scale detection method based on the YOLOv8 network described in the first aspect. The system includes:
[0033] An annotation module for obtaining an image of a mixture of coal and coal gangue and annotating the image of the mixture of coal and coal gangue to obtain a data set;
[0034] A training module for inputting the data set into the improved YOLOv8 network model for training to obtain a target network model;
[0035] An input module for inputting an image to be detected into the target network model to determine coal and coal gangue in the image to be detected;
[0036] The improved YOLOv8 network model includes a backbone network composed of 5 CBS modules, 4 CBC2f modules, and 1 fast spatial pyramid pooling module, a bidirectional feature fusion module composed of 2 CBS modules and 4 CBC2f modules in the neck, and finally a detection structure sent to the head composed of 3 detection heads of different sizes;
[0037] Among them, after an input stream in the CBC2f module undergoes a splitting operation, the original input stream is split into two equal parts. One part is sent to the remaining CBS modules in the CBC2f module, and the other part focuses on important features through a convolutional block attention module. Finally, the feature maps extracted through three convolutional blocks are concatenated with the feature maps generated by the convolutional block attention module for final output. The convolutional block attention module adopts an integrated channel and spatial attention mechanism;
[0038] After the fast spatial pyramid pooling module performs convolutional processing on the feature map output by the previous layer, the obtained feature map undergoes max pooling through three pooling layers of 5×5, 9×9, and 13×13, and is concatenated with the feature maps output by the three pooling layers in the channel dimension. At the same time, an additional convolutional layer is set to extract features and perform a skip connection between the finally obtained features and the input features;
[0039] The bidirectional feature fusion module increases the detection ability by setting skip connections between each feature pyramid;
[0040] The detection head includes 2 CBS modules connected in sequence, and convolutional blocks respectively connected to the 2 CBS modules.
[0041] A third aspect of the embodiments of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the coal gangue multi-scale detection method based on the YOLOv8 network provided in the first aspect.
[0042] A fourth aspect of the embodiments of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the coal gangue multi-scale detection method based on the YOLOv8 network provided in the first aspect.
[0043] A method and system for multi-scale detection of coal gangue based on the YOLOv8 network provided in the embodiments of the present invention improve and optimize the structure from the perspective of the detection ability of small targets under multi-scale detection on the basis of the YOLOv8 network structure. First, the CBAM attention mechanism is embedded in the core module C2f of the entire network, increasing the detection ability of small targets; secondly, a fast spatial pyramid pooling module is adopted in the backbone network, integrating more front and back features, enhancing the interaction of feature information, and further improving the multi-scale detection ability; then, a bidirectional feature fusion module with more skip connections is used to increase the multi-scale detection ability and prevent the loss of small target features due to the network depth; finally, the convolutional blocks used for feature fusion in the detection head are merged to be lightweight, and the regression loss is set to EIoU_Loss to avoid errors caused by the same aspect ratio of the predicted box and the true box. Through the above target network model, coal gangue with a particle size of 6 mm to 50 mm can be effectively detected, realizing cross-particle size detection, and can be effectively deployed in coal gangue separation equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 FIG. 6 is a flowchart of the implementation of a method for multi-scale detection of coal gangue based on the YOLOv8 network provided in Embodiment 1 of the present invention;
[0045] Figure 2 FIG. 7 is a schematic structural diagram of the CBC2f module;
[0046] Figure 3 FIG. 8 is a schematic structural diagram of the convolutional block attention module;
[0047] Figure 4 FIG. 9 is a schematic structural diagram of the fast spatial pyramid pooling module;
[0048] Figure 5 FIG. 10 is a schematic diagram of the bidirectional feature fusion module;
[0049] Figure 6 FIG. 11 is a schematic structural diagram of the detection head;
[0050] Figure 7 FIG. 12 is a structural block diagram of a system for multi-scale detection of coal gangue based on the YOLOv8 network provided in Embodiment 2 of the present invention;
[0051] Figure 8 FIG. 13 is a structural block diagram of an electronic device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant accompanying drawings. Several embodiments of the present invention are shown in the accompanying drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0053] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there can also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0055] Embodiment 1
[0056] According to an embodiment of the present invention, an embodiment of a multi-scale detection method for coal gangue based on the YOLOv8 network is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0057] In this first embodiment, a multi-scale detection method for coal gangue based on the YOLOv8 network is provided, which is used to detect coal gangue with a particle size of 6 mm to 50 mm and can be used in electronic devices such as computers. Please refer to Figure 1 , Figure 1 which shows the implementation flowchart of a multi-scale detection method for coal gangue based on the YOLOv8 network provided in the first embodiment of the present invention, specifically including steps S01 to S03.
[0058] Step S01, obtain an image of a mixture of coal and coal gangue, and label the image of the mixture of coal and coal gangue to obtain a data set.
[0059] Among them, the picture annotation tool Labelme can be used for annotation, and the data set can be divided according to a ratio, and divided into a training set, a validation set and a test set according to 8:1:1.
[0060] Step S02: Input the dataset into the improved YOLOv8 network model for training to obtain the target network model.
[0061] Specifically, the improved YOLOv8 network model includes a backbone network composed of 5 CBS modules, 4 CBC2f modules, and 1 Spatial Pyramid Pooling Fast Module (SPPFS module), a Bidirectional Feature Aggregation Module (BiPAN) composed of 2 CBS modules and 4 CBC2f modules in the neck, and finally a detection structure sent to the head consisting of 3 detection heads of different sizes.
[0062] Among them, please refer to Figure 2 for the schematic diagram of the CBC2f module structure. Among them, the sizes of the input feature map and the output feature map are the same. After the input stream in the CBC2f module undergoes a splitting operation, the original input stream is split into two equal parts. One part is sent to the remaining CBS module in the CBC2f module, and the other part focuses on important features through a Convolutional Block Attention Module (CBAM). Finally, the feature maps extracted by three convolutional blocks are concatenated with the feature map generated by the convolutional block attention module for final output. The convolutional block attention module adopts an integrated channel and spatial attention mechanism. In this embodiment, the CBAM attention mechanism is added to the basic building block C2f module of the entire network model to enhance the small target detection ability. The convolutional block attention module helps the network model more effectively focus on important features through an integrated channel and spatial attention mechanism, thereby improving the model's attention to the feature channels and location information of small targets.
[0063] Please refer to Figure 3 for the schematic diagram of the convolutional block attention module structure. Among them, c represents concatenation, a represents addition, and the function of Bottleneck is to add the part of the input that has passed through two 3×3 CBS modules to the other part of the input and then output. Specifically, the convolutional block attention module first performs average pooling and max pooling on the input feature map F(C×H×W) in the spatial dimension to obtain two C×1×1 first feature maps, which represent the average value and the maximum value of this channel respectively. The calculation formulas are as follows:
[0064]
[0065] Among them, Favg and Fmax are the pooling results after channel average pooling and max pooling respectively, is the feature description at the position in the c-th channel, is the scale range to which the original feature map belongs, C is the number of channels, H is the height, and W is the width. It is the size of the pooled feature map with the number of channels being C, and both the height and width being 1.
[0066] The two first feature maps are respectively input into two layers of fully connected networks to generate two channel attention vectors, which are added together and then normalized through the sigmoid function to generate the final channel attention weight M. The calculation formula is:
[0067]
[0068] Among them, M is the channel attention weight, σ is the sigmoid activation function, and MLP is a shared multi-layer perceptron, which consists of a learnable weight matrix and a ReLU activation function;
[0069] The original feature map F is multiplied element-wise with the channel attention weight M to obtain the weighted feature map F1. The calculation formula is:
[0070]
[0071] Among them, is the dot product;
[0072] The weighted feature map F1 is input into the spatial attention unit. First, average pooling and max pooling are also performed on the channel dimension to obtain two second feature maps of 1×H×W Favg(i,j) and Fmax(i,j) , and the calculation formula is:
[0073]
[0074] The two second feature maps are concatenated on the channel dimension to obtain a dual-channel feature map , and then a convolution operation is performed through a 7×7 convolution kernel to obtain a spatial attention map ;
[0075] The original feature map F is multiplied element-wise with the spatial attention map to obtain the weighted target feature map F2. The calculation formula is:
[0076]
[0077] After the fast spatial pyramid pooling module performs convolution processing on the feature map output by the previous layer, the obtained feature map is subjected to max pooling through three pooling layers of 5×5, 9×9, and 13×13, and is concatenated with the feature maps output by the three pooling layers on the channel. At the same time, an additional convolution layer is set to extract features and a skip connection is made between the finally obtained features and the input features. Please refer to Figure 4 , which is the structural schematic diagram of the fast spatial pyramid pooling module;
[0078] The bidirectional feature fusion module increases the detection ability by setting skip connections between the feature pyramids. Please refer to Figure 5 , which is a schematic diagram of the bidirectional feature fusion module. The specific processing flow using the bidirectional feature fusion module is as follows:
[0079] Original feature pyramid: P = {P1, P2, P3}, where P1 is a feature map of 80×80 with 256 channels; P2 is a feature map of 40×40 with 512 channels; P3 is a feature map of 20×20 with 1024 channels.
[0080] Upsampling and fusion:
[0081] P3 corresponds to the new feature map T3. After 1×1 convolution and normalization, it is upsampled by a factor of 2 to obtain a 40×40 feature map, which is horizontally connected to P2, and the corresponding elements are added to obtain a 40×40 feature map T2.
[0082] After convolution and normalization processing of T2, it is upsampled by a factor of 2 to obtain an 80×80 feature map, which is horizontally connected to P1, and the corresponding elements are added to obtain an 80×80 feature map T1.
[0083] Downsampling and skip connection:
[0084] Fuse T1 and P1 to obtain an 80×80 feature map F1.
[0085] After convolution and normalization processing of F1, it is downsampled by a factor of 2 to obtain a 40×40 feature map, which is skip-connected to P2, and the corresponding elements are added to obtain a 40×40 feature map F2.
[0086] After convolution and normalization processing of F2, it is downsampled by a factor of 2 to obtain a 20×20 feature map, which is skip-connected to P3, and the corresponding elements are added to obtain a 20×20 feature map F3.
[0087] Skip connection:
[0088] Perform a skip connection between F3 and T3 to obtain a 20×20 feature map N3.
[0089] Perform a skip connection between F2 and T2 to obtain a 40×40 feature map N2.
[0090] Perform a skip connection between F1 and T1 to obtain an 80×80 feature map N1.
[0091] Final output:
[0092] N1: 80×80
[0093] N2: 40×40
[0094] N3: 20×20;
[0095] The detection head includes two CBS modules connected in sequence, and a convolutional block connected to the two CBS modules respectively. The convolutional blocks for feature fusion in the detection head for regression loss and classification loss are merged together, and the convolutional layer for conversion to the output remains unchanged, so as to achieve the purpose of lightweight. Please refer to Figure 6 , which is the schematic diagram of the detection head structure.
[0096] It should be noted that the classification loss uses BCE_Loss for multi-label classification tasks, and the regression loss uses two parts: EIoU_Loss and Distribution Focal Loss. Among them, the calculation formula of EIoU_Loss is:
[0097]
[0098] Among them, IOU is the intersection over union, b , bgt are the center points of box A and box B respectively, c is the diagonal length of the smallest circumscribed rectangle that can contain both the predicted box and the ground truth box, A is the predicted box, B is the ground truth box, wgt and hgt are the width and height of the predicted box, w and h are the width and height of the target box, is the Euclidean distance between two points, , are the width and height of the smallest circumscribed rectangle that can contain both the predicted box and the ground truth box respectively.
[0099] In this embodiment, the evaluation metrics are Precision (P), Recall (R), and Mean Average Precision (mAP) on the test set, including mAP50 (IoU = 0.5), mAP50 - 95 (IoU = 0.5:0.05:0.95), detection speed (ms), and Parameters (the number of model parameters). These 6 performance metrics are used to evaluate the performance of the network model. The calculation formulas of P, R, and mAP are:
[0100]
[0101]
[0102]
[0103]
[0104] Among them, TP is the number of correct targets detected by the model, FP is the number of targets detected incorrectly by the model, FN is the number of missed detections and false negatives by the model, AP is the average precision of a category, and K is the number of categories.
[0105] In step S03, the image to be detected is input into the target network model to determine coal and coal gangue in the image to be detected.
[0106] In summary, the coal gangue multi-scale detection method based on the YOLOv8 network in the above embodiments of the present invention improves and optimizes the structure from the perspective of the small target detection ability under multi-scale detection on the basis of the YOLOv8 network structure. First, the CBAM attention mechanism is embedded in the entire network core module C2f, increasing the small target detection ability; secondly, a fast spatial pyramid pooling module is adopted in the backbone network, fusing more front and back features, enhancing the interaction of feature information, and further improving the multi-scale detection ability; then, a bidirectional feature fusion module with more skip connections is used to increase the multi-scale detection ability and prevent the loss of small target features due to the network depth; finally, the convolutional blocks for feature fusion in the detection head are merged to be lightweight, and the regression loss is set to EIoU_Loss to avoid errors due to the same aspect ratio of the predicted box and the ground truth box. Through the above target network model, coal gangue with a particle size of 6 mm to 50 mm can be effectively detected, cross-particle size detection can be achieved, and it can be effectively deployed in coal gangue separation equipment.
[0107] Embodiment 2
[0108] Please refer to Figure 7 , Figure 7 which is a structural block diagram of a coal gangue multi-scale detection system based on the YOLOv8 network provided by Embodiment 2 of the present invention. The coal gangue multi-scale detection system 200 based on the YOLOv8 network is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0109] Specifically, the coal gangue multi-scale detection system 200 based on the YOLOv8 network includes: an annotation module 21, a training module 22, and an input module 23, where:
[0110] The annotation module 21 is used to obtain an image of a mixture of coal and coal gangue, and annotate the image of the mixture of coal and coal gangue to obtain a data set;
[0111] A training module 22 for inputting the dataset into an improved YOLOv8 network model for training to obtain a target network model;
[0112] An input module 23 for inputting the image to be detected into the target network model to determine coal and gangue in the image to be detected;
[0113] The improved YOLOv8 network model includes a backbone network composed of 5 CBS modules, 4 CBC2f modules and 1 fast spatial pyramid pooling module, a bidirectional feature fusion module composed of 2 CBS modules and 4 CBC2f modules in the neck, and finally a detection structure sent to the head and composed of 3 detection heads of different sizes;
[0114] Among them, after an input stream in the CBC2f module undergoes a splitting operation, the original input stream is split into two equal parts. One part is sent to the remaining CBS module in the CBC2f module, and the other part focuses on important features through a convolutional block attention module. Finally, the feature maps extracted through three convolutional blocks are concatenated with the feature map generated by the convolutional block attention module and finally output. The convolutional block attention module adopts an integrated channel and spatial attention mechanism. The convolutional block attention module is used to perform average pooling and max pooling on the input feature map in the spatial dimension to obtain two first feature maps, which respectively represent the average value and the maximum value of this channel. The calculation formula is:
[0115]
[0116] Among them, F avg and F max are the pooling results after channel average pooling and max pooling respectively, is the feature description at the position of in the c-th channel, is the scale range to which the original feature map belongs, C is the number of channels, H is the height, and W is the width, is the size of the pooled feature map with the number of channels C and both height and width being 1;
[0117] The two first feature maps are respectively input into two-layer fully connected networks to generate two channel attention vectors and add them, and then normalized through a sigmoid function to generate the final channel attention weight;
[0118] The original feature map is multiplied element-wise by the channel attention weight to obtain a weighted feature map. The calculation formula is:
[0119]
[0120] Among them, M is the channel attention weight, σ is the sigmoid activation function, and MLP is a shared multi-layer perceptron, which consists of a learnable weight matrix and a ReLU activation function;
[0121] The weighted feature map is input into the spatial attention unit. First, average pooling and max pooling are performed on the channel dimension to obtain two second feature maps;
[0122] The two second feature maps are concatenated on the channel dimension to obtain a dual-channel feature map, and then a convolution operation is performed through a 7×7 convolution kernel to obtain a spatial attention map;
[0123] The original feature map is multiplied element-wise by the spatial attention map to obtain a weighted target feature map;
[0124] After the fast spatial pyramid pooling module performs convolution processing on the feature map output by the previous layer, the obtained feature map is max-pooled through three pooling layers of 5×5, 9×9, and 13×13, and is concatenated with the feature maps output by the three pooling layers on the channel. At the same time, an additional convolution layer is set to extract features and a skip connection is made between the finally obtained features and the input features;
[0125] The bidirectional feature fusion module increases the detection ability by setting skip connections between each feature pyramid;
[0126] The detection head includes 2 CBS modules connected in sequence, and convolution blocks respectively connected to the 2 CBS modules. The convolution blocks for feature fusion that perform regression loss and classification loss in the detection head are merged together, while the convolution for conversion to the output remains unchanged to achieve the purpose of lightweight. The classification loss uses BCE_Loss for multi-label classification tasks. The regression loss uses two parts: EIoU_Loss and Distribution Focal Loss. Among them, the calculation formula of EIoU_Loss is:
[0127]
[0128] Among them, IOU is the intersection over union, b 、 b gt are the center points of bounding box A and bounding box B respectively, c is the diagonal length of the smallest circumscribed rectangle that can simultaneously contain the predicted box and the ground truth box, A is the predicted box, B is the ground truth box, w gt and h gtare the width and height of the prediction box, w and h are the width and height of the target box, is the Euclidean distance between two points, C w 、 C h are respectively the width and height of the smallest bounding rectangle that simultaneously contains the prediction box and the ground truth box.
[0129] Embodiment III
[0130] On the other hand, the present invention also proposes an electronic device. Please refer to Figure 8 , which shows the electronic device in Embodiment III of the present invention, including a memory 20, a processor 10, and a computer program 30 stored in the memory and executable on the processor. When the processor 10 executes the computer program 30, it implements the above-mentioned coal gangue multi-scale detection method based on the YOLOv8 network.
[0131] Among them, the processor 10 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments, and is used to run the program code stored in the memory 20 or process data, such as executing an access restriction program, etc.
[0132] Among them, the memory 20 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a magnetic memory, a magnetic disk, an optical disk, etc. The memory 20 can be an internal storage unit of the electronic device in some embodiments, such as the hard disk of the electronic device. The memory 20 can also be an external storage device of the electronic device in other embodiments, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory 20 can also include both the internal storage unit and the external storage device of the electronic device. The memory 20 can be used not only to store application software and various data of the electronic device, but also to temporarily store data that has been output or will be output.
[0133] It should be noted that Figure 8 the structure shown does not limit the electronic device. In other embodiments, the electronic device may include fewer or more components than shown, or combine some components, or have different component arrangements.
[0134] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the above-mentioned multi-scale detection method of coal gangue based on the YOLOv8 network.
[0135] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0136] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.
[0137] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0138] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0139] The above embodiments only express several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several variations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.
Claims
1. A multi-scale detection method for coal gangue based on the YOLOv8 network, characterized in that For detecting coal gangue with a particle size of 6 mm to 50 mm, the method includes: Obtain an image of the mixture of coal and coal gangue, and annotate the image of the mixture of coal and coal gangue to obtain a data set; Input the data set into the improved YOLOv8 network model for training to obtain a target network model; Input the image to be detected into the target network model to determine the coal and coal gangue in the image to be detected; The improved YOLOv8 network model includes a backbone network composed of 5 CBS modules, 4 CBC2f modules, and 1 fast spatial pyramid pooling module, and a bidirectional feature fusion module composed of 2 CBS modules and 4 CBC2f modules in the neck, and finally a detection structure sent to the head composed of 3 detection heads of different sizes. Among them, the CBC2f module is a module obtained by embedding the bidirectional feature fusion module into the C2f module, and the CBS module is a module combined with a convolutional layer, a normalization layer, and an activation function; Among them, after the input stream in the CBC2f module undergoes a splitting operation, the original input stream is split into two equal parts. One part is sent to the remaining CBS modules in the CBC2f module, and the other part focuses on important features through a convolutional block attention module. Finally, the feature maps extracted by three convolutional blocks are spliced with the feature map generated by the convolutional block attention module, and finally output. The convolutional block attention module adopts an integrated channel and spatial attention mechanism; After the fast spatial pyramid pooling module performs convolutional processing on the feature map output by the previous layer, the obtained feature map is subjected to max pooling through three pooling layers of 5×5, 9×9, and 13×13, and is spliced with the feature maps output by the three pooling layers in the channel dimension. At the same time, an additional convolutional layer is set to extract features and perform a skip connection between the finally obtained features and the input features; The bidirectional feature fusion module increases the detection ability by setting skip connections between each feature pyramid; The detection head includes 2 CBS modules connected in sequence, and convolutional blocks respectively connected to the 2 CBS modules; The convolutional block attention module is used to perform average pooling and max pooling on the input feature map in the spatial dimension to obtain two first feature maps, which respectively represent the average value and the maximum value of this channel; Input the two first feature maps into two fully connected networks respectively to generate two channel attention vectors and add them, and then perform normalization through the sigmoid function to generate the final channel attention weight; Multiply the original feature map element by element with the channel attention weight to obtain a weighted feature map; Input the weighted feature map into the spatial attention unit, and also perform average pooling and max pooling in the channel dimension first to obtain two second feature maps; Splice the two second feature maps in the channel dimension to obtain a two-channel feature map, and then perform a convolution operation through a 7×7 convolutional kernel to obtain a spatial attention map; Multiply the original feature map element by element with the spatial attention map to obtain a weighted target feature map; Merge the convolutional blocks for feature fusion that perform regression loss and classification loss in the detection head, while keeping the convolutional layers for conversion to the output unchanged, for the purpose of lightweighting.
2. The multi-scale detection method of coal gangue based on the YOLOv8 network according to claim 1, wherein In the step of performing average pooling and max pooling on the input feature map in the spatial dimension to obtain two first feature maps, the calculation formula is: Among them, F avg and F max are the pooling results after channel average pooling and max pooling respectively, is the feature description at the position of in the c-th channel, is the scale range to which the original feature map belongs, C is the number of channels, H is the height, and W is the width, is the size of the pooled feature map with the number of channels C and both the height and width being 1.
3. The multi-scale detection method of coal gangue based on the YOLOv8 network according to claim 2, wherein, In the step of multiplying the original feature map element-wise with the channel attention weight to obtain a weighted feature map, the calculation formula is: Where M is the channel attention weight, σ is the sigmoid activation function, and MLP is a shared multi-layer perceptron, consisting of a learnable weight matrix and a ReLU activation function.
4. The gangue multi-scale detection method based on the YOLOv8 network according to claim 3, wherein, The classification loss uses BCE_Loss and is used for multi-label classification tasks.
5. The gangue multi-scale detection method based on the YOLOv8 network according to claim 4, wherein, The regression loss consists of EIoU_Loss and Distribution Focal Loss. Among them, the calculation formula of EIoU_Loss is: wherein, IOU is the intersection over union (IoU), b , b gt are the center points of bounding box A and bounding box B respectively, c is the diagonal length of the smallest bounding rectangle that can contain both the predicted box and the ground truth box, A is the predicted box, B is the ground truth box, w gt and h gt are the width and height of the predicted box, w and h are the width and height of the target box, is the Euclidean distance between two points, C w , C h are the width and height of the smallest bounding rectangle that can contain both the predicted box and the ground truth box respectively.
6. A multi-scale detection system for coal gangue based on the YOLOv8 network, characterized in that, For implementing the multi-scale coal gangue detection method based on the YOLOv8 network as described in any one of claims 1-5, the system includes: An annotation module for obtaining an image of a mixture of coal and coal gangue and annotating the image of the mixture of coal and coal gangue to obtain a data set; A training module for inputting the data set into the improved YOLOv8 network model for training to obtain a target network model; An input module for inputting the image to be detected into the target network model to determine the coal and coal gangue in the image to be detected; The improved YOLOv8 network model includes a backbone network composed of 5 CBS modules, 4 CBC2f modules, and 1 fast spatial pyramid pooling module, a bidirectional feature fusion module composed of 2 CBS modules and 4 CBC2f modules in the neck, and a detection structure finally sent to the head composed of 3 detection heads of different sizes; Among them, in the CBC2f module, after the input stream undergoes a splitting operation, the original input stream is split into two equal parts. One part is sent to the remaining CBS modules in the CBC2f module, and the other part focuses on important features through a convolutional block attention module. Finally, the feature maps extracted by three convolutional blocks are concatenated with the feature map generated by the convolutional block attention module and finally output. The convolutional block attention module adopts an integrated channel and spatial attention mechanism; The fast spatial pyramid pooling module performs convolutional processing on the feature map output by the previous layer, then performs max pooling on the obtained feature map through three pooling layers of 5×5, 9×9, and 13×13, and concatenates the feature maps output by the three pooling layers in the channel dimension. At the same time, an additional convolutional layer is set to extract features and perform a skip connection between the finally obtained features and the input features; The bidirectional feature fusion module increases the detection ability by setting skip connections between each feature pyramid; The detection head includes 2 CBS modules connected in sequence, and convolutional blocks respectively connected to the 2 CBS modules.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the multi-scale detection method of coal gangue based on the YOLOv8 network as described in any one of claims 1-5.
8. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the multi-scale detection method of coal gangue based on the YOLOv8 network as described in any one of claims 1-5.
Citation Information
Patent Citations
Water turbine top cover defect detection method based on improved YOLOv8 model
CN117541538A