Coal gangue rapid identification method based on lightweight improved YOLOv8 algorithm

By using the lightweight and improved YOLOv8 algorithm in coal gangue detection, the CSPDarkNet backbone network and PAN-FAN neck network are built, and the problems of low efficiency and poor effect of coal gangue detection in the existing technology are solved, and the rapid and accurate identification of coal gangue is achieved.

CN120032164APending Publication Date: 2025-05-23NORTHEAST DIANLI UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510001331.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In the prior art, coal gangue detection efficiency is low and has poor results, and relies on manual observation and analysis, which is time-consuming and labor-intensive, and is easily affected by subjective factors.

Method used

The rapid identification method of coal gangue based on the lightweight and improved YOLOv8 algorithm is adopted. By building a CSPDarkNet as the backbone network and PAN-FAN as the neck network, and combining a detection network including three detection heads, the target detection network is trained to achieve rapid identification of coal gangue.

Benefits of technology

It realizes efficient and accurate detection of coal gangue, improves detection efficiency and accuracy, and reduces dependence on manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032164A_ABST
    Figure CN120032164A_ABST
Patent Text Reader

Abstract

The invention belongs to a coal gangue detection and sorting technology in the coal industry, and relates to a rapid coal gangue recognition method based on a lightweight improved YOLOv8 algorithm, which comprises the following steps: performing simulation according to a specific scene, acquiring a coal gangue image, and performing frame selection and labeling to prepare a data set; establishing a YOLOv8 target detection network which takes CSPDarkNet as a trunk network and PAN-FAN as a neck network, and is combined with a detection network comprising three detection output ends to carry out lightweight improvement, training a coal gangue data set by using an improved YOLOv8 algorithm, and checking the performance of the coal gangue data set through a verification set and a test set; and the trained lightweight improved YOLOv8 target detection network is used for rapid coal gangue identification. The problems that a manual observation and analysis method is time-consuming and labor-consuming, and the coal gangue detection efficiency is low and the coal gangue detection effect is poor are solved, and a new solution is provided for efficient and accurate detection of the coal gangue. And the coal gangue in the coal pile can be effectively identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the detection and sorting technology of coal gangue in the coal industry, in particular to the technology of automatic identification and classification of coal gangue using a deep learning algorithm. More specifically, the present invention provides a method for rapid identification of coal gangue based on a lightweight improved YOLOv8 algorithm. Background Art

[0002] As a traditional fossil energy, coal occupies a pivotal position in the energy structure. According to analysis, my country will still be engaged in coal mining for a long time in the future, and the detection and sorting of coal gangue is a key step to improve coal quality and reduce environmental pollution. The traditional manual sorting method is not only inefficient, but also labor-intensive and easily affected by subjective factors. Summary of the invention

[0003] The technical problem to be solved by the present invention is to provide a method for quickly identifying coal gangue based on a lightweight improved YOLOv8 algorithm, so as to solve the problems that manual observation and analysis methods are time-consuming and labor-intensive, and the existing methods for detecting coal gangue are inefficient and have poor effects.

[0004] The present invention is achieved in this way.

[0005] A fast identification method for coal gangue based on the lightweight improved YOLOv8 algorithm is proposed. According to the specific scenario, the coal gangue images are collected and marked by box selection to form a data set.

[0006] A lightweight and improved YOLOv8 target detection network was built with CSPDarkNet as the backbone network and PAN-FAN as the neck network, combined with a detection network including three detection heads. The improved YOLOv8 algorithm was used to train the gangue dataset, and its performance was tested through the validation set and test set.

[0007] The lightweight improved YOLOv8 target detection network specifically includes:

[0008] The backbone network includes, in order of data input: a first convolutional layer, a second convolutional layer, a first cross-stage local partial convolutional module, a third convolutional layer, a second cross-stage local partial convolutional module, a fourth convolutional layer, a third cross-stage local partial convolutional module, a fifth convolutional layer, a fourth cross-stage local partial convolutional module, and a fast spatial pyramid pooling layer;

[0009] The neck network includes, in order of data input: a second upsampling layer, a second splicing layer, a sixth cross-stage local partial convolution module, a partial self-attention module, a first upsampling layer, a first splicing layer and a fifth cross-stage local partial convolution module, a seventh convolution layer, a fourth splicing layer, an eighth cross-stage local partial convolution module, a sixth convolution layer, a third splicing layer and a seventh cross-stage local partial convolution module;

[0010] The second cross-stage local partial convolution module also outputs to the first splicing layer, the third cross-stage local partial convolution module also outputs to the second splicing layer, and the fast spatial pyramid pooling layer outputs to the third splicing layer;

[0011] The fifth cross-stage local partial convolution module also outputs a first detection output terminal;

[0012] The eighth cross-stage local partial convolution module also outputs a second detection output terminal;

[0013] The seventh cross-stage local partial convolution module outputs a third detection output terminal;

[0014] The trained lightweight improved YOLOv8 target detection network is used for rapid identification of coal gangue.

[0015] Furthermore: the backbone network is divided into five parts according to P1~P5, wherein P1 includes the first convolution layer, which is a convolution kernel with a size of 3, a step size of 2 and a padding of 1, P2~P4 are respectively composed of a convolution layer and a cross-stage local partial convolution module, and P5 includes a fast spatial pyramid pooling layer.

[0016] Furthermore: each detection output end in the detection network is connected to two 3×3 Conv modules and one 1×1 Conv2d module through two paths, respectively, to perform Bbox.Loss detection and Cls.Loss detection, respectively.

[0017] Furthermore, in Cls.Loss detection, binary cross entropy is used as classification loss, each category is judged "whether it is this type", and the confidence is output, where the calculation formula is:

[0018]

[0019] Y represents the sample label, the positive sample label is 1, the negative sample label is 0, p represents the probability of predicting a positive sample, and N is the number of samples;

[0020] In Bbox.Loss detection, DFL loss is used as the regression loss. DFL loss focuses on the value near the label, so that the network quickly converges to the distribution of the target location and its neighboring areas. The DFL loss formula is:

[0021] (S i ,s i+1 )=-((y i+1 -y)log(s i )+(yy i )log(s i+1 ))

[0022] Si and Si+1 are the "predicted value" and "near predicted value" output by the network, and y, yi, and yi+1 are the "actual value", "label integral value", and "near label integral value" of the label;

[0023] At the same time, IOU Loss is used as the regression loss. IOU Loss measures the overlap between the predicted bounding box and the true bounding box to optimize the positioning accuracy. The IOU Loss formula is:

[0024]

[0025] A is the predicted box and B is the true box.

[0026] Further: the improved YOLOv8 target recognition network performs hyperparameter setting and data preprocessing before training the collected images, and the hyperparameters are set as follows: the learning rate is set to 0.01, the training rounds are 300 rounds, the training batch batchsize=16, and the optimizer type is SGD;

[0027] Data preprocessing: The original images were rotated 90° clockwise, counterclockwise and inverted, saturated by -30%-30%, and noise was added at 1.6% of the pixels. The data set was expanded by 3 times to increase the generalization ability of the model, and the images were adjusted to a pixel size of 640×640 to match the size required by the model. The data set was then divided into training set, validation set and test set in a ratio of 8:1:1.

[0028] Furthermore: in the cross-stage local partial convolution module structure, the following are included in the order of data input: convolution layer, cutting layer, multi-layer partial convolution layer and splicing layer, wherein the partial convolution layer is used to reduce the number of parameters in each cross-stage local partial convolution module, wherein the number of parameters of the partial convolution layer is much less than that of the convolution layer, and the cutting layer and the multi-layer partial convolution layer are also directly output to the splicing layer.

[0029] Furthermore: the partial self-attention module includes, in order of output and input: a convolutional layer at the input end, a multi-head self-attention module, a feedforward network including two convolutional layers, a splicing layer and a convolutional layer at the output end. The features obtained by cutting the output of the convolutional layer at the input end are evenly divided into two parts, one part is input into the multi-head self-attention module and then into the feedforward network, and the other part is directly input into the feedforward network and the splicing layer. The two parts of the features are then connected and fused through the convolutional layer at the output end. The convolutional layer at the input end also directly outputs all the output features to the splicing layer.

[0030] Compared with the prior art, the invention has the following beneficial effects: the invention provides a new solution for efficient and accurate detection of coal gangue and can effectively identify coal gangue in a coal pile. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 A schematic diagram of the network structure of the lightweight improved YOLOv8 target detection network provided by the method of the present invention;

[0032] Figure 2 A schematic diagram of the network structure of a cross-stage local partial convolution module provided by the method of the present invention;

[0033] Figure 3 A schematic diagram of the network structure of some self-attention modules provided by the method of the present invention;

[0034] Figure 4 In the figure, (a) is the original image, (b) is the result of the method using the prior art, and (c) is the result of the method of the present invention. DETAILED DESCRIPTION

[0035] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0036] A fast identification method for coal gangue based on the lightweight improved YOLOv8 algorithm is proposed. According to the specific scenario, the coal gangue images are collected and marked by box selection to form a data set.

[0037] A lightweight and improved YOLOv8 target detection network was built with CSPDarkNet as the backbone network and PAN-FAN as the neck network, combined with a detection network including three detection heads. The improved YOLOv8 algorithm was used to train the gangue dataset, and its performance was tested through the validation set and test set.

[0038] The lightweight improved YOLOv8 target detection network specifically includes:

[0039] The backbone network includes, in order of data input, a first convolutional layer, a second convolutional layer, a first cross-stage local partial convolutional module, a third convolutional layer, a second cross-stage local partial convolutional module, a fourth convolutional layer, a third cross-stage local partial convolutional module, a fifth convolutional layer, a fourth cross-stage local partial convolutional module, and a fast spatial pyramid pooling layer; the convolutional layer is represented by Conv, and the cross-stage local partial convolutional module is represented by Csppc.

[0040] The neck network includes, in the order of data input, a second upsampling layer, a second splicing layer, a sixth cross-stage local partial convolution module, a partial self-attention module, a first upsampling layer, a first splicing layer and a fifth cross-stage local partial convolution module, a seventh convolution layer, a fourth splicing layer, an eighth cross-stage local partial convolution module, a sixth convolution layer, a third splicing layer and a seventh cross-stage local partial convolution module; the upsampling layer is represented by Upsample, the splicing layer is represented by Concat, and the partial self-attention module is represented by PSA.

[0041] The second cross-stage local partial convolution module also outputs to the first splicing layer, the third cross-stage local partial convolution module also outputs to the second splicing layer, and the fast spatial pyramid pooling layer outputs to the third splicing layer;

[0042] The fifth cross-stage local partial convolution module also outputs a first detection head;

[0043] The eighth cross-stage local partial convolution module also outputs a second detection head;

[0044] The seventh cross-stage local partial convolution module outputs a third detection head;

[0045] The trained lightweight improved YOLOv8 target detection network is used for rapid identification of coal gangue.

[0046] Among them, according to the specific experimental scenario, collecting coal gangue images and marking them to make a data set includes: selecting an electric turntable of appropriate size to ensure that it can stably support the entire simulation device. The electric turntable is placed on a flat and stable work surface to ensure that it will not shake or shift during the experiment, and two rings are fixed on the aluminum disc as the inner ring and the outer ring. The two rings will be used to fix the black belt and simulate the boundary of the coal conveyor belt. Use appropriate fasteners to firmly fix the inner and outer rings on the aluminum disc to ensure that they will not fall off or shift during rotation. Connect the power cord of the electric turntable and ensure that it works properly. Start the electric turntable through the control switch and make it start to rotate slowly. During the continuous rotation of the electric turntable, use a camera to monitor and collect data on the simulated coal conveyor belt in real time.

[0047] Prepare coal and gangue for combustion, mix the coal and gangue for combustion and place them randomly on the turntable, use a camera to take pictures of the coal transportation, and ensure that the pictures are clear and contain complete coal or gangue. Label the gangue in the collected pictures to form a data set. In this way, when training the lightweight improved YOLOv8 object detection network, you can read the pictures and label files at the same time for supervised learning.

[0048] Repeat the above steps to take more coal transport pictures and create corresponding label files to increase the size and diversity of the dataset. This will help improve the accuracy and generalization ability of the model.

[0049] Building a YOLOv8 lightweight network structure, backbone network, neck network and detection network specifically includes:

[0050] The backbone network includes, in the order of data input: a first convolution layer, a second convolution layer, a first cross-stage local partial convolution module, a third convolution layer, a second cross-stage local partial convolution module, a fourth convolution layer, a third cross-stage local partial convolution module, a fifth convolution layer, a fourth cross-stage local partial convolution module, and a fast spatial pyramid pooling layer;

[0051] The backbone network uses a series of convolution and deconvolution layers to extract features. These layers capture local features in the image through different convolution operations and restore the spatial resolution of the image through deconvolution operations. At the same time, in order to reduce the size of the network and improve performance, the structure also uses residual connections and bottleneck structures. Residual connections connect the input directly to the output, allowing the network to learn the direct relationship between the input and output, thereby reducing the complexity of the network and the difficulty of training. In the CSPDarknet structure, residual connections are used to connect adjacent convolutional layers and deconvolution layers to better transfer feature information. The partial convolution layer (Pconv) is another lightweight convolutional network commonly used in neural networks. It reduces the number of parameters in the network by performing partial convolution on the feature map.

[0052] like Figure 2 As shown, in the cross-stage local partial convolution module (Csppc module) structure, the partial convolution layer (Pconv) is used to reduce the number of parameters in each cross-stage local partial convolution module. Specifically, each cross-stage local partial convolution module includes, in the order of data input: convolution layer, cutting layer (Split), multi-layer partial convolution layer (Pconv) and splicing layer, wherein the partial convolution layer is used to reduce the number of parameters in each cross-stage local partial convolution module, wherein the number of parameters of the partial convolution layer is much less than that of the convolution layer, and the cutting layer and the multi-layer partial convolution layer are also directly output to the splicing layer.

[0053] The core of the present invention is to optimize and adjust the CSPDarknet backbone network structure. Figure 1 The structure is divided into five key parts (P1 to P5), wherein P1 includes the first convolution layer, which is a convolution kernel of size 3, stride 2 and padding 1, P2 to P4 are respectively composed of a convolution layer and a cross-stage local partial convolution module, and P5 includes a fast spatial pyramid pooling layer. Each part is carefully designed for feature extraction and fusion to achieve more efficient and accurate target detection. The structural configuration of P1 helps to reduce the spatial dimension of the feature map while maintaining important spatial information, laying the foundation for subsequent feature extraction. From P2 to P4, these three parts use the same convolution kernel and a cross-stage local partial convolution module (Csppc module). The Csppc module is an important innovation in the algorithm. It enhances the model's ability to recognize targets by fusing information from low-level feature maps and high-level feature maps, especially in the case of complex backgrounds or occlusions, which can significantly improve the accuracy and robustness of detection. And partial convolution (Pconv) is used instead of ordinary convolution on the original basis to reduce model parameters and achieve lightweight. In the P5 section, the SPPF (Fast Spatial Pyramid Pooling Layer) module is introduced, which is an important supplement to the traditional convolutional neural network. SPPF provides a multi-scale representation of the input feature map by performing pooling operations at different scales, allowing the model to capture features at different levels of abstraction. This feature is particularly critical for detecting objects of different sizes because it allows the model to focus on both detail information and global context at the same time, thereby improving the flexibility and adaptability of detection.

[0054] The lightweight and improved YOLOv8 algorithm uses the CSPDarknet backbone structure, combined with efficient feature extraction, fusion strategy and multi-scale feature representation capabilities, to achieve fast and accurate recognition of targets such as gangue. It is particularly suitable for scenarios with limited resources or requiring real-time processing, such as gangue identification in coal.

[0055] The neck network includes, in the order of data input, a second upsampling layer, a second splicing layer, a sixth cross-stage local partial convolution module, a partial self-attention module, a first upsampling layer, a first splicing layer and a fifth cross-stage local partial convolution module, a seventh convolution layer, a fourth splicing layer, an eighth cross-stage local partial convolution module, a sixth convolution layer, a third splicing layer and a seventh cross-stage local partial convolution module. The second cross-stage local partial convolution module also outputs to the first splicing layer, the third cross-stage local partial convolution module also outputs to the second splicing layer, and the fast spatial pyramid pooling layer outputs to the third splicing layer;

[0056] The fifth cross-stage local partial convolution module also outputs a first detection head;

[0057] The eighth cross-stage local partial convolution module also outputs a second detection head;

[0058] The seventh cross-stage local partial convolution module outputs the third detection head

[0059] The neck network adopts multi-scale feature fusion technology to fuse feature maps from different stages of the backbone network (Backbone) to enhance the feature representation capability. Specifically, the neck part of YOLOv8 includes an FPN module and a PAN module. The FPN module includes: the second upsampling layer, the second splicing layer, the sixth cross-stage local partial convolution module, the partial self-attention module, the first upsampling layer, the first splicing layer and the fifth cross-stage local partial convolution module; the PAN module includes the seventh convolution layer, the fourth splicing layer, the eighth cross-stage local partial convolution module, the sixth convolution layer, the third splicing layer and the seventh cross-stage local partial convolution module. The neck network aims to improve the feature representation capability of the model through multi-scale feature fusion. This part consists of two main modules: Feature Pyramid Network (FPN) and Pyramid Attention Network (PAN). The FPN module is mainly used to fuse feature maps at different scales. In a top-down manner, the information of the high-level feature map is passed to the low-level feature map, so that the low-level feature map can contain more contextual information. At the same time, the FPN module also fuses the feature maps of different layers through lateral connections, further enhancing the expressiveness of the features. The PAN module is mainly used to strengthen the important information in the feature map through the attention mechanism. By modeling the global context of the feature map, the model can better focus on the key areas in the image. This attention mechanism helps to improve the model's recognition accuracy for targets, especially in complex scenarios.

[0060] like Figure 3 As shown in the figure, the partial self-attention module includes, in the order of input and output: a convolutional layer at the input end, a multi-head self-attention module (MHSA), a feedforward network (FFN) including two convolutional layers, a splicing layer and a convolutional layer at the output end. The features obtained by cutting (split) the output of the convolutional layer (conv) at the input end are evenly divided into two parts, one part is input into the multi-head self-attention module (MHSA), and then into the feedforward network (FFN), and the other part is directly input into the feedforward network and the splicing layer. The two parts of the features are then connected and fused through the convolutional layer at the output end. The convolutional layer at the input end also directly outputs all the output features to the splicing layer.

[0061] After completing two upsamplings to restore the spatial dimension, this part adds a partial self-attention module (PSA module) in the middle of the sampling, followed by two convolution operations to refine the features, and finally significantly increases the number of channels of the feature map through four feature fusions. This strategy not only enhances the model's ability to capture subtle features, but also ensures that the features extracted from different levels are fully utilized, thereby improving the recognition accuracy of complex textures and morphologies of coal gangue.

[0062] In order to achieve the two tasks of target detection and classification, a neural network model called "detection network" is used. This detection network includes multiple detection output ends, each of which consists of two parts: a detection head and a classification head. The main function of the detection head is to generate detection results. It contains a series of convolutional layers and deconvolutional layers. The convolutional layer is used to extract the features of the input image, while the deconvolutional layer is used to restore these features to the size of the original image. In this way, the detection head generates a detection result map with the same size as the input image, in which each pixel corresponds to the confidence score of an object. The classification head is used to classify each feature map. It uses global average pooling to downsample each feature map to a vector of a fixed size. Then, this vector is sent to a fully connected layer to calculate the score of each category. Finally, by comparing these scores, the category to which each feature map belongs can be determined.

[0063] Each detection output in the detection network is connected to two 3×3 Conv modules and a 1×1 Conv2d module through two paths, respectively, to perform Bbox.Loss detection and Cls.Loss detection. This multi-scale feature processing method enables the network to capture local detail information and global context information at the same time, effectively improving the accuracy of target positioning. Bbox.Loss is used to optimize bounding box regression to ensure that the detection box fits the target closely; while Cls.Loss focuses on classification tasks to distinguish between coal powder and coal gangue. The combination of the two achieves high-precision target detection.

[0064] In Cls.Loss detection, BCE (binary cross entropy) is used as the classification loss. Each category is judged as "whether it is this category" and the confidence is output. The calculation formula is:

[0065]

[0066] Where N is the number of samples, L i Represents the cross entropy loss function for each sample, y i Represents the sample label, the positive sample label is 1, the negative sample label is 0, p i Represents the probability predicted by the model.

[0067] In the Bbox.Loss detection, DFL loss is used as the regression loss. DFL loss focuses on the values near the labels, enabling the network to quickly converge to the distribution of the target position and its adjacent regions. The DFL formula is as follows:

[0068] (S i , s i+1 ) = -((y i+1 - y) log(s i ) + (y - y i ) log(s i+1 ))

[0069] Si and Si+1 are the "predicted values" and "adjacent predicted values" output by the network. y, yi, and yi+1 are the "actual values", "label integral values", and "adjacent label integral values" of the labels;

[0070] At the same time, IOU Loss is used as the regression loss. IOU Loss measures the overlap degree between the predicted bounding box and the true bounding box to optimize the localization accuracy. The IOU Loss formula is as follows:

[0071]

[0072] A is the predicted box and B is the true box.

[0073] Before training the improved YOLOv8 object recognition network on the collected pictures, hyperparameter settings and data preprocessing are carried out. The hyperparameter settings are as follows: the learning rate is set to 0.01, the number of training epochs is 300, the training batch size batchsize = 16, and the optimizer type is SGD.

[0074] Perform data preprocessing: Rotate the original pictures 90° clockwise and counterclockwise and invert them, perform saturation processing of -30% - 30%, and add noise processing of 1.6% pixels. Expand the dataset by 3 times to increase the generalization ability of the model, and adjust the pictures to a pixel size of 640×640 to adapt to the size required by the model; then divide the dataset into a training set, a validation set, and a test set according to the ratio of 8:1:1.

[0075] The accuracy, recall rate, and mAP50 value obtained after training the original YOLOv8 algorithm are 0.8, 0.585, and 0.723 respectively; the three types of indicators of the YOLOv8 algorithm based on lightweight improvement are 0.873, 0.647, and 0.785 respectively. Compared with them, there are respective increases of 9.1%, 10.6%, and 8.6%, and there is a relatively large improvement in accuracy. Moreover, the improved algorithm has 820,000 fewer parameters and 2G less computational volume compared with the original algorithm, achieving lightweight and better detection efficiency.

[0076] Such as Figure 4 As shown, Figure 4 (a) is the original image, and the results of the prior art method are shown in Figure 4 (b) and the result of the method of the present invention are shown in FIG. Figure 4 By comparing with (c) in the figure, it can be seen that the detection confidence of the method of the present invention for coal gangue is higher and has a higher recognition efficiency.

[0077] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for rapid identification of coal gangue based on a lightweight improved YOLOv8 algorithm, characterized in that: Simulate according to specific scenarios, collect coal gangue images and mark them by box selection to create a data set; A lightweight and improved YOLOv8 target detection network was built with CSPDarkNet as the backbone network and PAN-FAN as the neck network, combined with a detection network including three detection heads. The improved YOLOv8 algorithm was used to train the gangue dataset, and its performance was tested through the validation set and test set. The lightweight improved YOLOv8 target detection network specifically includes: The backbone network includes, in order of data input: a first convolutional layer, a second convolutional layer, a first cross-stage local partial convolutional module, a third convolutional layer, a second cross-stage local partial convolutional module, a fourth convolutional layer, a third cross-stage local partial convolutional module, a fifth convolutional layer, a fourth cross-stage local partial convolutional module, and a fast spatial pyramid pooling layer; The neck network includes, in order of data input: a second upsampling layer, a second splicing layer, a sixth cross-stage local partial convolution module, a partial self-attention module, a first upsampling layer, a first splicing layer and a fifth cross-stage local partial convolution module, a seventh convolution layer, a fourth splicing layer, an eighth cross-stage local partial convolution module, a sixth convolution layer, a third splicing layer and a seventh cross-stage local partial convolution module; The second cross-stage local partial convolution module also outputs to the first splicing layer, the third cross-stage local partial convolution module also outputs to the second splicing layer, and the fast spatial pyramid pooling layer outputs to the third splicing layer; The fifth cross-stage local partial convolution module also outputs a first detection output terminal; The eighth cross-stage local partial convolution module also outputs a second detection output terminal; The seventh cross-stage local partial convolution module outputs a third detection output terminal; The trained lightweight improved YOLOv8 target detection network is used for rapid identification of coal gangue.

2. The method for rapid identification of coal gangue based on the lightweight improved YOLOv8 algorithm according to claim 1 is characterized in that: The backbone network is divided into five parts according to P1 to P5, wherein P1 includes the first convolution layer, which is a convolution kernel with a size of 3, a step size of 2 and a padding of 1; P2 to P4 are respectively composed of a convolution layer and a cross-stage local partial convolution module; and P5 includes a fast spatial pyramid pooling layer.

3. The method for rapid identification of coal gangue based on the lightweight improved YOLOv8 algorithm according to claim 1 is characterized in that: Each detection output in the detection network is connected to two 3×3 Conv modules and one 1×1 Conv2d module through two paths to perform Bbox.Loss detection and Cls.Loss detection respectively.

4. The method for rapid identification of coal gangue based on the lightweight improved YOLOv8 algorithm according to claim 3 is characterized in that: In Cls.Loss detection, binary cross entropy is used as the classification loss. Each category is judged as "whether it is this category" and the confidence is output. The calculation formula is: Y represents the sample label, the positive sample label is 1, the negative sample label is 0, and p i Indicates the probability of predicting a positive sample, N is the number of samples, y i represents the sample label, L i Represents the cross entropy loss function for each sample; In Bbox.Loss detection, DFL loss is used as the regression loss. DFL loss focuses on the value near the label, so that the network quickly converges to the distribution of the target location and its neighboring areas. The DFL loss formula is: (S i ,s i+1 )=-((and i+1 -y)log(s i )+(yy i )log(s i+1 )) Si and Si+1 are the "predicted value" and "near predicted value" output by the network, and y, yi, and yi+1 are the "actual value", "label integral value", and "near label integral value" of the label; At the same time, IOU Loss is used as the regression loss. IOU Loss measures the overlap between the predicted bounding box and the true bounding box to optimize the positioning accuracy. The IOU Loss formula is: A is the predicted box and B is the true box.

5. The method for rapid identification of coal gangue based on the lightweight improved YOLOv8 algorithm according to claim 1 is characterized in that: The improved YOLOv8 target recognition network performs hyperparameter setting and data preprocessing before training the collected images. The hyperparameter settings are: learning rate is set to 0.01, training rounds are 300 rounds, training batch batchsize=16, and optimizer type is SGD; Data preprocessing: The original images were rotated 90° clockwise, counterclockwise and inverted, saturated by -30%-30%, and noise was added at 1.6% of the pixels. The data set was expanded by 3 times to increase the generalization ability of the model, and the images were adjusted to a pixel size of 640×640 to match the size required by the model. The data set was then divided into training set, validation set and test set in a ratio of 8:1:

1.

6. The method for rapid identification of coal gangue based on the lightweight improved YOLOv8 algorithm according to claim 1 is characterized in that: In the cross-stage local partial convolution module structure, the following are included in the order of data input: convolution layer, cutting layer, multi-layer partial convolution layer and splicing layer. The partial convolution layer is used to reduce the number of parameters in each cross-stage local partial convolution module. The number of parameters of the partial convolution layer is much less than that of the convolution layer. The cutting layer and the multi-layer partial convolution layer are also directly output to the splicing layer.

7. The method for rapid identification of coal gangue based on the lightweight improved YOLOv8 algorithm according to claim 1 is characterized in that: The partial self-attention module includes, in the order of output and input: a convolutional layer at the input end, a multi-head self-attention module, a feedforward network including two convolutional layers, a splicing layer and a convolutional layer at the output end. The features obtained by cutting the output of the convolutional layer at the input end are evenly divided into two parts. One part is input into the multi-head self-attention module and then into the feedforward network, and the other part is directly input into the feedforward network and the splicing layer. The two parts of the features are then connected and fused through the convolutional layer at the output end. The convolutional layer at the input end also directly outputs all the output features to the splicing layer.

Citation Information

Cited By

  • Coal sorting method, system, medium and equipment

    CN120673181A

  • A coal sorting method, system, medium and equipment

    CN120673181B