Lightweight non-coal foreign matter image recognition algorithm based on feature enhancement and Slim-Neck

The lightweight YOLOv7-tiny model with Slim-Neck architecture and feature enhancement addresses the inefficiencies of existing coal conveyor belt detection systems, enhancing detection speed and accuracy while reducing model complexity and resource requirements.

CN120318648APending Publication Date: 2025-07-15CHENGDU UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510379580.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing non-coal foreign matter detection algorithms have problems such as low detection efficiency, high cost, false detection and missed detection in coal mine conveying equipment. In addition, the YOLOv7-tiny model has many channels, high model complexity and large parameters, making it difficult to deploy efficiently on hardware equipment with limited resources.

Method used

The Slim-Neck structure is designed based on the GSConv module, combined with partial convolution and feature enhancement modules, and built a lightweight YOLO-Coal network model, including using PConv to replace backbone network convolution, designing MAB-Y module to enhance feature extraction, and optimizing the network structure through the VoV-GSCSP module and BAM mechanism.

Benefits of technology

It realizes efficient detection of non-coal foreign matter on hardware equipment with limited resources, improves detection accuracy and speed, reduces the amount of model parameters, and is suitable for deployment in coal mine conveying equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention provides a lightweight non-coal foreign matter image recognition algorithm based on feature enhancement and Slim-Neck, and belongs to the field of deep learning and image processing. In order to solve the problems that a YOLOv7-tiny model is too many in channel number, high in complexity, large in parameter quantity and the like, a Slim-Neck structure is designed to replace an original Neck structure so as to reduce the size of the model; a convolution layer in the backbone network is replaced by partial convolution, so that the calculation efficiency is improved, and the reasoning speed is increased; and an improved feature enhancement module is embedded between the backbone network and the Slim-Neck structure, so that the feature extraction and fusion capability is improved, and the problems of missing detection, false detection and low precision of foreign matters are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer vision and deep learning, and particularly relates to a lightweight YOLOv7-tiny network model based on feature enhancement and Slim-Neck for non-coal foreign object detection. Background Art

[0002] In coal mine conveying equipment, belt conveying is a common coal transportation method and plays a very important role in the production efficiency of coal. During the transmission process, the contact between the conveyor and external sharp objects (such as anchor bolts, gangue, iron blocks, etc.) may cause the belt to be scratched, thus seriously affecting the production and transportation of the entire project. In the actual transportation of coal, it is necessary to accurately and quickly detect foreign objects at the early stage when they fall onto the belt and remove them in time to minimize the damage of foreign objects to the belt.

[0003] Traditional detection algorithms mainly include ray method, spectral recognition method, image recognition method, etc. However, traditional detection algorithms have problems such as low detection efficiency, high cost, false detection, and missed detection. Deep learning technology has excellent feature extraction ability in object detection, with higher accuracy and faster detection speed, thus having obvious performance advantages in foreign object detection.

[0004] The VoV-GSCSP module inherits the advantages of the GSConv module and GS Bottleneck. Through the new skip connection branch, the VoV-GSCSP module has stronger non-linear representation, effectively solving the problem of gradient disappearance. At the same time, the split-channel method of VoV-GSCSP realizes rich gradient combinations, solves the problem of redundant gradient information, and improves the learning ability.

[0005] The present invention constructs a YOLO-Coal non-coal foreign object detection network model to solve problems such as the large number of channels, high model complexity, and large number of parameters of YOLOv7-tiny, while ensuring its recognition accuracy while improving the detection speed of the model for coal. Summary of the Invention

[0006] The present invention proposes a lightweight YOLOv7-tiny network model based on feature enhancement and Slim-Neck. The method has been improved in many aspects to achieve higher detection accuracy and faster detection speed, while reducing the volume of the model, making it more suitable for deployment on hardware devices with limited resources. The method includes the following steps:

[0007] S10, designing a Slim-Neck structure based on the GSConv module;

[0008] S20, lightweighting the backbone network using partial convolution;

[0009] S30. Design the design feature enhancement module MAB-Y to improve the detection accuracy of the lightweight model.

[0010] S40. Design the YOLO-Coal non-coal foreign object recognition algorithm model.

[0011] Specifically,

[0012] S101. Adopt a structure combining standard convolution and depthwise separable convolution - the GSConv module - to design the GS Bottleneck module in the neck network structure. The GSConv module combines standard convolution, depthwise separable convolution, and the Shuffle module. This module combines the advantages of the two perfect convolution modules, and its network structure is as Figure 1 shown. The structure of the GSBottleneck module is as Figure 2 shown.

[0013] S102. Build the VoV-GSCSP module on the basis of the GSConv module and the GS Bottleneck module by using the method of one-time aggregation, as Figure 3 shown.

[0014] S103. Design the Slim-Neck structure according to the VoVNet and CSPNet ideas. Apply the GSConv module and the VoV-GSCSP module to the structure design of the Silm-Neck of the model at the same time to maintain detection accuracy while reducing the number of parameters. The Silm-Neck structure is shown in Figure 4.

[0015] S201. The research content of the present invention needs to meet the requirements of real-time detection. In addition to reducing the number of parameters of the network model as much as possible, it also needs to have a relatively fast operation speed. To further reduce the number of parameters of the model and improve the operation speed of the model, partial convolution (PConv) is used to replace the ordinary convolution in the backbone network, which can achieve the purpose of high FLOPS (floating point operations per second) and low FLOPs. The structure of PConv is as Figure 5 shown.

[0016] S203. Use the PConv module and combine a convolution module, a BN module, and a LeakyReLu activation function behind the PConv to construct the PCBL module in the backbone network.

[0017] S204. Build the ELAN-P module in the backbone network based on the PCBL module in step S203.

[0018] S301. Although PConv can bring low parameter quantity and high computing power, it will also cause information loss, affecting the accuracy and generalization ability of the model. Therefore, in the present invention, a feature enhancement module is embedded before the feature output layer to strengthen the extraction of feature information. The "Bottleneck Attention Module" (BAM) has the ability to improve the network feature expression. This module adopts a parallel manner of spatial attention mechanism and channel attention mechanism. By adjusting the weights of the feature map, it retains useful information, eliminates redundancy, and improves the robustness and generalization performance of the model. As Figure 6 shown.

[0019] S302. In order to make more full use of global context information and enhance the feature extraction ability of occlusions and foreign objects in low-light images, the MAB-Y feature enhancement module is proposed to improve the ability of the MAB module.

[0020] The constructed MAB-Y module is as Figure 7 shown.

[0021] S401. In the backbone network part, partial convolution (PConv) is used to replace traditional convolution to reduce the network parameter quantity while maintaining high computing power. The PConv module, BN module, and LeakyReLu activation function are used to extract flame features. In the neck network, the Silm-Neck structure is adopted, which combines the advantages of ordinary convolution and depthwise separable convolution. While maintaining detection accuracy, it reduces the parameter quantity, achieving a lightweight effect and facilitating subsequent model deployment. In addition, the MAB-Y feature enhancement module is added during the feature extraction process to enhance the model's ability to extract and fuse feature information. After the above improvements, the final network structure of YOLO-Coal is as Figure 8 shown. Description of the Drawings

[0022] Figure 1 The GSConv structure diagram in the present invention.

[0023] Figure 2 The GS Bottleneck structure diagram in the present invention.

[0024] Figure 3 The VoV-GSCSP module structure diagram in the present invention.

[0025] Figure 4 The Silm-Neck structure diagram in the present invention.

[0026] Figure 5 The PConv structure diagram in the present invention.

[0027] Figure 6 The BAM module structure diagram in the present invention.

[0028] Figure 7 Structure diagram of the MAB-Y module in the present invention.

[0029] Figure 8 YOLO-Coal network structure obtained by improving YOLOv7-tiny in the present invention. Specific implementation manner

[0030] Next, in combination with the accompanying drawings in the present invention, the technical solutions of the present invention will be clearly and completely described. As Figure 8 shown, this embodiment provides a YOLO-Coal network structure model obtained by improving YOLOv7-tiny, including the following steps:

[0031] Referring to Figure 1 , specific descriptions are made for the above steps S101 and S102. The following is the specific process of constructing the GSConv structure. "Conv" is composed of a convolutional layer, a BN layer, and an activation function layer. First, the feature map is divided into two parts after passing through the standard convolutional module. Half of the feature map is used for depthwise separable convolution, and the remaining part is used for standard convolution. Then, the two parts after convolution are feature-connected. Finally, the Shuffle technique is used to merge the feature maps of the two channels for feature concatenation, so that the feature information generated by the Conv module is completely mixed into the feature information generated by the DWConv module through various channels. As a channel mixing technique, Shuffle transmits its feature information on each channel, so that the information from the Conv module is completely mixed with the output channels of DWConv, thus realizing channel information interaction.

[0032] In Figure 2 , specific descriptions are made for step S103. The two branches of the GS Bottleneck perform separate convolutions without sharing weights, and the channel information is propagated to different network paths by reducing the number of channels. Therefore, the information propagated through the GSBottleneck can obtain greater correlation and diversity, ensuring more accurate image information while reducing the computational amount.

[0033] In Figure 3 , specific descriptions are made for step S104. First, the input feature information is adjusted to 0.5 times the original number of channels through convolution operations. Then, the output result of the GS Bottleneck module is concatenated with the convolution operation result of the input feature map. Finally, it is output through a convolution module.

[0034] Figure 5Shown is the PConv structure described in step S201. PConv can reduce redundant calculations and memory accesses, thus better utilizing the computing power on the device. It only needs to apply ordinary convolution on a part of the input channels for spatial feature extraction while keeping the remaining channels unchanged. For continuous or regular memory access, the first or last continuous channels are regarded as representatives of the entire feature map for calculation, reducing the computational amount and memory space of the model. The calculation formula for the FLOPs of PConv is shown as follows.

[0035]

[0036] In the formula represents the height of the input feature map, represents the width of the input feature map, represents the number of convolution kernels in the convolutional layer, represents the number of convolutional channels for feature extraction.

[0037] For a typical ratio , is the number of channels for which ordinary convolution is used for feature extraction. The FLOPs of PConv are only 1 / 16 of those of ordinary convolution. In addition, PConv also has a smaller memory access amount, and the calculation formula is as follows:

[0038]

[0039] For , the memory access amount of PConv is only 1 / 4 of that of ordinary Conv, reducing the redundant calculations caused by memory access and improving .

[0040] As can be seen from the formula, PConv can reduce the computational amount while still maintaining the efficient operation speed of the model and enhancing the generation ability of the feature map. In addition, PConv can better utilize the computing device capabilities, thus bringing benefits to the training and inference of the foreign object detection model.

[0041] Figure 6 is the original BAM module structure described in step S301. On the channel attention branch, in order to aggregate the feature maps in each channel, first, global average pooling is performed on the feature map F to generate the channel vector , which encodes the global information in each channel. Then, a multi-layer perceptron (MLP) with one hidden layer is used to estimate the cross-channel attention of the channel vector. After the MLP, a batch normalization (BN) layer is added to adjust the scale of the output of the spatial branch.

[0042] The spatial attention branch emphasizes features at different spatial positions, utilizes a large receptive field range to effectively utilize context information, and generates spatial attention. . The "bottleneck structure" proposed by ResNet is adopted, which saves both the number of parameters and the computational overhead. First, the feature map is integrated and compressed in channels using 1×1 convolution. Then, two 3×3 dilated convolutions are used to expand the field of view and effectively utilize context information. Finally, the feature map is simplified again to using 1×1 convolution for the spatial attention map, and a BN layer is used at the end of the spatial branch to adjust the scale.

[0043] Figure 7 Step S302's MAB-Y module is obtained by improving the MAB module in step S301. In this module, in order to make more full use of global context information and enhance the feature extraction ability of foreign objects on occluded and low-light images, an MAB-Y feature enhancement module is proposed to improve the ability of the MAB module. The global average pooling operation in the channel attention branch of the MAB module only operates by averaging each channel feature map separately and uses this average value to represent the channel feature map information, without making full use of global information. In this paper, the global average pooling operation of MAB is improved. Instead of using global average pooling, one-dimensional convolution operation is used to obtain the dependency relationship of the pixel positions of each channel feature map, so as to obtain an output vector, that is, a probability value vector is obtained through the Softmax function operation and multiplied by the original input feature map. This improvement can make full use of the full text information of the feature map to obtain the connection between each channel pixel point, and the obtained probability value vector makes the key information on the feature map more prominent, which is more convenient for subsequent learning and enhances the feature extraction ability of foreign objects on partially occluded foreign objects and dust images.

[0044] In Figure 8 , the final network structure of YOLO-Coal is designed based on all the above steps. In this structure, the backbone network part uses partial convolution (PConv) to replace traditional convolution to reduce the number of network parameters while maintaining high computational efficiency, and PConv modules, BN modules, and LeakyReLu modules are used to extract flame features. The neck network adopts the Silm-Neck structure, which combines the advantages of ordinary convolution and depthwise separable convolution, reduces the number of parameters while maintaining detection accuracy, achieving a lightweight effect and facilitating subsequent model deployment. In addition, an MAB-Y feature enhancement module is added during the feature extraction process to enhance the model's feature information extraction and feature information fusion capabilities.

Claims

1. A lightweight non-coal foreign object image recognition algorithm based on feature enhancement and Slim-Neck, characterized in that The method includes the following steps: S10. Design the Slim-Neck structure of the neck network, and use a structure that combines standard convolution and depthwise separable convolution - the GSConv module to design and build the neck network structure; S20. Use partial convolution (PConv) to lightweight the backbone network; S30. Add the MAB-Y feature enhancement module during the feature extraction process; S40. Based on the above three steps, design the YOLO-Coal non-coal foreign object image recognition algorithm model.

2. The lightweight non-coal foreign object image recognition algorithm based on feature enhancement and Slim-Neck according to claim 1, characterized in that The design steps of the Slim-Neck structure are as follows: S101. Introduce the GS Bottleneck module based on the GSConv module, which ensures more accurate image information while reducing the computational amount; S102. Adopt a one-time aggregation method to build the VoV-GSCSP module based on the GSConv module and the GS Bottleneck module, effectively solving the problems of gradient disappearance and redundant gradient information, and improving the learning ability; S103. Apply the GSConv module and the VoV-GSCSP module to the structure design of the Silm-Neck of the model at the same time to maintain detection accuracy while reducing the number of parameters; For a lightweight non-coal foreign object image recognition algorithm based on feature enhancement and Slim-Neck according to claim 1, wherein step S20 includes the following steps: S201. Replace the ordinary convolution in the PCBL module of the backbone network with partial convolution (PConv), and then combine the convolution module after PConv to form the PCBL module in the backbone network; ​ S202. Use the PCBL module obtained in step S201 to construct the ELAN-P module in the backbone network.

3. A lightweight non-coal foreign object image recognition algorithm based on feature enhancement and Slim-Neck according to claim 1, characterized in that Step S30 includes the following steps: S301. Based on the original BAM module, use one-dimensional convolution operation to obtain the dependence relationship of the pixel positions of each channel feature map, so as to obtain the output vector; S302. Obtain the probability value vector through the Softmax function operation to improve the prediction result of the model; S303. Multiply the probability value vector by the original input feature map, S304. Obtain the new MAB-Y feature enhancement module.