A Method and System for Object Detection of Conveyor Belt Workpiece Images Based on Lightweight YOLOV4-tiny

By adopting CSPDarknet53-tiny and FPN networks in the YOLOV4-tiny algorithm and integrating the efficient channel attention mechanism ECA, the real-time object detection problem of the existing YOLO series algorithms on computing power and memory limited devices is solved, and efficient and accurate object detection effect is achieved.

CN114898200BActive Publication Date: 2025-06-24XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210577168.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-25
Publication Date
2025-06-24
Estimated Expiration
2042-05-25

AI Technical Summary

Technical Problem

The existing YOLO series algorithms require powerful GPU computing power in real-time object detection, making it difficult to achieve efficient object detection on devices with limited computing power and memory.

Method used

The lightweight YOLOV4-tiny algorithm is adopted to enhance the network by constructing the CSPDarknet53-tiny backbone feature extraction network and FPN feature enhancement network, and integrate the efficient channel attention mechanism ECA to reduce network parameters and computing complexity, and adapt to the computing power and memory of ordinary devices.

Benefits of technology

It realizes efficient and accurate object detection on devices with limited computing power and memory, avoids dependence on powerful GPU computing power, and meets the lightweight deployment needs of industrial scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114898200B_ABST
    Figure CN114898200B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for target detection of conveyor belt workpiece images based on lightweight YOLOV4-tiny, belonging to the technical field of target detection. The specific steps are as follows: constructing a training image set of conveyor belt workpieces; constructing a lightweight YOLOV4-tiny conveyor belt workpiece image target detection network, including a backbone feature extraction network CSPDarknet53-tiny, an enhanced feature extraction network FPN integrated with an efficient channel attention mechanism ECA, and a prediction feature layer YOLO Head; training the lightweight YOLOV4-tiny conveyor belt workpiece image target detection network with the training image set to obtain a lightweight YOLOV4-tiny conveyor belt workpiece image target detection network model; inputting the conveyor belt workpiece image to be detected into the lightweight YOLOV4-tiny conveyor belt workpiece image target detection network model to obtain the target detection result of the conveyor belt workpiece image. The detection method of the present invention combines the requirements of lightweight deployment in industrial scenarios to save computer resources and memory, can improve the detection speed while improving the detection accuracy, and further realizes efficient and accurate target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and more particularly to a method and system for detecting conveyor belt workpiece image targets based on lightweight YOLOV4-tiny. Background Art

[0002] With the rapid development of artificial intelligence, as a branch of the field of computer vision, object detection technology has achieved many breakthrough results. Thanks to the technological breakthroughs, object detection technology has gradually moved towards practical applications and has been widely used in many fields such as industrial manufacturing, video surveillance, and national defense and military. For industrial manufacturing, this technology can replace manual vision for automatic sorting, defect detection, counting and packaging, production quality assessment, etc. of various types of components on the production line. It not only avoids missed inspections, misinspections, and low production efficiency of components on the production line, but also can mark various components and analyze the production progress.

[0003] Existing object detection methods are mainly divided into One-Stage and Two-Stage algorithms. Comparatively speaking, One-Stage algorithms have better accuracy and real-time detection speed. The YOLO series of detection algorithms is a typical One-Stage algorithm model. However, the network structures of the YOLO series of algorithms and their improved algorithms are complex and have many network parameters. They require powerful GPU (graphics processing unit) computing power to achieve real-time object detection. Especially in real-world applications where real-time object detection needs to be performed on some mobile devices and industrial robots, the computing power and memory of ordinary devices are limited. Summary of the Invention

[0004] In order to solve the problems existing in the prior art, the present invention provides a method and system for detecting conveyor belt workpiece image targets based on lightweight YOLOV4-tiny. Combining the requirements of lightweight deployment in industrial scenarios to save computer resources and memory, it can improve the detection speed while improving the detection accuracy, and further achieve efficient and accurate object detection.

[0005] To achieve the above object, the present invention provides the following technical solution: A method for detecting conveyor belt workpiece image targets based on lightweight YOLOV4-tiny, the specific steps are as follows:

[0006] S1 Construct a training image set of conveyor belt workpieces;

[0007] S2 Construct a lightweight YOLOV4-tiny conveyor belt workpiece image target detection network, including a backbone feature extraction network CSPDarknet53-tiny, an enhanced feature extraction network FPN, and a prediction feature layer YOLO Head, wherein an efficient channel attention mechanism ECA is incorporated into the enhanced feature extraction network FPN;

[0008] S3 uses the training image set of the conveyor belt workpieces to train the lightweight YOLOV4-tiny conveyor belt workpiece image target detection network, and obtains the lightweight YOLOV4-tiny conveyor belt workpiece image target detection network model;

[0009] S4 inputs the conveyor belt workpiece image to be detected into the lightweight YOLOV4-tiny conveyor belt workpiece image target detection network model, and obtains the target detection result of the conveyor belt workpiece image.

[0010] Further, in step S2, the feature enhancement network FPN includes three efficient channel attention mechanisms ECA. After two effective feature layers of the backbone feature extraction network CSPDarknet53-tiny, an efficient channel attention mechanism ECA is connected, and an efficient channel attention mechanism ECA is connected after the upsampling of the feature enhancement network FPN.

[0011] Further, in step S2, the backbone feature extraction network CSPDarknet53-tiny extracts features from the conveyor belt workpiece image to be detected and outputs two-scale feature maps feat1 and feat2 of 13×13 and 26×26. The feature maps feat1 and feat2 are respectively input into the efficient channel attention mechanism ECA for channel weighting of the feature maps to obtain feature maps f1 and f2. The feature enhancement network FPN performs one upsampling on the feature map f2 and then passes through the efficient channel attention mechanism ECA to obtain the feature map f2'. The output feature map f2' is stacked with the feature map f1 and the feature map f2.

[0012] Further, in step S2, the efficient channel attention mechanism ECA uses a global average pooling layer to reduce the dimension of the original feature map, and the reduced-dimension original feature map is subjected to feature extraction through a one-dimensional convolution to obtain a feature extraction output map; the sigmoid activation function is used to obtain the corresponding channel probability value of the feature extraction output map, and the corresponding channel probability value is multiplied by the original feature map to obtain the output of the efficient channel attention mechanism ECA.

[0013] Further, in step S2, the convolution kernel size K of the one-dimensional convolution is:

[0014]

[0015] In the formula, |*| odd represents the nearest odd number, and C is the channel dimension.

[0016] Further, Matrix Non-Maximum Suppression (Matrix NMS) is used to perform parallel screening on the prediction boxes of the prediction feature layer YOLO Head in a linearly decaying manner.

[0017] Further, in step S2, the specific steps of the screening are as follows:

[0018] S2.1 Arrange the detection frames of the prediction feature layer YOLO Head in descending order according to the class confidence, and obtain a list of bounding boxes;

[0019] S2.2 Move the bounding box with the highest confidence score in the list to the output bounding box list;

[0020] S2.3 Calculate the areas of all the remaining detection frames in the output bounding box list;

[0021] S2.4 Calculate the intersection over union (IoU) between all the remaining bounding boxes in the bounding box list and the bounding box with the highest confidence score;

[0022] S2.5 Calculate the linear decay coefficient decay of the bounding boxes with an IoU greater than the threshold j , and change the confidence of the candidate bounding box. The score reset function is:

[0023]

[0024] where s i is the confidence score; M is the reference bounding box with the highest confidence; b i is the other candidate bounding box for calculating the IoU with the reference bounding box; N t represents the set IoU threshold;

[0025] The linear decay coefficient decay j has the formula:

[0026]

[0027] where: f(iou i,j ) = 1 - iou(i,j),

[0028] S2.6 Repeat steps S2.1 - S2.5 until the bounding box list is empty. At this time, the bounding boxes in the output bounding box list are the finally screened prediction bounding boxes.

[0029] Further, the intersection over union (IoU) represents the area ratio of the intersection and union parts between the generated rectangular bounding boxes. Specifically: IoU = (A ∩ B) / (A ∪ B), where A and B represent different rectangular bounding boxes.

[0030] Further, in step S1, the conveyor belt workpiece image labeled with the target to be detected is pre - processed to obtain a training image, and the training image is collected in the training image set.

[0031] The present invention also provides a conveyor belt workpiece image target detection system based on lightweight YOLOV4-tiny, including:

[0032] An image acquisition module for constructing a training image set of conveyor belt workpieces;

[0033] A network construction module for constructing a lightweight YOLOV4-tiny conveyor belt workpiece image target detection network, including a backbone feature extraction network CSPDarknet53-tiny, an enhanced feature extraction network FPN, and a prediction feature layer YOLO Head. An efficient channel attention mechanism ECA is incorporated into the enhanced feature extraction network FPN;

[0034] A network training module for training the lightweight YOLOV4-tiny conveyor belt workpiece image target detection network using the training image set of conveyor belt workpieces to obtain all trained lightweight YOLOV4-tiny conveyor belt workpiece image target detection network models;

[0035] A retrieval module for inputting the conveyor belt workpiece image to be detected into the lightweight YOLOV4-tiny conveyor belt workpiece image target detection network model to obtain the target detection result of the conveyor belt workpiece image.

[0036] Compared with the prior art, the present invention has at least the following beneficial effects:

[0037] The present invention provides a conveyor belt workpiece image target detection method based on lightweight YOLOV4-tiny. By constructing a lightweight YOLOV4-tiny conveyor belt workpiece image target detection network, compared with other YOLO series methods, the method of the present invention has a simple network structure, fewer network parameters, good real-time detection speed, and less memory occupied by the model. This method does not require powerful GPU (graphics processing unit) computing power to achieve real-time target detection, and is better adapted to general devices with limited computing power and memory, such as some mobile devices and industrial robots in practical applications. It fully meets the requirements of lightweight deployment in industrial scenarios for saving computer resources and memory.

[0038] Furthermore, the CSPDarknet53-tiny network is adopted as the backbone feature extraction network of the lightweight YOLOV4-tiny conveyor belt workpiece image target detection network. The efficient channel attention mechanism ECA is incorporated into the enhanced feature extraction network FPN. On the one hand, the efficient channel attention mechanism ECA is used to effectively extract the channel semantic information of the two feature layers output by the backbone feature extraction network, and efficiently utilize the feature information captured by the backbone network. On the other hand, the enhanced feature network FPN structure incorporating the efficient channel attention mechanism ECA is used to capture multi-scale features, enhance the expression of top-level feature information, explore the interdependence between channel graphs, enhance the feature representation of specific semantics, make the finally trained model more discriminative for conveyor belt images, and improve the detection efficiency.

[0039] Furthermore, the Matrix Non-Maximum Suppression (Matrix NMS) is adopted to parallelize the screening of the prediction boxes of the prediction feature layer YOLO Head in a linearly decaying manner, improving the problem that the traditional non-maximum suppression algorithm may misdelete prediction boxes with high confidence, avoiding inaccurate threshold setting and poor detection effects caused by overly dense conveyor belt workpieces, so as to realize the precise detection of reagent bottles, improve the detection speed while improving the detection accuracy, and further achieve efficient and accurate workpiece detection.

[0040] Furthermore, the training image set is used to iteratively train the target detection model, and the model with the highest detection accuracy is selected from all the trained target detection models as the optimal target detection model. The detection image is input into the optimal target detection model to obtain the target detection result of the detection image, completing the target detection of the conveyor belt workpiece. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is Figure 1 a schematic flowchart of a lightweight conveyor belt workpiece detection method based on YOLOV4-tiny according to the present invention;

[0042] Figure 2 is the network structure diagram of YOLOV4-tiny in the present invention;

[0043] Figure 3 is the structure diagram of the enhanced feature extraction network incorporating ECA in the present invention;

[0044] Figure 4 is the structure diagram of the efficient channel attention mechanism ECA;

[0045] Figure 5 is the IOU representation diagram;

[0046] Figure 6Data flow diagram of a lightweight conveyor belt workpiece detection method based on YOLOV4-tiny according to the second embodiment of the present invention;

[0047] Figure 7 Conveyor belt workpiece diagram of the present invention;

[0048] Figure 8 Conveyor belt workpiece detection result diagram of the present invention;

[0049] Figure 9 Process diagram of detecting and grasping conveyor belt workpieces using the method of the present invention; Detailed implementation manners

[0050] The present invention will be further described below in conjunction with the accompanying drawings and detailed implementation manners.

[0051] As Figure 1 shown, a conveyor belt workpiece image target detection method based on lightweight YOLOV4-tiny, the specific steps are as follows:

[0052] S1. Preprocess the conveyor belt workpiece image marked with the target to be detected to obtain a training image, and collect the training image in a training image set;

[0053] S2. Use the CSPDarknet53-tiny network as the backbone feature extraction network of YOLOV4-tiny, construct a target detection model of lightweight YOLOV4-tiny, incorporate efficient channel attention (ECA) into the enhanced feature extraction network FPN to enhance the expression of top-level feature information, and use matrix non-maximum suppression (Matrix NMS) in post-processing to perform parallel screening on the prediction boxes of the prediction feature layer YOLO Head in a linearly decaying manner;

[0054] S3. Iteratively train the target detection model using the training image set, and select the model with the highest detection accuracy among all the trained target detection models as the optimal target detection model;

[0055] S4. Input the detection image into the optimal target detection model to obtain the target detection result of the detection image.

[0056] In step S1, the conveyor belt artifact image is directly obtained, or the conveyor belt artifact image is extracted from the captured video. For example, the captured video is cut every 25 frames to obtain a conveyor belt artifact image with a resolution of 960×600 pixels. The LabelImg annotation tool is used to annotate the target to be detected in the artifact image, that is, the conveyor belt artifact target. The annotation used is an XML file in the PASCAL VOC format. The annotated conveyor belt artifact image is preprocessed by image cropping, image flipping, image enhancement, etc. to obtain a training image, and the training images are collected to obtain a training image set.

[0057] In step S2, the network structure of YOLOV4-tiny is as follows Figure 2 As shown in the figure, as a lightweight detection network that takes into account both detection speed and accuracy, it can be divided into three parts, namely, the backbone feature extraction network CSPDarknet53-tiny, the enhanced feature extraction network FPN and the prediction feature layer YOLO Head. The backbone feature extraction network CSPDarknet53-tiny is used to extract features of the input image and output feature maps feat1 and feat2 of two scales of 13×13 and 26×26. An efficient channel attention mechanism ECA is added to the two feature maps respectively, and each channel of the two feature maps is given a weight. Then, the feature maps f1 and f2 are outputted by multiplication channel by channel weighting. After the enhanced feature extraction network FPN upsamples the feature map f2 extracted by the backbone network CSPDarknet53-tiny once, it uses the efficient channel attention mechanism ECA to output the feature map f2′, and the stacking result of f2′ and the feature map f1 and the feature map f2 are used as the output of the feature extraction network FPN, which enhances the information extraction ability of the enhanced feature extraction network. The output result of the feature extraction network FPN is passed to YOLO Head performs image regression and classification, and outputs the prediction box of the conveyor belt workpiece; the prediction box is screened using matrix non-maximum suppression (Matrix NMS) to obtain the final conveyor belt workpiece prediction box.

[0058] Furthermore, in step S2, an efficient channel attention mechanism ECA is integrated into the enhanced feature extraction network FPN. On the one hand, an efficient channel attention mechanism ECA is added to the two effective feature layers of 13×13 and 26×26 output by the backbone feature extraction network, respectively, so as to efficiently utilize the feature information captured by the backbone network; on the other hand, an attention mechanism efficient channel attention mechanism ECA is added to the up-sampled result in the feature enhancement network FPN to mine the interdependence between channel graphs and enhance the feature representation of specific semantics, so that the finally trained model is more discriminative of the conveyor belt image, that is, the salient features that can distinguish categories are generally referred to, and then more attention resources are invested in this area. The specific integration position is as follows: Figure 3 shown.

[0059] The efficient channel attention mechanism ECA effectively avoids the side effects of dimensionality reduction on channel attention prediction through one-dimensional convolution, and realizes adaptive cross-channel interaction while maintaining performance with extremely few parameters. Its schematic diagram is as Figure 4 shown. The processing process of the efficient channel attention mechanism ECA is as follows:

[0060] Step 1: Use the global average pooling layer (GAP) to modify the size of each channel of the feature map χ extracted by the previous convolution module to achieve dimensionality reduction;

[0061] Step 2: Pass the dimensionality-reduced feature map through a one-dimensional convolution that can reduce the depth of the feature map for feature extraction to obtain a feature extraction output map. When performing feature extraction, the efficient channel attention mechanism ECA captures local cross-channel interaction information by considering all channels and k neighboring channels;

[0062] Step 3: Calculate the corresponding channel probability value of the feature extraction output map through the sigmoid activation function;

[0063] Step 4: Then use the result of multiplying the probability value by the original input feature map χ as the input for the subsequent layer. Among them, the hyperparameter K is the convolution kernel size of the one-dimensional convolution, and its value represents the coverage range of local cross-channel interaction.

[0064] The value of K can be determined in an adaptive manner, and its size is obtained from its proportional relationship with the channel dimension C:

[0065]

[0066] In the formula, |*| odd represents the nearest odd number, and C is the channel dimension. γ and b are set to 2 and 1 respectively in the experiments of this article.

[0067] Furthermore, in step S2, the object detection algorithm often outputs multiple overlapping prediction boxes for the same object, resulting in false detections. In order to accurately obtain accurate prediction boxes, the matrix non-maximum suppression (Matrix NMS) is used to replace the original NMS post-processing strategy, and the prediction boxes of the prediction feature layer are screened in parallel by a linear attenuation method to improve the problem that the traditional non-maximum suppression algorithm may misdelete the prediction boxes with higher confidence, resulting in unsatisfactory detection effects. The processing process is as follows:

[0068] Step 1: According to the class confidence, sort the output detection boxes in descending order and obtain a list of bounding boxes;

[0069] Step 2: Move the bounding box with the highest confidence score in the list to the output bounding box list;

[0070] Step 3: Calculate the areas of all the remaining detection boxes in the output bounding box list;

[0071] Step 4: Calculate the intersection over union (IOU) between all the remaining bounding boxes in the bounding box list and the bounding box with the highest confidence score. The IOU represents the ratio of the area of the intersection to the area of the union of the generated rectangular bounding boxes, as shown in Figure 5 the following figure. It usually reflects the degree of overlap between two bounding boxes. It is defined as: IOU = (A ∩ B) / (A ∪ B), where A and B represent different rectangular bounding boxes;

[0072] Step 5: Calculate the linear decay coefficient decay for the bounding boxes with an IOU greater than the threshold (usually this threshold is set to 0.5) in a matrix parallel manner j to change the confidence of this candidate box. Its score reset function is:

[0073]

[0074] where s i is the confidence score; M is the reference box with the highest confidence; b i is the other candidate box for calculating the IOU with the reference box; N t represents the set IOU threshold.

[0075]

[0076] where f(iou i,j ) = 1 - iou(i,j),

[0077] Step 6: Repeat the previous 5 steps until the bounding box list is empty. At this time, the bounding boxes in the output bounding box list are the final predicted boxes.

[0078] In step S3, set the training parameters of the object detection model, and use the training image set to iteratively train the object detection model until the number of training times reaches the preset number of iterations. Adopt the stochastic gradient descent method, and use the training image set to perform frozen iterative training and unfrozen iterative training on the fine-tuned object detection model to obtain the trained object detection model; select the model with the highest detection accuracy from all the obtained trained object detection models as the optimal object detection model, and its data flow diagram is Figure 6 as shown in the following figure.

[0079] Among them, the training parameters are set as follows:

[0080] Since the features of the backbone feature extraction network are general, using frozen iterative training can accelerate the training speed and prevent the weights from being damaged in the initial stage of training. Therefore, it is set to train for 100 epochs (iterations). In the first 50 epochs, the backbone feature extraction network is frozen, Batchsize (batch size) = 16, and the initial learning rate is 1e-3. Considering that when starting training, the weights of the object detection model are randomly initialized. At this time, if a relatively large learning rate is selected, it may cause instability (oscillation) of the object detection model. The Warmup preheated learning rate method is selected, so that the learning rate within the first 10 epochs of starting training is trained at the preheated small learning rate of 1e-4, and the object detection model can gradually become stable. After the object detection model is relatively stable, then select the preset initial learning rate of 1e-3 for training, and the lower limit of the learning rate is 1e-6. After unfreezing, set Batchsize = 8, the initial learning rate is 1e-4, and also select the Warmup preheated learning rate method, so that the learning rate within the first 10 epochs of starting training is trained at the preheated small learning rate of 1e-5, and then select the preset initial learning rate of 1e-4 for training after the object detection model is relatively stable. Select the model with the highest detection accuracy from all the trained object detection models as the optimal object detection model.

[0081] As Figures 7 - 9 shown, in step S4, preprocessing such as image cropping, image flipping, and image scaling is performed on the detection image to make the parameters such as the size of the detection image consistent with those of the training image. The preprocessed detection image is input into the optimal object detection model to obtain the object detection result of the detection image. Figure 8 The detection result diagram of conveyor belt workpiece detection using the method of the present invention is given. It can be seen from the figure that for various types of workpieces, the method of the present invention can accurately detect the workpiece position, achieving the best detection performance, and the detection accuracy reaches 99.5%.

[0082] The present invention also provides a conveyor belt workpiece image object detection system based on lightweight YOLOV4-tiny, which is characterized by including:

[0083] An image acquisition module for constructing a training image set of conveyor belt workpieces;

[0084] A network construction module for constructing a lightweight YOLOV4-tiny conveyor belt workpiece image object detection network, including a backbone feature extraction network CSPDarknet53-tiny, an enhanced feature extraction network FPN, and a prediction feature layer YOLO Head, wherein an efficient channel attention mechanism ECA is incorporated into the enhanced feature extraction network FPN;

[0085] A network training module, which is used to train a lightweight YOLOV4-tiny conveyor belt workpiece image target detection network by using a training image set of conveyor belt workpieces, so as to obtain all the trained lightweight YOLOV4-tiny conveyor belt workpiece image target detection network models;

[0086] A retrieval module, which is used to input the conveyor belt workpiece image to be detected into the lightweight YOLOV4-tiny conveyor belt workpiece image target detection network model to obtain the target detection result of the conveyor belt workpiece image.

Claims

1. A method for object detection of conveyor belt workpiece images based on lightweight YOLOV4-tiny, characterized in that, The specific steps are as follows: S1 Construct a training image set of conveyor belt workpieces; S2 Construct a lightweight YOLOV4-tiny conveyor belt workpiece image target detection network, including a backbone feature extraction network CSPDarknet53-tiny, a feature enhancement network FPN, and a prediction feature layer YOLO Head. An efficient channel attention mechanism ECA is incorporated into the feature enhancement network FPN; S3 Use the training image set of conveyor belt workpieces to train the lightweight YOLOV4-tiny conveyor belt workpiece image target detection network to obtain a lightweight YOLOV4-tiny conveyor belt workpiece image target detection network model; S4 Input the conveyor belt workpiece image to be detected into the lightweight YOLOV4-tiny conveyor belt workpiece image target detection network model to obtain the target detection result of the conveyor belt workpiece image; In step S2, matrix non-maximum suppression (Matrix NMS) is used to perform parallel screening on the prediction boxes of the prediction feature layer YOLOHead in a linearly decaying manner. The specific steps of the screening are as follows: S2.1 According to the class confidence, arrange the detection boxes of the prediction feature layer YOLO Head in descending order and obtain a list of bounding boxes; S2.2 Move the bounding box with the highest confidence score in the list to the output bounding box list; S2.3 Calculate the areas of all the remaining detection boxes in the output bounding box list; S2.4 Calculate the intersection over union (IoU) between all the remaining bounding boxes in the bounding box list and the bounding box with the highest confidence score; S2.5 Calculate the linear attenuation coefficient decay of the bounding boxes with intersection over union greater than the threshold j , change the confidence of the candidate bounding boxes, and the score reset function is as follows: Among them, s i is the confidence score; M is the reference box with the highest confidence; b i is the other candidate box for calculating the intersection over union with the reference box; N t represents the set intersection over union threshold; Linear attenuation coefficient decay j The formula is as follows: Among them: f(iou i,j ) = 1 - iou(i, j), S2.6 Repeat steps S2.1 - S2.5 until the bounding box list is empty. At this time, the bounding boxes in the output bounding box list are the finally screened prediction boxes.

2. The object detection method for conveyor belt workpiece images based on lightweight YOLOV4-tiny according to claim 1, wherein, In step S2, the feature enhancement network FPN includes three efficient channel attention mechanisms ECA. An efficient channel attention mechanism ECA is connected after each of the two effective feature layers of the backbone feature extraction network CSPDarknet53-tiny. An efficient channel attention mechanism ECA is connected after the upsampling of the feature enhancement network FPN.

3. A method for detecting target objects in conveyor belt workpiece images based on lightweight YOLOV4-tiny according to claim 2, characterized in that, In step S2, the backbone feature extraction network CSPDarknet53-tiny extracts features from the conveyor belt workpiece image to be detected and outputs two-scale feature maps feat1 and feat2 of 13×13 and 26×26. The feature maps feat1 and feat2 are respectively input into the efficient channel attention mechanism ECA for channel weighting of the feature maps to obtain feature maps f1 and f2. The feature enhancement network FPN performs one upsampling on the feature map f2 and then passes it through the efficient channel attention mechanism ECA to obtain the feature map f2'. The output feature map f2' is stacked with the result of the feature map f1 and the feature map f2.

4. A conveyor belt workpiece image target detection method based on lightweight YOLOV4-tiny according to claim 3, characterized in that, In step S2, the efficient channel attention mechanism ECA uses a global average pooling layer to reduce the dimension of the original feature map. The dimension-reduced original feature map is subjected to feature extraction through one-dimensional convolution to obtain a feature extraction output map. The sigmoid activation function is used to obtain the corresponding channel probability values of the feature extraction output map, and the corresponding channel probability values are multiplied by the original feature map to obtain the output of the efficient channel attention mechanism ECA.

5. A method for detecting target of conveyor belt workpiece image based on lightweight YOLOV4-tiny according to claim 4, characterized in that, In step S2, the convolution kernel size K of the one-dimensional convolution is: where |*| odd represents the nearest odd number, and C is the channel dimension.

6. A conveyor belt workpiece image target detection method based on lightweight YOLOV4-tiny according to claim 1, characterized in that, The intersection over union IOU represents the area ratio of the intersection and union parts between the generated rectangular borders. Specifically: IOU = (A ∩ B) / (A ∪ B), where A and B represent different rectangular borders.

7. A method for detecting conveyor belt workpiece image targets based on lightweight YOLOV4-tiny according to claim 1, characterized in that, In step S1, the conveyor belt workpiece image marked with the target to be detected is preprocessed to obtain a training image, and the training image is collected in the training image set.

8. A conveyor belt workpiece image target detection system based on lightweight YOLOV4-tiny, characterized in that, Including: An image acquisition module for constructing a training image set of conveyor belt workpieces; A network construction module for constructing a lightweight YOLOV4-tiny conveyor belt workpiece image target detection network, including a backbone feature extraction network CSPDarknet53-tiny, an enhanced feature extraction network FPN, and a prediction feature layer YOLO Head. An efficient channel attention mechanism ECA is incorporated into the enhanced feature extraction network FPN; A network training module for training the lightweight YOLOV4-tiny conveyor belt workpiece image target detection network using the training image set of conveyor belt workpieces to obtain all trained lightweight YOLOV4-tiny conveyor belt workpiece image target detection network models; A retrieval module for inputting the conveyor belt workpiece image to be detected into the lightweight YOLOV4-tiny conveyor belt workpiece image target detection network model to obtain the target detection result of the conveyor belt workpiece image; In the network construction module, matrix non-maximum suppression Matrix NMS is used to perform parallel screening on the prediction boxes of the prediction feature layer YOLO Head in a linearly decaying manner. The specific steps of the screening are as follows: S1 According to the class confidence, the detection boxes of the prediction feature layer YOLO Head are sorted in descending order, and a border list is obtained; S2 Move the border with the highest confidence score in the list to the output border list; S3 Calculate the areas of all the remaining detection boxes in the output border list; S4 Calculate the intersection over union between all the remaining borders in the border list and the border with the highest confidence score; S5 Calculate the linear decay coefficient decay of the bounding boxes with an intersection over union greater than the threshold j , change the confidence of the candidate bounding boxes, and the score reset function is as follows: Among them, s i is the confidence score; M is the reference box with the highest confidence; b i is the other candidate box for calculating the intersection over union with the reference box; N t represents the set intersection over union threshold; Linear attenuation coefficient decay j The formula is as follows: where: f(iou i,j ) = 1 - iou(i, j), S6 Repeat steps S2.1 - S2.5 until the border list is empty. At this time, the borders in the output border list are the finally screened prediction boxes.

Citation Information

Patent Citations

  • Electric power intelligent construction site violation behavior detection method based on YOLOv4 improved algorithm

    CN112149761A

  • Improved YOLO-v3 algorithm and application thereof in pulmonary nodule detection

    CN112184684A