Cigarette appearance defect detection method and system based on improved YOLOv5s architecture
By improving the YOLOv5s architecture, the use of Ghost+ACIN lightweight convolutional network and cross-layer cascade fusion network, combined with the hybrid attention mechanism, the accuracy and speed of smoke branch appearance defect detection are improved, the low accuracy and real-time problems of traditional detection technology are solved, and an efficient detection solution is provided for industrial quality inspection and mobile terminals.
Patent Information
- Application Number
- CN202510442912.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional tobacco appearance detection technology has low detection accuracy and high working intensity, and is difficult to meet the real-time requirements of industrial production. The existing models are complex and difficult to deploy on mobile terminals.
Adopting the improved YOLOv5s architecture, by introducing Ghost+ACIN lightweight convolutional network to replace ordinary convolutional layers, integrating a cross-layer cascaded fusion network and a hybrid attention mechanism, improving feature extraction and recognition accuracy.
It realizes efficient detection of appearance defects of cigarette posts, improves detection accuracy and speed, is suitable for industrial quality inspection equipment and mobile terminal equipment, and provides high-precision and low-latency visual inspection solutions.
Smart Images

Figure CN120339705A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for detecting appearance defects of cigarette sticks based on an improved YOLOv5s architecture, belonging to the field of computer vision technology. Background Art
[0002] In a cigarette automated production system, insufficient periodic maintenance of the cigarette making and tipping equipment is likely to induce multiple appearance defects. Surface defects of cigarette sticks not only directly affect the appearance consistency of products, but also have a negative impact on sensory quality and brand image. However, traditional cigarette stick appearance detection technologies not only have low detection accuracy and high working intensity, but also have hysteresis and contingency in sampling inspection, and the detection speed is difficult to meet the real-time detection requirements of industrial production. The dual defects of detection accuracy and real-time performance severely restrict the stability of cigarette products and brand credibility. To address these problems, based on the machine vision technology framework, the present invention aims to construct a more efficient detection model through deep learning methods to improve the accuracy and speed of cigarette stick appearance defect detection.
[0003] For the requirements of large-sample cigarette stick appearance defect detection and mobile deployment, the present invention constructs a lightweight detection model based on improved YOLOv5s. By introducing the Ghost+ACIN convolution module to replace the ordinary convolution layer, while retaining the multi-scale feature extraction ability, the model parameter quantity and computational complexity are synchronously compressed, significantly improving the detection efficiency; improving the feature pyramid structure in the PAN module of the YOLOv5s neck network, integrating the cross-level cascaded fusion network, fully fusing the deep and shallow layer information of the feature map, and improving the recognition accuracy of defect targets; at the same time, introducing the CBAM hybrid attention mechanism to strengthen the model's attention to the channel and spatial information of the feature map, enhancing the model's semantic feature and detail information extraction ability, and improving the model's recognition performance. Summary of the Invention
[0004] Aiming at the problems existing in traditional cigarette stick appearance detection technologies, the present invention provides a method and system for detecting appearance defects of cigarette sticks based on an improved YOLOv5s architecture, to solve the problems of insufficient detection accuracy, high model complexity, and difficult mobile deployment for surface defects of long and thin cigarette sticks in the prior art. The present invention constructs a more efficient detection model to improve the accuracy and speed of cigarette stick appearance defect detection.
[0005] The technical solution of the present invention is: a method for detecting appearance defects of cigarette sticks based on an improved YOLOv5s architecture, the method comprising:
[0006] Step 1, collecting cigarette stick appearance image data, and performing preprocessing and data augmentation on the cigarette stick appearance images;
[0007] Step 2, defining different types of cigarette stick appearance image defects and annotating the defect positions, and constructing a defect cigarette stick image dataset;
[0008] Step 3. Replace the traditional convolutional layer in the YOLOv5s model with the Ghost+ACIN lightweight convolutional network;
[0009] Step 4. Integrate the cross-layer cascade fusion network in the YOLOv5s model neck network PAN module;
[0010] Step 5: Introduce the hybrid attention mechanism module to strengthen the YOLOv5s model's attention to feature map space and channel information;
[0011] Step 6. Train the improved YOLOv5s model using the training set images;
[0012] Step 7: Classify and detect the test set images using the trained improved YOLOv5s model to complete the cigarette appearance image detection.
[0013] Furthermore, the Step 1 includes:
[0014] The original cigarette appearance image is grayed and image filtered for preprocessing. The image filtering preprocessing includes median filtering, mean filtering, and Gaussian filtering. The cigarette appearance image is randomly flipped and Gaussian noise data is added for enhancement.
[0015] Furthermore, the Step 2 includes:
[0016] Step 2.1. The types of cigarette appearance defects in the cigarette appearance image are uniformly divided into six categories: warping, misaligned teeth, wrinkles, spots, no filter, and loose tips according to the third edition of the national standard "Cigarette" and the guidance of relevant technical personnel from the cigarette factory. The Pycharm built-in toolbox Labelimg is used to mark the location and type of the cigarette image defects;
[0017] Step 2.2, construct a defective cigarette image dataset with the labeled data, and divide it into a training set and a test set in a ratio of 8:2.
[0018] Furthermore, the Step 3 includes:
[0019] Step 3.1, configure four groups of Ghost modules to replace the traditional convolutional layer for high-order feature extraction;
[0020] Step 3.2. Design the ACIN module based on the AC module of the asymmetric convolutional network ACNet, and use two ACIN modules with a convolution kernel size of 3, a step size of 2, and the same number of input and output channels to replace the traditional convolution layer during feature fusion in the neck network.
[0021] Furthermore, the Step 4 includes:
[0022] Introduce a cross-level cascading fusion network into the backbone network of the YOLOv5s model to integrate the low-level position information, contour information, and detail information in the feature map, and combine it with the high-level semantic information at the same time.
[0023] Furthermore, the Step5 includes:
[0024] Embed the Convolutional Block Attention Module (CBAM) of the hybrid attention mechanism into the neck network of the YOLOv5s model. Through the parallel channel attention and spatial attention mechanisms, it enables the model to capture the channel dependence and spatial context information of the feature map simultaneously, improve the recognition accuracy of the semantic features of small targets, and obtain an improved YOLOv5s object detection network model.
[0025] Furthermore, the Step6 includes:
[0026] Step6.1: Prepare the model configuration file and set the training parameters of the improved YOLOv5s model;
[0027] Step6.2: Input the training dataset into the improved YOLOv5s model for iterative training, and end the training after reaching the training iteration times;
[0028] Step6.3: Save the weight file corresponding to the trained improved YOLOv5s model; the weight file includes the knowledge learned by the improved YOLOv5s model during the training process, and retains the connection weights and biases between different layers of the improved YOLOv5s model.
[0029] By loading the weight file, reload the parameter values into the improved YOLOv5s model to obtain the trained improved YOLOv5s model.
[0030] The present invention also provides a cigarette appearance defect detection system based on the improved YOLOv5s architecture, and the system includes: a module for executing the cigarette appearance defect detection method based on the improved YOLOv5s architecture.
[0031] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the cigarette appearance defect detection method based on the improved YOLOv5s architecture.
[0032] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the cigarette appearance defect detection method based on the improved YOLOv5s architecture.
[0033] The beneficial effects of the present invention are as follows:
[0034] 1. The present invention uses the Ghost+ACIN lightweight network to replace ordinary convolution. By introducing linear transformation to generate redundant feature maps, it reduces the computational overhead of traditional convolution operations, providing theoretical support for model lightweighting.
[0035] 2. Through the multi-level feature fusion mechanism, the present invention fully integrates the semantic information of the high layer and the detailed features of the low layer, solving the problem of reduced detection accuracy caused by feature information loss in traditional methods.
[0036] 3. The model of the present invention embeds the CBAM module of the hybrid attention mechanism. Through the parallel channel attention and spatial attention mechanisms, the model can simultaneously capture the channel dependence and spatial context information of the feature map, thereby significantly improving the recognition accuracy of the semantic features of small targets.
[0037] 4. The present invention can provide a high-precision and low-latency visual detection solution for the deployment of industrial quality inspection equipment and mobile terminal equipment; the improved YOLOv5s model proposed by the present invention is superior to other comparison models in terms of detection accuracy and detection speed, providing a new method for the detection of cigarette appearance defects with high requirements for lightweighting and real-time performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a flowchart of the detection of cigarette appearance defects based on the improved YOLOV5s model of the present invention;
[0039] Figure 2 It is a network framework diagram of the improved YOLOv5s model of the present invention;
[0040] Figure 3 It is a comparison diagram of the average detection accuracy of the YOLOv5s model before and after improvement of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] Example 1: As Figures 1-3 shown, the present invention will be described in detail below with reference to the drawings and specific embodiments. The present invention provides a method for detecting cigarette appearance defects based on an improved YOLOv5s architecture. After collecting cigarette appearance images and performing preprocessing and image enhancement and other processing operations, unified classification of defect types and annotation and division of the data set are carried out. Then, the Ghost+ACIN lightweight network, multi-level feature fusion mechanism, and hybrid attention mechanism CBAM improvement strategy are used to improve the YOLOv5s model. Finally, the model is trained and verified. The method specifically includes:
[0042] Step 1. Collect the appearance image data of cigarette sticks, and perform preprocessing and data augmentation on the appearance images of cigarette sticks; perform grayscale conversion and image filtering preprocessing on the original appearance images of cigarette sticks. The image filtering preprocessing includes median filtering, mean filtering, and Gaussian filtering, and perform random flipping on the appearance images of cigarette sticks and add Gaussian noise data for enhancement.
[0043] Step 2. Define different defect types of cigarette stick appearance images and label the defect positions to construct a defective cigarette stick image dataset;
[0044] Step 2.1. Uniformly classify the defect types of cigarette stick appearances in the cigarette stick appearance images into six categories: warping edge, misaligned teeth, wrinkle, spot, no filter tip, and loose head according to the provisions of the third edition of the "Cigarettes" series of national standards and the guidance of relevant technical personnel in cigarette factories, and use the built-in toolbox Labelimg in Pycharm to label the defect positions and types of cigarette stick images;
[0045] Step 2.2. Construct a defective cigarette stick image dataset from the labeled data, and divide it into a training set and a test set in a ratio of 8:2.
[0046] Step 3. Replace the traditional convolutional layer in the YOLOv5s model with a Ghost+ACIN lightweight convolutional network;
[0047] Step 3.1. Replace the traditional convolutional layer with four groups of Ghost modules (channel mapping 64→128→256→512→1024, kernel size 3×3, stride 2) for high-order feature extraction, and realize the synchronous optimization of the network parameter quantity and computational complexity (GFLOPS);
[0048] Step 3.2. Design an ACIN module based on the AC module of the asymmetric convolutional network ACNet (Asymmetric Convolutional Network), and replace the traditional convolutional layer during feature fusion in the neck network with two ACIN modules with a convolutional kernel size of 3, a stride of 2, and the same number of input and output channels to strengthen the extraction and fusion of local feature information.
[0049] Step 4. Integrate a cross-level cascading fusion network into the PAN module of the neck network of the YOLOv5s model;
[0050] Introduce a cross-level cascading fusion network into the backbone network of the YOLOv5s model to integrate the low-level position information, contour information, and detail information in the feature map, and combine it with the high-level semantic information at the same time.
[0051] Step 5. Introduce a hybrid attention mechanism module to strengthen the attention of the YOLOv5s model to the spatial and channel information of the feature map; The Step 5 includes:
[0052] Embed the CBAM (Convolutional Block Attention Module) with a hybrid attention mechanism into the neck network of the YOLOv5s model. Through the parallel channel attention and spatial attention mechanisms, it enables the model to capture both the channel dependencies and spatial context information of the feature map simultaneously, significantly improving the recognition accuracy of the semantic features of small targets, and obtaining an improved YOLOv5s object detection network model.
[0053] Step 6. Train the improved YOLOv5s model using the training set images.
[0054] Step 6.1. Prepare the model configuration file and set the training parameters of the improved YOLOv5s model.
[0055] Step 6.2. Input the training data set into the improved YOLOv5s model for iterative training, and end the training after reaching the training iteration times.
[0056] Step 6.3. Save the weight file corresponding to the trained improved YOLOv5s model; the weight file includes the knowledge learned by the improved YOLOv5s model during training and retains the connection weights and biases between different layers of the improved YOLOv5s model.
[0057] By loading the weight file, reload the parameter values into the improved YOLOv5s model to obtain the trained improved YOLOv5s model.
[0058] The present invention also provides a cigarette appearance defect detection system based on the improved YOLOv5s architecture, and the system includes:
[0059] An acquisition and preprocessing module, which is used to acquire cigarette appearance image data, perform preprocessing and data augmentation on the cigarette appearance images.
[0060] A defective cigarette image data set construction module, which is used to define different types of cigarette appearance image defects and label the defect positions to construct a defective cigarette image data set.
[0061] A Ghost+ACIN lightweight convolutional network replacement module, which is used to replace the traditional convolutional layers in the YOLOv5s model with a Ghost+ACIN lightweight convolutional network.
[0062] A cross-level cascading fusion network integration module, which is used to integrate a cross-level cascading fusion network into the PAN module of the neck network of the YOLOv5s model.
[0063] A hybrid attention mechanism module, which is used to enhance the attention of the YOLOv5s model to the spatial and channel information of the feature map.
[0064] A training module for training the improved YOLOv5s model using training set images;
[0065] A detection module for classifying and detecting test set images using the trained improved YOLOv5s model to complete the detection of cigarette appearance images.
[0066] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method for detecting cigarette appearance defects based on the improved YOLOv5s architecture is implemented.
[0067] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for detecting cigarette appearance defects based on the improved YOLOv5s architecture is implemented.
[0068] The present invention constructs a basic feature extraction module using the Ghost network and the ACIN activation function, generates redundant feature maps through linear transformation, and compresses the model parameter quantity under the condition of maintaining the detection accuracy; designs a cross-level cascaded feature fusion network structure, and realizes the collaborative enhancement of shallow and deep semantic information through a bidirectional feature pyramid, effectively improving the positioning accuracy of small-scale defect targets; integrates the channel-space dual-domain hybrid attention mechanism (CBAM), establishes a multi-dimensional feature dynamic weighting mechanism, and strengthens the fine-grained characterization ability of defects with low pixel distribution.
[0069] The present invention classifies and detects test set images using the trained improved YOLOv5s model to complete the detection of cigarette appearance images. The present invention uses five indicators, namely the mean average precision (mAP) value, precision (P), recall (R), frames per second (FPS), and model size (Size), to evaluate the model performance. Among them, the calculation methods of the recall rate R and the precision rate P are as follows:
[0070]
[0071] In the formula, P is the precision rate, also known as the precision, indicating the proportion of correctly predicted positive samples in all predicted positive samples; R is the recall rate, also known as the recall, TP is the number of correctly predicted positive samples; FP is the number of incorrectly predicted positive samples; FN is the number of true positive samples that fail to be predicted as positive samples.
[0072] The calculation method of the mAP value: First, the result of the R value can be integrated within the range of [0, 1] according to the calculation results of the P and R values. That is:
[0073]
[0074] Subsequently, the mAP value is the average precision of all defect types, i.e.:
[0075]
[0076] In the formula, n is the number of detected target categories, corresponding to 6 types of appearance defects of defective cigarettes.
[0077] The experiment of the present invention was carried out in an environment of 11th Gen Intel(R) Core(,M) i5-1155G7, 2.50GHz, and 16GB of memory. Based on the cuda11.3 and pytorch1.10 frameworks, the hardware configuration was GPU: RTX3060 with 24G of video memory. It was carried out based on Pycharm as the development and training platform and Python3.8 as the development language. The initial learning rate of the YOLOv5 model was learning_rate of 0.01, the batch size was BatchSize of 8, and the training cycle epoch was 250.
[0078] After the data augmentation strategy, the number of dataset samples was significantly expanded from the original 1200 to 6000, fully meeting the requirements of the model training for the data scale. Using the Pycharm integrated toolbox LabelImg, the defective areas of the images in the dataset were accurately labeled and bounding box labeled, and the labeled files were saved in the TXT format compatible with YOLOv5. Finally, the dataset was divided into a training set and a test set in a ratio of 8:2. The experimental evaluation metrics were divided into two aspects: performance evaluation and complexity evaluation. The core metrics for model performance evaluation included recall rate (Recall) and mAP@0.5. The model complexity evaluation metrics included frames per second (FPS) and model size (Size), evaluating the computational efficiency and processing speed of the model.
[0079] To verify the effectiveness of the method proposed in the present invention, the present invention selected Faster R-CNN, SSD, YOLOv3, YOLOv4, CenterNet and the baseline model YOLOv5s as comparison algorithms to carry out comparative experiments. The comparative experiments were carried out on the self-built dataset of defective cigarettes, and the experimental environment and parameter configurations were the same.
[0080] The final experimental results are shown in Table 1 below.
[0081] Table 1 shows the comparison of classification results of different models
[0082]
[0083] As can be observed from Table 1, the Faster R-CNN, SSD, and YOLOv4 algorithms are inferior to the YOLOv5s algorithm in terms of metrics such as accuracy and recall. The reason for this phenomenon is that the models of the aforementioned algorithms are relatively large, with a large number of parameters and a deep network depth, which affects their processing speed and efficiency. In contrast, the accuracy and detection time consumption of YOLOv5 have been greatly improved. After the model is reduced, the accuracy is not affected because the model size becomes smaller and the corresponding parameters become fewer, resulting in a further increase in the inference speed. The algorithm of the present invention performs better than the listed algorithms in all metrics except the detection time consumption. The size of the parameter model is 15.9MB, mAp@0.5 is 95.6%, the recall rate R is 89.5%, and the detection time consumption meets the detection requirements of the dual-channel cigarette-making unit. This is because after adding other modules to the model in this article, the number of parameters and the amount of calculation of the model increase. Under the condition that the detection accuracy is comparable to that of YOLOv4 and CenterNet, the model volume of the algorithm proposed in the present invention is reduced to 6.3% of the original YOLOv4 and 12.8% of CenterNet respectively, and the advantage of model lightweight is significant. To sum up, it can be shown that the algorithm in this article is a lightweight cigarette appearance detection method with excellent performance.
[0084] Verified by experiments, while maintaining the average detection accuracy (mAP) at 95.6%, the model size of this solution is reduced by 0.9MB compared with the benchmark model YOLOv5s, and the inference speed is increased to 54.2FPS, which can provide a high-precision and low-latency visual detection solution for the deployment of industrial quality inspection equipment and mobile terminal equipment.
[0085] The comparison chart of the average detection accuracy of the YOLOv5s model before and after improvement is shown in Figure 3 . From Figure 3 It can be seen that after improving YOLOV5s, the original model can have better performance. Specifically, the model converges rapidly after 50 epochs and finally stabilizes after 150 epochs, and the average precision rate is much better than the YOLOv5 model before improvement.
[0086] The specific implementation manners of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above implementation manners, and various changes can be made without departing from the spirit of the present invention within the knowledge scope of those of ordinary skill in the art.
Claims
1. A method for detecting appearance defects of cigarette sticks based on an improved YOLOv5s architecture, characterized in that: The method includes: Step 1: Collect the appearance image data of cigarette sticks, and perform preprocessing and data enhancement on the appearance images of cigarette sticks. Step 2: Define different types of defects in the appearance images of cigarette sticks and label the defect positions to construct a dataset of defective cigarette stick images. Step 3: Replace the traditional convolutional layer in the YOLOv5s model with the Ghost+ACIN lightweight convolutional network. Step 4: Integrate the cross-level cascade fusion network in the PAN module of the neck network of the YOLOv5s model. Step 5: Introduce a hybrid attention mechanism module to strengthen the attention of the YOLOv5s model to the spatial and channel information of the feature map. Step 6: Train the improved YOLOv5s model using the training set images. Step 7: Perform classification detection on the test set images using the trained improved YOLOv5s model to complete the detection of the appearance images of cigarette sticks.
2. The method for detecting appearance defects of cigarette sticks based on an improved YOLOv5s architecture according to claim 1, wherein: The above Step 1 includes: Perform grayscale conversion and image filtering preprocessing on the original appearance images of cigarette sticks. The image filtering preprocessing includes median filtering, mean filtering, and Gaussian filtering, and perform random flipping on the appearance images of cigarette sticks and add Gaussian noise data for enhancement.
3. A method for detecting appearance defects of cigarette sticks based on an improved YOLOv5s architecture according to claim 1, characterized in that: The above Step 2 includes: Step 2.1: Uniformly classify the types of defects in the appearance of cigarette sticks in the appearance images of cigarette sticks into six categories: warping, misaligned teeth, wrinkles, spots, no filter tip, and loose head according to the provisions of the third edition of the "Cigarettes" series of national standards and the guidance of relevant technical personnel in the cigarette factory, and use the built-in toolbox Labelimg in Pycharm to label the defect positions and types of cigarette stick images. Step 2.2: Construct a dataset of defective cigarette stick images from the labeled data, and divide it into a training set and a test set in a ratio of 8:
2.
4. A method for detecting appearance defects of cigarette sticks based on an improved YOLOv5s architecture according to claim 1, characterized in that: The above Step 3 includes: Step 3.1: Replace the traditional convolutional layer with four groups of Ghost modules for high-order feature extraction. Step 3.2: Design the ACIN module based on the AC module of the asymmetric convolutional network ACNet, and use two ACIN modules with a convolutional kernel size of 3, a stride of 2, and the same number of input and output channels to replace the traditional convolutional layer during feature fusion in the neck network.
5. A method for detecting appearance defects of cigarette sticks based on an improved YOLOv5s architecture according to claim 1, characterized in that: The above Step 4 includes: Introduce a cross-level cascade fusion network in the backbone network of the YOLOv5s model to integrate the low-level position information, contour information, and detail information in the feature map, and combine it with the high-level semantic information at the same time.
6. A method for detecting appearance defects of cigarette sticks based on an improved YOLOv5s architecture according to claim 1, characterized in that: The above Step 5 includes: Embed the hybrid attention mechanism module CBAM in the neck network of the YOLOv5s model. Through the parallel channel attention and spatial attention mechanisms, it is used to enable the model to simultaneously capture the channel dependence and spatial context information of the feature map, improve the recognition accuracy of the semantic features of small targets, and obtain an improved YOLOv5s object detection network model.
7. A method for detecting appearance defects of cigarette sticks based on an improved YOLOv5s architecture according to claim 1, characterized in that: The above Step 6 includes: Step 6.1: Prepare the model configuration file and set the training parameters of the improved YOLOv5s model. Step 6.2: Input the training dataset into the improved YOLOv5s model for iterative training, and end the training after reaching the training iteration times. Step 6.3: Save the weight file corresponding to the improved YOLOv5s model after training; the weight file includes the knowledge learned by the improved YOLOv5s model during training and retains the connection weights and biases between different layers of the improved YOLOv5s model. By loading the weight file, reload the parameter values into the improved YOLOv5s model to obtain the trained improved YOLOv5s model.
8. A cigarette appearance defect detection system based on an improved YOLOv5s architecture, characterized in that, The system includes: a module for executing a method for detecting appearance defects of cigarette sticks based on an improved YOLOv5s architecture according to any one of claims 1 to 7.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements a method for detecting appearance defects of cigarette sticks based on an improved YOLOv5s architecture according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements a method for detecting appearance defects of cigarette sticks based on an improved YOLOv5s architecture according to any one of claims 1 to 7.
Citation Information
Cited By
Defect detection method and device based on improved YOLO model
CN121169795A