Cigarette piece cigarette box carton lacking detection method

The YOLOv8 target detection model, which utilizes multi-camera acquisition and image enhancement, solves the problems of false detection and missed detection in cigarette box defect detection, achieving efficient and accurate defect detection and real-time traceability, thus improving detection efficiency and reliability.

CN121962054APending Publication Date: 2026-05-01CHANGDE COMPANY OF CHINA TOBACCO HUNAN
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHANGDE COMPANY OF CHINA TOBACCO HUNAN
Filing Date
2026-01-13
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing cigarette box defect detection technologies suffer from high false detection and false negative rates under complex lighting and multi-angle views, and lack real-time recording and traceability mechanisms, resulting in low detection efficiency and commercial losses.

Method used

Image data is acquired using multiple cameras, and the YOLOv8 target detection model, which combines image enhancement and attention mechanisms, is used to adjust the lighting and extract features at multiple scales. The relevant data of missing bar events are recorded and transmitted to the data storage system.

Benefits of technology

It significantly improves the accuracy and reliability of detection, reduces false detection and false negative rates, enhances the ability to identify cigarette stick targets, and enables real-time data traceability and auditability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962054A_ABST
    Figure CN121962054A_ABST
Patent Text Reader

Abstract

The invention discloses a cigarette carton missing detection method for a cigarette carton box. The method comprises the following steps: S1, collecting image data of the cigarette carton through a plurality of cameras when the carton is opened; s2, performing illumination adjustment on the acquired image data, and applying image enhancement processing to generate an enhanced image; s3, inputting the enhanced image into a target detection model, detecting whether the cigarette bar is missing, outputting a cigarette bar bounding box and confidence coefficient by the target detection model through extracting multi-scale features of the image and applying attention weight calculation, and judging a bar missing state according to a confidence coefficient threshold value; and S4, when it is detected that the cigarette bar is missing, triggering a data recording operation, recording related data of a missing event, including a timestamp, a box body number and an image fragment, and packaging, transmitting and storing the data. Through the steps of collecting images by multiple cameras, adjusting illumination, integrating a target detection model of an attention mechanism and automatically recording and transmitting data, the accuracy of detecting the carton lacking of the cigarette box of the cigarette piece is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

A method for detecting missing strips in cigarette packs Technical Field

[0001] This invention relates to a method for detecting missing strips in cigarette packs, belonging to the field of cigarette pack inspection technology. Background Technology

[0002] With the increasing volume of cigarette logistics and distribution and the accelerated pace of sorting operations, the integrity verification process for cigarette cartons before they leave the warehouse faces greater pressure. For example, the Changde Tobacco Logistics and Distribution Center processes thousands of cigarette cartons a day, and manually checking for missing cigarettes requires a significant investment of manpower. If the problem of missing cigarette cartons is not detected in time, it will directly lead to economic losses for commercial enterprises and seriously affect retail customer satisfaction and the company's service image.

[0003] Traditional methods for detecting missing cigarette packs primarily include manual visual inspection and electronic weighing. Manual inspection relies on operators visually examining the contents of the packs. While simple, this method requires significant manpower, is inefficient, and prone to errors due to fatigue or subjective bias. Electronic weighing technology detects missing cigarette packs based on their weight differences, using a weighing sensor to measure the total weight of the pack. However, because cigarette pack weights can fluctuate naturally or vary with packaging, this method is prone to missed or false detections, especially in batch processing where accumulated errors further reduce reliability.

[0004] With the development of automation technology, machine vision and deep learning have been introduced to replace manual labor and improve the accuracy and efficiency of detection. Existing literature (CN118196050A) discloses a method for detecting missing cigarette packs based on an improved YOLOv5s. This method captures images of the cigarette pack using a camera and trains a model based on the improved YOLOv5s to identify the cigarette packs. However, the following problems still exist in actual detection: 1. Image quality is affected by lighting conditions; low light or overexposure can lead to inaccurate feature extraction. 2. Under multi-angle views, the cigarette packs are tightly packed, making features indistinct and easily causing false positives or false negatives. Furthermore, existing methods lack effective real-time recording and traceability mechanisms, making it difficult to capture detection process data in real time, thus hindering the tracing of missing cigarette pack incidents and making it difficult to determine the time, location, and cause of the incident. Summary of the Invention

[0005] Based on the above, the present invention provides a method for detecting missing strips in cigarette packs, in order to solve the problems of high false detection and false negative rates in the prior art under complex lighting and multi-angle views, and the fixed algorithm that cannot adapt to environmental changes.

[0006] The technical solution of the present invention is: a method for detecting missing strips in cigarette boxes, comprising the following steps:

[0007] S1. Collect image data of the cigarette box when it is opened using multiple cameras. The image data includes the cigarette box barcode image, the front image of the cigarette stack, and the back image of the cigarette stack.

[0008] S2. Adjust the illumination of the acquired image data, calculate parameters based on the illumination conditions of the image, and apply image enhancement processing to generate an enhanced image;

[0009] S3. Input the enhanced image into the target detection model to detect whether the cigarette bar is missing. The target detection model extracts multi-scale features of the image and applies attention weights to calculate and output the cigarette bar bounding box and confidence score. The missing bar status is determined based on the confidence score threshold.

[0010] S4. When a missing cigarette stick is detected, a data recording operation is triggered to record relevant data of the missing event, including timestamp, box number and image fragment, and the data is packaged and transmitted to the data storage system for evidence preservation and traceability.

[0011] Preferably, in step S1, the plurality of cameras include a first camera, a second camera, and a third camera, wherein the first camera captures the flow image of the tobacco box barcode, the second camera captures the front image of the tobacco stack, and the third camera captures the back image of the tobacco stack.

[0012] Preferably, in step S2, the illumination adjustment includes:

[0013] S21. Convert the image data into a floating-point tensor and normalize it;

[0014] S22. Estimate the illumination map using a pre-trained image enhancement network, wherein the image enhancement network includes convolutional layers and activation functions;

[0015] S23. Apply learnable curve functions to adjust the illumination map;

[0016] S24. Generate an enhanced image by pixel-by-pixel fusion and scale it to a preset value range.

[0017] Preferably, in step S22, the image enhancement network uses a lightweight convolutional neural network structure, takes an image as input, and outputs a lighting map, wherein the convolutional layer has a 3×3 convolutional kernel.

[0018] Preferably, in step S3, the object detection model is based on the YOLOv8 architecture and integrates an attention mechanism, which includes:

[0019] S31. Extract multi-scale feature maps of images using the CSPDarkNet backbone network;

[0020] S32. Insert an attention module at the end of the backbone network, calculate the channel attention weights and spatial attention weights, and fuse them to enhance features;

[0021] S33. Generate bounding boxes, confidence scores, and class probabilities through a prediction network, and filter the results using nonmaximum suppression.

[0022] Preferably, in step S32, the attention module uses global average pooling to calculate channel attention weights, uses 1×1 convolution to calculate spatial attention weights, and fuses them through weighted summation.

[0023] Preferably, the training method of the target detection model is as follows: the target detection model is trained using a labeled tobacco stack image dataset, the dataset including the tobacco stack positions labeled by rectangular boxes, and expanded by data augmentation techniques, including rotation, flipping, cutMix and random occlusion.

[0024] Preferably, in step S4, the relevant data for recording missing events includes: extracting video segments of preset duration before and after detection from the original video stream, and packaging the detection results, enhanced images, and video segments into JSON and binary file formats, and sending them to the data storage system via HTTP API.

[0025] The beneficial effects of this invention are as follows: By acquiring images through multiple cameras, adjusting illumination, integrating an attention-based target detection model, and automating data recording and transmission steps, this invention significantly improves the accuracy of detecting missing cigarette packs. Compared with existing technologies, this invention can dynamically adapt to complex lighting and multi-angle environmental changes, reducing false positives and false negatives through image enhancement and processing, increasing recall while decreasing false positives; the target detection model, combining attention mechanisms and multi-scale feature extraction, enhances the ability to identify cigarette pack targets, reduces reliance on a large number of labeled samples, and makes the training process more efficient; simultaneously, data storage and traceability functions ensure the auditability and reliability of the detection process. Attached Figure Description

[0026] Figure 1 is a schematic diagram of the method for detecting missing strips in cigarette packs;

[0027] Figure 2 shows the data annotations for the tobacco strips on the front of the tobacco stack;

[0028] Figure 3 shows the data annotations for the tobacco strips on the back of the tobacco stack. Detailed Implementation

[0029] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0030] Referring to Figure 1, an embodiment of the present invention provides a method for detecting missing strips in cigarette packs, comprising the following steps:

[0031] Step S1: Collect image data of the tobacco box when it is opened using multiple cameras. The image data includes the barcode image of the tobacco box, the front image of the tobacco stack, and the back image of the tobacco stack.

[0032] Specifically, three industrial-grade cameras are deployed at key locations along the cigarette box transport path: the first camera is installed above the conveyor belt of the unpacking machine to capture images of the flow of the cigarette box barcodes, with the capture range covering the entire front barcode area of ​​the box; the second camera is installed on the side of the pusher device to capture images of the front of the cigarette stack; and the third camera is set in the cigarette replenishment waiting area to capture images of the back of the cigarette stack.

[0033] In one example, the camera uses a global shutter CMOS sensor with a resolution of 1920×1080 pixels, acquiring image data at a fixed frame rate of 30fps. The three cameras are synchronously triggered via the GenICam protocol to ensure timestamp consistency. The acquired image data is transmitted in real-time to the central processing platform via the RTSP protocol, with transmission latency controlled within 100ms to ensure data real-time performance and synchronization. The transmitted data format can be an H.264 encoded video stream or an uncompressed RGB image array with dimensions of height (H) × width (W) × 3 channels.

[0034] Barcode images are used to identify the box number of the tobacco cartons. Front and back images are used for subsequent tobacco stick loss detection, thus establishing data association in the detection events and ensuring that each missing tobacco stick event accurately corresponds to a specific tobacco carton and process. Taking images of the front, back, and barcode flow direction is to comprehensively cover the visible surface of the tobacco stack, ensuring the accuracy of tobacco stick loss detection. The front image is used to detect missing tobacco sticks on the front of the stack, and the back image is used to detect missing tobacco sticks on the back, avoiding missed detections due to viewing angle limitations.

[0035] Step S2: Adjust the illumination of the acquired image data, calculate parameters based on the illumination conditions of the image, and apply image enhancement processing to generate an enhanced image.

[0036] Specifically, lighting adjustment includes the following steps:

[0037] S21, Data Preprocessing

[0038] The raw image data (e.g., uint8 format, value range [0, 255]) obtained from the RTSP stream is converted into a floating-point tensor and normalized to the value range [0, 1].

[0039] in The original image data has a value range of [0, 255]. is the normalized floating-point tensor with a range of [0,1].

[0040] S22, Illumination Map Estimation

[0041] Illumination maps are estimated using a pre-trained lightweight image augmentation network (IAT). in For illumination estimation network, These are the pre-trained parameters.

[0042] The image enhancement network uses a lightweight convolutional neural network architecture, specifically configured as follows:

[0043] Input layer: Receives 3-channel RGB images, with dimensions of 256×256×3.

[0044] Convolutional layer 1: 3×3 convolutional kernel, 32 filters, stride 1, ReLU activation.

[0045] Convolutional layer 2: 3×3 convolutional kernel, 64 filters, stride 2, ReLU activation.

[0046] Convolutional layer 3: 3×3 convolutional kernel, 32 filters, stride 1, ReLU activation.

[0047] Output layer: 3×3 convolutional kernel, 3 filters, stride 1, sigmoid activation.

[0048] Output lighting map Same size as the input image, with a value range of [0,1].

[0049] S23, Lighting Adjustment

[0050] Adjusting the illumination map using learnable curve functions:

[0051] in and These are learnable parameters, and their optimal values ​​are obtained through training.

[0052] S24, Image Reconstruction

[0053] Enhanced images are generated by pixel-by-pixel fusion:

[0054] The enhanced image value range is then scaled to [0, 255] and converted to uint8 format.

[0055] By adjusting the lighting, problems such as uneven lighting, overexposure, or underexposure in actual industrial environments are overcome, providing a stable input for subsequent target inspection.

[0056] Step S3: Input the enhanced image into the target detection model to detect whether the cigarette bar is missing. The target detection model extracts multi-scale features of the image and applies attention weights to calculate and output the cigarette bar bounding box and confidence score. The missing bar status is determined based on the confidence score threshold.

[0057] Specifically, the object detection model is based on the YOLOv8 architecture and integrates an attention mechanism, which includes:

[0058] S31: Feature Extraction

[0059] Multi-scale feature maps of images are extracted using the CSPDarkNet backbone network to detect targets of different sizes. The enhanced image will then be used. The input model uses the backbone network CSPDarkNet. This network efficiently extracts feature maps of images at different scales through a series of convolutions, CSP modules, and downsampling operations. .

[0060] For example, suppose the input image size is After five downsampling stages, feature maps at four scales are output:

[0061] P3 / 8: Size

[0062] P4 / 16: Size

[0063] P5 / 32: Size

[0064] P6 / 64: Size

[0065] S32: Attention Mechanism

[0066] To improve the false positives and false negatives caused by the tight adhesion of cigarette sticks and the lack of significant features, an attention module is inserted at the end of the backbone network to calculate channel attention weights and spatial attention weights and fuse them to enhance features.

[0067] Channel attention enables the model to focus on feature channels that contain more information. The attention module first calculates the channel attention weights. Specifically, for feature maps Each channel undergoes global average pooling (GAP) to obtain a vector representing the global information of each channel. Then, a learnable weight matrix is ​​used... The attention weights for each channel are calculated using the sigmoid function:

[0068] GAP stands for Global Average Pooling. This is a learnable weight matrix.

[0069] Spatial attention enables the model to focus on key regions in the image where smoke trails might appear. The attention module calculates the spatial attention weights. Use 1×1 convolutions on the feature maps The transformation is performed, and then the importance weight of each spatial location (pixel) is calculated using the Sigmoid function:

[0070]

[0071] in For 1×1 convolution, This is a learnable weight matrix.

[0072] Finally, the two attention weights are fused by weighted summation and applied to the original feature map. The enhanced feature map is obtained above. :

[0073]

[0074] This operation effectively integrates local and global, channel and spatial information, enhancing the model's ability to represent smoke bar targets.

[0075] S33: Predicted Output

[0076] Enhanced feature map After multi-scale feature fusion via PANet, the data is fed into the detection head. The detection head predicts the bounding box coordinates (center point, width, and height), a confidence score (representing the probability of an object being present within the box), and a class probability (divided into "normal cigarette sticks" and "missing cigarette sticks") for each preset anchor box. Subsequently, the non-maximum suppression (NMS) algorithm is applied to filter out redundant bounding boxes with high overlap and low confidence. Finally, the state of each detection box is determined based on a set confidence threshold (e.g., 0.5): if a bounding box is classified as "missing cigarette sticks" and its confidence score is greater than the threshold, it is determined that a cigarette stick is missing at that location; the number of detected "normal cigarette sticks" in the entire image is counted and compared with the standard number to determine whether a cigarette stick is missing.

[0077] The object detection model is constructed as follows:

[0078] 1. Model structural parameters

[0079] Input dimensions: 640×640×3

[0080] Backbone network: CSPDarkNet, depth coefficient 1.0, width coefficient 1.0

[0081] Attention module: 256 channels, 8 attention heads

[0082] Detection head: 3 anchor points per scale, 255 convolutional kernels

[0083] 2. Training Data Preparation

[0084] Data source: 2,000 images of the front and back of the tobacco stacks collected from the unpacking machine.

[0085] Data labeling: Use rectangles to label the position of each cigarette stick, and use YOLO format (normalized coordinates) for the labeling.

[0086] Dataset split: 70% training set (1400 images), 15% validation set (300 images), and 15% test set (300 images).

[0087] 3. Data Augmentation Technology

[0088] 3.1 Explicit Enhancement:

[0089] Rotation: Random rotation ±15 degrees

[0090] Flip: 50% probability of horizontal flip.

[0091] CutMix: Blends partial regions of two images at a blending ratio of 0.2-0.4.

[0092] Color jitter: Brightness, contrast, and saturation are randomly adjusted by ±20%.

[0093] 3.2 Implicit Enhancement:

[0094] Random obscuring: Randomly obscuring 20%-40% of the cigarette pack area.

[0095] Gaussian blur: kernel size 3×3, σ value 0.5-1.5

[0096] Noise addition: Gaussian noise, mean 0, variance 0.01

[0097] 4. Training parameter settings

[0098] Optimizer: AdamW, learning rate 0.001, weight decay 0.05

[0099] Batch size: 16, number of training rounds: 300

[0100] Learning rate scheduling: CosineAnnealingLR, T_max=300, η_min=0.0001

[0101] Loss function: CIoU loss + focus loss, weighted at a ratio of 1.0:0.8.

[0102] 5. Training process monitoring

[0103] Performance is evaluated on the validation set every 10 epochs.

[0104] Early stopping mechanism: Training is stopped if the loss on the validation set does not improve after 20 consecutive epochs.

[0105] Model selection: Select the model with the highest mAP on the validation set as the final model.

[0106] Step S4: When a missing cigarette stick is detected, a data recording operation is triggered to record relevant data of the missing event, including timestamp, box number and image fragment, and the data is packaged and transmitted to the data storage system for evidence preservation and traceability.

[0107] Specifically, the missing event data record includes:

[0108] Timestamp: UTC time accurate to milliseconds.

[0109] Box number: A unique identifier identified from the barcode image.

[0110] Detection results: bounding box coordinates, confidence level, number of missing bars.

[0111] Image data: Enhanced front and back images.

[0112] Video clips: 5 seconds of raw video stream before and after detection.

[0113] When the target detection model determines that the cigarette bar is missing in step S3, the system immediately triggers the data recording pipeline to extract a 10-second H.264 format video clip from the original video stream (5 seconds before detection + 5 seconds after detection); serialize the detection result into JSON format; package the JSON file, enhanced image, and video clip into a ZIP archive; transmit the data via the RESTful API interface of the data platform using the HTTPS protocol; and store the data using the MinIO object storage system.

[0114] In the process of detecting missing cigarettes after unpacking, the close proximity of the cigarette packs makes their features indistinct, leading to false positives and false negatives when using the improved YOLOv8 algorithm. This invention proposes an improved algorithm, YOLOv8-Attention. A Hybrid Enhanced Attention (MIX) module is inserted into the backbone network, simultaneously extracting global and local features of the image to generate spatial attention and channel attention respectively. This results in a hybrid enhanced attention module containing local, global, spatial, and channel information, thereby enhancing the network's ability to capture key target features. It can retain scale-related details that are easily lost during sampling, achieving efficient fusion of semantic and detailed information and improving the consistency of information across feature maps at different scales. This method achieves a recall rate of over 95% and a false positive rate of less than 3% on the test set. Compared to the baseline model, this method improves the average detection accuracy by 4.7%, and the precision and recall by 6.9%, respectively, achieving more reliable and accurate automated detection.

[0115] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. A method for detecting missing strips in cigarette packs, characterized in that, Includes the following steps: S1. Collect image data of the cigarette box when it is opened using multiple cameras. The image data includes the cigarette box barcode image, the front image of the cigarette stack, and the back image of the cigarette stack. S2. Adjust the illumination of the acquired image data, calculate parameters based on the illumination conditions of the image, and apply image enhancement processing to generate an enhanced image; S3. Input the enhanced image into the target detection model to detect whether the cigarette bar is missing. The target detection model extracts multi-scale features of the image and applies attention weights to calculate and output the cigarette bar bounding box and confidence score. The missing bar status is determined based on the confidence score threshold. S4. When a missing cigarette stick is detected, a data recording operation is triggered to record relevant data of the missing event, including timestamp, box number and image fragment, and the data is packaged and transmitted to the data storage system for evidence preservation and traceability.

2. The method for detecting missing strips in cigarette packs according to claim 1, characterized in that, In step S1, the multiple cameras include a first camera, a second camera, and a third camera, wherein the first camera captures the flow image of the tobacco box barcode, the second camera captures the front image of the tobacco stack, and the third camera captures the back image of the tobacco stack.

3. The method for detecting missing strips in cigarette packs according to claim 1, characterized in that, In step S2, the illumination adjustment includes: S21, converting the image data into floating-point tensors and normalizing them; S22, estimating the illumination map through a pre-trained image enhancement network, the image enhancement network including convolutional layers and activation functions; S23, applying a learnable curve function to adjust the illumination map; S24, generating an enhanced image by pixel-by-pixel fusion and scaling it to a preset value range.

4. The method for detecting missing strips in cigarette packs according to claim 3, characterized in that, In step S22, the image enhancement network uses a lightweight convolutional neural network structure, takes an image as input and outputs a lighting map, wherein the convolutional layer has a 3×3 convolutional kernel.

5. The method for detecting missing strips in cigarette packs according to claim 1, characterized in that, In step S3, the object detection model is based on the YOLOv8 architecture and integrates an attention mechanism, which includes: S31, extracting multi-scale feature maps of the image through the CSPDarkNet backbone network; S32, inserting an attention module at the end of the backbone network, calculating channel attention weights and spatial attention weights, and fusing them to enhance features; S33, generating bounding boxes, confidence scores, and class probabilities through a prediction network, and filtering the results using nonmaximum suppression.

6. The method for detecting missing strips in cigarette packs according to claim 5, characterized in that, In step S32, the attention module uses global average pooling to calculate channel attention weights, uses 1×1 convolution to calculate spatial attention weights, and fuses them through weighted summation.

7. The method for detecting missing strips in cigarette packs according to claim 1, characterized in that, The training method for the target detection model is as follows: the target detection model is trained using a labeled tobacco stack image dataset, which includes the tobacco stack positions labeled by rectangular boxes, and is expanded using data augmentation techniques, including rotation, flipping, cutMix, and random occlusion.

8. The method for detecting missing strips in cigarette packs according to claim 1, characterized in that, In step S4, the relevant data for recording missing events includes: extracting video clips of preset duration before and after detection from the original video stream, and packaging the detection results, enhanced images, and video clips into JSON and binary file formats, and sending them to the data storage system via HTTP API.

Citation Information

Patent Citations

  • Improved YOLOv5s-based cigarette box carton lacking detection method

    CN118196050A