Dense cigarette packet target detection method based on lightweight neural network

Through the improved C2fPA module, BIMAFPN and DyDhead architecture, the problem of insufficient accuracy of lightweight target detection models in dense target scenarios is solved, and efficient and real-time dense target detection is achieved, which is suitable for resource-constrained devices.

CN120635401APending Publication Date: 2025-09-12HUNAN NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510556133.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing lightweight object detection models lack accuracy in dense target scenes, especially in retail product displays and cigarette packaging images, making it difficult to achieve efficient real-time detection on resource-constrained devices.

Method used

The improved C2fPA module, BIMAFPN and DyDhead architecture are adopted, combined with weighted bidirectional multi-branch auxiliary feature pyramid network and dynamic deformable convolution to enhance feature extraction and target detection capabilities.

Benefits of technology

The accuracy and efficiency of dense target detection are improved, the number of parameters is reduced by 18.2%, real-time detection is achieved on resource-constrained devices, and the recall rate and precision are significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635401A_ABST
    Figure CN120635401A_ABST
Patent Text Reader

Abstract

The invention discloses a dense cigarette packet target detection method based on a lightweight neural network, and relates to the technical field of real-time image processing. The invention provides a lightweight DBA-YOLO based on YOLOv10 for the computing power bottleneck of dense target end side detection. The improvement points are as follows: C2fPA and ParNetAttention are embedded into the backbone, and foreground features are highlighted; a weighted bidirectional multi-branch auxiliary pyramid BIMAFPN is designed, and meanwhile shallow layer positioning and deep layer semantics are reserved; and a detection head DyDhead of the set dynamic deformable convolution and multiple attention is constructed. According to the method, the model is matched with an Anchar-Free framework, TAL positive sample distribution, CIoU and other loss functions, the model reaches mAP (0.5: 0.95) 46.9% and 99.4%, the parameter is 7.8 M, FLOPs is 3.7 G and the end-to-end delay is smaller than or equal to 45 ms in SKU-110K and cigarette data sets, and real-time deployment can be carried out on an OrinNX / RK3588 embedded platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of real-time image processing, and in particular to a dense cigarette packet target detection method based on a lightweight neural network. Background Art

[0002] In recent years, the application of deep learning technology has significantly improved the efficiency of object detection tasks, demonstrating superior performance compared to traditional methods and achieving remarkable success. This paradigm shift has not only demonstrated its strength in diverse fields such as autonomous driving, face recognition, and text recognition, but has also become an indispensable part of various downstream applications. Despite the transformative impact of deep learning on object detection, a notable challenge remains: computational resource intensity, which is particularly evident in leading object detectors. This computational demand poses a significant obstacle to the deployment of such detectors on resource-limited devices, a dilemma exacerbated by the rapid adoption, accessibility, and practicality of mobile devices such as body cameras and fixed-wing unmanned aerial vehicles (UAVs).

[0003] Defects and shortcomings of existing technology:

[0004] In real-world scenarios, detecting densely distributed objects is often challenging, especially in scenarios such as retail product displays, drone aerial imagery, and cigarette packaging images. These scenarios typically feature a large number of similar or identical objects densely packed together, often with occlusion instances where some objects obscure others. Current applications of dense object detection extend to similar areas such as product detection and dense crowd detection. Notably, the need for lightweight models in the field of dense object detection is becoming increasingly prominent in specific application scenarios such as assisted driving devices, drones, and wearable technologies.

[0005] Convolutional neural networks (CNNs) have demonstrated remarkable capabilities in image understanding and representation, making dense object detection more accurate. By extracting relevant image features, CNN models can improve performance and accuracy, making significant progress in this research field. Dense object detectors based on deep learning are mainly divided into two categories: two-stage detectors and single-stage detectors. The two-stage detector uses a region proposal stage, followed by feature extraction and object classification. Its representative method is the R-CNN series of algorithms. These methods perform well in average precision, but due to their reliance on candidate region generation and subsequent feature resampling, they are challenging in real-time dense object detection applications. In contrast, single-stage detectors adopt a dense prediction paradigm and directly perform anchor box classification and regression without first generating candidate boxes.

[0006] In many real-world applications, object detection tasks often need to be performed on platforms with limited computing resources. Current large-scale deep models are difficult to deploy on mobile devices. Therefore, there has been a surge of interest in developing efficient and lightweight network architectures. Model compression is one of the effective techniques for implementing lightweight neural networks. Weight pruning aims to compress the network by removing weights from the network architecture. Weight quantization reduces the number of fixed points in the network weights, thereby reducing network computation and parameter count. However, model compression may affect the accuracy of existing complex models. Artificially designed neural network architectures such as ShuffleNet and GhostNet use efficient convolutions to generate lightweight networks, reducing the number of parameters while maintaining satisfactory performance. The main problem with single-stage methods is that they are more likely to produce noisy pseudo-boxes, so further improvements are needed in terms of dense target detection accuracy.

[0007] Real-world dense object detection, such as images of supermarket merchandise and cigarette packaging, often involves numerous similar objects placed adjacent to each other, with similar colors and shapes, making it difficult to distinguish object boundaries. Furthermore, similar objects of varying sizes often exist, making it difficult to achieve high accuracy by directly deploying existing lightweight architecture models on mobile devices. Therefore, despite the flexible and scalable architecture and excellent accuracy of YOLOv10, its accuracy is still insufficient to meet practical requirements in real-world applications. Therefore, it is necessary to improve the model architecture, informative feature extraction strategies, and new object localization benchmarks. First, we select YOLOv10 as the base model. To extract informative features, we propose a novel plug-and-play framework: the Weighted Bidirectional Multi-branch Auxiliary Feature Pyramid Network (BIMAFPN). In BIMAFN, BiSAF surface-assisted fusion maintains shallow-layer backbone information through bidirectional connections, enhancing the network's ability to detect small objects. BiAAF advanced-assisted fusion enriches gradient information in the output layer through multidirectional connections. BiFPN also improves object detection efficiency and accuracy through bidirectional cross-scale connections and weighted feature fusion. Secondly, considering that the amount of information contained in feature regions varies, it is necessary to strengthen the influence of foreground features. Therefore, we propose the C2fPA module, which adds an attention mechanism to C2f. This module can adaptively adjust the weights of input features based on their varying importance, thereby improving feature extraction. Finally, to achieve accuracy in complex scenes, we propose a detection head with an attention mechanism, DyDhead. Based on the YOLOv10 detection head, we add a dynamic head (DynamicHead) mechanism, using dynamic deformable convolution (DyDCNv3) and multiple attention mechanisms to enhance the representation capability of feature maps. The effectiveness of the proposed method will be verified using the SKU110K and self-made cigarette packaging datasets. Summary of the Invention

[0008] In order to overcome the defects of the above-mentioned prior art, the present invention provides the following technical solutions: a dense cigarette package target detection method based on a lightweight neural network, comprising the following steps: S1, data acquisition and preprocessing: collecting images of the scene to be detected, normalizing the size, gamut consistency and Mosaic data enhancement of the images, and generating a training data set including SKU-110K images and cigarette packaging images; S2, backbone feature extraction: inputting the preprocessed images into the improved C2fPA module integrating the ParNetAttention mechanism, and extracting three groups of multi-scale feature maps of 80×80, 40×40 and 20×20 in turn; S3, multi-scale feature fusion: sending the multi-scale features into the weighted bidirectional multi-branch auxiliary feature pyramid network BIMAFPN, maintaining shallow positioning information through the BiSAF surface auxiliary fusion module, coupling deep semantic information through the BiAAF advanced auxiliary fusion module, and utilizing the bidirectional Cross-scale weighted fusion of BiFPN completes feature redistribution; S4, target detection head reasoning: the fused features are input into the multi-attention detection head DyDhead with dynamic deformable convolution DyDCNv3, and scale-aware attention, spatial-aware attention and task-aware attention are performed in sequence, decoupling the output position regression branch, category prediction branch and target confidence branch; S5, loss calculation and model training: CIoU bounding box regression loss, distribution focus loss and BCEWithLogits classification loss with fusion of Focal coefficient are used for One-to-OneHead and One-to-ManyHead outputs, and task alignment learning TAL dynamic positive sample allocation is jointly optimized; S6, online reasoning and result output: S2 to S4 are executed on the input image at the deployment end, and the bounding box coordinates, category label and confidence score of each target are output to achieve real-time detection of dense targets.

[0009] Preferably, the improved C2fPA module is composed of a channel attention branch, a 1×1 convolution branch and a 3×3 convolution branch in parallel, and the channel attention branch adopts a SpatialSqueeze-and-Excitation structure to perform Sigmoid scaling weighting on the extended feature map.

[0010] Preferably, the BIMAFPN consists of three parts in series: a BiFPN base layer, a BiSAF surface auxiliary fusion layer, and a BiAAF advanced auxiliary fusion layer; BiSAF performs upsampling on the feature map and aligns it with the 1×1 convolution channel before performing bidirectional weighted summation, and BiAAF adopts a cross-scale four-way parallel fusion structure to improve the recall rate of mid-scale targets.

[0011] Preferably, the DyDhead detection head includes: a ScaleAttn ​​module that uses depthwise separable convolution to generate a scale weight vector and adaptively recalibrates features of different resolutions; a SpatialAttn module that uses DyDCNv3 to obtain sparse spatial offsets and aggregate discriminative regional features; and a TaskAttn module that performs global average pooling and offset Sigmoid activation on channel features, dynamically turning channels on and off to adapt to classification and regression tasks.

[0012] Preferably, an SGD optimizer is used with an initial learning rate of 0.01, a momentum of 0.937, a weight decay of 0.0005, and the number of training rounds is set to 100 epochs for the SKU-110K dataset and 300 epochs for the cigarette packaging dataset.

[0013] A dense cigarette packet target detection system based on a lightweight neural network comprises: a data acquisition unit for acquiring and annotating original images containing dense targets; an image preprocessing unit for performing size adjustment and data enhancement; a lightweight target detection model construction unit for building and loading the network structure described in steps S2-S4; a training unit for performing offline training of the model according to the methods of claims 1-5; an inference unit for calling the trained detection model in embedded hardware for real-time target detection; and a result output unit for returning detection coordinates, categories, and confidence levels to an upper-level business system.

[0014] Preferably, the embedded hardware is an industrial camera terminal integrating NVIDIA Orin NX or RK3588 AI SoC, and the system end-to-end delay does not exceed 45ms (1280×1280 input resolution).

[0015] Preferably, the lightweight target detection model construction unit adopts an improved C2fPA-BIMAFPN-DyDhead architecture.

[0016] Preferably, the inference unit is configured on Jetson Orin NX or RK3588 AI SoC, supports INT8 quantized inference and outputs structured detection results.

[0017] Compared with the existing technology, the present invention has the following advantages: (1) Lightweight and efficient: the total number of parameters is 7.8M, which is 3.6% less than the original YOLOv10; FLOPs is 3.7G, which is 18.2% less than the baseline; accuracy is improved: the mAP (0.5:0.95) of the SKU-110K dataset is improved from 45.3% to 46.9%; the mAP of the cigarette packaging dataset is ≥99.4%, and the recall rate is ≥93.9%; (3) The present invention has strong robustness: multi-attention dynamic recalibration significantly reduces the missed detection rate in densely occluded scenes; DyDCNv3 spatial offset improves positioning accuracy in complex backgrounds; (4) The present invention is easy to deploy: Anchor-Free design does not require resetting the prior frame; the model can run in 1GB of video memory and is suitable for real-time applications on the end side. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is the architecture diagram of the DBA-YOLO method of the present invention.

[0019] Figure 2 This is a flow chart of the object detection of the present invention.

[0020] Figure 3 This is a C2fPA module diagram of the present invention.

[0021] Figure 4 This is the BIMAFPN diagram of the present invention.

[0022] Figure 5 Schematic diagram of the BiSAF structure of the present invention.

[0023] Figure 6 Schematic diagram of the BiAAF structure of the present invention.

[0024] Figure 7 Schematic diagram of DyDhead of the present invention.

[0025] Figure 8 Schematic diagram of the detailed structure of DyDhead of the present invention. DETAILED DESCRIPTION

[0026] The following is combined with Figures 1-8 , and further illustrate the technical solution of the present invention through specific implementation methods.

[0027] The present invention provides a dense cigarette packet target detection method based on a lightweight neural network, such as Figure 1 As shown in , DBA-YOLO is a lightweight model for dense object detection, such as Figure 1, the architecture consists of a backbone, a neck, and a head. The improved C2fPA module is integrated in the backbone to capture the normalized object size at resolutions of 80×80, 40×40, and 20×20, respectively. The neck adopts a streamlined multi-scale feature fusion network BIMAFPN to enhance PAN by removing the inefficient "Concat" module and adding four feature fusion modules to focus on the object area and alleviate the complex background effect. The head adopts an improved detection head DyDhead responsible for position, classification probability, and object score, which consists of three detection layers representing feature maps of different sizes. The flowchart of the object detection model is as follows Figure 2 shown.

[0028] exist Figure 3 C2fPA is an improved version of the C2f (CSPBottleneckwith2convolutions) module in YOLOv8. Its core modification is the introduction of the ParNetAttention mechanism, which enhances feature selection and multi-scale fusion efficiency. While maintaining lightweight, the attention mechanism dynamically calibrates the importance of feature channels, improving the model's ability to capture key information. Dynamic feature channel calibration inserts an attention module after cv1. Using the channel-wise attention mechanism (SSE) and multi-branch convolutions, it weights the expanded features, enhancing the response of important channels (such as target region features) and harmonizing redundant or redundant channels (such as background regions). ParNetAttention's parallel 1x1 and 3x3 convolution branches fuse spatial information at different scales within the attention mechanism. This significantly improves feature selection capabilities with virtually no increase in computational cost. Its design balances accuracy and speed, making it particularly suitable for vision tasks that demand both real-time performance and accuracy.

[0029] The module is divided into three branches, including channel attention branch, 1×1 convolution and 3×3 convolution. The channel attention (SSE, Spatial Squeeze and Excitation) is calculated as follows:

[0030] x sse =σ(f 1×1 (GAP(x)))·x;

[0031] x 1×1 =BN(f 1×1 (x));

[0032] x 3×3 =BN(f 3×3 (x));

[0033] in is Global Average Pooling. It adjusts the channel information, and σ is the Sigmoid activation function. Finally, the obtained attention weight is multiplied by the original input x. It extracts global information and adjusts the weight. 1×1 1×1 convolution is used for feature transformation to reduce computational complexity. BN is batch normalization. 3×3 It is a 3×3 convolution, which is mainly used to extract local features. Finally, feature fusion is performed. The calculation process is as follows:

[0034] y=SiLU(x 1×1 +x 3×3 +x sse );

[0035] The channel attention, 1×1 and 3×3 features are fused by element-by-element addition. Finally, SiLU (Sigmoid-Weighted Linear Unit) activation is used to enhance the nonlinear expression capability.

[0036] Multi-scale attention feature fusion network:

[0037] In object detection tasks, objects in an image may appear at different scales due to factors such as distance, angle, and occlusion. A single feature extraction method struggles to capture information at these different scales, resulting in missing information at other scales. The feature pyramid fusion framework provides a flexible framework for processing multi-resolution information to capture objects and features of varying sizes.

[0038] A BIMAFPN structure is proposed, which combines the multi-branch auxiliary feature pyramid network MAFPN and the bidirectional feature pyramid network BIFPN to analyze feature maps at different scales, thereby capturing different features at multiple levels of detail and emphasizing the role of foreground targets in detection.

[0039] BiFPN achieves multi-level feature fusion through a bidirectional path (top-down + bottom-up), where represents the output features of level l, It is the set of input layers connected to layer l (such as adjacent layers). Resize represents upsampling or downsampling operations, which are used to align resolutions. Conv represents depthwise separable convolution (including BN and activation functions) to reduce computational complexity. The bifpn formula is as follows:

[0040]

[0041] The main goal of BiSAF is to combine deep layer information with the features of the same level and high resolution shallow layer in the backbone network to preserve rich positioning details and enhance the spatial representation of the network. In addition, 1×1 convolution is used to control the number of channels in the shallow layer information to ensure that the number of channels of each input is the same, ensuring that BiSFPN operations can be performed without affecting subsequent learning. n-1 ,P n and P n+1 ∈R H×W×C Represents feature maps at different resolutions, where P n ,P' n and P″ n Denotes the feature layer of the backbone network and the two paths of the BIMAFPN. The symbol U(·) denotes an upsampling operation. Down represents a 3×3 downsampling convolution followed by a batch normalization layer, δ represents a silu function, and C represents a 1×1 convolution that controls the number of channels. The output after applying SAF is as follows:

[0042] P' n =bifpn(δ(C(Down(P n-1 ))),δ(C(P n )),U(P' n+1 ));

[0043] To further enhance the interactive utilization of feature layer information, we use the BiAAF module in the deeper layers of BIMAFPN to integrate multi-scale information. Specifically, Figure 3 Shows P″ n The AAF connection in the , which involves crossing the shallow high-resolution layer P ' n+1 , shallow low-resolution layer P' n-1 , Brothers shallow P' n and the previous layer P″ n-1 At this point, the final output layer can simultaneously merge information from 4 different layers, significantly improving the performance of medium-sized targets. According to the traditional FPN single-path architecture, we assume that the initial guidance information is already embedded in the shallow layers of BIMAFPN. Therefore, we balance the number of channels in each layer to ensure that the model obtains different outputs. The output results after applying BiAAF are as follows

[0044] P″ n =bifpn(Down(P' n-1 ),Down(P″ n-1 ),P' n ,U(P′ n+1 ));#

[0045] Improvement of detection head: schematic diagram of improved detection head DyDhead Figure 7 As shown in the figure, it mainly includes three modules: scale-aware attention, space-aware attention and task-aware attention.

[0046] Given a feature tensor F∈R L×S×C , the general formula for applying self-attention is:

[0047]

[0048] Where π(·) is an attention function, but using a fully connected layer for self-attention is computationally expensive, so we convert the attention function into three consecutive attentions:

[0049]

[0050] where π L (·),π S (·), and π C (·) are three different attention functions applied to dimensions L, S, and C respectively. π L (·) The scale-aware attention module dynamically fuses features of different scales according to their semantic importance. The specific process corresponds to Figure 8 The DyDhead in the design details the ScaleAttn ​​module.

[0051]

[0052] π S (·) corresponds to a feature-fused spatially aware attention module that continuously focuses on discriminative regions that coexist across spatial locations and feature layers. Given the high dimensionality of the spatial dimension S, we decompose this module into two steps: first, attention learning is made sparse by using DCNV3, and then cross-layer features are aggregated at the same spatial location. DCNV3 is the third-generation Deformable Convolutional Network (DCN), designed to improve the performance of convolutional neural networks (CNNs) when handling changes in the shape and position of objects in images. DCNV3 improves upon the previous two generations of DCN by introducing grouping operations and dynamic offsets, further enhancing the model's deformation-aware capabilities.

[0053]

[0054] Where K is the number of sparse sampling locations, p k +Δp k is the spatial offset Δp through self-learning k The position of the shift is used to focus on the discriminative area, and Δm k is the self-learning importance scalar at position pk.

[0055] In order to achieve joint learning and generalize the representation of different objects, π C (·) corresponds to a task-aware attention module that can dynamically open and close feature channels to support different tasks.

[0056]

[0057] Here, Fc represents the feature slice at the cth channel, and [α1,α2,β1,β2]T=θ(·) is a hyperfunction used to learn the activation threshold. The implementation of function θ(·) involves first performing global average pooling on the L×S dimension to reduce the dimensionality, then applying two fully connected layers and a normalization layer, and finally applying a biased sigmoid function to normalize the output to the range [-1, 1]. This process corresponds to the TaskAttn module in the DyDhead detailed design shown in the figure.

[0058] Loss function: YOLOv10 continues the anchor-free design of YOLOv8, adopts dual decoupled detection heads (one-to-one head and one-to-many head), and optimizes the distribution of positive and negative samples through the task alignment learning (TAL) mechanism to solve the problem of misalignment between classification and regression tasks. Its loss function consists of three parts:

[0059]

[0060]

[0061] Where ρ is the Euclidean distance of the center point, c is the diagonal length of the minimum closure area, v is the aspect ratio penalty term, and α is the adaptive balance coefficient.

[0062] Distributed focus loss improves regression accuracy through probability distribution modeling. 12 bins are set for the small target area (coordinates 0-0.3) and 8 bins are set for the large target area (0.3-1.0). The formula is:

[0063]

[0064] where y gt is the true label, y i and y i+1 are two adjacent points of the predicted value, P(y i ) and P(y i+1 ) is y i and y i+1 The probability of log(P(y i )) and log(P(y i+1 )) is the logarithmic probability, which is used to measure the uncertainty of the predicted value, and n is the total number of predicted values

[0065] For classification loss (ClassLoss), YOLOv10 uses binary cross entropy loss (BCEWithLogitsLoss) and handles multi-label classification problems through the Sigmoid activation function. This design avoids the limitations of Softmax in multi-label classification. Softmax is more suitable for single-label classification of mutually exclusive categories, while the combination of BCEWithLogitsLoss and Sigmoid can handle multi-label scenarios more flexibly. At the same time, the FocalLoss mechanism is introduced to reduce the weight of easy-to-classify samples by adjusting the γ parameter to alleviate the class imbalance problem. The formula is:

[0066]

[0067] where N max is the maximum number of category samples in the dataset:

[0068]

[0069] For the confidence loss (ObjectnessLoss), the task alignment confidence is used. The target confidence is dynamically generated through TAL, and the IoU and classification score of the predicted box and the real box are combined to achieve the alignment of classification and positioning:

[0070] Score=IoU β ClsScore 1-β ;#

[0071] Where β is the task balance coefficient, which defaults to 0.5.

[0072] Effect of BIMAFPN model: After replacing the original neck module of yolov10 with BIMAFPN module, the number of parameters is significantly reduced. SAF surface assisted fusion maintains shallow backbone information through bidirectional connection, enhancing the network's ability to detect small targets, and AAF advanced assisted fusion enriches the gradient information of the output layer through multi-directional connection. At the same time, BiFPN uses bidirectional cross-scale connection and weighted feature fusion to improve target detection efficiency and accuracy. As shown in the table, the number of parameters is reduced by 30%, mAP is improved by 0.5%, and AP 75 It also increased by 0.8%.

[0073] Effect of C2fPA: By improving the C2fPA feature fusion module, the weights of the input features are adaptively adjusted according to their importance, improving the effect of feature extraction, thereby enhancing the feature extraction capability in complex background environments. As shown in the table, due to the addition of the attention mechanism, even if the number of parameters is slightly increased, the mAP is improved by 0.8%, and the AP 75 It also increased by 1.2%.

[0074] Impact of DyDhead: By improving the detection head, using dynamic convolution and multiple attention mechanisms to enhance the representation of feature maps, the model performs better in complex scenarios. As shown in the table, the number of parameters is slightly increased, but the mAP and AP 75 Improved again, mAP increased by 0.9% over the baseline model, AP 75 The 0.3% increase further verifies the feasibility of the improved model.

[0075] DBA-YOLO is a new method for dense target detection in complex backgrounds, which uses the improved C2-PA module as the backbone feature extraction network. DBA-YOLO is a lightweight network that reduces model parameters and computation, making it suitable for deployment on mobile devices. Fusion of MAFPN and BiFPN as a multi-scale attention feature fusion network can enhance feature extraction and improve detection accuracy. Replacing the original detection head with DyDhead enables the model to perform better in complex scenes. Verification on a real cigarette pack dataset demonstrates the practical feasibility of DBA-YOLO, which achieves superior performance on the SKU-110K and cigarette pack datasets compared to YOLOv10. DBA-YOLO reduces parameters by 3.6%, improves mAP by 1.6% on the SKU-110 dataset, and AP 75 It is improved by 2.2%. On the actual cigarette pack dataset, mAP is improved by 1.6%. 50 Improved to 99.4%, AP 50 The accuracy is increased to 93.9%, meeting the detection needs of complex scenes. Experimental results confirm that DBA-YOLO outperforms existing dense target detection models in terms of accuracy and local bounding box prediction in complex environments.

[0076] Dataset:

[0077] The SKU-110K dataset (Eran Goldman, Roei Herzig, Aviv Eisenschtat, Oria Ratzon, Itsik Levi, Jacob Goldberger, and Tal Hassner, 2019) and a custom cigarette pack dataset were used. SKU-110K contains 11,762 images and 173,678 instances collected from various supermarket stores. The dataset is split into 8,821 training samples and 2,941 test samples. The custom cigarette pack dataset includes over 1,000 real images taken in various cities, weather conditions, and times of day, each with a size of 960 × 1,280 pixels. To increase diversity and prevent overfitting, 50 natural scene images without cigarette packs are included as negative samples. Labeling was performed using labelImg, resulting in 1,073 images and 50,173 labeled samples, of which 887 images are in the training set and 186 samples are in the test set. Because the original cigarette pack images are proprietary, all faces shown are mosaicked for privacy reasons.

[0078] Experimental environment and parameter settings

[0079] The experiment used Python 3.10.14, CUDA 12.1, RTX 4090D GPU (24GB) and Xeon 8474C CPU. The input image sizes of SKU-110K and cigarette pack datasets are 640×640 and 1280×1280 respectively, and mosaic data augmentation is used to increase diversity. Training involves 100 epochs for SKU-110K and 300 epochs for the cigarette pack dataset to achieve the optimal weights. Stochastic gradient descent (SGD) with momentum of 0.937 and weight decay coefficient of 0.0005 is used for gradient update. The initial learning rate is 10 -2 , the final learning rate is 10 -4 , the batch size is 16.

[0080] Evaluation Metrics

[0081] This experiment uses evaluation indicators in the field of deep learning object detection, such as precision, recall, and average precision (AP). The formula is as follows:

[0082]

[0083] Where TP is the number of positive samples correctly predicted as positive, FP is the number of positive samples incorrectly predicted as positive, and FN is the number of positive samples incorrectly predicted as negative. P stands for precision and R stands for recall. We used the same evaluation method as COCO (Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick, 2014), which reports the average precision (mAP) when IoU = 0.5:0.05:0.95 (IoU ranges from 0.5 to 0.95 with a step size of 0.05). In addition, we report AP 50 (IoU=0.5) and AP 75 (IoU=0.75).

Claims

1. A dense cigarette packet target detection method based on a lightweight neural network, characterized in that: The steps include: S1. Data acquisition and preprocessing: Collect images of the scene to be detected, perform size normalization, color gamut uniformity, and mosaic data enhancement on the images, and generate a training dataset containing SKU-110K images and cigarette packaging images; S2. Backbone feature extraction: The preprocessed image is input into the improved C2fPA module integrated with the ParNetAttention mechanism, and three groups of multi-scale feature maps of 80×80, 40×40, and 20×20 are extracted in sequence; S3. Multi-scale feature fusion: The multi-scale features are fed into the weighted bidirectional multi-branch auxiliary feature pyramid network (BIMAFPN). The shallow positioning information is maintained through the BiSAF surface auxiliary fusion module, the deep semantic information is coupled through the BiAAF advanced auxiliary fusion module, and the feature redistribution is completed using the bidirectional cross-scale weighted fusion BiFPN. S4, object detection head reasoning: The fused features are input into the multi-attention detection head DyDhead with dynamic deformable convolution DyDCNv3, which performs scale-aware attention, space-aware attention, and task-aware attention in sequence, decoupling the output position regression branch, category prediction branch, and object confidence branch; S5. Loss calculation and model training: The CIoU bounding box regression loss, distribution focus loss, and BCEWithLogits classification loss with fused Focal coefficient are used for One-to-OneHead and One-to-ManyHead outputs, and the dynamic positive sample allocation of TAL is jointly optimized with task alignment learning. S6, online reasoning and result output: Execute S2 to S4 on the input image at the deployment end, output the bounding box coordinates, category label, and confidence score of each target, and achieve real-time detection of dense targets.

2. The method for dense cigarette packet target detection based on a lightweight neural network according to claim 1, characterized in that: The improved C2fPA module consists of a channel attention branch, a 1×1 convolution branch and a 3×3 convolution branch in parallel. The channel attention branch adopts a SpatialSqueeze-and-Excitation structure and performs Sigmoid scaling weighting on the extended feature map.

3. The method for dense cigarette packet target detection based on a lightweight neural network according to claim 1, characterized in that: The BIMAFPN consists of three parts connected in series: the BiFPN base layer, the BiSAF surface auxiliary fusion layer, and the BiAAF advanced auxiliary fusion layer; BiSAF performs upsampling on the feature map and aligns the 1×1 convolution channels before performing bidirectional weighted summation, while BiAAF adopts a cross-scale four-way parallel fusion structure to improve the recall rate of mid-scale targets.

4. The method for dense cigarette packet target detection based on a lightweight neural network according to claim 1, characterized in that: The DyDhead detection head comprises: ScaleAttn ​​module - uses depthwise separable convolution to generate scale weight vectors and adaptively recalibrate features of different resolutions; SpatialAttn module - uses DyDCNv3 to obtain sparse spatial offsets and aggregate discriminative regional features; TaskAttn module - performs global average pooling and offset Sigmoid activation on channel features, dynamically turning channels on and off to adapt to classification and regression tasks.

5. The method for dense cigarette packet target detection based on a lightweight neural network according to claim 1, characterized in that: The training phase uses an SGD optimizer with an initial learning rate of 0.01, a momentum of 0.937, a weight decay of 0.0005, and the number of training rounds is set to 100 epochs for the SKU-110K dataset and 300 epochs for the cigarette packaging dataset.

6. A dense cigarette packet target detection system based on lightweight neural network, characterized by: include: A data acquisition unit, used to collect and annotate original images containing dense objects; Image preprocessing unit for performing size adjustment and data augmentation; A lightweight target detection model building unit, used to build and load the network structure described in steps S2-S4; A training unit, configured to perform offline training on the model according to the method of claims 1 to 5; The inference unit is used to call the trained detection model in the embedded hardware for real-time target detection; The result output unit is used to return the detection coordinates, categories and confidence levels to the upper-level business system.

7. The dense cigarette packet target detection system based on a lightweight neural network according to claim 6, characterized in that: The embedded hardware is an industrial camera terminal integrated with NVIDIA OrinNX or RK3588AI SoC, and the system end-to-end delay does not exceed 45ms.

8. The dense cigarette packet target detection system based on a lightweight neural network according to claim 7, characterized in that: The lightweight target detection model construction unit adopts the improved C2fPA-BIMAFPN-DyDhead architecture.

9. The dense cigarette packet target detection system based on a lightweight neural network according to claim 8, characterized in that: The inference unit is configured on Jetson Orin NX or RK3588 AI SoC, supports INT8 quantized inference and outputs structured detection results.