Retail commodity target detection method based on improved YOLOv8 model
By introducing BiFPN, Triplet Attention and Shape-IoU loss functions in the YOLOv8 model, the problem of insufficient accuracy of target overlap, small target detection and similar category discrimination in retail scenarios is solved, and efficient and accurate target detection of retail products is achieved.
Patent Information
- Application Number
- CN202510338670.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-13
AI Technical Summary
In retail scenarios, the YOLOv8 model has insufficient accuracy and robustness when dealing with target overlap, small target detection and similar category discrimination, and optimization methods often sacrifice detection speed, making it difficult to meet real-time requirements.
The weighted bidirectional feature pyramid network BiFPN and triple attention mechanism were introduced in the YOLOv8 model. The optimization loss function uses the Shape-IoU loss function to improve the accuracy of border regression.
It improves the accuracy and stability of target detection of retail products, enhances the detection ability of small targets, improves the model's ability to judge similar categories, and maintains efficient detection speed.
Smart Images

Figure CN120147752A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of object detection, and specifically to a method for detecting retail commodity objects based on an improved YOLOv8 model. Background Art
[0002] In modern retail scenarios, the efficiency bottleneck of the traditional manual checkout mode has become increasingly prominent. When customers queue up in front of the cashier, the cashier needs to scan the barcodes of each item one by one. This process not only takes time and effort but also often results in recognition failures due to problems such as barcode smudging, item stacking, or packaging reflection. Especially during promotional activities, the number of items in the shopping cart surges, and the limitations of manual operations become more prominent. Price disputes caused by incorrect or missed scanning of items occur frequently, seriously affecting the consumption experience and the operational efficiency of the store.
[0003] Automatic commodity recognition technology based on computer vision provides an innovative solution to this problem. This technology captures commodity images through cameras deployed in the cashier area and uses object detection algorithms to achieve rapid multi-commodity positioning and classification. Advanced detection models represented by YOLOv8, with their single-stage detection architecture and optimized network design, have achieved remarkable breakthroughs in real-time performance and accuracy. Compared with earlier versions, YOLOv8 can handle object detection tasks in complex scenarios more efficiently by improving the feature fusion mechanism and dynamic label assignment strategy.
[0004] However, when directly applying YOLOv8 to actual retail scenarios, many challenges still exist. The close arrangement of commodities often leads to overlapping objects, and the interference between detection frames causes missed detections; small-sized commodities only occupy a very small area in the image, and it is difficult for the model to capture their effective features; while the packaging of the same type of commodities in different specifications is highly similar in appearance, which is extremely likely to cause misclassification. Although YOLOv8 performs excellently on general datasets, its default configuration is insufficiently adapted to the particularity of retail scenarios - the feature extraction accuracy of dense objects is limited, the robustness of small object detection needs to be improved, and the discrimination ability of similar categories still needs to be strengthened. Existing improvement methods often sacrifice the detection speed while trying to optimize the accuracy by increasing the network depth or introducing complex modules, making it difficult to meet the strict real-time requirements of the checkout system. Therefore, how to conduct targeted optimization for the uniqueness of retail scenarios while maintaining the efficient inference characteristics of YOLOv8 has become a key research direction for promoting the implementation of intelligent cashier technology. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for detecting retail commodity objects based on an improved YOLOv8 model, aiming to improve the detection accuracy and detection stability of retail commodity objects.
[0006] To achieve the above purpose, the technical solutions adopted by the present invention include the following steps, including:
[0007] S1: Collect retail product images to establish a retail product object detection dataset;
[0008] S2: Build a retail product object detection model based on the improved YOLOv8;
[0009] S3: Optimize the loss function;
[0010] S4: Train the built improved YOLOv8 retail product object detection model;
[0011] S5: Evaluate the performance of the trained model.
[0012] Furthermore, in step S1, when collecting retail product images, a certain number of product images at the checkout counters are selected from the RPC dataset to establish a retail product object detection dataset, which is divided into a training set, a validation set, and a test set according to the ratio of 7:2:1;
[0013] Furthermore, in step S2, the specific steps for building a retail product object detection model based on the improved YOLOv8 include;
[0014] S201: Introduce a lightweight and efficient weighted bidirectional feature pyramid network BiFPN structure into the neck network;
[0015] S202: Introduce a triple attention mechanism (Triplet Attention) structure;
[0016] Furthermore, in step S3, when optimizing the loss function, the Shape-IoU loss function is used to replace the original CIoU loss function. By focusing on calculating the loss of the border's own shape and scale, the border regression becomes more accurate. Its definition is as follows:
[0017]
[0018] where scale is the scale factor, related to the scale of the objects in the dataset, and ww and hh are the weight coefficients in the horizontal and vertical directions respectively, and their values are related to the shape of the GT box. The corresponding border regression loss is as follows:
[0019] L Shape-IoU = 1 - IoU + distance shape + 0.5 × Ω shape
[0020] Furthermore, in step S4, the specific steps for training the built improved YOLOv8 retail product object detection model include;
[0021] S401: Set training parameters and use the SGD optimizer for training. The image input size is 640×640, the initial learning rate is 0.01, the number of iterations is 200, and the batch size is 16;
[0022] S402; Put the training set and validation set images in the dataset into the improved YOLOv8 retail commodity target detection model for training;
[0023] S403: Train the model according to the set parameters. By observing the change trend of the loss function, adjust the learning rate and the number of iterations of model training until the change of the loss function tends to be stable to obtain the final trained model;
[0024] Furthermore, in step S5, the performance of the trained model is evaluated. The network model is evaluated from aspects such as recall, precision, average precision, mean average precision, and model size. The formulas are as follows:
[0025]
[0026] Among them, TP represents the number of correctly identified positive samples, TN represents the number of correctly identified negative samples, FP is the number of incorrectly identified positive samples, and FN represents the number of incorrectly identified negative samples. Precision represents the accuracy of the prediction result. The higher the value, the fewer misdetections. Recall represents the comprehensiveness of target prediction. The higher the value, the fewer missed detections. The AP value is obtained by calculating the area under the PR curve with Recall as the abscissa and Precision as the ordinate. Finally, mAP is used as the evaluation index of precision to measure the comprehensive performance of the trained model on all categories;
[0027] Compared with the prior art, the present invention has the following beneficial effects:
[0028] The present invention introduces a lightweight and efficient weighted bidirectional feature pyramid structure BIFPN in the neck network to reduce model parameters while improving accuracy. Triplet Attention is also added to improve the model's detection ability for small targets. Shape-IOU is used to improve detection accuracy. Therefore, the improved YOLOv8 model can quickly and accurately detect retail commodity targets in images and play a huge role in practical applications. Description of the Drawings
[0029] Figure 1 Flowchart of a retail commodity target detection method based on an improved YOLOv8 model in a specific embodiment of the present invention;
[0030] Figure 2 Schematic diagram of the BiFPN structure in a retail commodity detection method based on an improved YOLOv8 model in a specific embodiment of the present invention;
[0031] Figure 3 Schematic diagram of the Triplet Attention in a retail commodity detection method based on an improved YOLOv8 model in a specific embodiment of the present invention; Specific embodiments
[0032] Examples
[0033] As Figure 1 shown, the technical solution of the present invention includes the following steps
[0034] S1: Collect retail commodity images to construct a retail commodity target detection dataset;
[0035] S2: Construct a retail commodity target detection model based on the improved YOLOv8;
[0036] S3: Optimize the loss function;
[0037] S4: Train the constructed improved YOLOv8 retail commodity target detection model;
[0038] S5: Evaluate the performance of the trained model.
[0039] Specifically, in step S1, when collecting retail commodity images, a certain number of checkout counter commodity images are selected from the RPC dataset to establish a retail commodity target detection dataset, which is divided into a training set, a validation set, and a test set according to a ratio of 7:2:1;
[0040] Specifically, in step S2, the specific steps for constructing a retail commodity target detection model based on the improved YOLOv8 include;
[0041] S201: Introduce a weighted bidirectional feature pyramid network (BiFPN) structure into the neck network;
[0042] Reconstruct the neck network using the lightweight and efficient weighted bidirectional feature pyramid (BiFPN) structure to enhance the model's feature extraction and fusion capabilities and improve the model's detection effect.
[0043] S202: Introduce a triple attention mechanism (Triplet Attention) structure;
[0044] Embed the triple attention mechanism (Triplet Attention) into the backbone network to enhance the feature learning ability and improve the model's perception level of targets at different scales
[0045] Specifically, in step S3, the optimized loss function uses the Shape-IoU loss function to replace the original CIoU loss function. By focusing on calculating the loss based on the shape and scale of the bounding box itself, the bounding box regression becomes more accurate. Its definition is as follows:
[0046]
[0047] Where scale is the scale factor, which is related to the scale of the objects in the dataset. ww and hh are the weight coefficients in the horizontal and vertical directions respectively, and their values are related to the shape of the GT box. The corresponding bounding box regression loss is as follows:
[0048] L Shape-IoU = 1 - IoU + distance shape + 0.5 × Ω shape
[0049] Specifically, in step S4, the improved YOLOv8 retail commodity object detection model constructed by training specifically includes the following steps:
[0050] S401: Set the training parameters and use the SGD optimizer for training. The image input size is 640×640, the initial learning rate is 0.01, the number of iterations is 200, and the batch size is 16;
[0051] S402; Put the training set and validation set images in the dataset into the improved YOLOv8 retail commodity object detection model for training;
[0052] S403: Train the model according to the set parameters. By observing the change trend of the loss function, adjust the learning rate and the number of iterations of the model training until the change of the loss function tends to be stable, and obtain the final trained model;
[0053] Specifically, in step S5, the performance of the trained model is evaluated. The network model is evaluated from aspects such as recall, precision, average precision, mean average precision, and model size. The formulas are as follows:
[0054]
[0055] Among them, TP represents the number of correctly identified positive samples, TN represents the number of correctly identified negative samples, FP is the number of incorrectly identified positive samples, and FN represents the number of incorrectly identified negative samples. Precision represents the accuracy of the prediction results, and the higher the value, the fewer the misdetections. Recall represents the comprehensiveness of the target prediction, and the higher the value, the fewer the missed detections. The AP value is obtained by calculating the area under the PR curve with Recall as the abscissa and Precision as the ordinate. Finally, mAP is used as the evaluation index of precision to measure the comprehensive performance of the trained model on all categories;
[0056] As described above, it is only the preferred specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent replacements or changes should be covered within the protection scope of the present invention.
Claims
1. A retail commodity target detection method based on improved YOLOv8, characterized in that: The process includes: S1: Collect retail product images to establish a retail product target detection dataset; S2: Build a retail commodity target detection model based on improved YOLOv8; S3: optimize the loss function; S4: Trained and constructed improved YOLOv8 retail commodity object detection model; S5: Perform performance evaluation on the trained model.
2. A retail commodity target detection method based on YOLOv8 according to claim 1, characterized in that: In step S1, the commodity settlement map in the RPC data set is selected to establish a retail commodity target detection data set, which is divided into a training set, a validation set and a test set according to a ratio of 7:2:
1.
3. A retail commodity target detection method based on YOLOv8 according to claim 1, characterized in that: In step S2, the specific steps of constructing a retail commodity target detection model based on improved YOLOv8 include: S201: Introduce a lightweight and efficient weighted bidirectional feature pyramid network BiFPN structure into the neck network; S202: Introduce the Triplet Attention structure and add a Triplet Attention module after the large backbone network of the original model to improve the detection accuracy of the model at the cost of a small number of parameters.
4. A retail commodity target detection method based on YOLOv8 according to claim 1, characterized in that: In step S3, the optimization loss function is to replace the original CIoU loss function of YOLOv8 with the Shape-IoU loss function which calculates the loss by focusing on the shape and scale of the border itself so as to make the border regression more accurate.
5. The method for detecting retail commodity targets based on YOLOv8 according to claim 1, characterized in that: In step S4, the training and construction of the improved YOLOv8 retail commodity target detection model specifically includes the following steps: S401: Set training parameters and use SGD optimizer for training. The image input size is 640×640, the initial learning rate is 0.01, the number of iterations is 200, and the batch size is 16; S402: putting the training set and validation set images in the data set into the improved YOLOv8 retail commodity target detection model for training; S403: Train the model according to the set parameters, and adjust the learning rate and number of iterations of the model training by observing the change trend of the loss function until the loss function changes tend to be stable, thereby obtaining the final training model.
6. A retail commodity target detection method based on YOLOv8 according to claim 1, characterized in that: In step S5, the performance of the training model is evaluated, and the network model is evaluated from aspects such as recall rate, precision rate, average precision, mean average precision and model size.