Dense shelf commodity target detection method based on improved YOLOv8
By introducing a global attention mechanism and multi-attention detection head in the YOLOv8 model and optimizing the loss function, the challenges of YOLOv8 in detection accuracy and stability in high-density commodity environments are solved, and more efficient detection of dense shelf commodity items is achieved.
Patent Information
- Application Number
- CN202510338668.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-13
AI Technical Summary
When deploying YOLOv8 to a real retail shelf scenario, there are still challenges in detection accuracy and stability, especially in high-density commodity environments.
By introducing global attention mechanism (GAM) and multi-attention detection head (DyHead), and optimizing the loss function as a WIoU loss function, the YOLOv8 model is improved to improve detection accuracy and stability.
The improved YOLOv8 model can detect dense shelf product targets more quickly and accurately, improving performance in practical applications.
Smart Images

Figure CN120147751A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of object detection, and specifically to a method for detecting dense commodity objects on a dense shelf based on improved YOLOv8. Background Art
[0002] Driven by the dual forces of the digital economy and consumption upgrading, the retail industry is undergoing unprecedented changes. Online e-commerce continues to squeeze the physical retail market with its convenience and price advantages, and the demand of consumers for services such as "instant delivery" and "omnichannel shopping" forces the supply chain to improve its response speed. At the same time, physical stores are facing multiple pressures such as rising rents, climbing labor costs, and insufficient inventory turnover efficiency. The traditional extensive management mode of the "huge workforce strategy" is no longer sustainable. Against this background, as the core link connecting the upstream and downstream of the supply chain, warehouse management urgently needs to achieve efficiency leap and cost optimization through intelligent upgrading. However, the traditional mode relying on manual inventory has become the key bottleneck restricting the development of the industry.
[0003] The automatic commodity recognition technology based on computer vision is rapidly changing the way of retail shelf management, providing a more intelligent and efficient solution for merchants. This technology arranges high-resolution cameras in the shelf area to capture commodity images in real time, and combines advanced object detection algorithms to quickly locate, classify, and monitor the status of commodities. The latest object detection models represented by YOLOv8, with their mature single-stage detection architecture and carefully optimized network design, have achieved an unprecedented balance between processing speed and detection accuracy. For example, the flexibility and efficiency of YOLOv8 enable it to quickly identify and process data in a high-density commodity environment, providing real-time commodity information for retailers, thereby enhancing the intelligence level of inventory management and sales strategies.
[0004] In addition, this technology not only improves the accuracy of commodity recognition, but also endows merchants with the ability to monitor the status of shelves in real time, helping them better manage inventory, reduce out-of-stock situations, and ensure the neatness and optimized display of shelves. Through data analysis, merchants can also gain insights into customer shopping behaviors, so as to formulate more precise marketing strategies. However, when deploying YOLOv8 to the actual retail shelf scenario, many complex technical challenges still need to be solved to further promote the implementation and development of intelligent retail. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for detecting dense shelf commodity objects based on improved YOLOv8, aiming to improve the detection accuracy and stability of dense shelf commodity objects.
[0006] To achieve the above purpose, the technical solutions adopted by the present invention include the following steps, including:
[0007] S1: Collect images of shelf goods to establish a dense shelf goods target detection dataset;
[0008] S2: Build a dense shelf goods target detection model based on the improved YOLOv8;
[0009] S3: Optimize the loss function;
[0010] S4: Train the built improved YOLOv8 dense shelf goods target detection model;
[0011] S5: Evaluate the performance of the trained model.
[0012] Furthermore, in step S1, the damaged pictures in the SKU-110K dataset are screened out, and the remaining ones are used to establish a dense shelf goods target detection dataset, which is divided into a training set, a validation set, and a test set according to the ratio of 7:2:1;
[0013] Furthermore, in step S2, the specific steps for building a dense shelf goods target detection model based on the improved YOLOv8 include;
[0014] S201: Introduce the global attention mechanism GAM structure, and add a GAM module behind the backbone network of the original YOLOv8 model to increase the global vision and dynamic adjustment ability of the model, and improve the recognition accuracy of small targets;
[0015] S202: Introduce the multi-attention detection head DyHead to replace the original detection head of YOLOv8;
[0016] Furthermore, in step S3, the optimized loss function uses the WIoU loss function to replace the original CIoU loss function. By comprehensively considering the orientation, centroid distance, and overlapping area, and introducing a dynamic non-monotonic focusing mechanism, the attention of the model to medium-quality samples is enhanced by adjusting the anchor box loss weight, thereby improving the overall performance. Its definition is as follows:
[0017]
[0018] IoU (Intersection over Union) represents the intersection ratio of the predicted bounding box and the ground truth bounding box, h and w represent the height and width of the predicted bounding box, b cx and b cy represent the center position of the predicted bounding box, and represent the center position of the ground truth bounding box, c h and c wIt represents the height and width of the minimum enclosing boundary formed by the predicted bounding box and the ground truth bounding box. Among them, r is the non-monotonic focusing factor in WIoU, and its formula is:
[0019]
[0020] β is defined as an outlier for measuring the quality of the bounding box, and δ and α are hyperparameters that can be adjusted to adapt to different models:
[0021] Furthermore, in the step S4, the improved YOLOv8 dense shelf commodity object detection model constructed by training specifically includes the following steps:
[0022] S401: Set the training parameters and use the SGD optimizer for training. The image input size is 640×640, the initial learning rate is 0.01, the number of iterations is 200, and the batch size is 16;
[0023] S402; Put the training set and validation set images in the dataset into the improved YOLOv8 dense shelf commodity object detection model for training;
[0024] S403: Train the model according to the set parameters, adjust the learning rate and the number of iterations of the model training by observing the change trend of the loss function until the change of the loss function tends to be stable, and obtain the final trained model;
[0025] Furthermore, in the step S5, the performance of the trained model is evaluated. The network model is evaluated from aspects such as recall, precision, average precision, mean average precision, and model size. The formulas are as follows:
[0026]
[0027] Among them, TP represents the number of correctly recognized positive samples, TN represents the number of correctly recognized negative samples, FP is the number of incorrectly recognized positive samples, and FN represents the number of incorrectly recognized negative samples. Precision represents the accuracy of the prediction result, and the higher the value, the fewer misdetections. Recall represents the comprehensiveness of the target prediction, and the higher the value, the fewer missed detections. The AP value is obtained by calculating the area under the PR curve with Recall as the abscissa and Precision as the ordinate. Finally, mAP is used as the evaluation index of accuracy to measure the comprehensive performance of the trained model on all categories;
[0028] Compared with the prior art, the present invention has the following beneficial effects:
[0029] The present invention introduces the self-attention detection head DyHead, which improves the model's feature learning ability and can perform adaptive weight adjustment, enabling the model to better handle objects of different scales and occlusions. GAM attention is also added to improve the model's detection ability for small targets. WIOU is used to further improve the detection accuracy. Therefore, the improved YOLOv8 model can quickly and accurately detect dense shelf commodity targets in images and play a huge role in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 Flowchart of a method for detecting dense shelf commodity targets based on an improved YOLOv8 model in a specific embodiment of the present invention;
[0031] Figure 2 Schematic diagram of the DyHead structure in a method for detecting dense shelf commodities based on an improved YOLOv8 model in a specific embodiment of the present invention;
[0032] Figure 3 Schematic diagram of GAM attention in a method for detecting retail commodities based on an improved YOLOv8 model in a specific embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0033] EXAMPLE
[0034] As Figure 1 shown, the technical solution of the present invention includes the following steps
[0035] S1: Collect shelf commodity images to establish a dense shelf commodity target detection dataset;
[0036] S2: Construct a dense shelf commodity target detection model based on the improved YOLOv8;
[0037] S3: Optimize the loss function;
[0038] S4: Train the constructed improved YOLOv8 dense shelf commodity target detection model;
[0039] S5: Evaluate the performance of the trained model.
[0040] Specifically, in step S1, when collecting shelf commodity images, damaged images in the SKU-110K dataset are screened out, and the remaining ones are used to establish a dense shelf commodity target detection dataset, which is divided into a training set, a validation set, and a test set according to a ratio of 7:2:1;
[0041] Specifically, in step S2, the specific steps for constructing a dense shelf commodity target detection model based on the improved YOLOv8 include;
[0042] S201: Introduce the Global Attention Mechanism (GAM) structure;
[0043] Add a GAM module behind the backbone network of the original YOLOv8 model to increase the global vision and dynamic adjustment ability of the model, and improve the recognition accuracy of the model for small targets;
[0044] S202: Introduce the Multi-Attention Detection Head (DyHead) to replace the original detection head of YOLOv8;
[0045] The DyHead unifies the attention mechanisms in three dimensions through the self-attention mechanism. Different attention mechanisms can be applied from three perspectives: scale-aware attention, space-aware attention, and task-aware attention. This integration improves the model's feature learning ability and enables adaptive weight adjustment, allowing the model to better handle objects of different scales and occlusions.
[0046] Specifically, in step S3, the optimized loss function uses the WIoU loss function to replace the original CIoU loss function. By comprehensively considering the orientation, centroid distance, and overlapping area, and introducing a dynamic non-monotonic focusing mechanism, the attention to medium-quality samples is enhanced by adjusting the anchor box loss weight, thus improving the overall performance. Its definition is as follows:
[0047]
[0048] IoU (Intersection over Union) represents the ratio of the intersection of the predicted bounding box and the ground truth bounding box. h and w represent the height and width of the predicted bounding box, b cx and b cy represent the center position of the predicted bounding box, and represent the center position of the ground truth bounding box, c h and c w represent the height and width of the smallest enclosing bounding box formed by the predicted box and the ground truth bounding box. Among them, r is the non-monotonic focusing factor in WIoU, and its formula is:
[0049]
[0050] β is defined as an outlier for measuring the bounding box quality, and δ and α are hyperparameters that can be adjusted to adapt to different models:
[0051] Specifically, in step S4, the training of the improved YOLOv8 dense shelf commodity target detection model constructed specifically includes the following steps;
[0052] S401: Set training parameters and use the SGD optimizer for training. The image input size is 640×640, the initial learning rate is 0.01, the number of iterations is 200, and the batch size is 16;
[0053] S402: Put the training set and validation set images in the dataset into the improved YOLOv8 dense shelf commodity target detection model for training;
[0054] S403: Train the model according to the set parameters. By observing the change trend of the loss function, adjust the learning rate and the number of iterations of model training until the change of the loss function tends to be stable to obtain the final trained model;
[0055] Specifically, in step S5, the performance of the trained model is evaluated. The network model is evaluated from aspects such as recall, precision, average precision, mean average precision, and model size. The formulas are as follows:
[0056]
[0057]
[0058] Among them, TP represents the number of correctly recognized positive samples, TN represents the number of correctly recognized negative samples, FP is the number of incorrectly recognized positive samples, and FN represents the number of incorrectly recognized negative samples. Precision represents the accuracy of the prediction result. The higher the value, the fewer misdetections. Recall represents the comprehensiveness of target prediction. The higher the value, the fewer missed detections. The AP value is obtained by calculating the area under the PR curve with Recall as the abscissa and Precision as the ordinate. Finally, mAP is used as the evaluation index of accuracy to measure the comprehensive performance of the trained model on all categories;
[0059] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered within the protection scope of the present invention.
Claims
1. A dense shelf commodity target detection method based on improved YOLOv8, characterized in that: The process includes: S1: Collect shelf product images to build a dense shelf product target detection dataset; S2: Build a dense shelf commodity target detection model based on improved YOLOv8; S3: optimize the loss function; S4: Trained and constructed improved YOLOv8 dense shelf commodity object detection model; S5: Perform performance evaluation on the trained model.
2. According to a method for detecting dense shelf commodity targets based on YOLOv8 according to claim 1, it is characterized in that: In step S1, damaged images in the SKU-110K dataset are screened out, and the remaining images are used to establish a dense shelf commodity object detection dataset, which is divided into a training set, a validation set, and a test set according to a ratio of 7:2:
1.
3. According to a method for detecting dense shelf commodity targets based on YOLOv8 according to claim 1, it is characterized in that: In step S2, the specific steps of constructing a dense shelf commodity target detection model based on improved YOLOv8 include: S201: Introduce the global attention mechanism GAM structure, and add the GAM module after the backbone network of the original YOLOv8 model to increase the global vision and dynamic adjustment ability of the model, and improve the recognition accuracy of the model for small targets; S202: Introduce the multi-attention detection head DyHead detection head to replace the original detection head of YOLOv8.
4. According to a method for detecting dense shelf commodity targets based on YOLOv8 according to claim 1, it is characterized in that: In step S3, the optimization loss function is to replace the original CIoU loss function of YOLOv8 with the WIoU loss function by comprehensively considering the orientation, centroid distance and overlapping area, and introducing a dynamic non-monotonic focusing mechanism. By adjusting the anchor box loss weight, the model's attention to medium-quality samples is enhanced to improve the overall performance.
5. According to a method for detecting dense shelf commodity targets based on YOLOv8 according to claim 1, it is characterized in that: In step S4, the training and construction of the improved YOLOv8 dense shelf commodity target detection model specifically includes the following steps: S401: Set training parameters and use SGD optimizer for training. The image input size is 640×640, the initial learning rate is 0.01, the number of iterations is 200, and the batch size is 16; S402: putting the training set and validation set images in the data set into the improved YOLOv8 dense shelf commodity target detection model for training; S403: Train the model according to the set parameters, and adjust the learning rate and number of iterations of the model training by observing the change trend of the loss function until the loss function changes tend to be stable, thereby obtaining the final training model.
6. A dense shelf commodity target detection method based on YOLOv8 according to claim 1, characterized in that: In step S5, the performance of the training model is evaluated, and the network model is evaluated from aspects such as recall rate, precision rate, average precision, mean average precision and model size.