Logistics cargo identification method and system based on dynamic segmentation
By optimizing the YOLOv8 model and combining the interactive segmentation network and attention mechanism, the problem of insufficient feature extraction of YOLOv8 in complex logistics scenarios is solved, efficient and accurate cargo recognition is achieved, and the level of intelligence of the logistics system is improved.
Patent Information
- Application Number
- CN202510484780.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-18
AI Technical Summary
The existing YOLOv8 algorithm has insufficient fusion of global and local information in the feature extraction process in complex logistics scenarios, resulting in limited attention to key goals and it is difficult to adapt to the needs of refined identification in changing environments.
By introducing interactive segmentation network and attention mechanism modules, the YOLOv8 model is optimized, the model's attention to key areas is enhanced, and the deep separable convolution and feature fusion network are combined to build an improved cargo recognition model.
Significantly improve the recognition accuracy in complex lighting and stacked occlusion scenarios, maintain efficient detection speed, adapt to changing environments, and provide efficient logistics and cargo identification support.
Smart Images

Figure CN120339591A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent identification of cargo packages, and in particular to a logistics cargo identification method and system based on dynamic segmentation. Background Art
[0002] With the booming development of e-commerce and the continuous growth of global trade, modern warehousing and logistics systems are facing unprecedented challenges. Especially in the process of goods out of the warehouse, how to efficiently and accurately identify and classify a large number of commodities has become a key factor in ensuring the smooth operation of the supply chain and improving logistics efficiency. When faced with massive amounts of goods, traditional manual identification methods can neither meet the needs of efficient processing nor be easily interfered by human factors, resulting in misidentification, increased logistics costs, affected customer satisfaction, and may even damage the company's reputation. Therefore, how to use advanced technology to improve the accuracy and efficiency of goods out of the warehouse has become a core issue that needs to be solved in the logistics industry.
[0003] In recent years, artificial intelligence technology, especially deep learning methods, has made significant progress in the field of image recognition and has become a key means to solve logistics identification problems. As the most popular real-time target detection algorithm, the YOLO (You Only Look Once) series has excellent detection accuracy and speed, which makes it show wide application potential in the logistics industry. However, as the latest generation of algorithms in the series, YOLOv8 has improved in speed and accuracy, but in complex logistics scenarios (such as cargo stacking, uneven lighting, and partial occlusion), its feature extraction process still has the problem of insufficient fusion of global and local information, resulting in limited attention to key targets. Traditional models lack a mechanism for dynamically adjusting feature weights, making it difficult to adapt to the needs of refined recognition in changing environments. Summary of the invention
[0004] In order to solve the above problems, the purpose of the present invention is to provide a logistics cargo identification technology based on dynamic segmentation, aiming to provide valuable technical support for the automation and intelligence of the logistics industry.
[0005] In order to achieve the above technical objectives, the present application provides a logistics cargo identification method based on dynamic segmentation, comprising the following steps: Based on the YOLOv8 model, an interactive segmentation network is introduced to segment a single target to be classified from multi-category images, and the model is improved to build an improved YOLOv8 model. In order to optimize the feature selection capability, the present invention further introduces an attention mechanism module into the YOLOv8 model, and enhances the model's attention to key areas by dynamically adjusting the feature map channels or spatial weights. Optional implementations include: Squeeze-and-Excitation (SE) block: It learns the dependencies between channels through compression and excitation operations, enhancing the weights of important feature channels. Channel-Spatial Attention Module (such as CBAM): It filters important feature channels through channel attention and combines spatial attention to focus on the target area. Efficient Channel Attention (ECA) block: It optimizes the channel feature response at low computational cost. The above-mentioned modules can be inserted into the backbone network, feature fusion network or detection head of the model. The specific position needs to be determined through ablation experiments.
[0006] Select a logistics parcel dataset to train the improved YOLOv8 model and build a cargo recognition model for identifying the parcel type and packaging condition of parcel images.
[0007] Preferably, when obtaining the YOLOv8 model, the YOLOv8 model includes: Backbone network: It is used to extract image features and adopts depthwise separable convolution and feature extraction methods. Neck module: It performs further fusion on the feature map, conducts feature fusion through multiple feature layers of different scales, and identifies targets of different scales. Head module: The detection part, which performs object detection by predicting the class, coordinates and confidence of the bounding box, and realizes accurate prediction of the object bounding box through optimization with a multi-task loss function.
[0008] Preferably, when obtaining the YOLOv8 model, the backbone network is constructed through convolution, C2f module, SPPF module and attention module, where: The C2f module fuses the features of the previous layer and the current layer through residual connection. The SPPF module first divides the input feature map into multiple scales through pooling windows of different sizes, performs pooling operations on each scale to generate multiple feature maps of different sizes; then, these pooled feature maps are concatenated to form a feature vector with rich multi-scale information. The attention module is used to perform global average pooling on the input feature map to generate a channel descriptor; learns channel weights through a fully connected layer and an activation function; multiplies the channel weights with the original feature map channel by channel to complete feature weighting.
[0009] Preferably, when obtaining the YOLOv8 model, the Neck module is used to strengthen the feature map output by the Backbone, and the object detection ability is improved through the PANet and FPN strategies.
[0010] Preferably, when obtaining the YOLOv8 model, the Head module is used to perform final target classification, bounding box regression and target confidence prediction on the feature map processed by Backbone and Neck, wherein the Head module adopts Anchor-Free design.
[0011] Preferably, when introducing the interactive segmentation network, the interactive segmentation network is utilized, ResNet101 is used as the feature extraction network, and a dilated convolution pooling pyramid module is subsequently connected to capture the multi-scale context of the object through dilated convolutions with different expansion rates to fuse information of different scales, and finally the segmentation result is generated through the decoder module.
[0012] Preferably, when introducing the interactive segmentation network, an initial interaction point attention FCA module is introduced, and features are extracted through 6 layers of 3×3 convolution to enhance the weight of the initial interaction point; the positive and negative Gaussian point maps are spliced with the original image in the channel direction and input into the ResNet101 network, the Gaussian point map generated by the initial interaction point is spliced with the first layer output of the ResNet101 network as the input of the FCA module, and the output of the FCA module is spliced with the output of the ResNet101 network as the input of the subsequent ASPP and decoder modules.
[0013] The present invention discloses a logistics cargo identification system based on dynamic segmentation, which is used to implement the above-mentioned logistics cargo identification method based on dynamic segmentation. The system comprises: The model building module is used to segment a single target to be classified from a multi-category image based on the YOLOv8 model by introducing an interactive segmentation network, improve the model, and build an improved YOLOv8 model; The classification and recognition module is used to select logistics package data sets, train the improved YOLOv8 model, and build a cargo recognition model to identify the package type and packaging condition of the package image.
[0014] The present invention discloses the following technical effects: The recognition accuracy of the present invention on the cargo category data set is significantly improved compared with the traditional method, and it maintains excellent performance in actual scenes such as complex lighting and stacking occlusion. Its detection speed reaches the real-time processing standard.
[0015] The present invention improves YOLOv8 and combines the comprehensive advantages of interactive segmentation technology in terms of accuracy, speed and robustness, providing key technical support for the intelligent upgrade of warehousing logistics that integrates efficient recognition, environmental adaptability and deployment feasibility. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 is the image classification network based on interactive image segmentation according to the present invention. Among them, the left side of the network is the interactive segmentation network, and the right side is the ViT network. ⊕ represents splicing in the channel direction; Figure 2 is the structural schematic diagram of the YOLOv8 model according to the present invention; Figure 3 is the schematic diagram of the method flow according to the present invention. Detailed implementation manners
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Usually, the components of the embodiments of the present application described and illustrated in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application to be protected, but only represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0019] As Figures 1-3 shown, the present invention provides a logistics goods recognition technology based on dynamic segmentation. By optimizing the YOLOv8 model structure and combining the interactive segmentation technology, the recognition accuracy and processing speed in the goods outbound link of warehousing logistics are improved, thereby promoting the intelligent upgrade of the logistics industry. The specific contents are as follows: The logistics recognition algorithm proposed by the present invention combines an interactive segmentation network and an attention classification network based on deep learning. Its overall framework is as Figure 1 shown. The interactive segmentation network is responsible for segmenting a single target to be classified from a multi-class image, and the classification network is based on YOLOv8 to classify the segmented item image.
[0020] The interactive segmentation network adopts an interactive image segmentation network based on initial interactive point attention. Based on the "point" interaction method, the user clicks on the image to output positive and negative interactive points as interactive information, and then generates corresponding positive and negative Gaussian point maps. The network uses ResNet101 as the feature extraction network and connects the Atrous Spatial Pyramid Pooling (ASPP) module behind it. Through atrous convolutions with different dilation rates, it captures the multi-scale context of the object to fuse information at different scales. Finally, the segmentation result is generated through the decoder module. In addition, the network also introduces the Feature Cross Attention (FCA) module, which extracts features through 6 layers of 3×3 convolutions to enhance the weight of the initial interactive points. The positive and negative Gaussian point maps are concatenated with the original image in the channel direction and then input into the ResNet101 network. The Gaussian point map generated by the initial interactive points is concatenated with the output of the first layer of the ResNet101 network as the input of the FCA module. The output of the FCA module is then concatenated with the output of the ResNet101 network and used as the input of the subsequent ASPP and decoder modules. In this way, the interactive segmentation network can accurately segment the target in the image and mark the segmentation result in the original image in the form of a mask. The segmented object region is used as the input of the subsequent classification network, thus transforming the classification problem of a single-category object in a multi-category image into a single-category image classification problem, which not only simplifies the classification task but also effectively removes the background noise in the image.
[0021] The YOLOv8 model structure and features proposed in this invention: YOLOv8 is the latest version of the YOLO (You Only Look Once) object detection model. It inherits the characteristics of the YOLO series and is further optimized to improve performance and efficiency. Compared with other algorithms in the YOLO series, the YOLOv8 model architecture is similar to YOLOv5, but there are some innovations and optimizations in details. The structure diagram of YOLOv8 is as Figure 2 shown and the model is usually divided into 2 - 3 components: backbone, neck, and head.
[0022] Backbone: Used to extract image features, usually composed of convolutional layers and other layers (such as CSPNet, PANet, etc.). YOLOv8 adopts more depthwise separable convolutions and efficient feature extraction methods here to improve the inference speed and accuracy.
[0023] Neck: Further fuses on the feature map, usually through multiple feature layers at different scales (such as FPN or PANet) for feature fusion to help the model better identify objects at different scales.
[0024] Head: The final detection part, which performs object detection based on the class, coordinates, and confidence of the prediction boxes. The head part of YOLOv8 is optimized by a multi-task loss function to achieve accurate prediction of object bounding boxes.
[0025] The backbone network of the YOLOv8 algorithm is mainly composed of convolutional layers, C2f modules, and SPPF modules. Among them, the convolutional layer is the most basic component in the neural network, responsible for extracting spatial features from the image. Compared with YOLOv5, YOLOv8 reduces the computational cost by using deeper convolutional modules or other techniques such as depthwise separable convolutions, thereby improving the efficiency of feature extraction. The C2f module usually fuses the features of the previous layer with the features of the current layer through residual connections. After fusion, the information of the feature map is more fully utilized, and not much important spatial information is lost during the transmission process. The SPPF (Spatial Pyramid Pooling) module is one of the important innovations in YOLOv8. First, the input feature map is divided into multiple scales through pooling windows of different sizes (such as 1x1, 5x5, 9x9, etc.), and pooling operations are performed on each scale to generate multiple feature maps of different sizes. Then, these pooled feature maps are concatenated to form a feature vector with rich multi-scale information.
[0026] The core goal of the neck module of YOLOv8 is to strengthen the feature maps output by the Backbone and fuse feature maps of different scales to provide richer and more accurate features for the subsequent Head module. Among them, strategies such as PANet and FPN are mainly used to improve the object detection ability. Both PANet (Path Aggregation Network) and FPN (Feature Pyramid Network) are network structures for multi-scale feature fusion. FPN combines high-level semantic features with low-level detailed features through top-down and lateral connections, thereby enhancing the representation ability of multi-scale features; while PANet further introduces bottom-up information flow on the basis of FPN, enhancing the upward transmission of low-level features and supplementing the position information of high-level features to improve the detection ability of small objects. The combination of the two can more effectively fuse features of different scales and improve the accuracy and robustness of object detection.
[0027] such as Figure 2As shown, the Head module in YOLOv8 is mainly responsible for performing final object classification, bounding box regression, and object confidence prediction on the feature maps processed by the Backbone and Neck. It decodes features at different scales and outputs detection results. Compared with the traditional Anchor-based method of YOLOv5, YOLOv8 adopts an Anchor-Free design, reducing the dependence on predefined anchor boxes. It can directly perform bounding box regression, avoiding the computational burden in the processes of anchor box selection, matching, and generation, thus improving the model's performance in complex scenarios. The Anchor-Free design enables YOLOv8 to achieve better detection results for objects of different sizes and shapes, especially outperforming YOLOv5 significantly in small object detection.
[0028] Example: A logistics goods recognition technology based on dynamic segmentation proposed by the present invention specifically includes the following processes: 1. Selection of the dataset: The COCO2017 dataset is a widely used dataset for image recognition, segmentation, and object detection, developed and released by the Microsoft team. This dataset aims to promote inventions in scene understanding, especially in applications related to object detection, segmentation, and semantic understanding. COCO2017 contains more than 200,000 labeled images, distributed across 80 categories, covering a variety of objects from humans and animals to everyday items. These images are sourced from complex daily scenes, where the objects appear in rich contexts, including not only instances of individual objects but also interaction information between objects.
[0029] The COCO2017 dataset is divided into a training set, a validation set, and a test set. Specifically, it contains 118,287 training images, 5,000 validation images, and a test set with unpublicized annotations for evaluating model performance. This dataset not only provides the bounding box annotations required for object detection but also includes pixel-level annotations for instance segmentation, which support the identification of the precise positions and shapes of each object in the image. Due to the diversity, complexity, and wide range of application scenarios of its data, COCO2017 has become one of the extremely important benchmarks in the field of computer vision.
[0030] The cargo logistics package dataset used in this invention is sourced from Shunhe Group and aims to support inventions related to package identification and classification in the logistics field. This dataset focuses on the image recognition problem of logistics packages and includes various types of package images, aiming to simulate the scenarios of handling packages in a real-world logistics center. The dataset includes a total of 355 samples, among which there are 243 training set samples, 57 validation set samples, and 55 test set samples. These images cover a variety of package types and packaging situations, with the aim of improving the robustness and accuracy of the model in practical applications.
[0031] 2. Transfer learning experiment process: To verify the effectiveness of the YOLO model in the target detection task of goods to be loaded, the experiment adopts a transfer learning strategy, using a pre-trained model trained on a large-scale dataset to improve the detection performance on a specific logistics package dataset.
[0032] First, the YOLO model was trained for 100 rounds on the MSCOCO2017 object detection dataset to obtain a general pre-trained model. The aim is to enable the model to learn rich feature representations and lay a foundation for subsequent transfer learning.
[0033] Then, the above pre-trained model was applied to the logistics package dataset for transfer learning to adapt to the specific features and detection tasks of the logistics package dataset.
[0034] Finally, to highlight the performance advantages of the YOLO model in the target goods detection task, the experiment compared it with the original YOLOV7 model and two YOLOV8 series models with similar numbers of parameters. All models were trained for the same number of rounds on the logistics package dataset and performance was evaluated using the same evaluation criteria.
[0035] The metrics include Precision, Recall, mAP.5 (mean Average Precision at IoU = 0.5), and mAP.75 (mean Average Precision at IoU = 0.75) to evaluate the progress and effectiveness of model learning.
[0036] This invention enhances the refined processing ability of targets in complex scenarios by combining manual and automated segmentation processes. Users can make interactive corrections when needed to accurately calibrate the object boundaries. Especially in cases where goods are stacked or severely occluded, it can effectively improve the segmentation accuracy, thereby further optimizing the recognition performance of the YOLOv8 model.
[0037] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in multiple blocks.
[0038] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined.
[0039] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A logistics goods recognition method based on dynamic segmentation, characterized in that, It includes the following steps: Based on the YOLOv8 model, by introducing an interactive segmentation network, a single target to be classified is segmented from multi-class images, the model is improved, and an improved YOLOv8 model is constructed; A logistics parcel dataset is selected to train the improved YOLOv8 model, and a cargo recognition model is constructed to identify the parcel type and packaging condition of the parcel image.
2. The method for logistics cargo recognition based on dynamic segmentation according to claim 1, wherein: When obtaining the YOLOv8 model, the YOLOv8 model includes: Backbone network: used to extract image features, adopting depthwise separable convolution and feature extraction methods; Neck module: further fusion is performed on the feature map, and feature fusion is carried out through multiple feature layers of different scales to identify targets of different scales; Head module: the detection part, performs target detection by predicting the category, coordinates, and confidence of the bounding box, and through the optimization of the multi-task loss function, accurate prediction of the target bounding box is achieved.
3. The method for logistics cargo recognition based on dynamic segmentation according to claim 2, wherein: When obtaining the YOLOv8 model, the Backbone network is constructed through convolution, C2f module, SPPF module, and attention module, where: The C2f module fuses the features of the previous layer and the current layer through a residual connection; For the SPPF module, first, the input feature map is divided into multiple scales through pooling windows of different sizes, pooling operations are performed on each scale to generate multiple feature maps of different sizes; then, these pooled feature maps are concatenated to form a feature vector with rich multi-scale information; The attention module is used to perform global average pooling on the input feature map to generate a channel descriptor; learn channel weights through a fully connected layer and an activation function; multiply the channel weights with the original feature map channel by channel to complete feature weighting.
4. The method for logistics cargo recognition based on dynamic segmentation according to claim 3, wherein: When obtaining the YOLOv8 model, the Neck module is used to strengthen the feature map output by the Backbone, and the target detection ability is improved through the PANet and FPN strategies.
5. The method for logistics cargo recognition based on dynamic segmentation according to claim 4, wherein: When obtaining the YOLOv8 model, the Head module is used to perform final target classification, bounding box regression, and target confidence prediction on the feature map processed by the Backbone and Neck, where the Head module adopts an Anchor-Free design.
6. The method for logistics cargo recognition based on dynamic segmentation according to claim 5, wherein: When introducing the interactive segmentation network, the interactive segmentation network is utilized, with ResNet101 as the feature extraction network, and the Atrous Spatial Pyramid Pooling (ASPP) module is connected behind it. Through atrous convolutions with different dilation rates, multi-scale context of the object is captured to fuse information at different scales. Finally, the segmentation result is generated through the decoder module.
7. The method for identifying logistics goods based on dynamic segmentation according to claim 6, characterized in that: When introducing the interactive segmentation network, the initial interactive point attention Feature Calibration Module (FCA) is introduced, and features are extracted through 6 layers of 3×3 convolutions to enhance the weight of the initial interactive point. The positive and negative Gaussian point maps are concatenated with the original image in the channel direction and then input into the ResNet101 network. The Gaussian point map generated by the initial interactive point is concatenated with the output of the first layer of the ResNet101 network as the input of the FCA module. The output of the FCA module is then concatenated with the output of the ResNet101 network and used as the input of the subsequent ASPP and decoder modules.
8. A logistics goods identification system based on dynamic segmentation, which is used to implement a logistics goods identification method based on dynamic segmentation as described in any one of claims 1-7, characterized in that, The system includes: A model construction module, which is used to improve the model by introducing an interactive segmentation network to segment a single target to be classified from multi-class images based on the YOLOv8 model, and construct an improved YOLOv8 model. A classification and recognition module, which is used to select a logistics parcel dataset, train the improved YOLOv8 model, and construct a goods recognition model for identifying the parcel type and packaging condition of the parcel image.
Citation Information
Cited By
Glacier area surface water resource distribution image extraction method based on deep learning
CN121033692A