Real-time Detection Method, Device, Equipment and Medium for Package Position Based on Rotating Object Detection
By designing a real-time package position detection method based on rotation target detection, and using deep learning and attention mechanism rotary target detection network model, the problem of multiple varieties and high-precision measurement of logistics packages is solved, fast, accurate and stable detection is achieved, hardware costs are reduced, and reliable guarantees are provided for logistics sorting.
Patent Information
- Application Number
- CN202210384167.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-04-13
AI Technical Summary
The prior art is difficult to meet the measurement needs of multiple varieties and high-precision logistics packages, and traditional methods require multiple 2D or 3D cameras to be arranged in complex and costly, making it impossible to achieve real-time detection of high resolution, high frame rate and low latency.
A real-time package position detection method based on rotation object detection is designed, a lightweight rotation object detection network model is designed using deep learning, and Yolov5 is used as the infrastructure. The PANet structure is modified as a bidirectional feature pyramid network Bi-FPN, and an attention mechanism is inserted into the feature fusion network, adjust the number of detection head branches and increase the angle classification output and angle loss function.
It realizes high-precision detection without limitations on parcel size, is fast, accurate, stable and maintainable, reduces the cost of hardware equipment, and provides reliable guarantees for the rapid and orderly sorting of logistics packages.
Smart Images

Figure CN114821408B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and particularly to a method, device, computer equipment and storage medium for real-time detection of package positions based on rotated object detection. Background Art
[0002] With the rapid development of the e-commerce and logistics industries, transportation has become increasingly convenient, leading to a sharp increase in the business volume of the express logistics industry. The types and quantities of express packages are constantly increasing, and the demand for real-time, accurate, and fast positioning of package recognition algorithms in the logistics process has become increasingly obvious. Enterprises need to improve the efficiency of package transportation and sorting, which poses great challenges to both the speed and management of logistics distribution.
[0003] Currently, manual sorting with low cost is usually adopted for package sorting in the logistics industry. However, due to the huge number of logistics packages, the workload and labor intensity of workers are large, and the efficiency is low.
[0004] With the wide application of computer vision technology in the single-piece separation of logistics packages and its cooperation with hardware devices such as image acquisition and conveyor belts, the efficiency of package sorting has been greatly improved, gradually replacing the manual sorting method. However, most of the existing technologies are traditional measurement methods, which are difficult to meet the measurement requirements of multiple varieties and high precision. Moreover, multiple 2D cameras or multiple 3D cameras need to be coordinated, with complex layout and high requirements. There is a lack of practical and effective solutions for the real-time and accurate detection of packages. In addition, to obtain high-resolution, high-frame-rate, and low-latency images on a conveyor belt with a transportation speed of 2 m / s, high requirements are imposed on 3D cameras, and the cameras that meet the requirements are expensive, which is not conducive to configuration on large-scale production lines and seriously affects the sorting efficiency. Summary of the Invention
[0005] In order to solve the above-mentioned deficiencies of the prior art, the present invention provides a method, device, computer equipment and storage medium for real-time detection of package positions based on rotated object detection. The method designs a lightweight rotated object detection network model based on deep learning, which is not limited by the size of packages, has high detection accuracy, and has the characteristics of rapidity, accuracy, stability and maintainability. Moreover, it has low requirements for cameras, reduces the cost of hardware devices, and provides a reliable guarantee for subsequent control of the rapid and orderly sorting of logistics packages.
[0006] The first object of the present invention is to provide a method for real-time detection of package positions based on rotated object detection.
[0007] The second object of the present invention is to provide a device for real-time detection of package positions based on rotated object detection.
[0008] The third object of the present invention is to provide a computer equipment.
[0009] The fourth object of the present invention is to provide a storage medium.
[0010] The first object of the present invention can be achieved by adopting the following technical solutions:
[0011] A real-time detection method for the position of a package based on rotated object detection, the method comprising:
[0012] Obtain real-time pictures of logistics packages, and obtain a data set according to the logistics package pictures;
[0013] Design a rotated object detection network model according to the characteristics of the data set, including: the rotated object detection network model uses the object detection network Yolov5 as the basic architecture. In the feature fusion Neck, modify the PANet structure into a feature fusion network Bi-FPN of a bidirectional feature pyramid, and insert an attention mechanism in the feature fusion network Bi-FPN; adjust the number of detection head branches in the detection layer Head and increase the angle classification output and the angle loss function, and predict the rotated rectangular box of the logistics package according to the angle classification output;
[0014] Train the rotated object detection network model using the data set;
[0015] Input the obtained video stream into the trained rotated object detection network model, and output the real-time status information of the package; separate the packages item by item according to the real-time status information of the package.
[0016] Further, the feature fusion network Bi-FPN adds skip connections between the input nodes and output nodes of the same scale, in order to fuse more features at the same layer without increasing additional calculations, and can perform top-down and bottom-up bidirectional feature fusion to achieve multi-scale feature fusion.
[0017] Further, inserting the attention mechanism in the feature fusion network Bi-FPN includes:
[0018] Insert multiple attention mechanisms CBAM in the feature fusion network Bi-FPN;
[0019] The attention mechanism CBAM is inserted between the CSP module and the basic convolution CBL module.
[0020] Further, adjusting the number of detection head branches in the detection layer Head includes:
[0021] Based on the characteristic that small target packages do not appear in the data set, remove the prediction modules in the detection layer Head responsible for predicting smaller sizes, and retain the prediction modules of appropriate sizes according to the size ratio of the real-time package in the picture.
[0022] Furthermore, for the prediction module retained in the detection layer Head, the output dimension of angle classification is increased to classify and output the angle value; where the angle is a set threshold.
[0023] Combine the output angle value with the information represented by the horizontal box to predict the rotated rectangular box of the package.
[0024] Furthermore, the angle loss function is as follows: Use binary cross-entropy and Logits loss function to calculate the loss of the output angle value.
[0025] The designed rotated object detection network model also includes modifying the confidence loss function, specifically:
[0026] Use the IOU of the rotated rectangular box to replace the IOU of the horizontal box as the weight coefficient in the confidence loss function, so that the confidence loss is associated with the output angle value.
[0027] Furthermore, the designed rotated object detection network model also includes improving the non-maximum suppression algorithm NMS, specifically:
[0028] Use the calculation of the IOU of the rotated rectangular box combined with angle information to replace the original IOU calculation based on the horizontal box, and filter out the redundant overlapping rotated prediction boxes.
[0029] Furthermore, use the dataset to train the rotated object detection network model, including:
[0030] Use the K-mean clustering algorithm to obtain the corresponding anchor box Anchor, and then update the anchor box of the feature map.
[0031] According to the characteristic of the small sample size of the dataset, the rotated object detection network model selects to use the Adam optimizer.
[0032] Perform multi-scale training on the rotated object detection network model. By setting different scales, randomly select an input image of one scale for training in each iteration cycle to enhance the robustness of the model, and finally obtain the network weights.
[0033] Furthermore, before using the dataset to train the rotated object detection network model, preprocess the dataset.
[0034] Use the preprocessed dataset to train the rotated object detection network model.
[0035] The preprocessing includes data cleaning and data augmentation, specifically including:
[0036] The data cleaning is to exclude and process the inferior and mislabeled pictures.
[0037] The data augmentation is to adopt a data augmentation method or technique according to the characteristics of the dataset and the number of samples in the dataset, so as to increase the number of samples in the dataset.
[0038] Further, obtaining the dataset according to the logistics package pictures includes:
[0039] Annotating each of the logistics package pictures to obtain the polygon corner point coordinate information of each logistics package;
[0040] According to the polygon corner point coordinate information, obtaining the four corner point coordinates of the minimum bounding rectangle;
[0041] Converting the four corner point coordinates into the long side representation method;
[0042] Taking the long side representation method of each logistics package as a sample, and the long side representation methods of all logistics packages constitute the dataset of samples.
[0043] The second objective of the present invention can be achieved by adopting the following technical solutions:
[0044] A real-time detection device for the position of a package based on rotated object detection, the device includes:
[0045] A dataset acquisition module, configured to acquire real-time logistics package pictures, and obtain a dataset according to the logistics package pictures;
[0046] A rotated object detection network model design module, configured to design a rotated object detection network model according to the characteristics of the dataset, including: the rotated object detection network model uses the object detection network Yolov5 as the basic architecture, in the feature fusion Neck, modifying the PANet structure into a bidirectional feature pyramid feature fusion network Bi-FPN, and inserting an attention mechanism into the feature fusion network Bi-FPN; adjusting the number of detection head branches and adding an angle classification output and an angle loss function in the detection layer Head, and predicting the rotated rectangle frame of the logistics package according to the angle classification output;
[0047] A rotated object detection network model training module, configured to train the rotated object detection network model by using the dataset;
[0048] A package real-time status detection module, configured to input the acquired video stream into the trained rotated object detection network model, and output the real-time status information of the package; and realizing single-piece separation of the packages according to the real-time status information of the packages.
[0049] The third objective of the present invention can be achieved by adopting the following technical solutions:
[0050] A computer device includes a processor and a memory for storing programs executable by the processor. When the processor executes the programs stored in the memory, the above-mentioned real-time detection method for the package position is implemented.
[0051] The fourth object of the present invention can be achieved by adopting the following technical solutions:
[0052] A storage medium stores a program. When the program is executed by a processor, the above-mentioned real-time detection method for the package position is implemented.
[0053] The present invention has the following beneficial effects compared with the prior art:
[0054] Based on the data set obtained from the package pictures, a lightweight rotation target detection network model is designed. This network model is not limited by the size of the package thickness, and has high detection accuracy. Therefore, the method provided by the present invention has rapidity, accuracy, stability and maintainability, and has low requirements for the camera, reducing the cost of hardware devices, and providing a reliable guarantee for the subsequent rapid and orderly sorting of logistics packages. Description of the Drawings
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0056] Figure 1 It is a flowchart of the real-time detection method for the package position based on rotation target detection in Embodiment 1 of the present invention.
[0057] Figure 2 It is a structural diagram of the rotation target detection network model in Embodiment 1 of the present invention.
[0058] Figure 3 It is a structural diagram of the multi-feature fusion Bi-FPN in Embodiment 1 of the present invention.
[0059] Figure 4 It is a structural block diagram of the real-time detection device for the package position based on rotation target detection in Embodiment 2 of the present invention.
[0060] Figure 5 It is a structural block diagram of the computer device in Embodiment 3 of the present invention. Detailed Embodiments
[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. It should be understood that the specific embodiments described are only for explaining the present application and not for limiting the present application.
[0062] Embodiment 1:
[0063] As Figure 1 shown, this embodiment provides a real-time detection method for the position of a package based on rotational object detection, including the following steps:
[0064] S101. Obtain real-time pictures of logistics packages, and obtain a data set based on the pictures of logistics packages.
[0065] Further, step S101 includes:
[0066] (1) Obtain real-time pictures of logistics packages.
[0067] Use an RGB camera to obtain real-time pictures of logistics packages.
[0068] Specifically, use an RGB camera fixed at a position above the conveyor belt to obtain real-time pictures of logistics packages.
[0069] (2) Obtain a data set based on the pictures of logistics packages.
[0070] Annotate each picture of a logistics package through an image annotation software to obtain the polygon corner point coordinate information of each logistics package; according to the polygon corner point coordinate information of each logistics package, obtain the four corner point coordinates of the minimum bounding rectangle of each logistics package; then convert the four corner point coordinates of each logistics package into the long side representation method (x c , y c , l s , s s , θ), use the long side representation method of each logistics package as a sample, and thus obtain a sample in the annotation format of a rectangle with angle direction information, where x c , y c are the center point coordinates of the minimum bounding rectangle, l s is the longest side of the minimum bounding rectangle, s s is the shortest side of the rectangle, and θ is the included angle from the horizontal axis counterclockwise to the long side, with a negative angle. All the samples constitute a data set, that is, the long side representation methods of all logistics packages constitute a data set.
[0071] Specifically, through an image annotation software Labelme with a graphical interface, use Create Polygon, one of the image data annotation methods in various forms, to annotate the dataset; according to the coordinate information of the polygon corner points of the logistics packages obtained from the aforementioned annotation, obtain the coordinate of the four corner points of the minimum bounding rectangle for each package; then convert it into the long side representation method (x c , y c , l s , s s , θ) as a sample, where x c , y c are the coordinates of the center point of the minimum bounding rectangle, l s is the longest side of the minimum bounding rectangle, s s is the other side of the rectangle, and θ is the angle from the horizontal axis counterclockwise to the long side, with a negative angle. In this way, a sample with the annotation format of a rectangle box with angle direction information is obtained, and all the samples constitute a dataset.
[0072] S102. Design a rotated object detection network model according to the characteristics of the dataset.
[0073] (1) Use the object detection network Yolov5 as the basic architecture of the rotated object detection network model and modify it on the basic architecture.
[0074] Further, step (1) specifically includes:
[0075] (1-1) Modify the original PANet structure combining the multi-scale feature fusion FPN and PAN to a bidirectional feature pyramid network Bi-FPN;
[0076] (1-2) Add a lightweight attention mechanism CBAM;
[0077] (1-3) Modify the detection head Head, add an angle θ dimension to the network prediction head, specifically add 180 angle classification channels, and increase the number of network output feature layers, so that the number of parameters responsible for prediction by each anchor box Anchor is 5 + num_classes + angle_classes, where the number 5 represents (x c , y c , longside, shortside, score);
[0078] (1-4) Modify the detection head branches, retain one detection branch, and reduce two detection branches;
[0079] (1-5) Apply the anchor box Anchor obtained from the package dataset through the K-mean clustering algorithm to the feature map;
[0080] (1-6) Improve the non-maximum suppression algorithm NMS. Replace the original IOU calculation based on horizontal boxes with the IOU calculation of rectangular boxes combined with angle information to filter out redundant overlapping rotated prediction boxes;
[0081] (1-7) In the loss function part, add the angle classification loss BCEWithLogitsLoss for the newly added angle classification output dimension, and modify the original horizontal box confidence loss to the rotated box confidence loss.
[0082] Specifically, modifying the multi-scale feature fusion network Bi-FPN includes: modifying the original PANet structure that combines the multi-scale feature fusion networks FPN and PAN into a bidirectional feature pyramid network Bi-FPN. When combined with Yolov5, only select three of the five nodes of Bi-FPN, and connect P5_in, P4_in, and P3_in in the backbone network Backbone to Bi-FPN to obtain P5_out, P4_out, and P3_out; Bi-FPN adds a skip connection between the input node Pn_in and the output node Pn_out at the same scale. The purpose is to enable the connection between the input node and the output node at the same layer to fuse more features without increasing additional computational costs, and repeatedly perform top-down and bottom-up bidirectional feature fusion to complete multi-scale feature fusion, and finally output to the network prediction head. Specifically, it includes: using the features of the 3rd, 4th, and 5th layers extracted by the backbone network as the input features of the 3 input nodes of the feature fusion network Bi-FPN from bottom to top; for the 3rd layer, the intermediate feature of the current layer is obtained by weighted fusion of the input feature of the current layer and the intermediate feature of the 4th layer, and the output feature of the current layer is obtained by weighted fusion of the input feature of the current layer and the intermediate feature of the current layer; for the 4th layer, the intermediate feature of the current layer is obtained by weighted fusion of the input feature of the current layer and the intermediate feature of the 5th layer, and the output feature of the current layer is obtained by weighted fusion of the input feature of the current layer, the intermediate feature of the current layer, and the output feature of the 3rd layer; for the 5th layer, the intermediate feature of the current layer is the input feature of the current layer, and the output feature is obtained by weighted fusion of the input feature of the current layer and the output feature of the 4th layer.
[0083] Specifically, adding the attention mechanism CBAM includes: adding lightweight CBAM modules between each cross-stage local network CSP module and the basic convolution CBL module in the above-mentioned feature fusion network Bi-FPN, a total of three places. CBAM first infers the attention map through the channel dimension and then through the spatial dimension in sequence and independently, and then multiplies the attention map with the input feature map to perform adaptive feature refinement modification, and cooperate with end-to-end training;
[0084] Specifically, adding the angle classification output dimension includes: at the last output part of the prediction head after the above multi-scale feature fusion module, adding the output of 180-dimensional angle information, and combining the angle values represented by the numerical values from 0 to 179 with the horizontal box representation information to predict the rotated rectangular box of the package.
[0085] Specifically, adjusting the number of branches of the detection head includes: according to the characteristics that small target packages generally do not appear in the logistics package dataset, removing the Detect prediction modules for small size and middle size in the original prediction head Head, and retaining the large size prediction module that adapts to the size according to the size ratio of the packages in the real scene pictures, so as to reduce the subsequent unnecessary non-maximum suppression (NMS) calculation and lower the calculation cost of actually predicting the package target.
[0086] Specifically, adding and modifying the loss function includes: adding the angle classification loss, and using the binary cross-entropy and Logits loss functions to calculate the loss of angle classification; using the rotated box IOU instead of the horizontal box IOU as the weight coefficient in the confidence loss function, so that the confidence loss is associated with the angle prediction result.
[0087] (2) Structure of the rotated object detection network model.
[0088] Specifically, as Figure 2 shown, the basic architecture of the rotated object detection network model is divided into three major parts, namely the backbone network Backbone, the feature fusion Neck, and the detection layer Head.
[0089] (2-1) Backbone network Backbone.
[0090] In the backbone network Backbone, CSPDarknet53, which is lightweight and has strong feature extraction ability, is used as the backbone network. Among them, the CSP structure is designed by referring to the design idea of CSPNet to enhance the convolutional learning ability and reduce the calculation cost.
[0091] (2-2) Feature fusion Neck.
[0092] In the feature fusion Neck, the original structure of the multi-scale feature fuser FPN and PAN combined PANet is modified to a bidirectional feature pyramid network Bi-FPN.
[0093] Specifically, as Figure 3As shown in the figure, the three input nodes of the multi-scale feature fusion Bi-FPN are the features P3_in, P4_in, and P5_in extracted from the backbone network at layers 3-5. P3_td is the intermediate feature of the third layer in the top-down path, which is obtained by weighted fusion of P3_in and the intermediate feature P4_td of the fourth layer. P3_out is the output feature of the third layer in the bottom-up path, which is obtained by weighted fusion of P3_in and the intermediate feature P3_td. The rule is that the input features will be repeatedly applied to the two-way feature fusion of top-down and bottom-up, and the inputs of the same scale will be directly connected to the output nodes, increasing feature fusion without increasing the computational cost. These fused features are fed into the classification regression sub-network.
[0094] In the feature fusion Neck, a lightweight attention mechanism CBAM is added. First, it passes through the channel attention module and then through the spatial attention module in sequence to perform Attention on the channel and space respectively. Specifically, the attention mechanism CBAM is inserted at three places in the feature fusion network Bi-FPN, and at each place, it is added between its CSP module and the basic convolution CBL module, as Figure 2 shown.
[0095] (2-3) Detection layer Head.
[0096] As Figure 2 shown, in the detection layer Head, the output of head P5 is retained, and the original two heads responsible for predicting smaller sizes are removed to reduce the computational cost; and in the output dimension of head P5, in order to increase the output of the angle information, 180 dimensions are added to classify and output the angle value, which is combined with the horizontal box representation information to predict the rotated rectangular box of the package; and the corresponding angle classification loss is added to learn the angle, and the subsequent IOU operation will use the rotated IOU for calculation.
[0097] S103. Use the dataset to train the rotated object detection network model.
[0098] Furthermore, step S103 includes:
[0099] (1) Use the dataset to train the rotated object detection network model.
[0100] First, use the K-mean clustering algorithm on the above-mentioned logistics package dataset to obtain the corresponding anchor boxes Anchor and update the anchor boxes applied to the feature map; according to the characteristics of the small sample size of the dataset, choose to use the Adam optimizer, which has advantages in training small datasets, for training; perform multi-scale training. By setting several different scales, randomly select one scale of input images for training every certain number of iteration cycles during training to enhance the robustness of the model; finally, train and obtain the network weights.
[0101] The above model training can be carried out on the high-performance GPUs of the server, with parameter settings as follows: Use K-mean clustering to obtain appropriate anchor boxes Anchor, and the Anchor responsible for head P5 is set to [89, 67]; The optimizer uses Adam, with the initial learning rate set to 0.0035 and momentum to 0.93; Enable multi-scale training; Batch_size is 64 and Epochs is 300.
[0102] (2) Preprocess the dataset, and use the preprocessed dataset to train the rotated object detection network model.
[0103] Preferably, using the preprocessed dataset to train the rotated object detection network model can improve the training efficiency of the network model. The process of using the preprocessed dataset to train the rotated object detection network model is the same as step (1) in S103.
[0104] The preprocessing includes data cleaning and data augmentation.
[0105] Data cleaning is to exclude and process inferior and mislabeled pictures.
[0106] Data augmentation is to adopt corresponding data augmentation methods or data augmentation methods according to the characteristics and quantity of the self-made logistics package dataset to increase the quantity of the dataset. Data augmentation includes dataset augmentation or augmentation methods such as Mosaic augmentation, Cutout mosaic augmentation, Mixup augmentation, HSV color gamut augmentation, and horizontal and vertical flipping augmentation.
[0107] Specifically, for Mosaic augmentation, four sample pictures are arranged in four directions and combined into a large sample picture, and then random rotation, scaling, translation, cropping, perspective and other affine transformations are performed on it, and finally the picture is stretched to the original sample size.
[0108] Specifically, for Cutout mosaic augmentation, randomly replace some areas in the sample picture with 0 pixels.
[0109] Specifically, for Mixup augmentation, randomly mix two sample pictures in a certain proportion; For random angle rotation augmentation, rotate the sample picture in a random angle direction, and stretch the picture size to ensure the integrity of the target information in the picture without loss, and improve the uneven angle situation in the dataset samples.
[0110] Specifically, dataset augmentation or augmentation methods such as HSV color gamut augmentation and horizontal and vertical flipping augmentation.
[0111] S104. Input the video stream obtained by the camera into the trained rotating object detection network model to output the real-time status information of the package. Control the running speed of the package according to the real-time status information of the package, so as to realize the separation of single packages.
[0112] (1) Use the trained rotating object detection network model to obtain the real-time status information of the package.
[0113] Use the weights of the above-trained rotating object detection network. The input is the video stream of the camera on the conveyor belt, and the output is the real-time package quantity information and position status information on the conveyor belt. Visualize the rotating rectangular box and corner point values of the package, calculate and screen out the first package at the forefront in the forward conveying direction of the conveyor belt for subsequent control and sorting.
[0114] (2) Control the running speed of the package according to the real-time status information of the package, so as to realize the separation of single packages.
[0115] The package position coordinates output from the above network model are coordinates in the camera coordinate system and need to be converted into real-world coordinates relative to the conveyor belt. Specifically, the position represented by the center point of the camera image is used as the origin, and the unit size formed is represented by the real-world physical size represented by the image unit pixel, such as 10 mm / pixel. Combine the camera focal length parameter to complete the conversion of camera coordinates to actual space coordinates. Finally, input the converted actual coordinates into the conveyor belt control system to control the belt of the first package at the forefront to run faster, and the belts where the other packages are located to run slower or stop, so as to orderly control the forward movement and separation of the packages.
[0116] Those skilled in the art can understand that all or part of the steps in the method of implementing the above embodiments can be completed by instructing relevant hardware through a program, and the corresponding program can be stored in a computer-readable storage medium.
[0117] It should be noted that although the method operations of the above embodiments are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the shown operations must be performed to achieve the desired result. On the contrary, the described steps can be changed in the execution order. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.
[0118] Embodiment 2:
[0119] As Figure 4 shown, this embodiment provides a device for real-time detection of package position based on rotating object detection. The device includes a dataset acquisition module 401, a rotating object detection network model design module 402, a rotating object detection network model training module 403, and a package real-time status detection module 404, where:
[0120] The dataset acquisition module 401 is used to acquire real-time pictures of logistics packages, and obtain a dataset according to the logistics package pictures;
[0121] The rotated object detection network model design module 402 is used to design a rotated object detection network model according to the characteristics of the dataset, including: the rotated object detection network model uses the object detection network Yolov5 as the basic architecture. In the feature fusion Neck, the PANet structure is modified into a bidirectional feature pyramid feature fusion network Bi-FPN, and an attention mechanism is inserted into the feature fusion network Bi-FPN; in the detection layer Head, the number of detection head branches is adjusted, and an angle classification output and an angle loss function are added, and the rotated rectangular box of the logistics package is predicted according to the angle classification output;
[0122] The rotated object detection network model training module 403 is used to train the rotated object detection network model by using the dataset;
[0123] The package real-time status detection module 404 is used to input the acquired video stream into the trained rotated object detection network model, and output the real-time status information of the package; according to the real-time status information of the package, single-piece separation of the package is realized.
[0124] For the specific implementation of each module in this embodiment, reference can be made to the above-mentioned Embodiment 1, which will not be elaborated here one by one; it should be noted that the device provided in this embodiment is only illustrated by the above-mentioned division of each functional module. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure is divided into different functional modules to complete all or part of the functions described above.
[0125] Embodiment 3:
[0126] This embodiment provides a computer device, which can be a computer, such as Figure 5 as shown, it is connected by a system bus 501 to a processor 502, a memory, an input device 503, a display 504, and a network interface 505. The processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium 506 and an internal memory 507. The non-volatile storage medium 506 stores an operating system, a computer program, and a database. The internal memory 507 provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. When the processor 502 executes the computer program stored in the memory, the real-time package position detection method of the above-mentioned Embodiment 1 is realized as follows:
[0127] Acquire real-time pictures of logistics packages, and obtain a dataset according to the logistics package pictures;
[0128] According to the characteristics of the dataset, a rotating object detection network model is designed, including: the rotating object detection network model uses the object detection network Yolov5 as the basic architecture. In the feature fusion Neck, the PANet structure is modified into a bidirectional feature pyramid feature fusion network Bi-FPN, and an attention mechanism is inserted into the feature fusion network Bi-FPN; in the detection layer Head, the number of detection head branches is adjusted, and the angle classification output and the angle loss function are added, and the rotated rectangular box of the logistics package is predicted according to the angle classification output;
[0129] Use the dataset to train the rotating object detection network model;
[0130] Input the obtained video stream into the trained rotating object detection network model to output the real-time status information of the package; according to the real-time status information of the package, single-piece separation of the package is realized.
[0131] Example 4:
[0132] This embodiment provides a storage medium, which is a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the real-time package position detection method of the above-mentioned Embodiment 1 is realized as follows:
[0133] Obtain real-time logistics package pictures, and obtain a dataset according to the logistics package pictures;
[0134] According to the characteristics of the dataset, a rotating object detection network model is designed, including: the rotating object detection network model uses the object detection network Yolov5 as the basic architecture. In the feature fusion Neck, the PANet structure is modified into a bidirectional feature pyramid feature fusion network Bi-FPN, and an attention mechanism is inserted into the feature fusion network Bi-FPN; in the detection layer Head, the number of detection head branches is adjusted, and the angle classification output and the angle loss function are added, and the rotated rectangular box of the logistics package is predicted according to the angle classification output;
[0135] Use the dataset to train the rotating object detection network model;
[0136] Input the obtained video stream into the trained rotating object detection network model to output the real-time status information of the package; according to the real-time status information of the package, single-piece separation of the package is realized.
[0137] It should be noted that the computer-readable storage medium in this embodiment can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0138] In summary, the detection method provided by the present invention obtains a dataset by acquiring real-time pictures of logistics parcels. According to the characteristics of the dataset, a lightweight rotated object detection network model is designed. This model uses the object detection network Yolov5 as the basic architecture, modifies the original PANet structure that combines the multi-scale feature fusion network FPN and PAN into a bidirectional feature pyramid network Bi-FPN, and adds a lightweight attention mechanism CBAM to the network Bi-FPN. In the detection head Head part, an angle θ dimension is added to the network prediction head, 180 angle classification channels are added, and the number of network output feature layers is increased, so that the number of parameters responsible for prediction by each anchor box Anchor is 5 + num_classes + angle_classes, where the number 5 represents (xc, yc, longside, shortside, score); and the detection head branches of the detection head Head part are modified: the detection branches are reduced, and the detection branches that adapt to the size are retained; the anchor boxes Anchor from the parcel dataset are obtained through the K-mean clustering algorithm and applied to the feature map; the non-maximum suppression algorithm NMS is improved, and the calculation of the original horizontal box IOU is replaced with the rotated box IOU combined with the angle information to filter out redundant and repeated rotated prediction boxes; in the loss function part, an angle classification loss BCEWithLogitsLoss is added for the newly added angle classification output dimension, and the original horizontal box confidence loss is modified to a rotated box confidence loss. Since the model designed by the present invention is not limited by the thickness and size of the parcel, has a fast calculation speed, high detection accuracy, and has low requirements for the camera, it provides a reliable guarantee for the subsequent rapid and orderly sorting of logistics parcels.
[0139] The above is only a preferred embodiment of the present invention for patents, but the protection scope of the present invention for patents is not limited thereto. Any person skilled in the art within the scope disclosed by the present invention for patents, according to the technical solution and inventive concept of the present invention for patents, makes equivalent substitutions or changes, all belong to the protection scope of the present invention for patents.
Claims
1. A real-time detection method for the position of packages based on rotated object detection, characterized in that, the method includes: Obtain real-time logistics package pictures, and obtain a data set according to the logistics package pictures; According to the characteristics of the data set, design a rotated object detection network model, including: the rotated object detection network model uses the object detection network Yolov5 as the basic architecture, modifies the PANet structure to a bidirectional feature pyramid feature fusion network Bi-FPN in the feature fusion Neck, uses the features of the 3rd to 5th layers extracted by the backbone network as the input of the feature fusion network Bi-FPN, and uses the output of the feature fusion network Bi-FPN for a detection head; in the feature fusion network Bi-FPN, insert an attention mechanism CBAM between each CSP module and the basic convolution CBL module; remove the prediction module responsible for predicting smaller sizes in the detection layer Head, retain the prediction module of the appropriate size according to the size ratio of the real-time package in the picture, and add an angle classification output and an angle loss function in the detection layer Head; predict the rotated rectangular box of the logistics package according to the angle classification output; Use the data set to train the rotated object detection network model; Input the obtained video stream into the trained rotated object detection network model, and output the real-time status information of the package; according to the real-time status information of the package, realize the separation of single packages.
2. The real-time detection method for the position of packages according to claim 1, characterized in that, The feature fusion network Bi-FPN adds skip connections between the input nodes and output nodes of the same scale, in order to fuse more features at the same layer without increasing additional calculations, and can perform top-down and bottom-up bidirectional feature fusion to achieve multi-scale feature fusion.
3. The real-time detection method for the position of packages according to claim 1, characterized in that, For the prediction module retained in the detection layer Head, increase the angle classification output dimension to classify and output the angle value; wherein, the angle is a set threshold; Combine the output angle value with the information represented by the horizontal box to predict the rotated rectangular box of the package.
4. The real-time detection method for the position of packages according to claim 3, characterized in that, The angle loss function is: use binary cross entropy and Logits loss function to calculate the loss of the output angle value; The design of the rotated object detection network model also includes modifying the confidence loss function, specifically: Use the rotated rectangle box IOU instead of the horizontal box IOU as the weight coefficient in the confidence loss function, so that the confidence loss is associated with the output angle value.
5. The real-time detection method for the position of packages according to claim 4, characterized in that, The design of the rotated object detection network model also includes improving the non-maximum suppression algorithm NMS, specifically: Use the rotated rectangle box IOU calculation combined with angle information to replace the original IOU calculation based on the horizontal box, and filter out redundant overlapping rotated prediction boxes.
6. The real-time detection method for the position of packages according to claim 1, characterized in that, Training the rotation target detection network model using the said data set, including: Using the K-mean clustering algorithm to obtain the corresponding anchor boxes Anchor, and then updating the anchor boxes of the feature map; According to the characteristic of small sample size of the said data set, the rotation target detection network model selects to use the Adam optimizer; Performing multi-scale training on the rotation target detection network model. By setting different scales, randomly select an input image of one scale for training in each iteration cycle during training to enhance the robustness of the model, and finally obtain the network weights.
7. The real-time detection method of package position according to claim 1, characterized in that, Before training the rotation target detection network model using the said data set, preprocess the said data set; Train the rotation target detection network model using the preprocessed data set; The said preprocessing includes data cleaning and data augmentation, specifically including: The data cleaning is to exclude and process inferior and mislabeled pictures; The data augmentation is to adopt data augmentation methods or data augmentation methods according to the characteristics of the said data set and the number of samples in the data set to increase the number of samples in the data set.
8. The real-time detection method of package position according to any one of claims 1-7, characterized in that, The obtaining of the data set according to the said logistics package picture includes: Annotating each said logistics package picture to obtain the polygon corner point coordinate information of each logistics package; According to the said polygon corner point coordinate information, obtaining the four corner point coordinates of the minimum bounding rectangle; Converting the said four corner point coordinates into the long side representation method; Taking the long side representation method of each logistics package as a sample, and the long side representation methods of all logistics packages constitute the data set of samples.
Citation Information
Patent Citations
Rotary frame positioning multi-form bottle-shaped article sorting target detection method
CN114266884A