Rapid positioning method for single-piece separated package based on object detection
By using RGBD cameras and YOLO object detection models in an automated sorting system, the fast and accurate positioning of the package is achieved, and the problems of errors and slow recognition of traditional two-dimensional cameras are solved, and the sorting efficiency and accuracy are improved.
Patent Information
- Application Number
- PCT/CN2024/099704
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-06-18
- Publication Date
- 2025-06-05
AI Technical Summary
The existing automated sorting system uses traditional two-dimensional cameras, which is prone to problems such as identification errors and misjudgment, which reduces the accuracy and reliability of the system, and the recognition speed is not ideal, affecting the sorting efficiency.
A single-piece separate package fast positioning method based on object detection is adopted, and the RGBD camera is used to collect the packaged RGB images and depth data, and the target detection model is combined with the YOLO-based lightweight object detection model to perform object detection, extract the object detection frame, and determine the position information of the package through depth data.
It realizes fast and accurate positioning of packages, improves detection efficiency and speed, reduces the possibility of misjudgment and missed inspections, and is suitable for packages of various types and shapes, suitable for applications in complex environments.
Smart Images

Figure CN2024099704_05062025_PF_FP_ABST
Abstract
Description
A fast positioning method for single separated packages based on target detection Technical Field
[0001] The present application relates to the field of logistics technology, and in particular to a method for quickly locating a single separated package based on target detection. Background Art
[0002] With the rapid development of the logistics industry, the number and variety of express parcels are increasing, posing significant challenges to parcel sorting and delivery. Traditional manual sorting methods are inefficient and costly, unable to meet the demands of the modern logistics market. Consequently, automated sorting and delivery systems are becoming a growing industry trend. Existing automated sorting systems typically use traditional two-dimensional cameras to track and locate moving objects. However, this approach is prone to recognition errors and misjudgments, reducing the accuracy and reliability of the automated sorting system. Furthermore, the recognition speed of current methods is suboptimal, directly impacting sorting efficiency. Technical issues
[0003] Existing automated sorting systems basically use traditional two-dimensional cameras to track and locate moving objects, which are prone to problems such as recognition errors and misjudgments, reducing the accuracy and reliability of the automated sorting system. In addition, the recognition speed of the current method is not ideal, which also directly affects the sorting efficiency. Technical Solutions
[0004] In response to the above-mentioned problems and technical requirements, this application proposes a method for rapid location of single separated packages based on target detection. The technical solution of this application is as follows:
[0005] A method for quickly locating a single separated package based on target detection, the method comprising:
[0006] An RGBD camera located above the single-piece separation area collects RGB images and depth data of the single-piece separation area to be detected. The single-piece separation area includes a separation belt and a plurality of packages placed on the separation belt.
[0007] Use a lightweight object detection model based on YOLO to detect the packages in the RGB image to be detected, and extract the object detection frame of each package in the RGB image to be detected;
[0008] Determine a reference depth value of the target detection frame according to the depth data to be detected in each target detection frame, and obtain the projection coordinates of the four vertices of the target detection frame on the separation belt according to the reference depth value of each target detection frame;
[0009] The location information of the corresponding package is determined based on the projection coordinates of the four vertices of each target detection frame on the separation belt.
[0010] A further technical solution is to obtain the projection coordinates of the four vertices of the target detection frame on the separation belt according to the reference depth value of each target detection frame, including:
[0011] The camera parameters of the RGBD camera are converted to obtain the two-dimensional coordinates of the four vertices of the target detection box in the camera coordinate system;
[0012] The reference depth value of the target detection frame is used as the third-dimensional coordinate of the four vertices of the target detection frame, and the three-dimensional coordinates of the corresponding vertex in the camera coordinate system are obtained by combining the two-dimensional coordinates of each vertex;
[0013] According to the three-dimensional coordinates of each vertex of the target detection frame, the projection coordinates of the vertex on the separation belt are obtained.
[0014] A further technical solution is to obtain the projection coordinates of each vertex on the separation belt according to the three-dimensional coordinates of each vertex of the target detection frame, including:
[0015] Based on the belt plane equation parameters of the separation belt, the three-dimensional coordinates of each vertex of the target detection frame are projected onto the separation belt to obtain the corresponding projection coordinates.
[0016] A further technical solution is that the method for quickly locating a single separated package includes:
[0017] Use the RGBD camera to collect calibrated RGB images and calibrated depth data of the single-piece separation area when no packages are set;
[0018] Based on the calibrated RGB image, the area where the separation belt is located in the single-piece separation area is determined as the area of interest;
[0019] The calibrated depth data in the region of interest is converted into 3D point cloud data of the belt according to the intrinsic parameter matrix of the RGBD camera;
[0020] The belt plane equation parameters are obtained by performing point cloud plane fitting on the three-dimensional point cloud data of the belt based on the random sampling consistency algorithm.
[0021] A further technical solution is that determining a reference depth value of a target detection frame according to the depth data to be detected in each target detection frame includes:
[0022] The median of all depth data to be detected in each target detection frame is used as the reference depth value of the target detection frame.
[0023] A further technical solution is to determine the location information of the corresponding package based on the projection coordinates of the four vertices of each target detection frame on the separation belt, including:
[0024] The maximum bounding rectangle of the two-dimensional coordinates of the projection coordinates of the four vertices of each target detection frame on the horizontal plane is fitted to obtain the maximum bounding rectangle of the landing area of the package corresponding to the target detection frame. The location information of the package is determined based on the maximum bounding rectangle of the landing area of the package.
[0025] Its further technical solution is that the backbone network of the lightweight target detection model uses CSPDarkNet53 for feature extraction. The target detection head includes three output layers of different scales to detect candidate boxes of different sizes. The target detection head gives prediction parameters for each candidate box. The prediction parameters include the center coordinates of the candidate box, the width and height of the candidate box, and the confidence level, where the confidence level represents the maximum value of the confidence level of the package and the background confidence level.
[0026] Its further technical solution is that the lightweight target detection model uses the Hard Swish activation function and also uses depth-wise separable convolution.
[0027] A further technical solution is that the method for quickly locating a single separated package includes:
[0028] Build the network structure of a lightweight object detection model based on YOLO;
[0029] Use multiple GPUs in parallel to use the COCO dataset to pre-train the model parameters based on the established network structure;
[0030] Get a sample RGB image of the package and annotate it;
[0031] The labeled sample RGB images are used to obtain model parameters based on pre-training for training, and the Greedy NMS algorithm is used for non-maximum suppression and the back propagation algorithm is used to update the model parameters.
[0032] A further technical solution is to obtain sample RGB images of packages of different sizes, packaging material types, and packaging colors respectively when obtaining the sample RGB images. Beneficial effects
[0033] This application discloses a method for rapid positioning of single-piece separated packages based on target detection. This method uses an RGBD camera to capture the RGB image and depth data of the package in real time, and uses a lightweight target detection model based on YOLO for target detection. The lightweight target detection model is suitable for industrial application scenarios with limited resources such as single-piece package separation, and can improve detection efficiency and speed. Then, the four vertices of the target detection frame are projected in combination with the depth data to achieve rapid positioning of the package. This method can quickly and in real time locate the position information of the package in the separation area, has low requirements for the quality of the depth data, low computational complexity, and high operating efficiency. It can improve the efficiency of logistics sorting and processing, and meet the needs of practical applications.
[0034] This method, based on an object detection algorithm, is applicable to packages of various types and shapes, demonstrating excellent detection results in complex environments. It also mitigates interference from side-by-side packages when locating their landing areas, effectively avoiding misjudgments and missed detections. Furthermore, the lightweight object detection model employed can be optimized and expanded as needed by incorporating additional features and algorithms to improve detection accuracy and robustness, demonstrating excellent scalability. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] FIG1 is a flow chart of a method for quickly locating a single separated package according to an embodiment of the present application.
[0036] FIG2 is a schematic diagram showing the three-dimensional coordinates of the four vertices obtained by transforming the target detection frame in an example, the corresponding projection coordinates, and the maximum circumscribed rectangle of the landing area of the package obtained by fitting. Modes for Carrying Out the Invention
[0037] The specific implementation of this application will be further described below with reference to the accompanying drawings.
[0038] This application discloses a method for quickly locating a single separated package based on target detection. Please refer to the flowchart shown in Figure 1. The method includes the following steps:
[0039] Step 1: Use an RGBD camera located above the single-piece separation area to collect the RGB image and depth data to be detected of the single-piece separation area.
[0040] The single piece separation area includes a separation belt and a plurality of packages arranged on the separation belt. In actual application, the separation belt includes a plurality of belt blocks arranged in an array, which are not further subdivided in this application and are collectively referred to as the separation belt.
[0041] The RGBD camera is generally installed at a preset height above the single-piece separation area and facing the single-piece separation area. The field of view of the RGBD camera at least covers the area where the separation belt is located.
[0042] Step 2: Use a lightweight object detection model based on YOLO to detect the packages in the RGB image to be detected, and extract the object detection frame of each package in the RGB image to be detected.
[0043] The YOLO-based lightweight target detection model used in this application is optimized and improved based on the YOLO series. It is lightweight and can run in real time on embedded devices and mobile terminals. It is very suitable for devices and application scenarios with limited resources and can meet the common usage requirements of single-piece separation.
[0044] The backbone network of this lightweight object detection model uses CSPDarkNet53 for feature extraction, which can capture global information in the image. Combined with various optimization techniques such as quantization and pruning, it helps accurately detect and locate various targets. The object detection head of the lightweight object detection model includes three output layers of different scales to detect candidate boxes of different sizes. The object detection head provides prediction parameters for each candidate box. The prediction parameters include the center coordinates of the candidate box, the width and height of the candidate box, and the confidence level. The confidence level represents the maximum value of the confidence level of whether the candidate box contains a package and the background confidence level of whether a package exists at that location.
[0045] In addition, some other improvements have been added to the lightweight object detection model, including the use of Hard Swish activation function and depth-wise separable convolution to improve the accuracy and efficiency of the lightweight object detection model.
[0046] Before using the lightweight target detection model, the training process of the lightweight target detection model is also included. After building the network structure of the lightweight target detection model based on YOLO, it includes two stages: pre-training and target detection training:
[0047] During the pre-training phase, the COCO dataset is used to pre-train the network structure. This step is carried out in parallel on multiple GPUs to improve data training efficiency and accelerate model convergence.
[0048] During the object detection training phase, sample RGB images of packages are first acquired and annotated. The number of categories can be set to detect and classify multiple targets of different categories, making the model suitable for real-world scenarios where multiple package categories must be detected simultaneously. For example, this allows for distinguishing between normal packages and black packages with poor depth data. Sample RGB images are acquired for packages of varying sizes, packaging material types, and packaging colors. Data augmentation is also performed on these images using random cropping, rotation, and scaling to improve the model's generalization capabilities.
[0049] Next, we use the labeled sample RGB images to train the model based on the pre-trained model parameters. We use the Greedy NMS algorithm for non-maximum suppression and backpropagation to update the model parameters. During this process, we also need to set an appropriate learning rate, loss calculation method, and iterations. In one example, we set the non-maximum suppression parameter to 0.6, the initial learning rate to 0.001, the category loss normalization factor to 1.0, and the IoU loss normalization factor to 0.07. For the IoU loss calculation method, we choose CIOU.
[0050] In step 3, based on the alignment of the RGB image to be detected and the depth data to be detected, the depth data to be detected within each target detection frame can be obtained, and a reference depth value for the target detection frame can be determined based on the depth data to be detected within each target detection frame. In one embodiment, all the depth data to be detected within each target detection frame are sorted, and the median of all the depth data to be detected within each target detection frame is used as the reference depth value for the target detection frame.
[0051] Step 4: Obtain the projection coordinates of the four vertices of each target detection frame on the separation belt according to the reference depth value of each target detection frame.
[0052] First, the camera parameters of the RGBD camera are converted to obtain the two-dimensional coordinates of the four vertices of the target detection frame in the camera coordinate system. Then, the reference depth value of the target detection frame is used as the third-dimensional coordinate of the four vertices of the target detection frame. Combined with the two-dimensional coordinates of each vertex, the coordinate conversion can be performed to obtain the three-dimensional coordinates of the corresponding vertex in the camera coordinate system.
[0053] Then, based on the 3D coordinates of each vertex in the target detection frame, the projection coordinates of the vertex on the separation belt are obtained. This step requires the plane of the separation belt to be determined in advance. Based on the parameters of the belt plane equation of the predetermined separation belt, the 3D coordinates of each vertex in the target detection frame can be projected onto the separation belt to obtain the corresponding projection coordinates.
[0054] The belt plane equation parameters used here are predetermined and are used to calibrate the plane where the separation belt is located. The method for determining the belt plane equation parameters includes: first, using an RGBD camera to collect a calibration RGB image and calibration depth data of the single-piece separation area when no package is set on the separation belt, and then extracting the belt plane equation parameters based on the calibration RGB image and calibration depth data. The method is as follows: (1) First, based on the calibration RGB image, the separation belt area in the single-piece separation area is determined as the area of interest. This is because the field of view of the RGBD camera generally includes the background image of the separation belt in addition to the area where the separation belt is located, so it is necessary to extract the area where the separation belt is located. (2) According to the internal parameter matrix of the RGBD camera, the calibration depth data in the area of interest is converted into belt three-dimensional point cloud data. (3) Based on the random sampling consistency algorithm, the point cloud plane fitting of the belt three-dimensional point cloud data is performed to obtain the belt plane equation parameters.
[0055] Step 5: Determine the location of the corresponding package based on the projected coordinates of the four vertices of each target detection frame onto the separation belt. This involves setting the height of the projected coordinates of the four vertices of each target detection frame to 0 and fitting the maximum bounding rectangle of the two-dimensional horizontal coordinates of the projected coordinates of the four vertices of each target detection frame. This step can be performed using the cv::minAreaRect function provided by the OpenCV library. This fits the maximum bounding rectangle of the landing area of the package corresponding to the target detection frame. The package's location can then be determined based on this maximum bounding rectangle.
[0056] For example, in one example, the four vertices of an object detection frame obtained in step 4 are converted to 3D coordinates, as shown in Figure 2: A, B, C, and D. It can be seen that the depth values of these four 3D coordinates are the same. The 3D coordinates A, B, C, and D are projected onto the separation belt as coordinates A', B', C', and D', respectively. Fitting these four projected coordinates to the maximum bounding rectangle of the package's landing area is shown in the shaded area.
[0057] The above description is only a preferred embodiment of the present application, and the present application is not limited to the above embodiments. It is understood that other improvements and variations directly derived or imagined by those skilled in the art without departing from the spirit and concept of the present application should be considered to be included in the scope of protection of the present application.
Claims
1. A method for rapid positioning of a single separated package based on target detection, characterized in that: The single-piece separated package rapid positioning method comprises: The RGBD camera located above the single-piece separation area collects the RGB image to be detected and the depth data to be detected of the single-piece separation area, wherein the single-piece separation area includes a separation belt and a plurality of packages arranged on the separation belt; Performing target detection on the packages in the RGB image to be detected by using a lightweight target detection model based on YOLO, and extracting target detection frames of each package in the RGB image to be detected; Determine a reference depth value of each target detection frame according to the depth data to be detected in each target detection frame, and obtain projection coordinates of four vertices of the target detection frame on the separation belt according to the reference depth value of each target detection frame; The location information of the corresponding package is determined according to the projection coordinates of the four vertices of each target detection frame on the separation belt.
2. The method for rapid positioning of a single separated package according to claim 1, characterized in that: The step of obtaining the projection coordinates of four vertices of each target detection frame on the separation belt according to the reference depth value of each target detection frame includes: The two-dimensional coordinates of the four vertices of the target detection frame in the camera coordinate system are obtained by using the camera parameter conversion of the RGBD camera; Using the reference depth value of the target detection frame as the third-dimensional coordinates of the four vertices of the target detection frame, and combining the two-dimensional coordinates of each vertex to obtain the three-dimensional coordinates of the corresponding vertex in the camera coordinate system; The projection coordinates of each vertex on the separation belt are obtained according to the three-dimensional coordinates of each vertex of the target detection frame.
3. The method for rapid positioning of a single separated package according to claim 2, characterized in that: The step of obtaining the projection coordinates of each vertex on the separation belt according to the three-dimensional coordinates of each vertex of the target detection frame comprises: Based on the belt plane equation parameters of the separation belt, the three-dimensional coordinates of each vertex of the target detection frame are projected onto the separation belt to obtain corresponding projection coordinates.
4. The method for rapid positioning of a single separated package according to claim 3, characterized in that: The single-piece separated package rapid positioning method comprises: The RGBD camera is used to collect the calibrated RGB image and calibrated depth data of the single-piece separation area when no package is set; Based on the calibrated RGB image, determine the area where the separation belt in the single-piece separation area is located as the area of interest; Converting the calibrated depth data in the region of interest into three-dimensional point cloud data of the belt according to the intrinsic parameter matrix of the RGBD camera; The belt plane equation parameters are obtained by performing point cloud plane fitting on the three-dimensional point cloud data of the belt based on a random sampling consistency algorithm.
5. The method for rapid positioning of a single separated package according to claim 1, characterized in that: The step of determining a reference depth value of each target detection frame according to the depth data to be detected in each target detection frame comprises: The median of all the depth data to be detected in each target detection frame is used as the reference depth value of the target detection frame.
6. The method for rapid positioning of a single separated package according to claim 1, characterized in that: The step of determining the location information of the corresponding package according to the projection coordinates of the four vertices of each target detection frame on the separation belt includes: The two-dimensional coordinates of the projection coordinates of the four vertices of each target detection frame on the horizontal plane are fitted with a maximum circumscribed rectangle to obtain the maximum circumscribed rectangle of the landing area of the package corresponding to the target detection frame, and the location information of the package is determined based on the maximum circumscribed rectangle of the landing area of the package.
7. The method for rapid positioning of a single separated package according to claim 1, characterized in that: The backbone network of the lightweight target detection model adopts CSPDarkNet53 for feature extraction. The target detection head includes three output layers of different scales to detect candidate boxes of different sizes. The target detection head gives prediction parameters for each candidate box. The prediction parameters include the center coordinates of the candidate box, the width and height of the candidate box, and the confidence, where the confidence represents the maximum value of the confidence of the package and the background confidence.
8. The method for rapid positioning of a single separated package according to claim 7, characterized in that: The lightweight object detection model uses the Hard Swish activation function and also uses depthwise separable convolution.
9. The method for rapid positioning of a single separated package according to claim 1, characterized in that: The single-piece separated package rapid positioning method comprises: Build the network structure of a lightweight target detection model based on YOLO; Use multiple GPUs in parallel to use the COCO dataset to pre-train the model parameters based on the established network structure; Get sample RGB images of packages and annotate them; The labeled sample RGB images are used to obtain model parameters based on pre-training for training, and the Greedy NMS algorithm is used for non-maximum suppression, and the back propagation algorithm is used to update the model parameters.
10. The method for rapid positioning of a single separated package according to claim 9, characterized in that: When acquiring the sample RGB images, sample RGB images of packages of different sizes, packaging material types, and packaging colors are acquired respectively.
Citation Information
Patent Citations
RGB image-based parcel detection method and system, medium and terminal
CN113344949A
Trolley parcel position detection method based on deep learning
CN113554706A
Automatic parcel sorting method based on deep learning
CN116228854A
Parcel static volume measurement method and system
CN116255912A
Single-piece separation parcel rapid positioning method based on target detection
CN117765065A