Multi-target parcel volume measurement method based on depth camera handheld inclined shooting
By using multiple screening fitting and RANSAC plane fitting algorithms when shooting with handheld tilt of the depth camera, combined with IR gradient information for edge detection, the problem of volume measurement accuracy of the depth camera at different angles is solved, and efficient and stable multi-objective packaging volume measurement is achieved.
Patent Information
- Application Number
- CN202510226067.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to meet the flexible use needs of depth cameras at different angles and scenes, resulting in point cloud data distortion when shooting with handheld tilt, affecting the accuracy of package volume measurement.
Multiple screening fitting and RANSAC plane fitting algorithms are used, combined with normal vector constraints, which significantly improves the accuracy of ground segmentation and target recognition. At the same time, IR gradient information is fused for edge detection and target segmentation to ensure the stability of multi-target package volume measurement at different angles and distances.
It breaks through the limitations of fixed angles of the camera, improves measurement efficiency and accuracy, and meets the demand for rapid multi-target volume detection in the fields of logistics, warehousing and other fields.
Smart Images

Figure CN120141298A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of volume measurement, and specifically relates to a multi-target package volume measurement method based on handheld inclined shooting with a depth camera. Background Art
[0002] With the rapid development of logistics, warehousing, and manufacturing, the demand for volume measurement of packages, goods, and other cuboid objects is increasing. Especially in automated warehousing, sorting centers, and intelligent logistics systems, accurate and efficient volume measurement is a key factor for rapid item classification, storage optimization, and transportation cost accounting. Most traditional volume measurement methods rely on manual operation or fixed laser measurement equipment, which have problems such as low measurement efficiency, large errors, and high equipment costs, and are difficult to meet the requirements of modern industry for high-efficiency and automated volume measurement.
[0003] In recent years, measurement technologies based on depth cameras have gradually been applied to the field of three-dimensional measurement. Depth cameras can quickly obtain the volume of the object to be measured by capturing the three-dimensional information of the object. In a Chinese invention patent with the publication number CN115077379B and the title "A Method for Measuring the Volumes of Multiple Packages Based on Vertical Shooting with a Depth Camera", a measurement method based on a fixed vertical shooting method is proposed, which uses a depth camera and an algorithm to perform real-time measurement of the volumes of multiple packages. This technology has successfully solved the need for simultaneous measurement of multiple targets and achieved high precision and efficiency in a fixed camera scenario. However, this method still has some limitations. This invention patent requires that the depth camera must be in a fixed vertical shooting position and is not applicable to non-vertical shooting scenarios.
[0004] Since the camera is tilted in the handheld state, the obtained point cloud data is prone to distortion, which in turn affects the accuracy of package volume measurement. To solve this problem, the present invention significantly improves the accuracy of ground segmentation and target recognition through multiple screening fittings and a RANSAC plane fitting algorithm combined with normal vector constraints. At the same time, IR gradient information is fused for edge detection and target segmentation to ensure the stability of multi-target package volume measurement at different angles and distances. Summary of the Invention
[0005] The technical problem to be solved by the present invention is: how to meet the flexible use requirements of a depth camera at different angles and in different scenarios during real-time measurement of the target volume. A multi-target package volume measurement method based on handheld inclined shooting with a depth camera is provided, which breaks through the limitation of the fixed angle of the camera, expands the application range of the depth camera in handheld and multi-angle scenarios, greatly improves flexibility and practicability, significantly improves the measurement efficiency, and meets the requirements for rapid detection of multi-target volumes in fields such as logistics and warehousing.
[0006] The present invention solves the above technical problems through the following technical solutions. The present invention includes the following steps:
[0007] S1: Obtain a frame of depth image, its corresponding IR image, and point cloud data of a depth camera. By screening and fitting the frame of depth image and point cloud data, determine the ground equation of the scene;
[0008] S2: Segment the point cloud data and depth image through the ground equation, separate the ground area from other areas, and determine the areas of all potential measured targets by screening the non-ground areas;
[0009] S3: Further screen out all qualified measured targets by fitting the top surface equation of potential measured targets, and accurately segment the top surface contour of the measured targets on the depth image according to the fitted top surface equation;
[0010] S4: Combine the top surface contour of the measured targets with the gradient information of the IR image or depth image to obtain the four vertex coordinates of the top surface of each measured target, and calculate the length and width of the measured targets accordingly;
[0011] S5: Calculate the height of all measured targets based on the intercept difference between the top surface equations and ground equations of all measured targets;
[0012] S6: According to the obtained length, width, and height of the measured targets, apply the volume formula to calculate the volume of each measured target;
[0013] S7: Apply the processing flow of steps S1 to S6 to each frame of depth image and IR image to realize the real-time measurement of the volumes of multiple targets based on the depth camera.
[0014] Furthermore, in the step S1, the specific processing process is as follows:
[0015] S11: Perform downsampling processing on a frame of depth image, store the depth values in an array, and sort the depth values in the array in ascending order; set a predetermined percentile value, and screen out the pixel points with depth values greater than the percentile value; store the point cloud coordinates corresponding to these pixel points in an array;
[0016] S12: Apply the RANSAC algorithm, set the inlier determination condition and the number of iterations, perform plane fitting on the point cloud coordinates in the array, calculate a plane equation and the inlier set, and then apply the least squares method to the inlier set of this plane for plane fitting, so as to obtain the plane equation of the ground, that is, obtain the ground equation of the scene.
[0017] Furthermore, in the step S2, the specific processing process is as follows:
[0018] S21: Create a two-dimensional array target_pixel_coordinate_1 and a three-dimensional array point_cloud_data_1;
[0019] S22: Traverse the point cloud data. According to the obtained ground equation, calculate the vertical distance from each point to the ground and set a threshold threshold1. For points with a vertical distance greater than this threshold, record their point cloud coordinates and store them in the three-dimensional array point_cloud_data_1, and at the same time store their corresponding pixel coordinates in the two-dimensional array target_pixel_coordinate_1;
[0020] S23: Create a grayscale image gray1 with the same resolution as the original depth map. Traverse the pixel coordinates in the two-dimensional array target_pixel_coordinate_1, set the pixel value corresponding to the pixel in the grayscale image gray1 to the maximum value, and set other unrecorded pixel points to the minimum value, thereby generating a binary image representing the segmentation of the ground area and other areas;
[0021] S24: Perform a closing operation on this binary image to fill small holes in it. Subsequently, apply the Canny edge detection algorithm to perform edge detection on the binary image, and use the edge tracking algorithm to extract all the contours in the image. Set a contour size screening threshold threshold2, and retain the contours with the number of contour pixels greater than the threshold threshold2. The remaining contours are regarded as the contours corresponding to the potential targets to be measured at this time. At this time, the number of contours is N1, indicating that there are at most N1 potential targets to be measured in this frame of scene;
[0022] S25: For the N1 effective contours obtained after screening in step S24, further perform the extraction work of the three-dimensional point cloud region to construct the point cloud data of the potential target region to be measured. The specific processing process is as follows:
[0023] Based on the boundary point set of each effective contour in the pixel coordinate system, calculate its corresponding three-dimensional camera coordinates through the coordinate transformation algorithm and obtain the point cloud region corresponding to each contour. Based on the geometric envelope characteristics of the contour, use the spatial interpolation algorithm to process the contour region to extract and integrate all the inlier data contained in the contour, thereby filling the vacant region inside the contour to form a complete point cloud region. Through this process, the point cloud data in the three-dimensional array point_cloud_data_1 is finally divided into N1 subsets, and each subset corresponds to the point cloud data set of a potential target to be measured.
[0024] Furthermore, in step S3, the specific processing process is as follows:
[0025] S31: For the N1 subsets of the three-dimensional array target_pixel_coordinate_1 in step S24, apply the RANSAC algorithm for plane fitting respectively to obtain the inlier sets of each subset;
[0026] S32: Subsequently, apply the least squares method to the inlier sets of each subset for precise plane fitting to obtain N1 plane equations, that is, obtain the top surface equations of each potential target to be measured;
[0027] S33: By calculating the angle between the top surface equation of each potential target to be measured and the ground equation, set the angle threshold threshold3, and eliminate the potential targets corresponding to the top surface equations with an angle lower than this threshold. At this time, the number of remaining top surface equations and subsets of the three-dimensional array is N2;
[0028] S34: Create a two-dimensional array target_pixel_coordinates_2, traverse the point cloud data of the remaining N2 subsets, calculate the perpendicular distance from each point to its corresponding top surface equation, set the threshold threshold4, and the points less than this threshold are considered inliers, and store their corresponding pixel coordinates into the two-dimensional array target_pixel_coordinates_2;
[0029] S35: Create a grayscale image gray2 with the same resolution as the original depth map. Traverse the pixel coordinates in the two-dimensional array target_pixel_coordinates_2, set the pixel value corresponding to the grayscale image gray2 to the maximum value, and set other unrecorded pixel points to 0. The generated binary image can segment the top surface area of the remaining potential targets to be measured; then perform a closing operation on this binary image to fill the holes, then apply an edge tracking algorithm to detect the edges of the image, and finally use a contour detection algorithm to extract the top surface contours of all remaining potential targets to be measured.
[0030] Furthermore, in the step S4, the specific processing process is as follows:
[0031] S41: For the top surface contours of the remaining N2 potential targets to be measured in step S35, first calculate their perimeters, and determine the accuracy threshold for polygon approximation according to the perimeter multiplied by a coefficient; then use the Douglas-Peucker algorithm to perform polygon approximation on these top surface contours. After the processing is completed, check the number of vertices of each generated polygon; if the polygon has four vertices, retain the contour and its corresponding four vertices and the top surface equation. If the number of vertices of the polygon is less than four, delete the contour and its top surface equation to further screen out potential targets to be measured. The number of remaining qualified potential targets to be measured is N3;
[0032] S42: Calculate the gradient maps in the X and Y directions respectively for the IR image of the depth camera, take their absolute values and add them to obtain the IR amplitude image; for the N3 contours retained in step S41, use the four line segments formed by the four vertices corresponding to each contour, and traverse each of the four line segments of each contour one by one; within the set range of the positive and negative directions of the line normal vector, find the point with the largest pixel value in the IR amplitude image as the two-dimensional image position of the top edge of the target to be measured; each contour corresponds to four groups of pixel points of the top edge. Perform least squares fitting on each group of pixel points to obtain four straight line equations; then calculate the included angle between every two opposite straight lines, and set an angle threshold threshold5. If both included angles are less than this angle threshold, then identify this contour as the actual target to be measured, otherwise delete this contour, its vertices, and the top surface equation; the final number of retained contours is N4, which is the number of actual targets to be measured;
[0033] S43: For the N4 actual targets to be measured retained in step S42, first calculate the intersection points of the four edge straight lines on the top surface of each target to determine the pixel coordinates of the four upper vertices of each target to be measured in the depth image; then combine these upper vertex pixel coordinates with the inverse perspective transformation formula and its top surface equation to calculate the three-dimensional coordinates of the upper vertices of each target to be measured in the camera coordinate system; finally, use the Euclidean distance formula to calculate the physical distances of the four upper vertices on the top surface of each target to be measured in the camera three-dimensional coordinate system, and then determine the length and width of each target to be measured.
[0034] Furthermore, in the said step S5, the specific processing process is as follows:
[0035] For the N4 targets to be measured finally retained in step S42, their top surface equations and the ground equations supporting them are both known; since the two planes represented by the top surface equation and the ground equation are parallel, by calculating the perpendicular distance between the top surface equation and the ground equation, determine the height of each target to be measured, and this perpendicular distance is calculated from the absolute value of the difference between the intercepts of their plane equations.
[0036] The present invention has the following advantages compared with the prior art:
[0037] 1. Improve measurement efficiency: It can measure multiple cuboid targets at one time in the same scene, greatly reducing the time for individual measurements and improving the overall measurement efficiency.
[0038] 2. Not limited by angles: It supports real-time multi-target measurement with a hand-held depth camera at different shooting angles and is not limited by fixed vertical shooting.
[0039] 3. Short time-consuming: Adopt lightweight point cloud processing and image processing algorithms, effectively avoiding the high computational overhead brought by deep learning network inference and ensuring the high-efficiency real-time performance of the measurement process.
[0040] 4. Strong adaptability: It can perform accurate measurements under various environmental conditions (such as different lighting, complex backgrounds, etc.), has strong adaptability, and meets the volume measurement requirements in different scenarios.
[0041] 5. Simplify the measurement process: Abandon the cumbersome process of traditional multiple shootings to collect data. Only a single shooting is required to complete the volume calculation through one frame of data, greatly simplifying the measurement process and improving the efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic flow chart of the multi-target package volume measurement method based on handheld tilt shooting with a depth camera in the first embodiment of the present invention;
[0043] Figure 2 It is a depth image of a frame of the depth camera and the point cloud data generated by combining the internal parameters in the first embodiment of the present invention, where (a) is the depth image and (b) is the corresponding generated point cloud data;
[0044] Figure 3 It is the corresponding IR image collected by the depth camera in the first embodiment of the present invention;
[0045] Figure 4 It is a binary image generated by segmenting the ground area and the non-ground area in the first embodiment of the present invention;
[0046] Figure 5 It is a schematic diagram of the contour of the non-ground area after screening in the first embodiment of the present invention;
[0047] Figure 6 It is the three-dimensional point cloud data of three potential measured targets extracted in the first embodiment of the present invention;
[0048] Figure 7 It is a schematic diagram of the segmentation result of the top surface area of the potential measured target in the first embodiment of the present invention;
[0049] Figure 8 It is a schematic diagram of the top surface contour of the measured target that meets the conditions after screening through the top surface contour in the first embodiment of the present invention;
[0050] Figure 9 It is the IR amplitude diagram in the first embodiment of the present invention;
[0051] Figure 10 It is a schematic diagram of the straight line fitted by the four groups of vertex coordinates on the top surface of the measured target in the first embodiment of the present invention, where (a) is the first measured target, (b) is the second measured target, and (c) is the third measured target;
[0052] Figure 11They are the pixel coordinate positions of the four vertices on the top surface of each measured target in the first embodiment of the present invention;
[0053] Figure 12 They are the spatial positions and shape contours of the measured targets reconstructed by perspective projection in the first embodiment of the present invention;
[0054] Figure 13 They are a frame of depth image and the corresponding IR image collected by the depth camera in the second embodiment of the present invention, where (a) is the depth image and (b) is the corresponding IR image;
[0055] Figure 14 They are the schematic diagram of the comparison between the measurement results and the actual sizes of the measured targets in the second embodiment of the present invention;
[0056] Figure 15 They are the example diagrams of the packages used in the experiment in the third embodiment of the present invention;
[0057] Figure 16 They are the schematic diagrams of the camera tilting angles for shooting in the fourth embodiment of the present invention, where (a) is 0° and (b) is 30°. Detailed implementation manners
[0058] The embodiments of the present invention will be described in detail below. These embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments.
[0059] Embodiment 1
[0060] As Figure 1 shown, this embodiment provides a method for measuring the volumes of multiple target packages by handheld tilting shooting with a depth camera. The measured target refers to any package or object with a cuboid shape, and the depth camera refers to a device that can generate a depth image or generate a depth image and an IR image simultaneously. This method is applicable to the efficient and real-time volume measurement of multiple cuboid targets.
[0061] S1: Obtain a frame of depth image of the depth camera, its corresponding IR image, and point cloud data; determine the ground equation of the scene by screening and fitting the frame of depth image and point cloud data;
[0062] In this embodiment, the image resolution of the depth camera is selected as 640×480, and the internal parameters of the camera f x , f y , c x , c y . Figure 2 shows the obtained frame of depth image and its corresponding point cloud data, Figure 3It is shown as the IR image corresponding to this frame. To improve the calculation efficiency, the depth image is downsampled with a row-column stride of 4, and the depth values are stored in an array. Subsequently, the depth values in this array are sorted in ascending order, a predetermined percentile of 50% is set, and the pixel points with depth values greater than the depth value at the 50% position of this array are screened out.
[0063] The point cloud coordinates corresponding to these screened pixel points are stored in a three-dimensional array. Next, the RANSAC algorithm is applied for plane fitting, with the inlier threshold set to 5 mm and the number of iterations set to 1000. By processing the point cloud coordinates in this three-dimensional array, we can calculate a preliminary plane equation and its corresponding inlier set. To improve the fitting accuracy, the least squares method is further applied to the inlier set of the plane for plane fitting, so as to obtain a more accurate ground plane equation A g x + B g y + C g z + D g = 0.
[0064] S2: Segment the point cloud data and the depth image through the ground equation, separate the ground area from other areas, and determine the areas of all potential targets to be measured by screening the non-ground areas;
[0065] First, create a two-dimensional array target_pixel_coordinates_1 to store pixel coordinates, and a three-dimensional array point_cloud_data_1 to store point cloud coordinates. Traverse the entire point cloud data, calculate the vertical distance from each point to the ground according to the ground plane equation calculated in the previous step, and set the threshold to 10 mm. For the points with a vertical distance greater than this threshold, record their point cloud coordinates and store them in the three-dimensional array point_cloud_data_1, and at the same time store their corresponding pixel coordinates in the two-dimensional array target_pixel_coordinates_1.
[0066] Next, create a grayscale image gray1 with the same resolution as the original depth image. Traverse all the pixel coordinates in the two-dimensional array target_pixel_coordinates_1, set the pixel value at the corresponding position in the grayscale image gray1 to the maximum value, and set the unrecorded pixel points to the minimum value. This process generates a binary image, as Figure 4 shown, representing the segmentation of the ground area and the non-ground area. Subsequently, perform a closing operation on this binary image to fill small holes, and apply the Canny edge detection algorithm to detect its edges.
[0067] Then, the edge tracking algorithm is used to extract all the contours in the image. The contour filtering threshold is set to 300 pixels, and only the contours with a pixel count greater than 300 are retained to exclude the interference of smaller contours. The remaining contours are regarded as the contours corresponding to the potential targets to be measured at this time, as Figure 5 shown. The number of retained contours is 3, indicating that there are at most 3 potential targets to be measured in the current frame scene. For these contours, first, based on the set of boundary points of each valid contour in the pixel coordinate system, its corresponding three-dimensional camera coordinates are accurately calculated through the coordinate transformation algorithm, and the point cloud region corresponding to each contour is obtained. Then, based on the geometric envelope characteristics of the contour, the spatial interpolation algorithm is used to process the contour region to extract and integrate all the inlier data contained in the contour, thereby filling the vacant region inside the contour to form a complete point cloud region. Through the above method, the point cloud coordinates in point_cloud_data_1 are respectively stored in the three subset three-dimensional arrays array_point_cloud[3], and the three groups of point clouds are as Figure 6 shown.
[0068] S3: By fitting the top surface equation of the potential target to be measured, all the qualified targets to be measured are further screened out, and the top surface contours of these targets to be measured are accurately segmented on the depth image according to the fitted top surface equation;
[0069] In step S3, for the three three-dimensional arrays in array_point_cloud[3], since the handheld camera is shooting from top to bottom, the top surface information of the target to be measured is dominant. First, the RANSAC algorithm is applied for plane fitting, with the inlier threshold set to 5 mm and the number of iterations set to 500 times, and the inlier sets of these three three-dimensional arrays are calculated respectively. Then, the least squares plane fitting is performed on these three groups of inlier sets to obtain three plane equations.
[0070] Subsequently, the angle between these planes and the ground is calculated. For example, if the ground equation is A 1 x + B 1 y + C 1 z + D 1 = 0, and a certain plane equation is A 2 x + B 2 y + C 2 z + D 2 = 0, then the angle between this plane and the ground can be calculated through Since the target to be measured is a cuboid and its top surface should be parallel to the ground, if the angle between the plane and the ground is greater than 3°, then this plane does not meet the requirements of the target to be measured and should be deleted. Finally, the remaining plane equation represents the top surface equation of the potential target to be measured at this time.
[0071] Next, create a two-dimensional array target_pixel_coordinates_2, traverse the point cloud data, and calculate the distance from each point in the scene to the three plane equations mentioned above. If the distance from a point to the plane is less than the threshold of 5mm, the pixel coordinates of the point in the depth image are stored in target_pixel_coordinates_2. Next, create a grayscale image gray2 with the same resolution as the original depth image, traverse the two-dimensional array target_pixel_coordinates_2, set the corresponding pixel values in the grayscale image to the maximum value, and set the remaining pixels to the minimum value. This operation generates a binary image containing the top surface area of each potential target to be measured, such as Figure 7 Then the binary image is closed to fill the holes, the Canny edge detection algorithm is applied, and then the edge tracking algorithm is applied to extract the top surface contours of the three actual measured targets.
[0072] S4: combining the top surface contour of the target to be measured with the gradient information of the IR image or the depth image, obtaining the coordinates of the four vertices of the top surface of each target to be measured, and calculating the length and width of the target accordingly;
[0073] For the top surface contours of the three potential targets, the perimeter of each contour is calculated in turn, and the perimeter is multiplied by a coefficient of 0.1 as the approximation accuracy. Then the Douglas-Peucker algorithm is used to perform polygonal approximation on each contour to determine the number of vertices of each contour. If the number of vertices of a contour is 4, its shape meets the requirements of the target (rectangular block), so the contour is retained; otherwise, it is removed. The contours that are finally retained are as follows: Figure 8 shown.
[0074] Next, calculate the X and Y direction gradient maps of the current frame IR image, take their absolute values and add them to obtain the IR amplitude map that can represent the edge strength, such as Figure 9 As shown. For the top surface contours of the three potential targets retained above, four end-to-end line segments are constructed for the four vertices of each contour. Pixel traversal with a stride of 5 is performed along each line segment, and the maximum pixel value in the IR amplitude image is searched within a set range in the positive and negative directions of its normal vector, and the point is taken as the possible edge position on the line. Four sets of top surface edge points are obtained for each contour, as shown Figure 10 As shown. The least squares method is used to fit these edge points to obtain the four straight line equations of each contour. Then, the angle detection is performed on the four straight lines of each contour. For example, the straight line v 1 =m 1 u+b 1 and v 2 =m 2 u+b 2 , the angle between the two straight lines Determine whether two opposite straight lines are parallel (the included angle is less than the threshold of 5°). If the two sets of opposite straight lines of each contour meet this included angle condition, then retain this contour; otherwise, eliminate it. The finally retained ones are the actual measured targets.
[0075] Next, calculate the intersection points of the four boundary straight lines on the top surface of each measured target to determine the four vertex pixel coordinates of each measured target in the depth image. For example, the intersection point of the straight line L 1 :v 1 =m 1 u + b 1 and L 2 :v 2 =m 2 u + b 2 is v=m 1 u + b 1 etc. By calculating the intersection points of the four sides, obtain the pixel coordinates of the four vertices of the top surface of the measured target in the depth image, as shown in Figure 11 . Next, according to the inverse perspective projection theorem, combine these pixel coordinates with the plane equation Ax + By + Cz + D = 0 of the top surface of the measured target, and calculate the three-dimensional point cloud coordinates of each vertex through the following formula:
[0076] Z=-D / (A*k_x + B*k_y + C)
[0077] X=Z*k_x
[0078] Y=Z*k_y
[0079] where:
[0080] k_x=(u - c x ) / f x
[0081] k_y=(v - c y ) / f y
[0082] Through the above method, the point cloud coordinates (X, Y, Z) of the four vertices of the top surface of each measured target can be obtained. Then, use the Euclidean distance formula to calculate the distances between the vertices to obtain the actual lengths of the four sides of the top surface. For the top surface of each measured target, calculate the lengths of the two pairs of opposite sides respectively and take their average values to determine the actual length and width of each measured target.
[0083] S5: Calculate the intercept difference between the top surface equation and the ground equation of all measured targets to obtain the heights of all measured targets;
[0084] For each target to be measured, the height of the target can be obtained by calculating the vertical distance between its top surface and the ground. Given that the ground equation and the top surface equation of each target to be measured are theoretically parallel, the standard plane equation Ax + By + Cz + D i = 0 after normalizing the normal vector can be used to determine the vertical distance between the two planes through the absolute difference |D1 - D2| of the intercepts of the two plane equations, thereby obtaining the height of each target to be measured.
[0085] After obtaining the length, width, and height of each target to be measured through the above method, the volume calculation formula V = length × width × height can be directly applied to calculate the volume of each target to be measured.
[0086] After obtaining the point cloud coordinates of the four vertices of the top surface of each target to be measured and the corresponding height, based on the coordinates of the four vertices of the top surface, the height distance can be displaced downward along the normal vector direction of the top surface plane equation to further calculate the point cloud coordinates of the four vertices of its bottom surface. Subsequently, combined with the perspective projection formula, the four vertices of the bottom surface of each target to be measured are converted into pixel coordinates. In this way, the spatial position and shape contour of each target to be measured can be accurately drawn by connecting the pixel coordinates of the eight vertices of the top surface and the bottom surface, as Figure 12 shown.
[0087] The figure shows the measurement results of the length, width, and height of each target to be measured, marked as package1, package2, and package3 in turn, arranged in a counterclockwise order starting from the bottommost target to be measured. The actual sizes of the three targets to be measured are 240mm × 220mm × 260mm, 333mm × 322mm × 190mm, and 333mm × 322mm × 350mm respectively, and the corresponding test results are 236mm × 221mm × 263mm, 336mm × 318mm × 186mm, and 329mm × 325mm × 357mm respectively. The measurement accuracy is relatively high, verifying the effectiveness and accuracy of the measurement method.
[0088] S6: Apply the processing flow of S1 - S5 to each frame of depth and IR images to achieve real-time measurement of the volumes of multiple targets based on a depth camera.
[0089] Embodiment 2
[0090] This embodiment introduces a method for measuring the volumes of multiple target packages by handheld tilting shooting with a depth camera. The specific steps are the same as those in Embodiment 1.
[0091] As Figure 13As shown, a certain frame collected by the depth camera includes a depth image and the corresponding IR image. Two measured targets from right to left are shown in the figure, with actual sizes of 333mm×322mm×190mm and 333mm×322mm×350mm respectively. The measurement results obtained by the method of the present invention, as Figure 14 shown, are 329mm×322mm×192mm and 336mm×317mm×345mm respectively. This measurement method can still achieve high-precision acquisition of the sizes of each measured target in different environments, showing good measurement stability and reliability, and fully verifying the robustness of the algorithm.
[0092] Embodiment 3
[0093] This embodiment introduces a method for real-time measurement of the volumes of multiple targets based on a TOF depth camera, and verifies the influence of the number of packages on the measurement accuracy. Four cuboid cardboard boxes were used for testing in the experiment, namely Carton F2 (250mm×200mm×150mm), Carton F3 (300mm×250mm×200mm), Carton F4 (400mm×300mm×200mm), and Carton F5 (400mm×300mm×300mm). The experimental design was divided into three groups to test the influence of different numbers of packages on the volume measurement accuracy.
[0094] As Figure 15 shown, are the packages used in this experiment. The experiment was conducted to measure the volumes of multiple targets by handholding a TOF depth camera. A free hand-held sampling method was adopted, with a sampling height of about 1 meter, and the angle between the camera and the vertical direction was within 10°. Two (F2, F3), three (F2, F3, F4), and four (F2, F3, F4, F5) packages were measured ten times respectively. The test data of the experiment are shown in Tables 1, 2, and 3, where the units of length, width, and height are all mm, and the volume unit is mm 3 , recording the average sizes and volumes of the packages in each group of experiments, as well as the corresponding relative volume errors.
[0095] Table 1 Test data when measuring two packages simultaneously
[0096]
[0097] Table 2 Test data when measuring three packages simultaneously
[0098]
[0099]
[0100] Table 3 Test data when measuring four packages simultaneously
[0101]
[0102] The experimental results show that in the face of different numbers of packages, the measurement accuracy remains at a relatively high level. The average relative volume error of the first group (cartons F2 and F3) tested was 2.36%, the average relative volume error of the second group (cartons F2, F3, and F4) tested was 3.40%, and the average relative volume error of the third group (cartons F2, F3, F4, and F5) tested was 4.42%.
[0103] This embodiment verifies the effectiveness and accuracy of the multi-objective package volume measurement method through experiments. Even when the number of packages increases, this method can still provide high-precision and stable volume measurement results, fully demonstrating the advantages of the method of the present invention in practical applications.
[0104] Embodiment 4
[0105] To verify the volume measurement accuracy of the method of the present invention at different angles, this embodiment conducts multi-angle tests on carton F2 (250mm×200mm×150mm) and carton F3 (300mm×250mm×200mm) to explore the influence of the camera tilt angle on the measurement accuracy. As Figure 16 shown, the tests were carried out from a tilt of 0° to 30°, with measurements taken every 5°, aiming to verify that the proposed method can still maintain a high accuracy even at different shooting angles. Each test was carried out in a free-hand manner, and the distance between the camera and the ground was approximately 1 meter during shooting. The experimental data is shown in Table 4 below, and the unit of all volume data is mm 3 :
[0106] Table 4 Test data for simultaneously measuring two packages at different angles
[0107]
[0108]
[0109] As can be seen from the above table, regardless of how the shooting angle of the camera changes, the measured values of the volumes of cartons F2 and F3 always remain within a range close to the true values, and the relative volume error fluctuates within 4%, with relatively high accuracy. This shows that even at different angles, the method of the present invention can still effectively maintain a high measurement accuracy and has good stability and adaptability.
[0110] In summary, the multi-object package volume measurement method based on handheld inclined shooting with a depth camera in the above embodiments breaks through the limitations of the traditional fixed shooting mode, supports measurements at any angle and distance, and significantly improves the environmental adaptability and operation flexibility of the measurement system; abandons the cumbersome process of collecting data through multiple shootings in the past, and only requires a single shooting to complete the volume calculation through one frame of data, greatly simplifying the measurement process and improving the efficiency; adopts lightweight point cloud processing algorithms and geometric calculation methods, effectively avoiding the high computational overhead brought by deep learning network inference, and ensuring the efficiency and real-time performance of the measurement process.
[0111] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A multi-target package volume measurement method based on handheld tilted shooting of a depth camera, characterized in that: The following steps are involved: S1: Obtain a depth image frame from a depth camera and its corresponding IR image and point cloud data, and determine the ground equation of the scene by screening and fitting the depth image frame and point cloud data; S2: Segment the point cloud data and depth image through the ground equation to separate the ground area from other areas, and determine the area of all potential targets by screening the non-ground area; S3: By fitting the top surface equation of the potential target to be measured, all qualified targets to be measured are further screened out, and the top surface contour of the target to be measured is accurately segmented on the depth image according to the fitted top surface equation; S4: combining the top surface contour of the measured target with the gradient information of the IR image or the depth image, obtaining the coordinates of the four vertices of the top surface of each measured target, and calculating the length and width of the measured target accordingly; S5: Calculate the heights of all measured targets based on the intercept differences between the top surface equations and the ground surface equations of all measured targets; S6: Calculate the volume of each measured object using a volume formula according to the obtained length, width and height of the measured object; S7: Apply the processing flow of steps S1 to S6 to each frame of the depth image and the IR image to achieve real-time measurement of multi-target volumes based on the depth camera.
2. The multi-target package volume measurement method based on handheld oblique shooting of a depth camera according to claim 1 is characterized in that: In step S1, the specific processing process is as follows: S11: down-sample a frame of depth image, store the depth value in an array, and sort the depth values in the array in ascending order; set a predetermined percentile value, filter out pixels whose depth value is greater than the percentile value; store the point cloud coordinates corresponding to these pixels in an array; S12: Apply the RANSAC algorithm, set the internal point judgment conditions and the number of iterations, perform plane fitting on the point cloud coordinates in the array, calculate a plane equation and an internal point set, and then apply the least squares method to the internal point set of the plane for plane fitting, so as to obtain the plane equation of the ground, that is, the ground equation of the scene.
3. The multi-target package volume measurement method based on handheld tilted shooting of a depth camera according to claim 1 is characterized in that: In step S2, the specific processing process is as follows: S21: Create a two-dimensional array target_pixel_coordinate_1 and a three-dimensional array point_cloud_data_1; S22: traverse the point cloud data, calculate the vertical distance from each point to the ground according to the obtained ground equation, and set the threshold threshold1; for points whose vertical distance is greater than the threshold, record their point cloud coordinates and store them in the three-dimensional array point_cloud_data_1, and store their corresponding pixel coordinates in the two-dimensional array target_pixel_coordinate_1; S23: create a grayscale image gray1 with the same resolution as the original depth image, traverse the pixel coordinates in the two-dimensional array target_pixel_coordinate_1, set the corresponding pixel value in the grayscale image gray1 to the maximum value, and set other unrecorded pixel points to the minimum value, thereby generating a binary image representing the segmentation of the ground area and other areas; S24: performing a closing operation on the binary image to fill a small hole set therein; then, applying the Canny edge detection algorithm to perform edge detection on the binary image, and using the edge tracking algorithm to extract all contours in the image; setting a contour size screening threshold threshold2, retaining contours with a number of contour pixels greater than the threshold threshold2, and the remaining contours are regarded as contours corresponding to potential targets to be detected at this time; at this time, the number of contours is N1, indicating that there are at most N1 potential targets to be detected in this frame scene; S25: For the N1 valid contours obtained after screening in step S24, further extracting the three-dimensional point cloud area is performed to construct point cloud data of the potential target area to be measured. The specific processing process is as follows: Based on the boundary point set of each valid contour in the pixel coordinate system, the corresponding three-dimensional camera coordinates are calculated through the coordinate transformation algorithm, and the point cloud area corresponding to each contour is obtained; based on the geometric envelope characteristics of the contour, the spatial interpolation algorithm is used to process the contour area to extract and integrate all the inner point data contained in the contour, so as to fill the vacant area in the contour and form a complete point cloud area; through this process, the point cloud data in the three-dimensional array point_cloud_data_1 is finally divided into N1 subsets, each subset corresponding to a point cloud data set of a potential target to be measured.
4. The multi-target package volume measurement method based on handheld oblique shooting of a depth camera according to claim 1 is characterized in that: In step S3, the specific processing process is as follows: S31: for the N1 subsets of the three-dimensional array target_pixel_coordinate_1 in step S24, apply the RANSAC algorithm to perform plane fitting respectively to obtain the internal point set of each subset; S32: Then, the least square method is applied to the internal point set of each subset to perform accurate plane fitting, and N1 plane equations are obtained, that is, the top surface equation of each potential target to be measured is obtained; S33: by calculating the angle between the top surface equation and the ground equation of each potential target to be detected, setting the angle threshold threshold3, and eliminating the potential targets corresponding to the top surface equations with angles lower than the threshold. At this time, the number of remaining top surface equations and subsets of the three-dimensional array is N2; S34: Create a two-dimensional array target_pixel_coordinates_2, traverse the point cloud data of the remaining N2 subsets, calculate the vertical distance from each point to its corresponding top surface equation, set a threshold threshold4, points smaller than the threshold are considered to be internal points, and store their corresponding pixel coordinates in the two-dimensional array target_pixel_coordinates_2; S35: Create a grayscale image gray2 with the same resolution as the original depth map, traverse the pixel coordinates in the two-dimensional array target_pixel_coordinates_2, set the corresponding pixel value in the grayscale image gray2 to the maximum value, and set other unrecorded pixel points to 0. The generated binary image can segment the top surface area of the remaining potential targets to be measured; then close the binary image to fill the holes, and then apply the edge tracking algorithm to perform edge detection on the image, and finally use the contour detection algorithm to extract the top surface contours of all remaining potential targets to be measured.
5. The multi-target package volume measurement method based on handheld oblique shooting of a depth camera according to claim 1 is characterized in that: In step S4, the specific processing process is as follows: S41: For the top surface contours of the remaining N2 potential targets to be detected in step S35, first calculate their perimeters, and determine the accuracy threshold of polygon approximation by multiplying the perimeters by a coefficient; then use the Douglas-Peucker algorithm to perform polygon approximation on these top surface contours, and after the processing is completed, check the number of vertices of each generated polygon; If the polygon has four vertices, the contour and its corresponding four vertices and top surface equation are retained. If the number of vertices of the polygon is less than four, the contour and its top surface equation are deleted to further screen out potential targets. The number of remaining potential targets that meet the conditions is N3. S42: Calculate the gradient images in the X and Y directions for the IR image of the depth camera, take their absolute values and add them to obtain the IR amplitude image; for the N3 contours retained in step S41, use the four end-to-end connected line segments formed by the four vertices corresponding to each contour to traverse the four line segments of each contour one by one; Within the set range of the positive and negative directions of the line segment normal vector, find the point with the largest pixel value in the IR amplitude image as the two-dimensional image position of the top edge of the target to be measured; each contour corresponds to four groups of top edge pixel points, and each group of pixel points is fitted by the least squares method to obtain four straight line equations; then calculate the angle between each two opposite straight lines, and set the angle threshold threshold5. If both angles are less than the angle threshold, the contour is identified as the actual target to be measured, otherwise the contour and its vertices and top equation are deleted; the number of contours retained in the end is N4, which is the number of actual targets to be measured; S43: For the N4 actual measured targets retained in step S42, first calculate the intersection points of the four edge straight lines on the top surface of each target to determine the four upper vertex pixel coordinates of each measured target on the depth image; then combine these upper vertex pixel coordinates with the inverse perspective transformation formula and its top surface equation to calculate the three-dimensional coordinates of each upper vertex of the measured target in the camera coordinate system; finally, use the Euclidean distance formula to calculate the physical distance between the four upper vertices of the top surface of each measured target in the camera three-dimensional coordinate system, and then determine the length and width of each measured target.
6. The multi-target package volume measurement method based on handheld tilted shooting of a depth camera according to claim 5 is characterized in that: In step S5, the specific processing process is as follows: For the N4 measured targets that are finally retained in step S42, their top surface equations and the ground equations supporting them are known; since the two planes represented by the top surface equation and the ground equation are parallel, the height of each measured target is determined by calculating the vertical distance between the top surface equation and the ground equation, and the vertical distance is calculated by the absolute value of the intercept difference of their plane equations.