Box clamping point estimation method and device based on depth camera and electronic equipment
By processing data from a depth camera and performing geometric calculations, irrelevant point clouds on the top surface of the box are removed and a horizontal plane is fitted. Combined with the rectangular base dimensions and straight line fitting, the convenience and stability of the box gripping points in robot handling are solved, and efficient gripping point estimation is achieved.
Patent Information
- Application Number
- CN202511705917.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-20
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2045-11-20
AI Technical Summary
Existing technologies for estimating the gripping points of boxes in robotic handling operations suffer from insufficient convenience and poor stability. In particular, when the box is incomplete, deep learning algorithms are complex and require high computing power.
Data is acquired and preprocessed using a depth camera. Point clouds unrelated to the top surface of the box are removed, a horizontal plane is fitted, and the point clouds are projected onto this plane. The clamping point position of the box is determined by combining the size of the rectangular base and the line fitting.
It achieves convenient and stable estimation of box clamping points, avoids the cumbersome operation of QR code pasting, reduces the dependence on the integrity of box images, reduces the instability of results, and has low computational complexity and low requirements for equipment computing power.
Smart Images

Figure CN121170017B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, and more specifically, relates to a method, apparatus and electronic device for estimating box gripping points based on a depth camera. Background Technology
[0002] When robots perform handling operations in environments such as factories, they rely on accurate gripping point locations to complete automatic gripping operations. Common methods include: (1) pasting QR codes for auxiliary positioning, estimating the gripping point location based on the state of the QR code (the state of the QR code corresponds to the posture of the box); (2) deep learning directly outputs the gripping point location, estimating the target pose and then calculating the gripping point location. For the first method, there is a problem that the target being handled may not be easy to paste QR codes on, which lacks convenience; for the second method, deep learning results are unstable when the box is incomplete in the image. In addition, deep learning algorithms are relatively complex and require high computing power from the algorithm execution device. How to conveniently and stably estimate the gripping point of the box is a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0003] In view of the shortcomings of the prior art, the purpose of this application is to achieve convenient and stable estimation of the box clamping point.
[0004] To achieve the above objectives, in a first aspect, this application provides a method for estimating box gripping points based on a depth camera, comprising:
[0005] Data acquisition and preprocessing are performed using a depth camera to obtain the preprocessed image and preprocessed point cloud of the box. The box is placed on a support platform, which supports the box through a horizontal plane.
[0006] Based on the preprocessed point cloud of the box, plane fitting is performed to determine the horizontal plane, and the projected point cloud is obtained by projecting the point cloud onto the horizontal plane.
[0007] Based on the projected point cloud and the dimensions of the rectangular base of the box, obtain the positions of the four intersection points of the rectangular base of the box;
[0008] The clamping point of the box is estimated based on the height of the clamping point and the position of the four intersection points of the rectangular bottom surface of the box.
[0009] In one possible implementation, the projected point cloud is obtained by projecting the point cloud onto a horizontal plane, including:
[0010] Based on the positional relationship between the horizontal plane and the depth camera, point clouds that are not related to the top surface of the box are removed from the preprocessed point cloud of the box to obtain the point cloud after removal.
[0011] Based on the point cloud after the removal process, the point cloud is projected onto a horizontal plane to obtain the projected point cloud;
[0012] Based on the positional relationship between the horizontal plane and the depth camera, point clouds unrelated to the top surface of the box are removed from the preprocessed point cloud of the box, resulting in a processed point cloud including:
[0013] The point clouds to be removed from the preprocessed point cloud of the box are determined. The point clouds to be removed include the point clouds located on the non-camera side and the point clouds to be removed also include the point clouds located on the camera side and whose distance from the horizontal plane is less than a preset distance. The camera side is the side where the depth camera is located on both sides of the horizontal plane, and the non-camera side is the side opposite to the camera side on both sides of the horizontal plane.
[0014] Remove the point cloud to be removed from the preprocessed point cloud of the box to obtain the point cloud after removal.
[0015] In one possible implementation, the positions of the four intersection points of the rectangular base of the box are obtained based on the projected point cloud and the dimensions of the rectangular base of the box, including:
[0016] Based on the projected point cloud, three lines with the highest probability are obtained through line fitting. The positions of the first and second intersection points are determined based on the three lines with the highest probability. The first and second intersection points are two of the four intersection points of the rectangular bottom surface of the box. The other two intersection points are defined as the third and fourth intersection points.
[0017] Based on the first intersection point, the second intersection point, the dimensions of the rectangular bottom surface of the box, and the preprocessed image of the box, the positions of the third and fourth intersection points are determined.
[0018] In one possible implementation, based on the projected point cloud, three lines with the highest probability are obtained through line fitting, including:
[0019] Based on the projected point cloud, perform the first line fitting to obtain the first fitted line;
[0020] Remove the point cloud distributed on the first fitted line from the projected point cloud to obtain the point cloud used for the second line fitting;
[0021] Based on the point cloud used for the second line fitting, a second line fitting is performed to obtain the second fitted line.
[0022] Remove the point cloud distributed on the second fitted line from the point cloud used for the second line fitting, and obtain the point cloud used for the third line fitting.
[0023] Based on the point cloud used for the third line fitting, a third line fitting is performed to obtain the third fitted line;
[0024] Among them, the first fitted line, the second fitted line, and the third fitted line are the three lines with the highest probability.
[0025] In one possible implementation, the positions of the third and fourth intersection points are determined based on the first intersection point, the second intersection point, the dimensions of the rectangular base of the box, and the preprocessed image of the box, including:
[0026] Based on the length of the reference edge and the dimensions of the rectangular base of the box, the positions of two candidate points are obtained on one side of the reference edge in the horizontal plane, and the positions of two candidate points are obtained on the other side of the reference edge in the horizontal plane. The line segment between the first intersection point and the second intersection point is used as the reference edge.
[0027] Based on the color of the box and the pre-processed image of the box, one side is determined as the target side from the two sides of the reference edge, and the colors of the two candidate points on the target side in the pre-processed image of the box are consistent with the color of the box.
[0028] The positions of the two candidate points on the target side are determined as the positions of the third and fourth intersection points.
[0029] In one possible implementation, the lengths of two adjacent sides in the rectangular bottom surface of the box are X units and Y units, respectively. Side A and side B are the two sides of the reference side in the horizontal plane. The direction perpendicular to the reference side and pointing to side A in the horizontal plane is the first reference direction, and the direction perpendicular to the reference side and pointing to side B in the horizontal plane is the second reference direction.
[0030] Based on the length of the reference side and the dimensions of the rectangular base of the box, the positions of two candidate points are obtained on one side of the reference side in the horizontal plane, and the positions of two candidate points are obtained on the other side of the reference side in the horizontal plane, including:
[0031] With the reference edge having a length of X units: starting from the first intersection point and the second intersection point, move Y units along the first reference direction in the horizontal plane to obtain two candidate points on side A; starting from the first intersection point and the second intersection point, move Y units along the second reference direction in the horizontal plane to obtain two candidate points on side B.
[0032] With the reference edge having a length of Y units: starting from the first intersection point and the second intersection point, move X units along the first reference direction on the horizontal plane to obtain two candidate points on side A; starting from the first intersection point and the second intersection point, move X units along the second reference direction on the horizontal plane to obtain two candidate points on side B.
[0033] In a possible implementation, the lengths of two adjacent sides of the rectangular bottom surface of the box body are X unit lengths and Y unit lengths respectively, the clamping point height is h unit lengths, 0 < h < H, and the vertical distance between the top surface and the bottom surface of the box body is H unit lengths;
[0034] Estimate the clamping points of the box body based on the clamping point height and the positions of the four intersection points of the rectangular bottom surface of the box body, including:
[0035] Based on the positions of the four intersection points of the rectangular bottom surface of the box body, determine the first bottom side and the second bottom side of the box body on the horizontal plane. The first bottom side and the second bottom side are two opposite bottom sides of the rectangular bottom surface of the box body, and the lengths of the first bottom side and the second bottom side are L unit lengths, and ;
[0036] Taking the midpoint of the first bottom side as the starting point, move h unit lengths in the vertical direction pointing to the camera side to obtain the position of the first clamping point; taking the midpoint of the second bottom side as the starting point, move h unit lengths in the vertical direction pointing to the camera side to obtain the position of the second clamping point;
[0037] Among them, the first clamping point and the second clamping point are used as the clamping points of the box body.
[0038] In a possible implementation, through a depth camera, data collection and preprocessing are performed to obtain the preprocessed image and preprocessed point cloud of the box body, including:
[0039] Through the depth camera, collect the first image and the first point cloud of the box body. The first image is an RGB-format image;
[0040] Based on the first image of the box body, perform preprocessing of mapping from RGB three-channel values to HSV three components to obtain the second image of the box body. The second image is an HSV-format image;
[0041] Based on the first point cloud of the box body, perform preprocessing of removing outlier noise points to obtain the second point cloud of the box body;
[0042] Among them, the second image is used as the preprocessed image of the box body, and the second point cloud is used as the preprocessed point cloud of the box body.
[0043] In a second aspect, the present application provides a device for estimating the clamping points of a box body based on a depth camera, including:
[0044] A collection and preprocessing module, configured to perform data collection and preprocessing through a depth camera to obtain the preprocessed image and preprocessed point cloud of the box body. The box body is placed on a carrier, and the carrier carries the box body through a horizontal plane;
[0045] A plane fitting module, configured to perform plane fitting based on the preprocessed point cloud of the box body to determine the horizontal plane;
[0046] The plane fitting and projection module is used to perform plane fitting based on the preprocessed point cloud of the box, determine the horizontal plane, and obtain the projected point cloud by projecting the point cloud onto the horizontal plane.
[0047] The intersection point determination module is used to obtain the positions of the four intersection points of the rectangular bottom surface of the box based on the projected point cloud and the dimensions of the rectangular bottom surface of the box.
[0048] The grip point estimation module is used to estimate the grip point of the box based on the grip point height and the position of the four intersection points of the rectangular bottom surface of the box.
[0049] Thirdly, this application provides an electronic device, including: a memory and one or more processors; the memory is coupled to one or more processors, the memory is used to store computer program code, the computer program code including computer instructions; one or more processors invoke the computer instructions to cause the electronic device to perform the method described in the first aspect or any possible implementation of the first aspect.
[0050] Fourthly, this application provides a computer-readable storage medium including instructions that, when executed on an electronic device, cause the electronic device to perform the method described in the first aspect or any possible implementation thereof.
[0051] Overall, the technical solutions conceived in this application have the following beneficial effects compared with the prior art:
[0052] The horizontal plane of the supporting box is determined by point cloud plane fitting, and point clouds irrelevant to the top surface of the box are removed to ensure that the point cloud data mainly contains the structural features required for subsequent projection. Then, the processed point cloud is projected onto the horizontal plane, and the positions of the four intersection points of the rectangular bottom surface of the box are obtained based on the dimensions of the projected point cloud and the rectangular bottom surface of the box (for example, the three lines with the highest probability are obtained by line fitting, and the two intersection points of the rectangular bottom surface of the box are determined by combining the intersection points of the lines, and the remaining two intersection points are located with the help of the known bottom surface dimensions and image information). Finally, the clamping point is estimated based on the height of the clamping point and the coordinates of the bottom surface intersection points. This method eliminates the need for pasting QR codes, directly calculating gripping points through point cloud geometric features and dimensional constraints, thus avoiding the cumbersome operation of pasting QR codes and improving convenience. Simultaneously, by using plane fitting, point cloud projection, and line intersection to determine the calculations, it reduces the dependence on the integrity of the box image. Compared to deep learning-based pose output methods, it exhibits better robustness and reduces the instability of results caused by incomplete images. Furthermore, the method in this application is primarily based on geometric calculations, resulting in low complexity and lower requirements for device computing power, thereby achieving convenient and stable box gripping point estimation. Attached Figure Description
[0053] Figure 1 This is a flowchart illustrating the box gripping point estimation method based on a depth camera provided in an embodiment of this application.
[0054] Figure 2 This is a schematic diagram illustrating data acquisition via a depth camera, provided in an embodiment of this application.
[0055] Figure 3 This is a schematic diagram of removing point clouds that are not related to the top surface of the box, provided in an embodiment of this application;
[0056] Figure 4 This is a schematic diagram of linear fitting provided in an embodiment of this application;
[0057] Figure 5 This is a schematic diagram of the first placement state provided in an embodiment of this application;
[0058] Figure 6 This is a schematic diagram of the second placement state provided in the embodiments of this application;
[0059] Figure 7 This is a schematic diagram of the third placement state provided in the embodiments of this application;
[0060] Figure 8 This is a schematic diagram of the fourth placement state provided in the embodiments of this application;
[0061] Figure 9 This is a schematic diagram of the estimated box clamping point provided in the embodiments of this application;
[0062] Figure 10 This is a schematic diagram of the structure of the box gripping point estimation device based on a depth camera provided in the embodiments of this application;
[0063] Figure 11 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0064] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0065] In this application, the terms "first" and "second," etc., are used to distinguish different objects, not to describe a specific order of objects. For example, "first intersection point" and "second intersection point," etc., are used to distinguish different intersection points, not to describe a specific order of intersection points.
[0066] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0067] In the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more, for example, multiple processing units means two or more processing units, multiple elements means two or more elements, etc.
[0068] The embodiments of this application are described below with reference to the accompanying drawings.
[0069] Figure 1 This is a flowchart illustrating the box gripping point estimation method based on a depth camera provided in an embodiment of this application, as shown below. Figure 1 As shown, the method includes steps S101 to S107.
[0070] Step S101: Data acquisition and preprocessing are performed using a depth camera to obtain the preprocessed image and preprocessed point cloud of the box.
[0071] Figure 2 This is a schematic diagram illustrating data acquisition via a depth camera, as provided in an embodiment of this application. Figure 2 As shown, the box is placed on a support platform, which supports the box through a horizontal plane, and data can be collected by a depth camera.
[0072] For example, the enclosure is typically a cuboid or cube, with a uniform color throughout. The enclosure is usually placed on a support platform for gripping by a robotic arm. The support platform supports the enclosure via a horizontal plane. The color of the support platform is generally different from the color of the enclosure. Intrinsic and extrinsic parameters are calibrated for the RGB data and point cloud data from the depth camera.
[0073] A depth camera is used to simultaneously acquire RGB images and point cloud information of the enclosure, which are denoted as the first image and the first point cloud, respectively. The first image is an RGB format image. Based on the first image of the enclosure, preprocessing is performed to map the RGB three-channel values to the HSV three-components to obtain the second image of the enclosure, which is an HSV format image. Based on the first point cloud of the enclosure, preprocessing is performed to remove outlier noise points to obtain the second point cloud of the enclosure. The second image serves as the preprocessed image of the enclosure, and the second point cloud serves as the preprocessed point cloud of the enclosure.
[0074] Here is an example of the preprocessing to remove outlier noise points. For instance, Euclidean clustering is performed on the collected point cloud data, and points below a threshold are identified as outlier noise points and removed. The processed point cloud is then recorded as the second point cloud.
[0075] This section provides an illustrative example of the preprocessing described above, which maps RGB three-channel values to HSV three-components. This preprocessing transforms an image from a color model based on red, green, and blue (RGB) to a color model based on hue (H), saturation (S), and lightness (V). This preprocessing separates the color information, color purity information, and brightness information of an image, simplifying color-based image analysis and processing tasks. For example, HSV-formatted images can efficiently identify colors at various locations within the image.
[0076] In the HSV color gamut, it is easy to distinguish the enclosure from its environment (including the platform) by color. It is worth noting that if the enclosure is directly distinguished from its environment using RGB images, it generally requires the use of deep learning technology, which is quite complex.
[0077] Step S102: Based on the preprocessed point cloud of the box, perform plane fitting to determine the horizontal plane.
[0078] Step S103: Based on the positional relationship between the horizontal plane and the depth camera, remove point clouds that are not related to the top surface of the box from the preprocessed point cloud of the box, and obtain the point cloud after removal.
[0079] Figure 3 This is a schematic diagram of removing point clouds unrelated to the top surface of the box, provided in an embodiment of this application. Figure 3 As shown, the point cloud near the top surface of the box is the point cloud related to the top surface of the box, while the point cloud far from the top surface of the box is the point cloud unrelated to the top surface of the box.
[0080] For example, the box is generally placed on a horizontal plane. The pre-processed point cloud is fitted to a plane. Based on the origin position of the pre-processed point cloud (the optical center position of the camera structured light), it is determined which side of the plane the camera is on. Point clouds that are less than 15cm away from the camera on one side (the height of interfering objects on the horizontal plane is generally less than 15cm) and all point clouds on the other side of the camera (non-camera side) are removed to eliminate interference. The processed point cloud is recorded as "point cloud after removal processing".
[0081] Ideally, the preprocessed point cloud will mostly be distributed on a horizontal plane (representing the plane of the support platform), which can be obtained by fitting the data. In non-ideal cases, such as when there is a wall nearby, the preprocessed point cloud may mostly be distributed on the horizontal plane (the desired plane) and a vertical plane. Multiple planes can be fitted, and the horizontal plane can be obtained after removing the vertical plane.
[0082] Step S104: Based on the point cloud after the removal process, project the point cloud onto a horizontal plane to obtain the projected point cloud.
[0083] Step S105: Based on the projected point cloud, obtain three lines with the highest probability through line fitting, and determine the positions of the first and second intersection points based on the three lines with the highest probability (by finding the intersection points between the lines). The first and second intersection points are two of the four intersection points of the rectangular bottom surface of the box (the four intersection points of the rectangular bottom surface of the box are the four intersection points obtained by the intersection of the four sides of the rectangular bottom surface of the box). The other two intersection points are defined as the third and fourth intersection points.
[0084] Figure 4 This is a schematic diagram of linear fitting provided in an embodiment of this application, as shown below. Figure 4 As shown, P1 and P2 are the first and second intersection points, and P3 and P4 are the third and fourth intersection points.
[0085] It should be noted that due to factors such as the uncertainty of camera orientation, the projected image obtained by projecting the point cloud after the removal process onto a horizontal plane may not be a complete rectangle. Based on this, a straight line fitting is performed on the projected point cloud data to fit three lines with the highest probability, which are denoted as the first line, the second line, and the third line. These three lines must have two intersection points, which are denoted as the first intersection point and the second intersection point.
[0086] Step S106: Based on the first intersection point, the second intersection point, the dimensions of the rectangular bottom surface of the box, and the preprocessed image of the box, determine the positions of the third intersection point and the fourth intersection point.
[0087] Step S107: Estimate the clamping point of the box based on the height of the clamping point and the position of the four intersection points of the rectangular bottom surface of the box.
[0088] Understandably, the horizontal plane of the supporting box is determined by point cloud plane fitting, and point clouds irrelevant to the top surface of the box are removed accordingly, ensuring that the point cloud data mainly contains the structural features required for subsequent projection. The processed point cloud is then projected onto the horizontal plane, and three lines with the highest probability are obtained through line fitting. The intersection points of these lines are used to determine the two intersection points of the rectangular bottom surface of the box. The remaining two intersection points are located using known bottom surface dimensions and image information. Finally, the clamping point is estimated based on the height of the clamping point and the coordinates of the bottom surface intersection points. This method eliminates the need for pasting QR codes, directly calculating the clamping points through point cloud geometric features and dimensional constraints, avoiding the cumbersome operation of pasting QR codes and improving convenience. Simultaneously, the calculation through plane fitting, point cloud projection, and line intersection point determination reduces the dependence on the integrity of the box image. Compared to deep learning-based pose output methods, it has better robustness and reduces the instability of results caused by incomplete images. Furthermore, the method in this application is mainly based on geometric operations, with low complexity and lower requirements for device computing power, thus achieving convenient and stable clamping point estimation for the box.
[0089] In one possible implementation, step S103 above, based on the positional relationship between the horizontal plane and the depth camera, removes point clouds unrelated to the top surface of the box from the preprocessed point cloud of the box, obtaining the point cloud after removal, including:
[0090] The point clouds to be removed from the pre-processed point cloud of the box are determined. The point clouds to be removed include those located on the non-camera side and those located on the camera side and whose distance from the horizontal plane is less than a preset distance (e.g., 15cm). The camera side is the side of the horizontal plane where the depth camera is located, and the non-camera side is the side of the horizontal plane that is opposite to the camera side.
[0091] Remove the point cloud to be removed from the preprocessed point cloud of the box to obtain the point cloud after removal.
[0092] like Figure 3 As shown, the camera side is the side where the depth camera is located, and the non-camera side is the side opposite to the camera side.
[0093] Specifically, the space is divided into the camera side (the side where the depth camera is located) and the non-camera side (opposite side) based on the horizontal plane. All point clouds located on the non-camera side are removed to avoid introducing irrelevant data due to box occlusion or background interference. At the same time, the camera side is further filtered to remove point clouds that are less than a preset distance from the horizontal plane. These point clouds usually belong to low structures such as the support platform or the ground, rather than the top surface of the box. Through the double removal mechanism, only point clouds located on the camera side and whose height is significantly higher than the horizontal plane (greater than the preset distance) are retained to ensure that the dataset processed later includes the effective structure of the box top surface, thereby improving the accuracy of point cloud projection and line fitting.
[0094] In one possible implementation, step S105 above, based on the projected point cloud, obtains three lines with the highest probability through line fitting, including:
[0095] Based on the projected point cloud, perform the first line fitting to obtain the first fitted line;
[0096] Remove the point cloud distributed on the first fitted line from the projected point cloud to obtain the point cloud used for the second line fitting;
[0097] Based on the point cloud used for the second line fitting, a second line fitting is performed to obtain the second fitted line.
[0098] Remove the point cloud distributed on the second fitted line from the point cloud used for the second line fitting, and obtain the point cloud used for the third line fitting.
[0099] Based on the point cloud used for the third line fitting, a third line fitting is performed to obtain the third fitted line;
[0100] Among them, the first fitted line, the second fitted line, and the third fitted line are the three lines with the highest probability.
[0101] Specifically, by performing a first straight-line fitting on the projected point cloud, the first fitted line with the highest probability can be obtained, and the points on this line are removed from the original point cloud to avoid duplicate fitting. Then, a second straight-line fitting is performed based on the remaining point cloud to obtain a second fitted line with the second highest probability, and the points on this line are removed again. Finally, a third straight-line fitting is performed on the remaining point cloud to obtain a third fitted line. Through this process of successive fitting and removal, it is ensured that each fitting is based on point clouds that have not participated in previous fittings, thereby avoiding the reuse of the same point cloud by multiple lines, effectively improving the independence and accuracy of the straight-line fitting. The three fitted lines obtained in the end can stably represent the main edge features of the rectangular bottom surface of the box, providing a reliable geometric basis for subsequent intersection point location and grip point estimation.
[0102] In one possible implementation, step S106 above, based on the first intersection point, the second intersection point, the dimensions of the rectangular bottom surface of the box, and the preprocessed image of the box, determines the positions of the third and fourth intersection points, including:
[0103] Based on the length of the reference edge and the dimensions of the rectangular base of the box, the positions of two candidate points are obtained on one side of the reference edge in the horizontal plane, and the positions of two candidate points are obtained on the other side of the reference edge in the horizontal plane. The line segment between the first intersection point and the second intersection point is used as the reference edge.
[0104] Based on the color of the box body and the pre - processed image of the box body, determine one side as the target side among the two sides of the reference edge, and the colors of the two alternative points on the target side are consistent with the color of the box body in the pre - processed image of the box body;
[0105] Determine the positions of the two alternative points on the target side as the positions of the third intersection point and the fourth intersection point.
[0106] It should be noted that taking the line segment between the first intersection point and the second intersection point as the reference edge, according to the known dimensions (such as length and width) of the rectangular bottom surface, the theoretical positions of the two alternative points can be calculated on both sides of the reference edge respectively, ensuring that these points satisfy the geometric relationship of parallel opposite sides and perpendicular adjacent sides of the rectangle; furthermore, by using the color information of the box body and the pixel characteristics of its pre - processed image, compare the consistency between the colors of the areas where the alternative points are located on both sides of the reference edge and the standard color of the box body, so as to screen out the "target side" with matching colors; furthermore, determine the two alternative points on the target side as the actual positions of the third intersection point and the fourth intersection point. This implementation method provides an initial positioning framework through geometric dimensions, and then uses the image color characteristics as verification conditions, effectively avoiding misjudgment of intersection points, enhancing the robustness to the change of the box body's placement posture, and ensuring that the positioning of the four intersection points of the rectangular bottom surface can be completed stably and accurately in a complex environment.
[0107] In a possible implementation, the lengths of two adjacent sides of the rectangular bottom surface of the box body are X unit lengths (for example, 1 unit length is 1 cm) and Y unit lengths respectively (possible situations include X > Y, X < Y, X = Y), A side and B side are the two sides of the reference edge in the horizontal plane, the direction perpendicular to the reference edge and pointing to the A side in the horizontal plane is the first reference direction, and the direction perpendicular to the reference edge and pointing to the B side in the horizontal plane is the second reference direction;
[0108] The above - mentioned method of obtaining the positions of two alternative points on one side of the reference edge in the horizontal plane and the positions of two alternative points on the other side of the reference edge based on the length of the reference edge and the dimensions of the rectangular bottom surface of the box body includes:
[0109] When the length of the reference edge is X unit lengths: taking the first intersection point and the second intersection point as starting points, move Y unit lengths along the first reference direction in the horizontal plane to obtain two alternative points on the A side; taking the first intersection point and the second intersection point as starting points, move Y unit lengths along the second reference direction in the horizontal plane to obtain two alternative points on the B side;
[0110] When the length of the reference edge is Y unit lengths: taking the first intersection point and the second intersection point as starting points, move X unit lengths along the first reference direction in the horizontal plane to obtain two alternative points on the A side; taking the first intersection point and the second intersection point as starting points, move X unit lengths along the second reference direction in the horizontal plane to obtain two alternative points on the B side.
[0111] The reference side mentioned above may be either the long side or the short side of the rectangular base. The box may be located on side A or side B of the reference side. These possible situations correspond to different placement states of the box: first placement state, second placement state, third placement state, and fourth placement state. Examples are provided below.
[0112] For example, suppose X > Y, and the length of the reference side is X units (i.e., the reference side is the longer side of the base of the rectangle), such as Figure 5 and 6 As shown, in this case, the box may be located on side A of the reference edge (first placement state) or side B of the reference edge (second placement state). It is necessary to determine whether the box is located on side A or side B. If the colors of the two candidate points on side A in the preprocessed image of the box are consistent with the color of the box, then side A is the target side, and the positions of the two candidate points on side A are determined as the positions of the third and fourth intersection points. If the colors of the two candidate points on side B in the preprocessed image of the box are consistent with the color of the box, then side B is the target side, and the positions of the two candidate points on side B are determined as the positions of the third and fourth intersection points.
[0113] For example, suppose X>Y, and the length of the reference side is Y units (i.e., the reference side is the shorter side of the rectangle's base), such as Figure 7 and 8 As shown, in this case, the box may be located on side A of the reference edge (third placement state) or side B of the reference edge (fourth placement state). It is necessary to determine whether the box is located on side A or side B. If the colors of the two candidate points on side A in the preprocessed image of the box are consistent with the color of the box, then side A is the target side, and the positions of the two candidate points on side A are determined as the positions of the third and fourth intersection points. If the colors of the two candidate points on side B in the preprocessed image of the box are consistent with the color of the box, then side B is the target side, and the positions of the two candidate points on side B are determined as the positions of the third and fourth intersection points.
[0114] It can be understood that according to the length relationship of adjacent sides of the rectangular bottom surface (the magnitudes of X and Y), the moving distance is adaptively determined: If the length of the reference side is X, starting from the first intersection point and the second intersection point, move Y unit lengths along the first reference direction (pointing to side A) and the second reference direction (pointing to side B) perpendicular to the reference side respectively, to generate two alternative points on each of side A and side B; conversely, if the length of the reference side is Y, the moving distance is adjusted to X unit lengths. This implementation method adaptively selects the moving distance according to the size, ensuring that the alternative points always satisfy the geometric characteristics that the adjacent sides of the rectangle are perpendicular and the opposite sides are parallel. This dynamic calculation method based on size and direction can effectively adapt to boxes with different aspect ratios and different placement states, providing an accurate geometric candidate set for subsequent screening of the target side through color features.
[0115] In a possible implementation, the lengths of two adjacent sides of the rectangular bottom surface of the box are X unit lengths and Y unit lengths respectively, the clamping point height is h unit lengths, 0 < h < H, and the vertical distance between the top surface and the bottom surface of the box (i.e., the height of the box) is H unit lengths (optionally, )
[0116] The above step S107 estimates the clamping points of the box based on the clamping point height and the positions of the four intersection points of the rectangular bottom surface of the box, including:
[0117] Based on the positions of the four intersection points of the rectangular bottom surface of the box, determine the first bottom edge and the second bottom edge of the box on the horizontal plane. The first bottom edge and the second bottom edge are two opposite bottom edges of the rectangular bottom surface of the box, and the lengths of the first bottom edge and the second bottom edge are L unit lengths, and ;
[0118] Taking the midpoint of the first bottom edge as the starting point, move h unit lengths along the vertical direction pointing to the camera side to obtain the position of the first clamping point; taking the midpoint of the second bottom edge as the starting point, move h unit lengths along the vertical direction pointing to the camera side to obtain the position of the second clamping point;
[0119] Among them, the first clamping point and the second clamping point are used as the clamping points of the box.
[0120] It can be understood that based on the positions of the four intersection points of the rectangular bottom surface, identify the length as L ( The two opposite bottom edges of the box are used as the first and second bottom edges to ensure that the gripping point is located on the narrower side of the box (the shorter side of the box) to improve gripping stability. Then, starting from the midpoint of the two bottom edges, a preset gripping height h (usually h = H / 2, where H is the height of the box) is moved in a direction perpendicular to the bottom surface and pointing towards the camera to generate the first and second gripping points respectively. This method reduces the risk of box tilting during gripping by selecting the midpoint of the shortest side as the reference, and at the same time achieves standardized positioning of the gripping point based on height offset. It can adapt to the gripping needs of boxes of different sizes and provide the robot with a stable, symmetrical gripping position that meets the actual operation requirements.
[0121] It should be noted that the positions of the first gripping point and the second gripping point can be the coordinates of the first gripping point and the second gripping point in the camera point cloud coordinate system. Based on the external parameter calibration results of the camera and the robotic arm, the coordinates can be transformed to the robotic arm coordinate system to obtain the coordinates of the first gripping point and the second gripping point in the robotic arm coordinate system.
[0122] Figure 9 This is a schematic diagram of the estimated clamping point of the box provided in the embodiments of this application, such as... Figure 9 As shown, the first and second bottom edges are the shorter sides of the rectangular bottom surface of the box, and the red center point is the midpoint of the shorter side. Starting from the midpoint of the shorter side, move the preset clamping height h in a direction perpendicular to the bottom surface and pointing towards the camera side to obtain the clamping point, that is... Figure 9 The green heart dot in the middle.
[0123] The following describes the box gripping point estimation device based on a depth camera provided in this application. The box gripping point estimation device based on a depth camera described below can be referred to in correspondence with the box gripping point estimation method based on a depth camera described above.
[0124] Figure 10 This is a schematic diagram of the structure of the box gripping point estimation device based on a depth camera provided in the embodiments of this application, as shown below. Figure 10 As shown, the device includes: an acquisition and preprocessing module 10, a plane fitting and projection module 20, an intersection point determination module 30, and a gripping point estimation module 40. Wherein:
[0125] The acquisition and preprocessing module 10 is used to acquire and preprocess data through a depth camera, obtain the preprocessed image and preprocessed point cloud of the box, and place the box on the support platform, which supports the box through a horizontal plane.
[0126] The plane fitting and projection module 20 is used to perform plane fitting based on the preprocessed point cloud of the box, determine the horizontal plane, and obtain the projected point cloud by projecting the point cloud onto the horizontal plane.
[0127] The intersection point determination module 30 is used to obtain the positions of the four intersection points of the rectangular bottom surface of the box based on the projected point cloud and the dimensions of the rectangular bottom surface of the box.
[0128] The gripping point estimation module 40 is used to estimate the gripping point of the box based on the gripping point height and the position of the four intersection points of the rectangular bottom surface of the box.
[0129] It is understood that the detailed functional implementation of each of the above units / modules can be found in the description in the aforementioned method embodiments, and will not be repeated here.
[0130] It should be understood that the above-described device is used to execute the methods in the above embodiments. The implementation principle and technical effect of the corresponding program modules in the device are similar to those described in the above methods. The working process of the device can be referred to the corresponding process in the above methods, and will not be repeated here.
[0131] Based on the methods in the above embodiments, this application provides an electronic device. Figure 11 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application, such as... Figure 11 As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute the methods in the above embodiments.
[0132] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.
[0133] Based on the methods in the above embodiments, this application provides a computer-readable storage medium storing a computer program that, when run on a processor, causes the processor to execute the methods in the above embodiments.
[0134] Based on the methods in the above embodiments, this application provides a computer program product that, when run on a processor, causes the processor to execute the methods in the above embodiments.
[0135] It is understood that the processor in the embodiments of this application can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor can be a microprocessor or any conventional processor.
[0136] The method steps in this application embodiment can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.
[0137] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0138] It is understood that the various numerical designations used in the embodiments of this application are merely for the convenience of description and are not intended to limit the scope of the embodiments of this application.
[0139] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for estimating box gripping points based on a depth camera, characterized in that, include: Data acquisition and preprocessing are performed using a depth camera to obtain the preprocessed image and preprocessed point cloud of the box. The box is placed on a support platform, which supports the box through a horizontal plane. Based on the preprocessed point cloud of the box, plane fitting is performed to determine the horizontal plane, and the projected point cloud is obtained by projecting the point cloud onto the horizontal plane. Based on the projected point cloud and the dimensions of the rectangular base of the box, obtain the positions of the four intersection points of the rectangular base of the box; Estimate the clamping point of the box based on the height of the clamping point and the position of the four intersection points of the rectangular bottom surface of the box. The method of obtaining the positions of the four intersection points of the rectangular base of the box based on the projected point cloud and the dimensions of the rectangular base of the box includes: Based on the projected point cloud, three lines with the highest probability are obtained through line fitting. The positions of the first and second intersection points are determined based on the three lines with the highest probability. The first and second intersection points are two of the four intersection points of the rectangular bottom surface of the box. The other two intersection points are defined as the third and fourth intersection points. Based on the first intersection point, the second intersection point, the dimensions of the rectangular bottom surface of the box, and the preprocessed image of the box, determine the positions of the third and fourth intersection points; The determination of the positions of the third and fourth intersection points based on the first intersection point, the second intersection point, the dimensions of the rectangular bottom surface of the box, and the preprocessed image of the box includes: Based on the length of the reference edge and the dimensions of the rectangular base of the box, the positions of two candidate points are obtained on one side of the reference edge in the horizontal plane, and the positions of two candidate points are obtained on the other side of the reference edge in the horizontal plane. The line segment between the first intersection point and the second intersection point is used as the reference edge. Based on the color of the box and the pre-processed image of the box, one side is determined as the target side from the two sides of the reference edge, and the colors of the two candidate points on the target side in the pre-processed image of the box are consistent with the color of the box. The positions of the two candidate points on the target side are determined as the positions of the third and fourth intersection points; The lengths of two adjacent sides in the rectangular bottom surface of the box are X units and Y units, respectively. Side A and side B are the two sides of the reference side in the horizontal plane. The direction perpendicular to the reference side and pointing to side A in the horizontal plane is the first reference direction, and the direction perpendicular to the reference side and pointing to side B in the horizontal plane is the second reference direction. The method of obtaining the positions of two candidate points on one side of the reference edge and the rectangular base of the box, based on the length of the reference edge and the dimensions of the rectangular base of the box, includes: With the reference edge having a length of X units: starting from the first intersection point and the second intersection point, move Y units along the first reference direction in the horizontal plane to obtain two candidate points on side A; starting from the first intersection point and the second intersection point, move Y units along the second reference direction in the horizontal plane to obtain two candidate points on side B. When the length of the reference edge is Y unit lengths: Taking the first intersection point and the second intersection point as starting points, move X unit lengths along the first reference direction on the horizontal plane to obtain two alternative points on the A side; Taking the first intersection point and the second intersection point as starting points, move X unit lengths along the second reference direction on the horizontal plane to obtain two alternative points on the B side.
2. The box gripping point estimation method based on a depth camera according to claim 1, characterized in that, The obtaining the projected point cloud by projecting the point cloud onto the horizontal plane includes: Based on the positional relationship between the horizontal plane and the depth camera, removing the point cloud irrelevant to the top surface of the box from the pre-processed point cloud of the box to obtain the point cloud after the removal process; Based on the point cloud after the removal process, projecting the point cloud onto the horizontal plane to obtain the projected point cloud; The based on the positional relationship between the horizontal plane and the depth camera, removing the point cloud irrelevant to the top surface of the box from the pre-processed point cloud of the box to obtain the point cloud after the removal process includes: Determining the point cloud to be removed in the pre-processed point cloud of the box. The point cloud to be removed includes the point cloud located on the non-camera side, and the point cloud to be removed also includes the point cloud located on the camera side and with a distance less than the preset distance from the horizontal plane. The camera side is the side where the depth camera is located among the two sides of the horizontal plane, and the non-camera side is the side opposite to the camera side among the two sides of the horizontal plane; Removing the point cloud to be removed from the pre-processed point cloud of the box to obtain the point cloud after the removal process.
3. The box gripping point estimation method based on a depth camera according to claim 1, characterized in that, The obtaining three lines with the highest probability by linear fitting based on the projected point cloud includes: Based on the projected point cloud, performing the first linear fitting to obtain the first fitted line; Removing the point cloud distributed on the first fitted line from the projected point cloud to obtain the point cloud used for the second linear fitting; Based on the point cloud used for the second linear fitting, performing the second linear fitting to obtain the second fitted line; Removing the point cloud distributed on the second fitted line from the point cloud used for the second linear fitting to obtain the point cloud used for the third linear fitting; Based on the point cloud used for the third linear fitting, performing the third linear fitting to obtain the third fitted line; Among them, the first fitted line, the second fitted line, and the third fitted line are used as the three lines with the highest probability.
4. The box gripping point estimation method based on a depth camera according to claim 1, characterized in that, The lengths of two adjacent sides of the rectangular bottom surface of the box are X unit lengths and Y unit lengths respectively, the clamping point height is h unit lengths, 0 < h < H, and the vertical distance between the top surface and the bottom surface of the box is H unit lengths; The estimating the clamping points of the box based on the clamping point height and the positions of the four intersection points of the rectangular bottom surface of the box includes: Based on the positions of the four intersection points of the rectangular base of the box, the first and second base edges of the box are determined on the horizontal plane. The first and second base edges are two opposite base edges of the rectangular base of the box, and the length of the first and second base edges is L units. ; Taking the midpoint of the first bottom side as the starting point, moving h unit lengths along the vertical direction pointing to the camera side to obtain the position of the first clamping point; Taking the midpoint of the second bottom side as the starting point, moving h unit lengths along the vertical direction pointing to the camera side to obtain the position of the second clamping point; Among them, the first clamping point and the second clamping point are used as the clamping points of the box.
5. A device for estimating box gripping points based on a depth camera, characterized in that, It includes: The acquisition and pre-processing module is used to perform data acquisition and pre-processing through the depth camera to obtain the pre-processed image and pre-processed point cloud of the box. The box is placed on the carrier table, and the carrier table bears the box through a horizontal plane; The plane fitting and projection module is used to perform plane fitting based on the preprocessed point cloud of the box, determine the horizontal plane, and obtain the projected point cloud by projecting the point cloud onto the horizontal plane. The intersection point determination module is used to obtain the positions of the four intersection points of the rectangular bottom surface of the box based on the projected point cloud and the dimensions of the rectangular bottom surface of the box. The gripping point estimation module is used to estimate the gripping point of the box based on the gripping point height and the position of the four intersection points of the rectangular bottom surface of the box. The method of obtaining the positions of the four intersection points of the rectangular base of the box based on the projected point cloud and the dimensions of the rectangular base of the box includes: Based on the projected point cloud, three lines with the highest probability are obtained through line fitting. The positions of the first and second intersection points are determined based on the three lines with the highest probability. The first and second intersection points are two of the four intersection points of the rectangular bottom surface of the box. The other two intersection points are defined as the third and fourth intersection points. Based on the first intersection point, the second intersection point, the dimensions of the rectangular bottom surface of the box, and the preprocessed image of the box, determine the positions of the third and fourth intersection points; The determination of the positions of the third and fourth intersection points based on the first intersection point, the second intersection point, the dimensions of the rectangular bottom surface of the box, and the preprocessed image of the box includes: Based on the length of the reference edge and the dimensions of the rectangular base of the box, the positions of two candidate points are obtained on one side of the reference edge in the horizontal plane, and the positions of two candidate points are obtained on the other side of the reference edge in the horizontal plane. The line segment between the first intersection point and the second intersection point is used as the reference edge. Based on the color of the box and the pre-processed image of the box, one side is determined as the target side from the two sides of the reference edge, and the colors of the two candidate points on the target side in the pre-processed image of the box are consistent with the color of the box. The positions of the two candidate points on the target side are determined as the positions of the third and fourth intersection points; The lengths of two adjacent sides in the rectangular bottom surface of the box are X units and Y units, respectively. Side A and side B are the two sides of the reference side in the horizontal plane. The direction perpendicular to the reference side and pointing to side A in the horizontal plane is the first reference direction, and the direction perpendicular to the reference side and pointing to side B in the horizontal plane is the second reference direction. The method of obtaining the positions of two candidate points on one side of the reference edge and the rectangular base of the box, based on the length of the reference edge and the dimensions of the rectangular base of the box, includes: With the reference edge having a length of X units: starting from the first intersection point and the second intersection point, move Y units along the first reference direction in the horizontal plane to obtain two candidate points on side A; starting from the first intersection point and the second intersection point, move Y units along the second reference direction in the horizontal plane to obtain two candidate points on side B. With the reference edge having a length of Y units: starting from the first intersection point and the second intersection point, move X units along the first reference direction on the horizontal plane to obtain two candidate points on side A; starting from the first intersection point and the second intersection point, move X units along the second reference direction on the horizontal plane to obtain two candidate points on side B.
6. An electronic device, characterized in that, include: Memory and one or more processors; The memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code including computer instructions; The one or more processors invoke the computer instructions to cause the electronic device to perform the method as described in any one of claims 1-4.
7. A computer-readable storage medium comprising instructions, characterized in that: When the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the method as described in any one of claims 1-4.
Citation Information
Patent Citations
Industrial material box volume measurement method based on 3D camera
CN117422756A
Pose calculation method and device based on point cloud and article grabbing system
CN119625063A
Carton stack identifying and positioning method and grabbing point determining method based on RGB image and point cloud data
CN120580401A