Carton stack identifying and positioning method and grabbing point determining method based on RGB image and point cloud data
By combining RGB images and point cloud data, using algorithms such as RANSAC and Canny edge detection, the high-precision identification and capture of carton stacks in complex environments is solved, and an efficient and stable automatic unloading system is realized.
Patent Information
- Application Number
- CN202510673323.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-02
AI Technical Summary
Existing visual detection algorithms are poorly adaptable in complex environments, especially sensitive to lighting conditions, background complexity and target occlusion. Learning-based detection algorithms rely on a large amount of training data and computing resources, making it difficult to apply in actual engineering.
Combining RGB images and point cloud data, the point cloud plane is extracted through the RANSAC method, and the camera internal reference is mapped to the RGB image. Combining Canny edge detection and surface growth clustering segmentation algorithm, the front surface information of the independent carton is divided, and the grab point is determined by combining the distance laser to achieve high-precision positioning and grabbing of the carton.
It realizes high-precision identification and stable grabbing of carton stacks in complex scenarios, improves the efficiency and safety of the automatic unloading system, and reduces the computing resource requirements.
Smart Images

Figure CN120580401A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of machine vision. Background Art
[0002] With the rapid development of the logistics industry and the rise of global e-commerce, the scale and complexity of freight transportation have increased exponentially. This change has placed new demands on every aspect of the logistics system, especially in unloading operations, where improving efficiency, accuracy, and safety has become a key focus for companies. For example, box trucks are one of the most common transport vehicles. Boxed cargo is often stacked in complex configurations, and irregular placement and limited cargo space further complicate unloading. Traditional manual unloading methods require a large workforce, are slow, and are prone to safety issues such as damage to cargo and personal injury. This operating model no longer meets the modern logistics industry's demands for efficiency, cost-effectiveness, and safety.
[0003] Faced with these challenges, developing efficient and intelligent unloading solutions has become a pressing task for the logistics industry. Existing hardware solutions for automated unloading technologies typically center around robotic arms. Through multi-degree-of-freedom design, flexible adaptation of end-effectors, and force feedback technology, these solutions can achieve high unloading efficiency in standardized scenarios. However, these solutions have limited adaptability and often struggle to cope with complex stacking or dynamically changing scenarios. Therefore, intelligent algorithms based on computer vision, due to their flexibility and efficiency, have become an important direction for improving the adaptability of automated unloading technology.
[0004] Traditional visual inspection algorithms primarily rely on hand-crafted features and geometric methods, with their core focus on extracting feature information about target objects through image processing techniques. Classic methods include the combination of HOG (Histogram of Oriented Gradients) and support vector machines (SVM), as well as template matching techniques. These algorithms analyze image gradients, edges, and textures to detect and demarcate targets. In logistics scenarios, template matching can effectively detect cartons with regular shapes and distinct features. The advantages of traditional algorithms lie in their relatively low computational complexity, reduced reliance on small datasets, and high interpretability. However, these algorithms have poor adaptability in complex scenarios, particularly when lighting conditions change, targets are obscured, or backgrounds are complex. Detection performance degrades significantly, failing to meet the high precision and robustness requirements of modern logistics applications.
[0005] Learning-based visual inspection algorithms, which have emerged in recent years, automatically learn target features through deep neural networks, significantly improving detection performance and applicability. These methods, based on convolutional neural networks (CNNs), have spawned efficient frameworks such as YOLO (You Only Look Once) and Faster R-CNN. YOLO completes target detection with a single forward propagation, making it suitable for real-time response scenarios, while Faster R-CNN improves target detection accuracy through a region proposal network (RPN). In the inspection of complex stacked cartons, YOLO can efficiently identify the boundaries and categories of multiple targets without the need for pre-built templates. However, these methods rely heavily on large-scale, high-quality annotated data, and model training and inference require significant computing resources. In practical applications, their generalization capabilities may be limited when the distribution of training data differs significantly from that of the target scene.
[0006] Furthermore, based on the input data type, the aforementioned visual detection algorithms can be further categorized as those that use RGB images as input and those that use point cloud data as input. Each has its own advantages in terms of input data type and application scenarios. Detection methods based on RGB images use 2D images as input and can capture rich texture, color, and lighting information. Traditional algorithms such as HOG-SVM achieve object detection by extracting gradient information, while the deep learning-based YOLO achieves efficient detection in complex scenes through end-to-end training. These methods perform well in conditions with good lighting and distinct target texture features. However, RGB images lack depth information, making it difficult to accurately resolve the 3D position of targets, especially in occluded and complex stacked scenes. Detection methods based on point cloud data are based on 3D spatial data and can directly utilize the geometric shape and spatial distribution of target objects. Traditional point cloud processing algorithms such as RANSAC segment target regions by fitting geometric models, while the deep learning-based PointNet directly extracts features from point cloud data, achieving efficient 3D detection. Point cloud algorithms have significant advantages in stacked scenes with severe occlusion, and can restore the 3D structure of targets through multi-view fusion. However, point cloud data processing has high computational complexity, relatively expensive hardware costs, and cannot provide detailed information such as texture and color, which limits its application in certain specific scenarios.
[0007] In summary, the existing visual inspection algorithms have the following shortcomings:
[0008] Traditional detection algorithms have limited adaptability in dynamic and complex environments and are sensitive to lighting conditions, background complexity, and object occlusion. Learning-based detection algorithms rely on large amounts of training data and require high computing resources, making them imperfect for practical engineering applications.
[0009] In terms of input data types, RGB images lack depth information, making it difficult to accurately interpret the three-dimensional spatial position of goods. They also provide insufficient information in scenes with severe occlusion or complex stacking. Point cloud data, on the other hand, lacks texture and color information, significantly limiting the algorithm's ability to identify and detect goods. Summary of the Invention
[0010] The present invention aims to solve the problems of poor accuracy in the existing three-dimensional spatial position analysis of goods and poor stability and accuracy in carton grabbing. A carton stack recognition and positioning method based on RGB images and point cloud data is now provided.
[0011] The carton stack recognition and positioning method based on RGB images and point cloud data of the present invention includes:
[0012] Step 1: Use a depth camera to obtain RGB image information and point cloud information of the mixed cardboard stack, and preprocess the point cloud information to obtain preprocessed point cloud data;
[0013] Step 2: Use the parameter-limited RANSAC method to extract the plane in the depth direction of the pre-processed point cloud data to obtain the point cloud plane closest to the camera, and then pre-process the point cloud plane; obtain the pre-processed point cloud plane P first_inlier ;
[0014] Step 3: Based on the camera intrinsic parameter matrix K, the point cloud plane P first_inlier Map it to the two-dimensional pixel coordinates of the RGB image to obtain the point set S0; based on the boundary range of the point set S0, establish a rectangular region ROI, perform ROI range clipping in the RGB image, use the Canny edge detection algorithm to extract the boundary point set S1 of the clipped RGB image, and inversely map the point set S1 back to the plane point cloud P first_inier After filtering out the boundary of the point cloud plane mapped back to the point cloud, we can obtain the point cloud P containing the front surface information of multiple independent cartons. cam ;
[0015] Step 4: Use a clustering segmentation algorithm based on surface growth to cluster the point cloud P containing the front surface information of multiple independent cartons. cam Split into multiple independent point cloud clusters;
[0016] Step 5: Create a minimum cube outside each independent point cloud cluster that can contain all the points in the point cloud cluster; use the spatial coordinates and side length of the minimum cube as the plane closest to the camera as the spatial position information of the carton.
[0017] Furthermore, in the present invention, in step 1, the point cloud information is preprocessed, and the method for the preprocessed point cloud information is:
[0018] The point cloud information is sequentially subjected to point cloud preprocessing operations including voxel filtering downsampling, pass filtering, statistical filtering and Kriging interpolation.
[0019] Furthermore, in the present invention, in step 2, the RANSAC method based on parameter limitation is used to extract the plane in the depth direction of the point cloud information, and the process of obtaining the point cloud plane closest to the camera is as follows:
[0020] Step S1: Use the point cloud data preprocessed in step 1 as a point cloud model, and randomly select three points in the point cloud model to construct a plane;
[0021] Step S2: Calculate all points p in the point cloud model i Distance to the plane constructed in step S1 The distance Compare with the preset distance threshold ∈; when it meets Point p i If the number of points is greater than M, a point cloud plane is determined using M points;
[0022] Step S3: Remove the point cloud plane determined in step S2 from the point cloud data preprocessed in step 1, use the remaining point cloud data as the point cloud model, randomly select three points from the point cloud model to construct a plane, and return to step S2 until all points in the point cloud model are traversed to obtain N point cloud planes.
[0023] Step S4: Calculate the distances from the center point of the depth camera to the N point cloud planes, and obtain the point cloud plane P closest to the center point of the depth camera. first_inlier .
[0024] Furthermore, in the present invention, in step 3, based on the camera intrinsic parameter matrix K, the point cloud plane P first_inier Mapping to the pixel coordinates of the two-dimensional RGB image, the method to obtain the point set S0 is:
[0025] Using the camera intrinsic parameter matrix K, the point cloud plane P first_inlier Map the three-dimensional points of the camera coordinate system to the pixel coordinates of the two-dimensional RGB image to obtain the point set S0;
[0026] S0={(u i ,v i )|i=1,2,3...}
[0027]
[0028] Among them, f x ,f y is the focal length of the camera in the x and y axis directions of the camera coordinate system, c x ,c yis the principal point coordinate of the image, u i ,v i They represent the horizontal and vertical coordinates of the i-th point cloud point in the two-dimensional RGB image in the point cloud coordinate system, respectively. i 、y i 、z i They represent the three-dimensional coordinates of the i-th point cloud point in the camera coordinate system.
[0029] Furthermore, in the present invention, in step 3, the Canny edge detection algorithm is used to extract the boundary of the cropped RGB image as follows:
[0030] First, detect the boundary of the point set S0 and obtain the region ROI = [u min ,u max ]×[v min ,v max ];
[0031] Among them, u min Indicates the minimum horizontal coordinate value of the point set S0 in the point cloud coordinate system, u max Indicates the maximum horizontal coordinate of the point set S0 in the point cloud coordinate system, v min Represents the minimum vertical coordinate v of the point set S0 in the point cloud coordinate system max Indicates the maximum longitudinal coordinate of the point set S0 in the point cloud coordinate system;
[0032] Reuse the formula:
[0033] E(u,v)=Cany(I ROI (u,v))
[0034]
[0035] Calculate the binarization result E(u,v), and use the binarization result E(u,v) to extract the boundary point set S1, I of the RGB image ROI (u,v) represents the RGB image after cropping based on the boundary range.
[0036] Furthermore, in the present invention, in step 3, the boundary is inversely mapped to the plane point cloud P first_inlier After filtering out the boundary point cloud, we obtain the point cloud P containing the front surface information of multiple independent cartons. cam The process is:
[0037] Set the camera center to the point cloud plane P first_inlier The distance is used as the plane depth Distance first_inlier , combined with the camera intrinsic parameter matrix, the boundary point set S1 of the RGB image is inversely mapped back to the camera coordinate system of the three-dimensional point cloud to obtain the three-dimensional point set P boundary , with a three-dimensional point set Pboundary Each point in is taken as the origin and the radius r is established b The spherical space of the point cloud plane P preprocessed in step 2 first_inlier The points whose coordinates are in the spherical space are marked as boundary points. first_inlier Eliminate all boundary points and obtain the point cloud P containing the front surface information of multiple independent cartons cam .
[0038] Furthermore, in the present invention, in step 4, a clustering segmentation algorithm based on surface growth is used to cluster the point cloud P containing the front surface information of multiple independent cartons. cam The method of segmenting into multiple independent point cloud clusters is:
[0039] Step S41: From point cloud P cam Select any point P i , search for point P i The neighborhood centered on is calculated by calculating the ratio of the minimum eigenvalue of the neighborhood covariance matrix to the sum of the eigenvalues, and the ratio is used as the curvature κ i ;
[0040] At the same time, the covariance matrix of the neighborhood is used to calculate the point P i Normal vector n at the center i , and then calculate the point P i Any point p in the neighborhood j The covariance matrix of the neighborhood centered on point p is then calculated to obtain j Normal vector n at the center j ; Calculate the normal vector n i and normal vector n j The angle between
[0041] Step S42: Set parameter smoothness threshold θ th and curvature threshold κ th ; When θ is satisfied ij <θ th , κ i <κ th When P i Neighborhood point p j Point P i Class, traverse P i All points in the neighborhood, get the point p i Starting point cloud cluster;
[0042] Step S43: From the point cloud P cam Select the unclassified point P k , execute step S41 and step S42 until the point cloud P cam All points are classified, and the point cloud P containing the front surface information of multiple independent cartons is completed.cam Split into multiple independent point cloud clusters.
[0043] Furthermore, in the present invention, in step S41, point P i Normal vector n at the center i The method is similar to the point p j Normal vector n at the center j The same method is used, click P i Normal vector n at the center i The method is:
[0044] Point P i Perform the nearest search for the center and obtain the neighborhood N i , calculate the neighborhood N i The center of mass Neighborhood covariance matrix C i ;
[0045] Using the neighborhood N i The center of mass For neighborhood N i Perform eigenvalue decomposition on the covariance matrix of the neighborhood N i The eigenvalues and eigenvectors of the covariance matrix;
[0046] The eigenvector of the minimum eigenvalue is taken as the i Normal vector n i .
[0047] A method for determining grasping points of a carton stack based on RGB images and point cloud data, which is implemented based on the above-mentioned carton stack recognition and positioning method based on RGB images and point cloud data; specifically, it includes:
[0048] A distance laser is set at the end of the robot's grabbing arm, and the distance D from the end of the grabbing arm to the carton surface is measured using the distance laser. α ;
[0049] Using the distance D α And the height of the carton, calculate the center point of the carton;
[0050] When the height of the carton meets the height y of the carton in the world coordinate system b +y cam >Preset parameter D ε When the center point is: center(x b ,y b , z b );
[0051] When the height of the carton meets the height y of the carton in the world coordinate system b +y cam <Default parameter D εWhen, the center point is: center_high(x b , y b + D_set, z b + D0 / 2);
[0052] Where, x b , y b , z b represent the spatial coordinates of the center point in the camera coordinate system, y cam is the camera height, D_set is the preset distance for top suction planning, and D0 is the distance between the nearest plane of the camera and the second point cloud plane;
[0053] Through the matrix R between the camera and the robotic arm, the center point or the center point height is converted into the spatial point P in the base coordinate system of the robotic arm a , and the spatial point P in the robotic arm coordinate system a is used as the grasping point.
[0054] Further, in the present invention, the spatial point P in the robotic arm coordinate system a is:
[0055] P a = R·P center + t
[0056] Where, R is the rotation matrix, t is the translation vector, and P center is the grasping point coordinate in the world coordinate system.
[0057] Further, in the present invention, the calculation method of the distance D0 between the nearest plane of the camera and the second point cloud plane:
[0058] Based on the camera internal parameter matrix K and the point cloud plane closest to the camera, the point cloud plane closest to the camera is deleted from the preprocessed point cloud data to obtain the remaining point cloud data, and the RANSAC method based on parameter limitation is used again to extract the second point cloud plane in the depth direction from the remaining point cloud data, and the distance between the extracted second point cloud plane and the camera closest plane satisfies: Pmin < Distance i - Min_Distance < Pmax, where Pmin and Pmax are the preset maximum distance threshold and minimum threshold, and at the same time, the number of points in the second point cloud plane is greater than the number threshold, and the distance D0 between the point cloud plane closest to the camera and the second point cloud plane is calculated.
[0059] This invention uses RGB images and point cloud data as input, leveraging both the surface feature information of the carton stack provided by the RGB images and the three-dimensional spatial depth information provided by the point cloud data. By combining the advantages of both, it achieves a more comprehensive characterization of the target object and effectively ensures the accuracy of positioning information recognition. Combined with the corresponding automatic unloading robot hardware system and software platform, this invention provides a stable, high-precision, and efficient solution for detecting and grasping carton cargo targets, enabling intelligent control of the entire system. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 The figure is a flow chart of the method of the present invention. DETAILED DESCRIPTION
[0061] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. It should be noted that the embodiments of the present invention and the features in the embodiments can be combined with each other in the absence of conflict.
[0062] Specific implementation method 1: refer to Figure 1 Specifically describing this embodiment, the carton stack recognition and positioning method based on RGB images and point cloud data described in this embodiment includes:
[0063] Step 1: Use a depth camera to obtain RGB image information and point cloud information of the mixed cardboard stack, and preprocess the point cloud information to obtain preprocessed point cloud data;
[0064] Step 2: Use the parameter-limited RANSAC method to extract the plane in the depth direction of the pre-processed point cloud data to obtain the point cloud plane closest to the camera, and then pre-process the point cloud plane; obtain the pre-processed point cloud plane P first_inlier ;
[0065] Step 3: Based on the camera intrinsic parameter matrix K, the point cloud plane P first_inlier Map it to the two-dimensional pixel coordinates of the RGB image to obtain the point set S0; based on the boundary range of the point set S0, establish a rectangular region ROI, perform ROI range clipping in the RGB image, use the Canny edge detection algorithm to extract the boundary point set S1 of the clipped RGB image, and inversely map the point set S1 back to the plane point cloud P first_inlier After filtering out the boundary of the point cloud plane mapped back to the point cloud, we can obtain the point cloud P containing the front surface information of multiple independent cartons. cam ;
[0066] Step 4: Use a clustering segmentation algorithm based on surface growth to cluster the point cloud P containing the front surface information of multiple independent cartons. cam Split into multiple independent point cloud clusters;
[0067] Step 5: Create a minimum cube outside each independent point cloud cluster that can contain all the points in the point cloud cluster; use the spatial coordinates and side length of the minimum cube as the plane closest to the camera as the spatial position information of the carton.
[0068] Furthermore, in this embodiment, in step 1, the point cloud information is preprocessed, and the method for the preprocessed point cloud information is:
[0069] The point cloud information is sequentially subjected to point cloud preprocessing operations including voxel filtering downsampling, pass filtering, statistical filtering and Kriging interpolation.
[0070] Furthermore, in this embodiment, in step 2, the RANSAC method based on parameter limitation is used to extract the plane in the depth direction of the point cloud information, and the process of obtaining the point cloud plane closest to the camera is as follows:
[0071] Step S1: Use the point cloud data preprocessed in step 1 as a point cloud model, and randomly select three points in the point cloud model to construct a plane;
[0072] Step S2: Use the pre-processed point cloud data as a point cloud model, and transform all points p in the point cloud model into i Distance to the plane The distance Compare with the preset distance threshold ∈; when it meets Point p i If the number of is greater than M, a point cloud plane is determined; M is adjusted according to the number of points in the point cloud data, and is generally greater than 3% of the number of points in the current point cloud model;
[0073] Step S3: Remove the point cloud plane determined in step S2 from the point cloud data preprocessed in step 1, use the remaining point cloud data as the point cloud model, randomly select three points from the point cloud model to construct a plane, and return to step S2 until all points in the point cloud model are traversed to obtain N point cloud planes.
[0074] Step S4: Calculate the distances from the center point of the depth camera to the N point cloud planes, and obtain the point cloud plane P closest to the center point of the depth camera. first_inlier .
[0075] Furthermore, in this embodiment, in step 3, based on the camera intrinsic parameter matrix K, the point cloud plane P first_inlierMapping to the pixel coordinates of the two-dimensional RGB image, the method to obtain the point set S0 is:
[0076] Using the camera intrinsic parameter matrix K, the point cloud plane is mapped from the three-dimensional point of the camera coordinate system to the pixel coordinates of the two-dimensional RGB image to obtain the point set S0;
[0077] S0={(u i ,v i )|i=1,2,3...}
[0078]
[0079] Among them, f x ,f y is the focal length of the camera in the x and y axis directions of the camera coordinate system, c x ,c y is the principal point coordinate of the image, u i ,v i They represent the horizontal and vertical coordinates of the i-th point cloud point in the two-dimensional RGB image in the point cloud coordinate system, respectively. i 、y i 、z i They represent the three-dimensional coordinates of the i-th point cloud point in the camera coordinate system.
[0080] Furthermore, in this embodiment, in step 3, the Canny edge detection algorithm is used to extract the boundary of the cropped RGB image as follows:
[0081] First, detect the boundary of the point set S0 and obtain the region ROI = [u min ,u max ]×[v min ,v max ];
[0082] Among them, u min Indicates the minimum horizontal coordinate value of the point set S0 in the point cloud coordinate system, u max Indicates the maximum horizontal coordinate of the point set S0 in the point cloud coordinate system, v min Represents the minimum vertical coordinate v of the point set S0 in the point cloud coordinate system max Indicates the maximum longitudinal coordinate of the point set S0 in the point cloud coordinate system;
[0083] Reuse the formula:
[0084] E(u,v)=Canny(I ROI (u,v))
[0085]
[0086] Calculate the binarization result E(u,v), and use the binarization result E(u,v) to extract the boundary point set S1, I of the RGB image ROI (u,v) represents the RGB image after cropping based on the boundary range.
[0087] Furthermore, in this embodiment, in step 3, the boundary is inversely mapped to the plane point cloud P firsr_inlier After filtering out the boundary point cloud, we obtain the point cloud P containing the front surface information of multiple independent cartons. cam The process is:
[0088] Set the camera center to the point cloud plane P first_inlier The distance is used as the plane depth Distance first_inlier , combined with the camera intrinsic parameter matrix, the boundary point set S1 of the RGB image is inversely mapped back to the camera coordinate system of the three-dimensional point cloud to obtain the three-dimensional point set P boundary , with a three-dimensional point set P boundary Each point in is taken as the origin and the radius r is established b The spherical space of the point cloud plane P preprocessed in step 2 first_inlier The points whose coordinates are in the spherical space are marked as boundary points. first_inlier Eliminate all boundary points and obtain the point cloud P containing the front surface information of multiple independent cartons cam .
[0089] Furthermore, in this embodiment, in step 4, a clustering segmentation algorithm based on surface growth is used to cluster the point cloud P containing the front surface information of multiple independent cartons. cam The method of segmenting into multiple independent point cloud clusters is:
[0090] Step S41: From point cloud P cam Select any point P i , search for point P i The ratio of the minimum eigenvalue of the covariance matrix of the neighborhood to the sum of the eigenvalues is calculated, and the ratio is used as the curvature κ i ;
[0091] At the same time, the covariance matrix of the neighborhood is used to calculate the point P i Normal vector n at the center i , and then calculate the point P i Any point p in the neighborhood j The covariance matrix of the neighborhood centered on point p is then calculated to obtain j Normal vector n at the center j ; Calculate the normal vector n i and normal vector n j The angle between
[0092] Step S42: Set parameter smoothness threshold θ th and curvature threshold κ th ; When θ is satisfied ij <θ th ,κ i <κ th When P i Neighborhood point p j Point P i Class, traverse P i All points in the neighborhood, get the point p i Starting point cloud cluster;
[0093] Step S43: From the point cloud P cam Select the unclassified point P k , execute step S41 and step S42 until the point cloud P cam All points are classified, and the point cloud P containing the front surface information of multiple independent cartons is completed. cam Split into multiple independent point cloud clusters.
[0094] Furthermore, in this embodiment, point P i Normal vector n at the center i The method is similar to the point p j Normal vector n at the center j The same method is used, click P i Normal vector n at the center i The method is:
[0095] Point P i Perform the nearest search for the center and obtain the neighborhood N i , calculate the neighborhood N i The center of mass Neighborhood covariance matrix C i ;
[0096] Using the neighborhood N i The center of mass For neighborhood N i Perform eigenvalue decomposition on the covariance matrix of the neighborhood N i The eigenvalues and eigenvectors of the covariance matrix;
[0097] The eigenvector of the minimum eigenvalue is taken as the i Normal vector n i .
[0098] Furthermore, in this embodiment, the normal vector n is calculated i and normal vector n j The formula for the angle is:
[0099] θ ij =arccos(ni ·n j )
[0100] Among them, θ ij is the normal vector n i and normal vector n j Angle.
[0101] Furthermore, in this embodiment, the curvature κ i :
[0102] κ i =λ3 / (λ1+λ2+λ3)
[0103] Among them, λ1, λ2, λ3 are the covariance matrix C i The three eigenvalues λ1, λ2, and λ3 obtained by eigenvalue decomposition, among which λ3 is the minimum eigenvalue.
[0104] A method for determining grasping points of a carton stack based on RGB images and point cloud data, which is implemented based on the above-mentioned carton stack recognition and positioning method based on RGB images and point cloud data; specifically, it includes:
[0105] A distance laser is set at the end of the grabbing arm, and the distance D from the end of the grabbing arm to the surface of the carton is measured using the distance laser. α ;
[0106] Using the distance D α And the height of the carton, calculate the center point of the carton;
[0107] When the height of the carton meets the height y of the carton in the world coordinate system b +y cam >Default parameter D ε When the center point is: center(x b ,y b ,z b );
[0108] When the height of the carton meets the height y of the carton in the world coordinate system b +y cam <Default parameter D ε When the center point is: center_high(x b ,y b +D_set,z b +D0 / 2);
[0109] Among them, x b ,y b ,z b Represents the spatial coordinates of the center point in the camera coordinate system, y camis the camera height, D_set is the preset distance for top suction planning, and D0 is the distance between the nearest plane of the camera and the second point cloud plane;
[0110] Convert the center point or the center point height into the space point P in the base coordinate system of the robotic arm through the matrix R between the camera and the robotic arm a , and use the space point P in the coordinate system of the robotic arm a as the grasping point.
[0111] Furthermore, in this embodiment, the space point P in the coordinate system of the robotic arm a is:
[0112] P a = R·P center + t
[0113] where R is the rotation matrix, t is the translation vector, and P center is the grasping point coordinate in the world coordinate system.
[0114] Furthermore, in this embodiment, the calculation method of the distance D0 between the nearest plane of the camera and the second point cloud plane:
[0115] Based on the camera intrinsic matrix K and the point cloud plane closest to the camera, delete the point cloud plane closest to the camera from the preprocessed point cloud data, obtain the remaining point cloud data, and then use the RANSAC method based on parameter limitation to extract the second point cloud plane in the depth direction from the remaining point cloud data, and the distance between the extracted second point cloud plane and the camera nearest plane satisfies: Pmin < Distance i - Min_Distance < Pmax, where Pmin and Pmax are the preset maximum and minimum distance thresholds, and at the same time, the number of points in the second point cloud plane is greater than the number threshold, and calculate the distance D0 between the point cloud plane closest to the camera and the second point cloud plane. The number threshold is generally 3% of the current point cloud number.
[0116] The present invention is applied to the equipment system of an automatic unloading robot for carton goods in a van truck, and proposes an accurate and efficient visual algorithm for detecting carton goods with universality. This algorithm provides a way of identifying, positioning, and grasping planning for carton goods based on the vision system for the automatic unloading robot, helping the automatic unloading robot in the working environment of the carriage to face the complex mixed stack of carton goods, filter out background information and noise, and accurately perform spatial positioning and reasonable grasping of the carton goods to be grasped at different positions and poses.
[0117] The present invention uses a TOF (Time of Flight) depth camera capable of capturing global information to obtain the RGB image information I of the mixed stack of cartons RAnd point cloud information P = {p i |i=1.2.3...}, indicating I R The pixel coordinates are (u, v), and the point cloud coordinates are p i =(X i ,Y i ,Z i It is worth noting that since the point cloud acquisition FOV of the camera model used is larger than the RGB image acquisition FOV, range clipping is required in the subsequent coordinate conversion process. In addition, the present invention installs a laser sensor at the end gripping device of the robot arm to measure the distance D from the end gripping device to the carton surface. α .
[0118] The traditional target detection algorithm uses RGB image information I R , point cloud information P and laser distance value D α The system takes the mixed stack of cartons inside the carriage as input, performs carton contour detection and spatial position calculation on the mixed stack of cartons inside the carriage, and outputs the spatial position of the final carton target to be grasped and the corresponding grasping method through the preset discrimination algorithm.
[0119] Affected by the complex spatial environment and noise, the present invention, based on the OpenCV and PCL visual libraries, first performs point cloud preprocessing operations on the input information, including voxel filtering downsampling, pass-through filtering, statistical filtering, and Kriging interpolation. To a certain extent, the preprocessing reduces the computational cost while retaining the basic initial features of the input information, and greatly reduces the complex patterns on the surface of the carton. Packaging information such as highlighted objects has an adverse effect on the algorithm's discrimination results. Next, the present invention uses the improved RANSAC method to perform plane extraction in the depth direction of the point cloud information, and performs secondary preprocessing operations such as radius filtering. Then, by combining the point cloud plane information with the detailed features in the same plane of the RGB image, the boundary lines between the cartons are identified, and with the help of repeated mapping of the 2D image space and the 3D point cloud space, the spatial position information of all the cartons on the outermost side of the current carton stack (i.e., the column to be grasped) is extracted. Finally, according to the grasping priority and grasping mode discrimination method, the target detection results required by the automatic unloading robot system are obtained.
[0120] The specific research contents are as follows:
[0121] Step 1: preprocessing;
[0122] Since the imaging range of the TOF depth camera is larger than the target range of the carton goods, the collected raw point cloud data will contain a large amount of sparse, disordered and useless information, mainly including:
[0123] (1) Background clutter information: the inner wall of the car, part of the robot body structure and the surface information of the rear carton.
[0124] (2) Box side information: Affected by the camera posture and viewing angle, the side information of some cartons will be collected.
[0125] (3) Noise point cloud information: The scanning of the depth camera is based on laser reflection. Due to the instability of the TOF depth camera itself and the influence of the complex working environment, a certain amount of chaotic and disordered noise point cloud will be collected.
[0126] Furthermore, complex patterns on the carton surface and reflective objects like tape can cause some point cloud data to be lost, creating a certain degree of point cloud voids. The object detection algorithm's speed needs to be controlled, otherwise it will affect the unloading robot's efficiency. To address these issues, the present invention preprocesses the raw data sequentially.
[0127] Voxel filtering and downsampling: The original point cloud information P collected by the camera is mainly the front surface of the mixed cardboard stack, which is mostly plane information and can be simplified to improve the calculation speed. Set the downsampled point cloud data size n1 and divide the point cloud space into multiple small cubes (voxels). All points in each voxel are represented by a representative point O(x o ,y o ,z o ) is used instead, where the representative point is obtained using the centroid method:
[0128] n1≤t=I(x)·I(y)·I(z)
[0129]
[0130] Among them, x max ,x min ,y max ,y min ,z max ,z min are the extreme coordinates of the three dimensions of the point cloud data, a, b, c are the side lengths of the divided voxels, t is the total prime number, INT is the upward rounding function, and k is the number of data points contained in each voxel.
[0131] Through filtering: Introduce six parameters in the filter processing algorithm: X min ,X max ,Y min ,Y max ,Z min ,Z max , set the range in the three dimensions of the spatial coordinate system, traverse all points through the threshold of the corresponding dimension, and filter the point cloud data in the spatial position:
[0132]
[0133] This range setting and data filtering can significantly reduce the impact of background environmental information in point cloud data, ensuring that new data is mainly within the target working range.
[0134] Statistical filtering: After the above processing, in order to further remove the residual background information and discrete noise information in the point cloud data P2, the optimization processing is performed based on the statistical filter. Construct the Kd-Tree model of the point cloud data P2, and for each point p i (x i ,y i ,z i ) Find the k points closest to it {p j ;j=1.2.3..k};p j (x j ,y j ,z j ), calculate p i The average distance d between neighboring points i :
[0135]
[0136] Traverse all points in the point cloud and calculate d i Average value And calculate the standard deviation σ based on them, based on the calculation results and the preset threshold T t1 Filter the point cloud data P2, where Adopt the trust region idea to ensure that the main data is preserved. and σ can be calculated:
[0137]
[0138] Kriging interpolation: Factors such as complex patterns on the carton surface and high-brightness items can cause camera laser divergence and point cloud loss, resulting in point cloud holes. Assuming that there is some correlation between the values of adjacent points, for each pair of points x in the point cloud data i and x j , calculate the Euclidean distance h between them e And their corresponding values z(x i ) and z(x j ) between the squares of the difference [z(x i )-z(x j )] 2 (Use the depth value as the point x i The attribute value z(x i )). Based on the relatively smooth characteristics of the carton surface point cloud data, the present invention constructs the empirical semivariogram function γ e (h e ), and select the Gaussian model for the semivariogram γg (h g ) Fitting:
[0139]
[0140] Where N(h0) is the number of point pairs with distance h0, C0 represents the nugget value, C represents the sill value, and a represents the range.
[0141] Next, we construct the Kriging equation calculation and use the linear equation system to determine the linear weight λ of each known point around the point to be interpolated. j , and then get the value of the interpolation point. Among them, the confidence of the interpolation result can be evaluated by the error variance, which facilitates the subsequent adjustment of the algorithm.
[0142] Step 2. Plane extraction and secondary preprocessing;
[0143] The pre-processed point cloud information P4 is a high-quality spatial point cloud containing information on multiple plane surfaces of a mixed stack of cartons. In order to distinguish the front surface information and side surface information of cartons at different depth planes, the present invention performs plane segmentation extraction on the point cloud based on the depth direction. RANSAC is a commonly used robust estimation algorithm suitable for estimating model parameters in a data set containing a large number of outliers. The algorithm extracts planes or other geometric models from point cloud data through random sampling. The RANSAC algorithm is optimized to extract information on the front surface of cartons at multiple depth dimensions and the distance between planes.
[0144] Randomly select three local points p1, p2, and p3 from the point cloud data P4 as initial values and define a unique plane a i x+b i y+c i z+d i =0(a i ,b i ,c i is the component of the normal vector of the i-th plane, d i is the distance from the i-th plane to the origin). Use this plane to evaluate the point cloud data P4 and calculate all points p i Distance to the plane Compare and judge with the preset distance threshold ∈. When the local points of the plane model to be estimated meet a certain number and conditions, the plane can be used as a result to be output. After multiple iterative calculations, multiple fitting plane point clouds of different depths are obtained. This invention introduces a new parameter Distance i Calculate the distance from the fitting plane to the origin (center of the camera body) (with the camera coordinate system as the reference, the camera position is at the origin of the coordinate system), and find the minimum distance Min_Distanc and the plane point cloud P closest to the camera in continuous iterations first_inlierExtract and segment it. After the extraction is completed, plane P first_inlier will be removed from the point cloud data P4 to become data P 4_delete . Input point cloud data P 4_delete Perform the RANSAC algorithm iteration again to extract the plane P closest to the camera second_inlier , and in this step, the present invention adds two additional constraint conditions: (a) the distance constraint between this plane and the fitting plane where P first_inlier is located: Pmin < Distance i -Min_Distance < Pmax, where Pmin and Pmax are preset distance thresholds. (b) This plane contains the minimum number of point clouds to ensure that the current plane is the information plane of the front surface of the carton rather than a noise plane. After obtaining the plane point clouds P first_inlier and P second_inlier , calculate the depth distance difference D0 between them to provide data for the grasping scheme of the upper surface of the carton:
[0145] D0 = Distance first_inlier -Distance second_inlier
[0146] After the plane processing, the present invention performs secondary preprocessing on the plane point cloud P first_inlier : (a) Perform the second statistical filtering process as shown above to filter the residual noise point clouds and make the data smoother. (b) Introduce two domain point number threshold parameter ranges: n min and n max , use the optimized radius filtering to process the plane point cloud, delete sparse points and over-dense points, and expand the point cloud gap at the position of the boundary line between cartons. The remaining point set forms a new point cloud data P filtered = {p i ∣n min ≤n i ≤n max}.
[0147] Step 3: Align the point cloud with RGB to identify the boundary;
[0148] Based on the experimental test results, due to the unstable point cloud acquisition effect and the large error in carton boundary recognition, the present invention uses the method of using the RGB image as the input to identify the carton boundary and map it to the point cloud space. The depth camera simultaneously acquires the point cloud and RGB, and the internal parameters are unified, but there are differences in the FOV (wide-angle range).
[0149] Obtain the camera internal parameter matrix K, and the point cloud data P filtered = {(X i , Y i , Z i)|i=1,2,3...} The three-dimensional points of the camera coordinate system are mapped to the pixel coordinates of the two-dimensional RGB image, and the point set S0={(u i ,v i )|i=1,2,3...}, S0 can be calculated:
[0150]
[0151] where f x ,f y is the focal length of the camera in the x and y axis directions, c x ,c y are the principal point coordinates of the image.
[0152] Detect the boundary range of point set S0 and obtain the parameters: u min =min(u i ),u max =max(u i ),v min =min(v i ),v max =max(v i ),i=1,2,3..., define a rectangular area ROI=[u min ,u max ]×[v min ,v max ]. For RGB image I R The cropping is performed based on the ROI area, and the Canny edge detection algorithm is used on the cropped image to extract the boundary line in RGB. The process is expressed as:
[0153]
[0154] E(u,v)=Canny(I ROI (u,v))
[0155] Where E(u,v) is the binarization result, which is used to extract the boundary point set S1={(u i ,v i )|i=1,2,3...}. For point set S1, use the fitting plane depth DIstance first_inlier The camera intrinsic parameter matrix is used to inversely map the RGB boundary line information back to the camera coordinate system of the three-dimensional point cloud to obtain the three-dimensional point set P boundary ={(X i ,Y i ,Distance first_inlier )|i=1.2.3...}:
[0156]
[0157] Construct point set P in point cloud space boundary The Kd-Tree model performs a Kd-Tree operation on each point based on the radius r. b Spherical space search (r b is the preset parameter), the point cloud P filtered The points whose mid-space coordinates are inside the sphere are marked as boundary points. filtered After removing all boundary points, we get a point cloud P containing the front surface information of multiple independent cartons. cam .
[0158] Step 4: BoundingBox segmentation;
[0159] In order to extract the spatial location information of a single carton, the present invention uses a region segmentation algorithm based on surface growth. cam Perform nearest neighbor search, distance point p i (x i ,y i ,z i ) The nearest K points form its neighborhood N i ={p i1 ,p i2 ,…,p iK}, then calculate the centroid of the neighborhood points (average of all point coordinates) and neighborhood covariance matrix C i :
[0160]
[0161] Covariance matrix C i Perform eigenvalue decomposition to obtain three eigenvalues λ1, λ2, λ3 and their corresponding eigenvectors v1, v2, v3: C i v k =λ k v k (k=1,2,3), where the eigenvector v3 corresponding to the minimum eigenvalue is the normal vector n i . For point p i and its neighboring point p j :(a) Calculate the normal vector n i and n j The angle θ ij As a measure of the smoothness between two points: cos(θ ij )=n i ·n j (b) Use the ratio of the smallest eigenvalue to the sum of the eigenvalues to approximate the curvature: κ i =λ3 / (λ1+λ2+λ3). Introducing the preset parameter smoothness threshold θ th and curvature threshold κ th, from the seed point p i Start by initializing a new point cloud cluster C k , when θ ij <θ th ,κ i <κ th When , the point is considered to belong to cluster C k By traversing all points and then performing iterative calculations from a new unclassified seed point, multiple independent point cloud clusters with plane features can be obtained, that is, multiple independent cardboard front surface information point sets. ij Belong to P cam .
[0162] Based on the characteristics of the carton surface, the present invention defines a three-dimensional rectangular structure (which can enclose the smallest cuboid of the point cloud cluster) BoundingBox = (x b ,y b ,z b ,l x ,l y ,l z ), that is, to find a minimum cuboid that can enclose all points in the three-dimensional space of a given point set. b ,y b ,z b Represents the spatial coordinates of the center point in the camera coordinate system, l x ,l y ,l z Represent the side lengths of BoundingBox in the X, Y, and Z directions respectively. i , calculate the corresponding BoundingBox according to the maximum value of the cluster point in the three coordinate axis directions i , get the spatial position information of all cartons on the plane.
[0163] Step 5: Crawl order and method;
[0164] Grasping order: In the Z direction, grasping begins at the plane closest to the unloading robot; in the Y direction, grasping begins at the row highest from the ground; in the X direction, grasping begins at the maximum value in the positive X-axis direction. The Z direction is prioritized, followed by the Y direction, and finally the X direction. For the X and Y directions, judgment can be made by comparing the BoundingBox parameters.
[0165] Grabbing method: This invention introduces a preset parameter D ε The height y of the carton in the world coordinate system b +y cam Compare (y cam is the camera height, in this system y camis a fixed value), the position of the cartons to be grabbed is divided into "high" and "low". For the cartons in the "high" position, the center point center (x b ,y b ,z b ) to realize the front surface grabbing of the carton; for the carton in the “low” position, the output point center_high(x b ,y b +D_set,z b +D0 / 2), where D_set is a manually set parameter.
[0166] When the industrial computer determines the position and grasping method of the next carton to be grasped, the output point coordinate center or center_high will be converted into the robot arm base coordinate system space point P through the matrix [R|t] between the camera and the robot arm. a , which helps the robot to locate, where R is the rotation matrix and t is the translation vector. Conversion formula:
[0167] P a =R·P center +t
[0168] Although the present invention is described herein with reference to specific embodiments, it should be understood that these embodiments are merely illustrative of the principles and applications of the invention. It should be understood that many modifications may be made to the illustrative embodiments, and that other arrangements may be devised, without departing from the spirit and scope of the invention as defined by the appended claims. It should be understood that the various dependent claims and features described herein may be combined in ways other than those described in the original claims. It should also be understood that features described in conjunction with individual embodiments may be used in conjunction with other described embodiments.
Claims
1. A cardboard stack recognition and positioning method based on RGB images and point cloud data, characterized in that: include: Step 1: Use a depth camera to obtain RGB image information and point cloud information of the mixed cardboard stack, and preprocess the point cloud information to obtain preprocessed point cloud data; Step 2: Use the parameter-limited RANSAC method to extract the plane in the depth direction of the pre-processed point cloud data to obtain the point cloud plane closest to the camera, and then pre-process the point cloud plane; obtain the pre-processed point cloud plane P first_inlier ; Step 3: Based on the camera intrinsic parameter matrix K, the point cloud plane P first_inlier Map it to the two-dimensional pixel coordinates of the RGB image to obtain the point set S0; based on the boundary range of the point set S0, establish a rectangular region ROI, perform ROI range clipping in the RGB image, use the Canny edge detection algorithm to extract the boundary point set S1 of the clipped RGB image, and inversely map the point set S1 back to the plane point cloud P first_inlier After filtering out the boundary of the point cloud plane mapped back to the point cloud, we can obtain the point cloud P containing the front surface information of multiple independent cartons. cam ; Step 4: Use a clustering segmentation algorithm based on surface growth to cluster the point cloud P containing the front surface information of multiple independent cartons. cam Split into multiple independent point cloud clusters; Step 5: Create a minimum cube outside each independent point cloud cluster that can contain all the points in the point cloud cluster; use the spatial coordinates and side length of the minimum cube as the plane closest to the camera as the spatial position information of the carton.
2. The method for identifying and locating carton stacks based on RGB images and point cloud data according to claim 1, characterized in that: In step 1, the point cloud information is preprocessed. The method of preprocessing the point cloud information is as follows: The point cloud information is sequentially subjected to point cloud preprocessing operations including voxel filtering downsampling, pass filtering, statistical filtering and Kriging interpolation.
3. The carton stack recognition and positioning method based on RGB image and point cloud data according to claim 1, characterized in that: In step 2, the RANSAC method based on parameter limitation is used to extract the plane in the depth direction of the point cloud information. The process of obtaining the point cloud plane closest to the camera is as follows: Step S1: Use the point cloud data preprocessed in step 1 as a point cloud model, and randomly select three points in the point cloud model to construct a plane; Step S2: Calculate all points p in the point cloud model i Distance to the plane constructed in step S1 The distance Compare with the preset distance threshold ∈; when it meets Point p i If the number of points is greater than M, a point cloud plane is determined using M points; Step S3: Remove the point cloud plane determined in step S2 from the point cloud data preprocessed in step 1, use the remaining point cloud data as the point cloud model, randomly select three points from the point cloud model to construct a plane, and return to step S2 until all points in the point cloud model are traversed to obtain N point cloud planes. Step S4: Calculate the distances from the center point of the depth camera to the N point cloud planes, and obtain the point cloud plane P closest to the center point of the depth camera. first_inlier .
4. The carton stack recognition and positioning method based on RGB image and point cloud data according to claim 1, characterized in that: In step 3, based on the camera intrinsic parameter matrix K, the point cloud plane P first_inlier Mapping to the pixel coordinates of the two-dimensional RGB image, the method to obtain the point set S0 is: Using the camera intrinsic parameter matrix K, the point cloud plane P first_inlier Map the three-dimensional points of the camera coordinate system to the pixel coordinates of the two-dimensional RGB image to obtain the point set S0; S0={(u i ,v i )|i=1,2,3...} Among them, f x ,f y is the focal length of the camera in the x and y axis directions of the camera coordinate system, c x ,c y is the principal point coordinate of the image, u i ,v i They represent the horizontal and vertical coordinates of the i-th point cloud point in the two-dimensional RGB image in the point cloud coordinate system, respectively. i 、y i 、z i They represent the three-dimensional coordinates of the i-th point cloud point in the camera coordinate system.
5. The carton stack recognition and positioning method based on RGB image and point cloud data according to claim 1, characterized in that: In step 3, the Canny edge detection algorithm is used to extract the boundary of the cropped RGB image as follows: First, detect the boundary of the point set S0 and obtain the region ROI = [u min ,u max ]×[v min ,v max ]; Among them, u min Indicates the minimum horizontal coordinate value of the point set S0 in the point cloud coordinate system, u max Indicates the maximum horizontal coordinate of the point set S0 in the point cloud coordinate system, v min Represents the minimum vertical coordinate v of the point set S0 in the point cloud coordinate system max Indicates the maximum longitudinal coordinate of the point set S0 in the point cloud coordinate system; Reuse the formula: E(u,v)=Canny(I ROI (u,v)) Calculate the binarization result E(u,v), and use the binarization result E(u,v) to extract the boundary point set S1, I of the RGB image ROI (u,v) represents I ROI (u,v) represents the RGB image after cropping based on the boundary range.
6. The method for identifying and locating carton stacks based on RGB images and point cloud data according to claim 1, characterized in that: In step 3, the boundary is inversely mapped to the plane point cloud P first_inlier After filtering out the boundary point cloud, we obtain the point cloud P containing the front surface information of multiple independent cartons. cam The process is: Set the camera center to the point cloud plane P first_inlier The distance is used as the plane depth Distance first_inlier , combined with the camera intrinsic parameter matrix, the boundary point set S1 of the RGB image is inversely mapped back to the camera coordinate system of the three-dimensional point cloud to obtain the three-dimensional point set P boundary , with a three-dimensional point set P boundary Each point in is taken as the origin and the radius r is established b The spherical space of the point cloud plane P preprocessed in step 2 first_inlier The points whose coordinates are in the spherical space are marked as boundary points. first_inlier Eliminate all boundary points and obtain the point cloud P containing the front surface information of multiple independent cartons cam .
7. The method for identifying and locating carton stacks based on RGB images and point cloud data according to claim 1, characterized in that: In step 4, a clustering segmentation algorithm based on surface growth is used to cluster the point cloud P containing the front surface information of multiple independent cartons. cam The method of segmenting into multiple independent point cloud clusters is: Step S41: From point cloud P cam Select any point P i , search for point P i The ratio of the minimum eigenvalue of the neighborhood covariance matrix to the sum of the eigenvalues is calculated, and the ratio is used as the curvature κ i ; At the same time, the covariance matrix of the neighborhood is used to calculate the point P i Normal vector n at the center i , and then calculate the point P i Any point p in the neighborhood j The covariance matrix of the neighborhood centered on point p is then calculated to obtain j Normal vector n at the center j ; Calculate the normal vector n i and normal vector n j The angle between Step S42: Set parameter smoothness threshold θ th and curvature threshold κ th ; When θ is satisfied ij <θ th ,κ i <κ th When P i Neighborhood point p j Point P i Class, traverse P i All points in the neighborhood, get the point p i Starting point cloud cluster; Step S43: From point cloud P cam Select the unclassified point P k , execute step S41 and step S42 until the point cloud P cam All points are classified, and the point cloud P containing the front surface information of multiple independent cartons is completed. cam Split into multiple independent point cloud clusters.
8. The method for identifying and locating carton stacks based on RGB images and point cloud data according to claim 1, characterized in that: In step S41, point P i Normal vector n at the center i The method is similar to the point p j Normal vector n at the center j The same method is used, click P i Normal vector n at the center i The method is: Point P i Perform the nearest search for the center and obtain the neighborhood N i , calculate the neighborhood N i The center of mass Neighborhood covariance matrix C i ; Using the neighborhood N i The center of mass For neighborhood N i Perform eigenvalue decomposition on the covariance matrix of the neighborhood N i The eigenvalues and eigenvectors of the covariance matrix; The eigenvector of the minimum eigenvalue is taken as the i Normal vector n i .
9. A method for determining grasping points of a carton stack based on RGB images and point cloud data, the method being implemented using the method for identifying and locating a carton stack based on RGB images and point cloud data as claimed in any one of claims 1 to 8; characterized in that: Specifically include: A distance laser is set at the end of the robot's grabbing arm, and the distance D from the end of the grabbing arm to the carton surface is measured using the distance laser. α ; Using the distance D α And the height of the carton, calculate the center point of the carton; When the height of the carton meets the height y of the carton in the world coordinate system b +y cam >Default parameter D ε When the center point is: center(x b ,y b ,z b ); When the height of the carton meets the height y of the carton in the world coordinate system b +y cam <Default parameter D ε When the center point is: center_high(x b ,y b +D_set,z b +D0 / 2); Among them, x b ,y b ,z b Represents the spatial coordinates of the center point in the camera coordinate system, y cam is the camera height, D_set is the preset distance for top suction planning, and D0 is the distance between the closest plane of the camera and the second point cloud plane; The center point or the center point height is converted into the space point P of the robot base coordinate system through the matrix R between the camera and the robot arm. a , the robot coordinate system space point P a As a grabbing point.
10. The method for determining grabbing points of a carton stack based on RGB images and point cloud data according to claim 9, characterized in that: Robotic arm coordinate system space point P a for: P a =R·P center +t Among them, R is the rotation matrix, t is the translation vector, P center The coordinates of the grab point in the world coordinate system.
Citation Information
Cited By
Box clamping point estimation method and device based on depth camera and electronic equipment
CN121170017A
Box clamping point estimation method and device based on depth camera and electronic equipment
CN121170017B
Self-adaptive control method and device for workbin sorting robot
CN122324452A