Target object recognition and pose estimation method based on 3D vision
Through a 3D vision-based method, using the point cloud data and images of the object to analyze the geometric features and historical pose data of the object, the problem of large pose estimation errors in the existing technology is solved, intelligent recognition and real-time adjustment are realized, and the accuracy and reliability of pose estimation of the object is improved.
Patent Information
- Application Number
- CN202510446280.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-04-10
AI Technical Summary
In the prior art, the pose estimation of the target object relies on a single sensor, resulting in large errors and the inability to judge whether the object is in a normal pose state in real time.
The target object recognition and pose estimation method based on 3D vision is used to obtain the point cloud data and images of the object, analyze the geometric center point, three-dimensional contour and voxel grid of the object, and combine historical pose data to determine whether the object position needs to be adjusted.
It realizes intelligent recognition of the position and posture state of the target object, reduces position estimation errors, and can timely detect abnormal position and posture state of the object and make timely adjustments.
Smart Images

Figure CN119964145B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and in particular to a target object recognition and posture estimation method based on 3D vision. Background Art
[0002] Object recognition and pose estimation are key tasks in computer vision and robotics, which aim to automatically identify objects in a scene through sensory data (such as images, point clouds, etc.) and accurately estimate the position and pose (i.e., rotation angle) of objects in three-dimensional space. First, object recognition extracts the features of the object (such as color, shape, texture, etc.), separates the object from the background, classifies it, and determines its type. Then, combining the geometric information of the object with the sensor data, pose estimation infers the position and orientation of the object relative to the reference coordinate system. The combination of the two enables the system to not only recognize the object, but also accurately determine the spatial position and pose of the object, providing a basis for subsequent operations (such as grasping, tracking, etc.).
[0003] In the prior art, the pose estimation of the target object often relies on a single sensor, which results in errors in the pose estimation of the target object, and it is impossible to determine whether the object is in a normal pose state according to different object categories, resulting in the inability to timely detect and adjust the object when it is in an abnormal pose state;
[0004] To this end, the present invention proposes a target object recognition and pose estimation method based on 3D vision. Summary of the invention
[0005] The purpose of the present invention is to propose a target object recognition and pose estimation method based on 3D vision to solve the problems raised in the above background technology.
[0006] The technical problems to be solved by the present invention are:
[0007] How to realize intelligent recognition of the corresponding posture state of the target object based on 3D vision.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions:
[0009] The target object recognition and pose estimation method based on 3D vision is as follows:
[0010] Step S1, obtaining an object image and point cloud data of an object in a warehouse;
[0011] Step S2, analyzing the point cloud data of the objects in the warehouse to obtain the center coordinates corresponding to the geometric center points of the objects in the warehouse;
[0012] Step S3, detecting the object images in the warehouse to obtain the object categories;
[0013] Step S4, obtaining historical posture data of the object, and analyzing the object based on the object's three-dimensional contour, object category and object voxel grid to determine whether the object's posture needs to be adjusted.
[0014] Preferably, the point cloud data of the object includes the coordinates of any point on the surface of all objects in the warehouse in the three-dimensional coordinate system.
[0015] Preferably, the step S1 includes the following sub-steps:
[0016] Step S11, taking any vertex of the warehouse as the origin, the east direction as the positive direction of the X axis, the north direction as the positive direction of the Y axis, and the direction perpendicular to the ground as the positive direction of the Z axis, constructing a three-dimensional coordinate system OXYZ of the warehouse;
[0017] Step S12, collecting point cloud data of the object and the image of the object in the warehouse.
[0018] Preferably, step S2 includes the following sub-steps:
[0019] Step S21, obtaining the coordinates of any coordinate point in the point cloud data of the object in the warehouse, calculating the distance between the corresponding coordinate point and all other coordinate points, and analyzing to obtain the voxel grid of the warehouse;
[0020] Step S22, analyzing the coordinates of all coordinate points in the voxel grid to obtain the warehouse sub-area;
[0021] Step S23, marking all coordinate points in the voxel grid, and constructing the coordinate points with the same object mark into an object voxel grid, specifically:
[0022] Traverse from any coordinate point in the voxel grid, traverse the coordinate points adjacent to the current coordinate point and not marked, and mark the corresponding two coordinate points with the same object identifier, and stop traversing when the coordinate point of the object cannot be traversed, and construct the object voxel grid of the object based on the coordinate points marked with the same object identifier;
[0023] Step S24, obtaining any coordinate point in the object voxel grid and its object identifier, taking the corresponding coordinate point as the detection coordinate point, and traversing the coordinate points adjacent to the detection coordinate point;
[0024] If the object identifiers of all coordinate points are the same as the object identifier of the detection coordinate point, it is determined that the detection coordinate point is not an edge coordinate point;
[0025] If the object identifier of any coordinate point is different from the object identifier of the detection coordinate point, or the detection coordinate point is adjacent to the blank area, the detection coordinate point is determined to be an edge coordinate point;
[0026] Step S25, obtaining edge coordinate points in the voxel grid of the object, connecting the edge coordinate points one by one to obtain a three-dimensional contour model of the object, and constructing a three-dimensional contour of the object based on the three-dimensional contour model;
[0027] Step S26, obtaining the geometric center point of the object according to the three-dimensional contour, and using the coordinates of the geometric center point of the object as the center coordinates of the object.
[0028] Preferably, the process of obtaining the voxel grid is as follows:
[0029] Get the coordinates (xi, yi, zi) of any coordinate point in the point cloud data of the object in the warehouse, where i is the number of the coordinate point, i=1, 2, ..., n, n is a positive integer, record the corresponding coordinate point as the voxel coordinate point, and calculate the distance between the voxel coordinate point and all other coordinate points;
[0030] If the distance between the voxel coordinate point and any coordinate point is greater than a first distance threshold, it is determined that the voxel coordinate point and the coordinate point do not belong to the same voxel grid;
[0031] If the distance between the voxel coordinate point and any coordinate point is less than or equal to the first distance threshold, it is determined that the voxel coordinate point and the coordinate point belong to the same voxel grid.
[0032] Preferably, the process of obtaining the warehouse sub-area is:
[0033] Traverse the coordinates of all coordinate points in the same voxel grid, perform calculations based on the traversal results, calculate the length, width and height of the voxel grid, and calculate the volume of the voxel grid;
[0034] Then, the number of coordinate points in the voxel grid is obtained to calculate the voxel grid density;
[0035] If the voxel grid density is less than or equal to the density threshold, it is determined that the corresponding voxel grid is not a warehouse sub-area, and the corresponding voxel grid is removed;
[0036] If the voxel grid density is greater than the density threshold, the corresponding voxel grid is determined to be a warehouse sub-region.
[0037] Preferably, the process of obtaining the three-dimensional contour is:
[0038] Draw six sets of planes tangent to the outermost edge points of the three-dimensional contour model from the six directions of up and down, left and right, front and back of the three-dimensional contour model of the object;
[0039] The three-dimensional area enclosed by the six groups of planes is the three-dimensional outline of the object.
[0040] Preferably, step S3 includes the following sub-steps:
[0041] Step S31, denoising the object image in the warehouse, and then cropping the object image, the resolution of the cropped object image is 800×600;
[0042] Step S32, inputting the cropped object image into the deep learning model for detection, and detecting and obtaining the object category corresponding to the object image;
[0043] Step S33, obtaining the shooting coordinates of the object image in the three-dimensional coordinate system according to the shooting position of the object image, and then calculating the distance between the shooting coordinates and the center coordinates corresponding to the geometric center points of all objects and traversing, traversing to obtain the center coordinates corresponding to the geometric center point of the object with the smallest distance from the shooting coordinates, and matching the object to which the center coordinates belong with the object category corresponding to the object image.
[0044] Preferably, step S4 includes the following sub-steps:
[0045] Step S41, obtaining the object category of any object, and obtaining the historical position and posture data of the corresponding object according to the object category; the historical position and posture data of the object includes the standard length, standard width and standard height of the object when it is normally placed;
[0046] Step S42, traversing the coordinates of the coordinate points in the three-dimensional outline of the object, and calculating and obtaining the length, width and height of the object based on the traversal results;
[0047] Step S43, comparing the calculated length of the object with the length in the historical pose data of the object, the calculated width of the object with the width in the historical pose data of the object, and the calculated height of the object with the height in the historical pose data of the object one by one;
[0048] If there is a difference in the comparison, it is determined that the posture of the object is abnormal. If the comparison is the same, go to the next step.
[0049] Preferably, the step S4 further includes the following sub-steps:
[0050] Step S44, obtain the center coordinates P0 (x0, y0, z0) corresponding to the geometric center point of the object, and the coordinates Pi (xi, yi, zi) of all coordinate points in the object voxel grid of the object, and calculate the covariance matrix C of the object by the formula, which is as follows:
[0051] , where T is the transposition symbol;
[0052] Step S45, decomposing the covariance matrix of the object, and calculating the eigenvector and eigenvalue of the object by the formula, the specific formula is as follows:
[0053] C·vi=Di·vi, where “·” is the dot product operation between vectors, vi is the eigenvector of the object, and Di is the eigenvalue corresponding to the eigenvector;
[0054] Step S46, taking the eigenvector corresponding to the calculated maximum eigenvalue as the main direction vector of the object;
[0055] Step S47, obtaining the standard direction vector BZ (0, 0, 1) of the object, and calculating the inclination angle JD of the object by the formula, the specific formula is as follows:
[0056] cos(JD)=(vi·BZ) / (|vi|·|BZ|), where |vi| is the modulus of the main direction vector of the object, |BZ| is the modulus of the standard direction vector, and |BZ|=1;
[0057] Step S48, when the tilt angle of the object is (0°, 180°), it is determined that the abnormal state of the object is a tilt state;
[0058] When the tilt angle of the object is 0°, it is determined that the object is not in an abnormal state;
[0059] When the tilt angle of the object is 180°, the abnormal state of the object is determined to be an inverted state.
[0060] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:
[0061] The present invention first obtains the object image and point cloud data of the object in the warehouse, and then analyzes the point cloud data of the object in the warehouse to obtain the center coordinates corresponding to the geometric center points of the objects in the warehouse. At the same time, the object images in the warehouse are detected to obtain the object categories, and finally the historical posture data of the object is obtained. Based on the historical posture data, three-dimensional contour, object category and object voxel grid analysis of the object, it is determined whether the posture of the object needs to be adjusted. The present invention realizes intelligent recognition of the corresponding posture state of the target object based on 3D vision. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to facilitate understanding by those skilled in the art, the present invention is further described below with reference to the accompanying drawings.
[0063] Figure 1 is the overall system block diagram of the present invention;
[0064] Figure 2 A top view of the warehouse in the present invention;
[0065] Figure 3 An example diagram of an object to which a coordinate point in a voxel grid belongs in the present invention;
[0066] Figure 4It is a schematic diagram of the structure of the computing device in the present invention. DETAILED DESCRIPTION
[0067] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0068] Example 1: Please refer to Figure 1-Figure 3 As shown, the technical solution provided by the present invention is: a target object recognition and posture estimation method based on 3D vision. This method is used to identify the object type of any object in the warehouse and detect whether the posture of the object is abnormal. The method is as follows:
[0069] Step S1, obtaining an object image and point cloud data of an object in a warehouse;
[0070] The laser radar is used to collect point cloud data of objects, and the high-definition camera is used to collect images of objects in the warehouse. The point cloud data of objects and the images of objects in the warehouse are combined to form the data required for 3D visual recognition of objects and their position and posture estimation.
[0071] It should be specifically explained that the laser radar generates point cloud data of the object by emitting a laser beam and receiving the laser beam reflected from the surface of the object. The point cloud data of the object includes the coordinates of any point on the surface of all objects in the warehouse in the three-dimensional coordinate system;
[0072] In this embodiment, step S1 includes the following sub-steps:
[0073] Step S11, taking any vertex of the warehouse as the origin, the east direction as the positive direction of the X axis, the north direction as the positive direction of the Y axis, and the direction perpendicular to the ground as the positive direction of the Z axis, constructing a three-dimensional coordinate system OXYZ of the warehouse;
[0074] Step S12, collecting point cloud data of the object and the image of the object in the warehouse.
[0075] Step S2, analyzing the point cloud data of the objects in the warehouse to obtain the center coordinates corresponding to the geometric center points of the objects in the warehouse;
[0076] In this embodiment, step S2 includes the following sub-steps:
[0077] Step S21, obtaining the coordinates of any coordinate point in the point cloud data of the object in the warehouse, calculating the distance between the corresponding coordinate point and all other coordinate points, and analyzing to obtain the voxel grid of the warehouse;
[0078] It should be specifically noted that in the point cloud data of an object, a voxel grid is a small unit in a three-dimensional space, which is constructed by a number of coordinates and is used to discretize the space;
[0079] Specifically, the coordinates (xi, yi, zi) of any coordinate point in the point cloud data of the object in the warehouse are obtained, where i is the number of the coordinate point, i=1, 2, ..., n, and n is a positive integer, and the corresponding coordinate point is recorded as a voxel coordinate point, and the distance between the voxel coordinate point and all other coordinate points is calculated;
[0080] If the distance between the voxel coordinate point and any coordinate point is greater than a first distance threshold, it is determined that the voxel coordinate point and the coordinate point do not belong to the same voxel grid;
[0081] If the distance between the voxel coordinate point and any coordinate point is less than or equal to a first distance threshold, it is determined that the voxel coordinate point and the coordinate point belong to the same voxel grid;
[0082] Step S22, analyzing the coordinates of all coordinate points in the voxel grid to obtain the warehouse sub-area;
[0083] Specifically, the coordinates of all coordinate points in the same voxel grid are traversed, and calculations are performed based on the traversal results to obtain the length CHA, width KUA, and height GAO of the voxel grid, specifically:
[0084] xmin=min(x1,x2,…,xn), xmax=max(x1,x2,…,xn);
[0085] ymin=min(y1,y2,…,yn), ymax=max(y1,y2,…,yn);
[0086] zmin=min(z1,z2,…,zn), zmax=max(z1,z2,…,zn);
[0087] CHA = xmax-xmin;
[0088] KUA=ymax-ymin;
[0089] GAO=zmax-zmin;
[0090] Then the voxel grid volume V is calculated by the formula. The calculation process is as follows:
[0091] V = CHA × KUA × GAO;
[0092] Further, the number of coordinate points in the voxel grid ZBS is obtained, and the voxel grid density TSM is calculated by the formula TSM=ZBS / V;
[0093] If the voxel grid density is less than or equal to the density threshold, it is determined that the corresponding voxel grid is not a warehouse sub-area, and the corresponding voxel grid is removed;
[0094] If the voxel grid density is greater than the density threshold, the corresponding voxel grid is determined to be a warehouse sub-area;
[0095] In this embodiment, each warehouse sub-area corresponds to a shelf area in the warehouse, wherein the shelf area without objects is taken as the standard shelf area, the number of coordinate points in the standard shelf area is obtained, and the voxel grid density of the standard shelf area is calculated, and the calculated voxel grid density is used as the density threshold;
[0096] Step S23, as Figure 3 As shown, all coordinate points in the voxel grid are marked as objects, and the coordinate points with the same object mark are constructed into an object voxel grid;
[0097] Specifically, traverse from any coordinate point in the voxel grid, traverse the coordinate points adjacent to the current coordinate point and not marked, and mark the above two coordinate points with the same object identifier, and stop traversing when the coordinate point of the object cannot be traversed, and construct the object voxel grid of the object based on the coordinate points marked with the same object identifier;
[0098] It should be specifically explained that, by using BFS / DFS search, the same object identifier is assigned to the coordinate points adjacent to the unmarked coordinate point and not yet marked. The object identifier can be the number corresponding to the object. If BFS / DFS traverses to the coordinate point of another object, a new BFS / DFS search is triggered, and the coordinate point is marked with the object identifier of the corresponding object; BFS / DFS search prior art, BFS traverses nodes layer by layer, suitable for finding the shortest path and breadth-first processing tasks, DFS traverses deeply through recursion or stack, and is often used for path search, connectivity detection and depth-first processing;
[0099] For example, there is currently a 3×3×3 voxel grid containing 27 coordinate points, and the coordinates are (0, 0, 0), (0, 0, 1), (0, 0, 2), ..., (2, 2, 2), and each coordinate point is set to be adjacent to its six faces. For example, the coordinate point (0, 0, 0) is adjacent to the coordinate point (0, 0, 1) and the coordinate point (0, 1, 0), and the coordinate point (0, 1, 0) is adjacent to the coordinate point (0, 1, 1);
[0100] Starting from the coordinate point (0, 0, 0), use BFS to search and obtain the adjacent coordinate points (0, 0, 1), (0, 1, 0) and (0, 1, 1), and mark the above coordinate points as object A. When BFS searches for the coordinate point (1, 0, 0) that does not belong to object A, create a new object B and mark the coordinate point that does not belong to A as object B. This is done by analogy to obtain the coordinate points occupied by any object. Finally, the coordinates of the coordinate points occupied by object A are (0, 0, 0), (0, 0, 1), (0, 1, 0) and (0, 1, 1), the coordinates of the coordinate points occupied by object B are (1, 0, 0), (1, 0, 1), (1, 1, 0) and (1, 1, 1), and the coordinates of the coordinate points occupied by object C are (2, 2, 0), (2, 2, 1) and (2, 1, 0);
[0101] Step S24, obtaining any coordinate point in the object voxel grid and its object identifier, taking the corresponding coordinate point as the detection coordinate point, and traversing the coordinate points adjacent to the detection coordinate point;
[0102] If the object identifiers of all coordinate points are the same as the object identifier of the detection coordinate point, it is determined that the detection coordinate point is not an edge coordinate point;
[0103] If the object identifier of any coordinate point is different from the object identifier of the detection coordinate point, or the detection coordinate point is adjacent to the blank area, the detection coordinate point is determined to be an edge coordinate point;
[0104] It should be specifically noted that in the warehouse sub-area, there is an area that is not occupied by any object, and there is no coordinate point in the area, which is regarded as a blank area;
[0105] Step S25, obtaining edge coordinate points in the voxel grid of the object, connecting the edge coordinate points one by one to obtain a three-dimensional contour model of the object, and constructing a three-dimensional contour of the object based on the three-dimensional contour model, specifically:
[0106] Draw six sets of planes tangent to the outermost edge points of the three-dimensional contour model from the six directions of up and down, left and right, front and back of the three-dimensional contour model of the object. The three-dimensional area enclosed by the six sets of planes is the three-dimensional contour of the object.
[0107] Step S26, obtaining the geometric center point of the object according to the three-dimensional contour, and using the coordinates of the geometric center point of the object as the center coordinates of the object.
[0108] Step S3, detecting the object images in the warehouse to obtain the object categories;
[0109] In this embodiment, step S3 includes the following sub-steps:
[0110] Step S31, denoising the object image in the warehouse, and then cropping the object image, the resolution of the cropped object image is 800×600;
[0111] Step S32, inputting the cropped object image into the deep learning model for detection, and detecting and obtaining the object category corresponding to the object image;
[0112] It should be specifically noted that the deep learning model used to detect the object category is a trained deep learning model, and the training process is: matching the object category with the object image, and then inputting it into the deep learning model for training to obtain the trained deep learning model, wherein the object image includes a front view, a side view, and a top view of the object;
[0113] Step S33, obtaining the shooting coordinates of the object image in the three-dimensional coordinate system according to the shooting position of the object image, and then calculating the distance between the shooting coordinates and the center coordinates corresponding to the geometric center points of all objects and traversing, traversing to obtain the center coordinates corresponding to the geometric center point of the object with the smallest distance from the shooting coordinates, and matching the object to which the center coordinates belong with the object category corresponding to the object image.
[0114] Step S4, obtaining historical posture data of the object, and determining whether the posture of the object needs to be adjusted based on the three-dimensional contour of the object, the object category, and the object voxel grid analysis;
[0115] The historical position data of the object includes the standard length, standard width and standard height of the object when it is normally placed;
[0116] It should be specifically noted that the object category of any object uniquely corresponds to the historical pose data of an object;
[0117] In this embodiment, step S4 includes the following sub-steps:
[0118] Step S41, obtaining the object category of any object, and obtaining the historical position and posture data of the corresponding object according to the object category;
[0119] Step S42, traversing the coordinates of the coordinate points in the three-dimensional outline of the object, and calculating and obtaining the length, width and height of the object based on the traversal results;
[0120] The process of calculating the length, width and height of the object is the same as that described in step S23;
[0121] Step S43, comparing the calculated length of the object with the length in the historical pose data of the object, the calculated width of the object with the width in the historical pose data of the object, and the calculated height of the object with the height in the historical pose data of the object one by one;
[0122] If there is a difference in the above comparison, it is determined that the posture of the object is abnormal and needs to be adjusted;
[0123] If the above comparisons are the same, proceed to the next step;
[0124] Step S44, obtain the center coordinates P0 (x0, y0, z0) corresponding to the geometric center point of the object, and the coordinates Pi (xi, yi, zi) of all coordinate points in the object voxel grid of the object, and calculate the covariance matrix C of the object by the formula, which is as follows:
[0125] , where T is the transposition symbol, i=1, 2, ..., n;
[0126] Step S45, decomposing the covariance matrix of the object, and calculating the eigenvector and eigenvalue of the object by the formula, the specific formula is as follows:
[0127] C·vi=Di·vi, where “·” is the dot product operation between vectors, vi is the eigenvector of the object, and Di is the eigenvalue corresponding to the eigenvector;
[0128] It should be specifically noted that the calculated eigenvalues correspond one-to-one to the eigenvectors;
[0129] Step S46, taking the eigenvector corresponding to the calculated maximum eigenvalue as the main direction vector of the object;
[0130] Step S47, obtaining the standard direction vector BZ (0, 0, 1) of the object, and calculating the inclination angle JD of the object by the formula, the specific formula is as follows:
[0131] cos(JD)=(vi·BZ) / (|vi|·|BZ|), where |vi| is the modulus of the main direction vector of the object, |BZ| is the modulus of the standard direction vector, and |BZ|=1;
[0132] Step S48, when the tilt angle of the object is (0°, 180°), it is determined that the abnormal state of the object is a tilt state, and the position and posture of the object needs to be adjusted;
[0133] When the tilt angle of the object is 0°, it is determined that the object is not in an abnormal state and there is no need to adjust the position of the object;
[0134] When the tilt angle of the object is 180°, the abnormal state of the object is determined to be an inverted state, and the position and posture of the object needs to be adjusted.
[0135] In this application, if corresponding calculation formulas appear, the above calculation formulas are all dimensionless and take their numerical calculations. The weight coefficients, proportional coefficients and other coefficients in the formulas are set to a result value obtained by quantifying each parameter. The size of the weight coefficient and the proportional coefficient can be determined as long as it does not affect the proportional relationship between the parameter and the result value.
[0136] Embodiment 2: Figure 4 The following is a schematic diagram of the structure of a computer device, such as Figure 4 As shown, the computer device may include: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus. The processor may call the logic instructions in the memory to execute a target object recognition and posture estimation method based on 3D vision, the method comprising: obtaining an object image and point cloud data of an object in a warehouse; analyzing the point cloud data of the object in the warehouse to obtain the center coordinates corresponding to the geometric center point of the object in the warehouse; detecting the image of the object in the warehouse to obtain the object category; obtaining the historical posture data of the object, and analyzing whether the posture of the object needs to be adjusted based on the three-dimensional contour of the object, the object category, and the object voxel grid.
[0137] In addition, the logic instructions in the above-mentioned memory can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0138] On the other hand, the present application also provides a computer program product, which includes a computer program stored on a computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the 3D vision-based target object recognition and posture estimation method provided by the above methods, and the method includes: obtaining object images and point cloud data of objects in the warehouse; analyzing the point cloud data of objects in the warehouse to obtain the center coordinates corresponding to the geometric center points of the objects in the warehouse; detecting the images of objects in the warehouse to obtain the object categories; obtaining the historical posture data of the object, and analyzing whether the posture of the object needs to be adjusted based on the three-dimensional contour of the object, the object category and the object voxel grid.
[0139] On the other hand, the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the above-mentioned 3D vision-based target object recognition and pose estimation methods, the methods comprising: obtaining object images and point cloud data of objects in a warehouse; analyzing the point cloud data of objects in the warehouse to obtain the center coordinates corresponding to the geometric center points of the objects in the warehouse; detecting the images of objects in the warehouse to obtain the object categories; obtaining the historical pose data of the object, and determining whether the pose of the object needs to be adjusted based on the object's three-dimensional contour, object category and object voxel grid analysis.
[0140] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0141] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0142] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for target object recognition and pose estimation based on 3D vision, characterized in that: Here’s how: Step S1, obtaining an object image and point cloud data of an object in a warehouse; Step S2, analyzing the point cloud data of the objects in the warehouse to obtain the center coordinates corresponding to the geometric center points of the objects in the warehouse; Step S3, detecting the object images in the warehouse to obtain the object categories; Step S4, obtaining historical posture data of the object, and determining whether the posture of the object needs to be adjusted based on the three-dimensional contour of the object, the object category, and the object voxel grid analysis; Wherein, the step S4 includes the following sub-steps: Step S41, obtaining the object category of any object, and obtaining the historical position and posture data of the corresponding object according to the object category; the historical position and posture data of the object includes the standard length, standard width and standard height of the object when it is normally placed; Step S42, traversing the coordinates of the coordinate points in the three-dimensional outline of the object, and calculating and obtaining the length, width and height of the object based on the traversal results; Step S43, comparing the calculated length of the object with the length in the historical pose data of the object, the calculated width of the object with the width in the historical pose data of the object, and the calculated height of the object with the height in the historical pose data of the object one by one; If there is a difference in the comparison, the position and posture of the object is determined to be abnormal. If the comparisons are the same, proceed to the next step; Step S44, obtain the center coordinates P0 (x0, y0, z0) corresponding to the geometric center point of the object, and the coordinates Pi (xi, yi, zi) of all coordinate points in the object voxel grid of the object, and calculate the covariance matrix C of the object by the formula, which is as follows: Where T is the transposition symbol; Step S45, decomposing the covariance matrix of the object, and calculating the eigenvector and eigenvalue of the object by the formula, the specific formula is as follows: C·vi=Di·vi, where "·" is the dot multiplication operation between vectors, vi is the eigenvector of the object, and Di is the eigenvalue corresponding to the eigenvector; Step S46, taking the eigenvector corresponding to the calculated maximum eigenvalue as the main direction vector of the object; Step S47, obtain the standard direction vector BZ (0, 0, 1) of the object, and calculate the inclination angle JD of the object by the formula, the specific formula is as follows: cos(JD)=(vi·BZ) / (|vi|·|BZ|), where |vi| is the modulus of the main direction vector of the object, |BZ| is the modulus of the standard direction vector, and |BZ|=1; Step S48, when the tilt angle of the object is (0°, 180°), it is determined that the abnormal state of the object is a tilt state; When the tilt angle of the object is 0°, it is determined that the object is not in an abnormal state; When the tilt angle of the object is 180°, the abnormal state of the object is determined to be an inverted state.
2. The method for target object recognition and pose estimation based on 3D vision according to claim 1, characterized in that: The point cloud data of the object includes the coordinates of any point on the surface of all objects in the warehouse in the three-dimensional coordinate system.
3. The method for target object recognition and pose estimation based on 3D vision according to claim 2, characterized in that: The step S1 includes the following sub-steps: Step S11, taking any vertex of the warehouse as the origin, the east direction as the positive direction of the X axis, the north direction as the positive direction of the Y axis, and the direction perpendicular to the ground as the positive direction of the Z axis, constructing a three-dimensional coordinate system OXYZ of the warehouse; Step S12, collecting point cloud data of the object and the image of the object in the warehouse.
4. The method for target object recognition and pose estimation based on 3D vision according to claim 1, characterized in that: The step S2 includes the following sub-steps: Step S21, obtaining the coordinates of any coordinate point in the point cloud data of the object in the warehouse, calculating the distance between the corresponding coordinate point and all other coordinate points, and analyzing to obtain the voxel grid of the warehouse; Step S22, analyzing the coordinates of all coordinate points in the voxel grid to obtain the warehouse sub-area; Step S23, marking all coordinate points in the voxel grid, and constructing the coordinate points with the same object mark into an object voxel grid, specifically: Traverse from any coordinate point in the voxel grid, traverse the coordinate points adjacent to the current coordinate point and not marked, and mark the corresponding two coordinate points with the same object identifier, and stop traversing when the coordinate point of the object cannot be traversed, and construct the object voxel grid of the object based on the coordinate points marked with the same object identifier; Step S24, obtaining any coordinate point in the object voxel grid and its object identifier, taking the corresponding coordinate point as the detection coordinate point, and traversing the coordinate points adjacent to the detection coordinate point; If the object identifiers of all coordinate points are the same as the object identifier of the detection coordinate point, it is determined that the detection coordinate point is not an edge coordinate point; If the object identifier of any coordinate point is different from the object identifier of the detection coordinate point, or the detection coordinate point is adjacent to the blank area, the detection coordinate point is determined to be an edge coordinate point; Step S25, obtaining edge coordinate points in the voxel grid of the object, connecting the edge coordinate points one by one to obtain a three-dimensional contour model of the object, and constructing a three-dimensional contour of the object based on the three-dimensional contour model; Step S26, obtaining the geometric center point of the object according to the three-dimensional contour, and using the coordinates of the geometric center point of the object as the center coordinates of the object.
5. The method for target object recognition and pose estimation based on 3D vision according to claim 4, characterized in that: The process of obtaining the voxel grid is as follows; Obtain the coordinates (xi, yi, zi) of any coordinate point in the point cloud data of the object in the warehouse, where i is the number of the coordinate point, i=1, 2, ..., n, and n is a positive integer. Record the corresponding coordinate point as a voxel coordinate point, and calculate the distance between the voxel coordinate point and all other coordinate points; If the distance between the voxel coordinate point and any coordinate point is greater than a first distance threshold, it is determined that the voxel coordinate point and the coordinate point do not belong to the same voxel grid; If the distance between the voxel coordinate point and any coordinate point is less than or equal to the first distance threshold, it is determined that the voxel coordinate point and the coordinate point belong to the same voxel grid.
6. The method for target object recognition and pose estimation based on 3D vision according to claim 4, characterized in that: The process of obtaining the warehouse sub-area is as follows: Traverse the coordinates of all coordinate points in the same voxel grid, perform calculations based on the traversal results, calculate the length, width and height of the voxel grid, and calculate the volume of the voxel grid; Then, the number of coordinate points in the voxel grid is obtained to calculate the voxel grid density; If the voxel grid density is less than or equal to the density threshold, it is determined that the corresponding voxel grid is not a warehouse sub-area, and the corresponding voxel grid is removed; If the voxel grid density is greater than the density threshold, the corresponding voxel grid is determined to be a warehouse sub-region.
7. The method for target object recognition and pose estimation based on 3D vision according to claim 4, characterized in that: The process of obtaining the three-dimensional contour is: Draw six sets of planes tangent to the outermost edge points of the three-dimensional contour model from the six directions of up and down, left and right, front and back of the three-dimensional contour model of the object; The three-dimensional area enclosed by the six groups of planes is the three-dimensional outline of the object.
8. The method for target object recognition and pose estimation based on 3D vision according to claim 1, characterized in that: The step S3 includes the following sub-steps: Step S31, denoising the object image in the warehouse, and then cropping the object image, the resolution of the cropped object image is 800×600; Step S32, inputting the cropped object image into the deep learning model for detection, and detecting and obtaining the object category corresponding to the object image; Step S33, obtaining the shooting coordinates of the object image in the three-dimensional coordinate system according to the shooting position of the object image, and then calculating the distance between the shooting coordinates and the center coordinates corresponding to the geometric center points of all objects and traversing, traversing to obtain the center coordinates corresponding to the geometric center point of the object with the smallest distance from the shooting coordinates, and matching the object to which the center coordinates belong with the object category corresponding to the object image.
Citation Information
Patent Citations
6D pose estimation method using contour information
CN110322510A
Target object recognition and pose estimation method and device based on 3D vision
CN111127556A