Class-level 6D attitude estimation method and system, medium and program product
The target object in the image is processed through neural networks and pose estimation is performed using the features of external rectangles and contour points, which solves the disadvantages of relying on point clouds or key points in the prior art, and achieves faster pose reasoning and stronger generalization capabilities.
Patent Information
- Application Number
- CN202510186620.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-20
AI Technical Summary
The existing 6D pose estimation method relies on high-quality point cloud data or pre-set key points, and has slow processing speed and high requirements for depth information, making it difficult to meet the needs of real-time applications.
The target object in the image is processed through a neural network, and the real external rectangle and contour points are used for branch mapping, standard rectangle parameters and contour potential characteristics are obtained, and the posture of the object is estimated.
There is no need to rely on point clouds or pre-set key points, which reduces the difficulty and speed of data acquisition, speeds up the pose reasoning process, and improves the generalization ability and robustness of objects that have not been seen in the same type.
Smart Images

Figure CN120182365A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a category-level 6D pose estimation method, system, computer-readable storage medium, and computer program product. Background Art
[0002] Six-dimensional (6D) pose estimation is a fundamental problem in computer vision and is widely applied in multiple fields such as industrial robot grasping, unmanned aerial vehicle flight control, virtual reality, augmented reality, and autonomous driving. Existing 6D pose estimation methods are mainly divided into instance-level and category-level. Specifically, instance-level pose estimation can only estimate the pose of a specific instance that has been trained, while category-level pose estimation can estimate the pose of an unseen instance among objects of the same category. Currently, the existing technologies mainly include:
[0003] (1) Method based on prior point cloud deformation: By constructing a point cloud prior model for each category of objects, using an RGB image (color image) and a depth image (D image) to guide the generation of a deformed instance point cloud from the category point cloud, and finally achieving the alignment of the depth image constructed point cloud and the deformed point cloud through rigid matching to obtain the pose of the object. However, the disadvantages of this method are that it relies on high-quality point cloud data, has a slow processing speed (the rigid matching process between point clouds), and has high requirements for depth information, which all limit the possibility of real-time applications.
[0004] (2) Method based on key point matching: By obtaining 3D key points of an object for compact representation and performing tracking and matching in consecutive RGB-D frames. The disadvantage of this method is that the acquisition and matching of key points are not stable enough and are easily affected by noise and occlusion. Summary of the Invention
[0005] The purpose of the embodiments of the present invention is to provide a category-level 6D pose estimation method, system, computer-readable storage medium, and computer program product, which do not rely on point clouds or pre-set key points, nor on rigid matching between point clouds, but process the target object in the image through a neural network, effectively accelerating the pose inference process.
[0006] The first aspect embodiment of the present invention provides a category-level 6D pose estimation method, including:
[0007] According to the image of the target object, obtain the true circumscribed rectangle and true contour points of each component;
[0008] Input the rectangle parameters of the true circumscribed rectangle and the true contour points into the corresponding common mapping models for branch mapping respectively to obtain standard rectangle parameters and contour latent features; wherein, the common mapping model is used to map the target object to a standard object model to obtain common geometric information;
[0009] Convert the standard rectangular parameters into circumscribed rectangle features, and input the circumscribed rectangle features and the contour potential features into a pose estimation model to obtain the central rotation angle and depth of the target object;
[0010] Based on the camera internal parameter matrix, the depth, and the position of the center point of the target object on the image, calculate the horizontal offset, vertical offset, and offset rotation matrix of the target object;
[0011] Based on the offset rotation matrix and the central rotation angle, calculate the true rotation angle of the target object.
[0012] Optionally, the rectangular parameters include: the abscissa of the center point of the circumscribed rectangle, the ordinate of the center point of the circumscribed rectangle, the length of the long side of the circumscribed rectangle, the length of the short side of the circumscribed rectangle, and the rotation angle of the long side.
[0013] Optionally, the step of inputting the rectangular parameters of the true circumscribed rectangle and the true contour points into corresponding common mapping models for branch mapping to obtain standard rectangular parameters and contour potential features includes:
[0014] Input the rectangular parameters of all the true circumscribed rectangles into a first mapping model to obtain the standard rectangular parameters of all the components;
[0015] Input all the true contour points into a second mapping model to obtain the standard contour points and contour potential features of all the components; wherein, the second mapping model includes: an encoder, a latent feature mean embedding layer, and a decoder; the contour potential feature is the output of the latent feature mean embedding layer.
[0016] Optionally, the step of converting the standard rectangular parameters into circumscribed rectangle features includes:
[0017] Input all the standard rectangular parameters into a preset neural network model to obtain the output of the hidden layer of the last layer;
[0018] Use the output of the hidden layer as the circumscribed rectangle feature.
[0019] Optionally, the horizontal offset and vertical offset of the target object are calculated by the following formula:
[0020]
[0021] where zc is the depth; (u, v) are the coordinates corresponding to the position of the center point of the target object on the image; fx and fy are the focal lengths of the camera; (cx, cy) are the coordinates of the principal point of the camera; xc is the horizontal offset; and yc is the vertical offset.
[0022] Optionally, the offset rotation matrix is obtained through the following steps:
[0023] Obtain the position vector [xc, yc, zc] of the target object T ; where xc is the horizontal offset, yc is the vertical offset, and zc is the depth;
[0024] Calculate the position vector [xc, yc, zc] T and the included angle between the axis vector [0, 0, zc] T ;
[0025] Convert the included angle into a three-dimensional rotation matrix to obtain the offset rotation matrix.
[0026] Optionally, calculating the true rotation angle of the target object based on the offset rotation matrix and the central rotation angle includes:
[0027] Convert the central rotation angle into a three-dimensional rotation matrix to obtain a central rotation matrix;
[0028] Multiply the offset rotation matrix by the central rotation matrix to obtain a true rotation matrix;
[0029] Convert the true rotation matrix into Euler angles to obtain the true rotation angle.
[0030] An embodiment of the second aspect of the present invention provides a category-level 6D pose estimation system, including:
[0031] A data acquisition module, configured to obtain the true circumscribed rectangle and true contour points of each component according to the image of the target object;
[0032] A standard mapping module, configured to input the rectangle parameters of the true circumscribed rectangle and the true contour points into corresponding common mapping models for branch mapping respectively, to obtain standard rectangle parameters and contour latent features; wherein, the common mapping model is used to map the target object to a standard object model to obtain common geometric information;
[0033] A pose estimation module, configured to convert the standard rectangle parameters into circumscribed rectangle features, and input the circumscribed rectangle features and the contour latent features into a pose estimation model to obtain the central rotation angle and depth of the target object;
[0034] An offset calculation module, configured to calculate the horizontal offset, vertical offset and offset rotation matrix of the target object based on the camera internal parameter matrix, the depth and the center point position of the target object on the image;
[0035] A rotation angle calculation module, configured to calculate the true rotation angle of the target object based on the offset rotation matrix and the central rotation angle.
[0036] In a third aspect of the embodiments of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium includes a stored computer program. Wherein, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the category-level 6D pose estimation method described in any of the embodiments of the first aspect above.
[0037] In a fourth aspect of the embodiments of the present invention, a computer program product is provided, including computer instructions. When the computer instructions are executed by a processor, the category-level 6D pose estimation method described in any of the embodiments of the first aspect above is implemented.
[0038] Compared with the prior art, the embodiments of the present invention provide a category-level 6D pose estimation method, system, computer-readable storage medium, and computer program product. On the one hand, the contour sampling points and circumscribed rectangles of the two-dimensional imaging of each component in the target object are used as prior knowledge, so that there is no need to rely on point clouds or pre-set key points, reducing the difficulty and requirements for obtaining raw data. On the other hand, there is no need for rigid matching from point cloud to point cloud, but the target object in the image is processed through a neural network, accelerating the pose inference speed. On the other hand, through the consistency of the component structure (i.e., the mapping of the standard object model), when dealing with unseen objects of the same type, it can show strong generalization ability and robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a flowchart of an embodiment of the category-level 6D pose estimation method provided by the present invention;
[0040] Figure 2 is an imaging example diagram of an object instance rotating in space provided by the present invention;
[0041] Figure 3 is an example diagram of the central imaging and offset imaging of an object instance provided by the present invention;
[0042] Figure 4 is a flowchart of another embodiment of the category-level 6D pose estimation method provided by the present invention;
[0043] Figure 5 is a structural diagram of an embodiment of the category-level 6D pose estimation system provided by the present invention;
[0044] Figure 6 is a structural diagram of another embodiment of the category-level 6D pose estimation system provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art in the technical field of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0046] See Figure 1 , which is a schematic flowchart of an embodiment of the category-level 6D pose estimation method provided by the present invention.
[0047] The first aspect of the embodiments of the present invention provides a category-level 6D pose estimation method, including steps S1 to S5, specifically as follows:
[0048] Step S1: According to the image of the target object, obtain the true circumscribed rectangle and true contour points of each component;
[0049] Step S2: Input the rectangle parameters of the true circumscribed rectangle and the true contour points into the corresponding common mapping models for branch mapping respectively to obtain standard rectangle parameters and contour latent features; wherein, the common mapping model is used to map the target object to a standard object model to obtain common geometric information;
[0050] Step S3: Convert the standard rectangle parameters into circumscribed rectangle features, and input the circumscribed rectangle features and the contour latent features into a pose estimation model to obtain the central rotation angle and depth of the target object;
[0051] Step S4: Based on the camera internal parameter matrix, the depth, and the center point position of the target object on the image, calculate the horizontal offset, vertical offset, and offset rotation matrix of the target object;
[0052] Step S5: Based on the offset rotation matrix and the central rotation angle, calculate the true rotation angle of the target object.
[0053] It should be noted that the true circumscribed rectangle refers to the smallest rectangular boundary around the component in the image; the true contour points refer to the discrete point set on the edge of the component.
[0054] Specifically, in step S1, through a target detection algorithm, the image containing the target object is recognized, and the true circumscribed rectangle and true contour points of each component in the target object are extracted; wherein, the true circumscribed rectangle is used to describe the external boundary of the object, and the true contour points are helpful for capturing the shape information of the object.
[0055] Step S2 can eliminate the differences in appearance among different objects, thus ensuring the wide applicability of the algorithm. In practical applications, different object instances may have some geometric differences, although they belong to the same category. For example, two mugs of different brands may have slightly different shapes and sizes of the cup body and the handle; if the geometric information of these two objects is directly compared, inaccurate estimates may be obtained. For this reason, the embodiment of the present invention creates a "standard object model" to represent the typical geometric shape / information of a mug at the same rotation angle; then, the geometric information of each instance (i.e., mugs of different brands) is mapped onto this standard object model, thereby eliminating the error caused by geometric differences.
[0056] Step S3 is used to predict the basic pose information of the target object, that is, the central rotation angle and depth of the target object; among them, the central rotation angle is used to describe the rotation angle when the target object is located at the center of the camera field of view; the depth is the depth distance of the target object relative to the camera lens, usually along the optical axis of the camera.
[0057] The horizontal offset in Step S4 is the offset of the target object in the horizontal direction relative to the camera optical axis, and a positive value indicates that the object is on the right side of the camera field of view, while a negative value indicates that the object is on the left side of the camera field of view. The vertical offset is the offset of the target object in the vertical direction relative to the camera optical axis, and a positive value indicates that the object is below the camera field of view, while a negative value indicates that the object is above the camera field of view. The offset rotation matrix is the rotation change generated after the object is offset from the camera center position.
[0058] The true rotation angle in Step S5 is the total rotation matrix obtained after applying the offset rotation to the target object, representing the final pose of the target object, which is a combined pose formed by the rotation of the target object after offset (i.e., the rotation after deviating from the camera center position) and the rotation at the camera center position.
[0059] On the one hand, the embodiment of the present invention uses the contour sampling points and the circumscribed rectangle of the two-dimensional imaging of each component in the target object as prior knowledge, thus eliminating the need to rely on point clouds or pre-set key points, reducing the difficulty and requirements for obtaining raw data; on the other hand, there is no need for rigid matching from point cloud to point cloud, but the target object in the image is processed through a neural network, accelerating the pose inference speed; on the other hand, through the consistency of the component structure (i.e., the mapping of the standard object model), when dealing with unseen objects of the same category, it can show strong generalization ability and robustness.
[0060] In an alternative embodiment, the rectangle parameters include: the abscissa of the center point of the circumscribed rectangle, the ordinate of the center point of the circumscribed rectangle, the length of the long side of the circumscribed rectangle, the length of the short side of the circumscribed rectangle, and the rotation angle of the long side.
[0061] Further, the conversion of the standard rectangle parameters into the circumscribed rectangle features includes:
[0062] Input all the standard rectangle parameters into a preset neural network model to obtain the output of the hidden layer of the last layer;
[0063] Use the output of the hidden layer as the circumscribed rectangle features.
[0064] It should be noted that in the embodiments of the present invention, the sampling points (i.e., real contour points) of the outer contour of the object component and the component rotation detection frame (i.e., real circumscribed rectangle) are used as the imaging geometric information of the object component.
[0065] The specific mathematical representation is as follows:
[0066] I(ci, R) = f(ci, R);
[0067] Wherein, R is the self-rotation angle of the object; ci is the i-th component of the object; f is an information extraction function for extracting the imaging geometric information after the object rotates; I(ci, R) is the imaging geometric information of the component ci at the rotation angle R.
[0068] In the embodiments of the present invention, by analyzing the imaging geometric information of each component, it can be used to estimate the pose of the object.
[0069] In an alternative embodiment, the step of inputting the rectangle parameters of the real circumscribed rectangle and the real contour points into the corresponding common mapping models for branch mapping to obtain the standard rectangle parameters and the contour latent features includes:
[0070] Input all the rectangle parameters of the real circumscribed rectangle into the first mapping model to obtain the standard rectangle parameters of all the components;
[0071] Input all the real contour points into the second mapping model to obtain the standard contour points and the contour latent features of all the components; wherein, the second mapping model includes: an encoder, a latent feature mean embedding layer, and a decoder; the contour latent feature is the output of the latent feature mean embedding layer.
[0072] It should be noted that the main principle of the embodiments of the present invention is as follows:
[0073] (1) The similarity of objects of the same type.
[0074] The objects existing in reality are always composed of a fixed number of components with similar structures; for example, a thermos cup is usually composed of two components, a cup body and a cup lid, and a mug is usually composed of two components, a cup body and a cup handle; the structures of thermos cups of different brands are similar, and the relative positions and shapes of the cup body and the cup lid are also similar.
[0075] (2) Similarities and differences in imaging with different self-rotation angles.
[0076] When the self-rotation angle of an object is different, the geometric information contained in the imaging of each component in the object will also be different. For example Figure 2 shown is an imaging example diagram of an object instance provided by the present invention when rotating in space.
[0077] Figure 2 In it, at different rotation angles, the imaging geometric information of components in the same instance is different; at the same rotation angle, the imaging geometric information of components in different instances of the same type is similar. This similarity and difference provide effective constraints for estimating the self-rotation angle of objects of the same type.
[0078] (3) Mapping of instances to a standard object model.
[0079] At the same self-rotation angle, the imaging geometric information of components in different instances of the same type is similar but not exactly the same. In an embodiment of the present invention, by setting the imaging geometric information of a standard object model, the geometric information of an instance is first mapped to the geometric information of the standard object model for more accurate pose estimation.
[0080] The specific mathematical representation is as follows:
[0081]
[0082] Where R is the self-rotation angle of the object; ci is the i-th component of the instance; cni is the i-th component of the standard object model; m is the mapping function of the geometric information of the instance to the standard object model; I(ci, R) is the imaging geometric information of component ci in the instance at rotation angle R; is the imaging geometric information of component cni in the standard object model at rotation angle R.
[0083] In an embodiment of the present invention, the mapping function m is actually a commonality mapping model; among them, the commonality mapping model is a first mapping model or a second mapping model. Through the mapping function m, common geometric features can be found between different object instances, reducing the subtle differences between different instances of the same type, thereby improving the accuracy and generalization ability of category-level pose estimation.
[0084] In an optional embodiment, the horizontal offset and vertical offset of the target object are calculated by the following formula:
[0085]
[0086] where zc is the depth; (u, v) are the coordinates corresponding to the center point position of the target object in the image; fx and fy are the focal lengths of the camera; (cx, cy) are the principal point coordinates of the camera; xc is the horizontal offset; and yc is the vertical offset.
[0087] It should be noted that is the camera internal parameter matrix, which is usually a known matrix; the depth zc is obtained through a pose estimation model. In the embodiments of the present invention, the center point pixel coordinates (u, v) of the target object on the image can be converted into the three-dimensional position (xc, yc, zc) in the camera coordinate system through the above formula, which describes the position of the object in the camera's field of view.
[0088] In the camera coordinate system, xc is the horizontal offset of the object relative to the camera optical axis (left - right direction); yc is the vertical offset of the object relative to the camera optical axis (up - down direction); zc is the depth distance of the object relative to the camera lens (along the positive direction of the camera optical axis); xc and yc determine the offset of the object on the camera imaging plane, and zc determines the distance of the object from the camera.
[0089] It is worth noting that if it is necessary to further convert from the camera coordinate system to the world coordinate system, only the external parameters of the camera (such as Rw and T) need to be simply measured, and the specific conversion formula is as follows:
[0090] where Rw is the rotation matrix of the camera in the world coordinate system; T is the translation vector of the camera in the world coordinate system; and (xw, yw, zw) are the object coordinates in the world coordinate system.
[0091] In an optional embodiment, calculating the true rotation angle of the target object based on the offset rotation matrix and the center rotation angle includes:
[0092] Performing a three - dimensional rotation matrix conversion on the center rotation angle to obtain a center rotation matrix;
[0093] Multiplying the offset rotation matrix by the center rotation matrix to obtain a true rotation matrix;
[0094] Performing an Euler angle conversion on the true rotation matrix to obtain the true rotation angle.
[0095] Furthermore, the offset rotation matrix is obtained through the following steps:
[0096] Obtaining the position vector [xc, yc, zc] of the target object T ; where xc is the horizontal offset, yc is the vertical offset, and zc is the depth;
[0097] Calculate the position vector [xc, yc, zc] T and the axis vector [0, 0, zc] T to obtain the included angle therebetween;
[0098] Perform a three-dimensional rotation matrix transformation on the included angle to obtain the offset rotation matrix.
[0099] It should be noted that 6D pose estimation requires estimating the position and orientation of an object in three-dimensional space. If the pose estimations at various positions of the object in space are calculated, the solution space is too large. In other words, for a certain object, it can freely move and rotate in space, which will result in countless possible imaging results (i.e., the images captured by the camera).
[0100] As Figure 3 shown, it is an example diagram of central imaging and offset imaging of an object instance provided by the present invention. The inventor found that the imaging of the object at any position and any self-angle in space is always approximately the same as the imaging of the object when it is at a certain self-rotation angle and on the camera center line. In the embodiments of the present invention, by finding the rotation angle of the object when the imaging on the center line is the same, and then calculating the angle offset value of the object when it is not on the center line, the true rotation angle (actual rotation angle) of the object can be obtained.
[0101] The specific mathematical representation is as follows:
[0102] Rtotal = Roffset · Rcenter;
[0103] wherein, Rtotal is the true rotation matrix, that is, the total rotation matrix obtained after the object is applied with the offset rotation; Rcenter is the self-rotation matrix of the object when it is on the center line, which is estimated through a pose estimation model; Roffset is the offset rotation matrix, which is used to describe the rotation of the object from the center line position to the offset position.
[0104] The central imaging is the imaging result when the object is on the camera optical axis. At this time, the pose of the object can be described by the central rotation matrix Rcenter. The offset imaging refers to the imaging result after the object deviates from the camera optical axis. In the camera field of view, when the object on the center line is deflected, additional rotation changes will be introduced, and this part of the change is described by Roffset. The total rotation matrix / true rotation matrix is a complete description of the actual pose of the object, which is calculated by the combination of Roffset and Rcenter.
[0105] In the embodiments of the present invention, through the decomposition of central imaging and offset imaging, the complex three-dimensional pose problem is transformed into two relatively simple rotation problems, which can significantly reduce the computational complexity and improve the computational efficiency.
[0106] To further describe the category-level 6D pose estimation method provided by the embodiments of the present invention more clearly, a specific application example of a technical solution in the R & D process of the inventor will be elaborated in detail or referred to below.
[0107] It should be noted that each neural network model involved in the embodiments of the present invention (such as the first mapping model and the second mapping model in the common mapping model, the preset neural network model, and the pose estimation model) can be any neural network model in the prior art. The specific model to be adopted shall be determined when the embodiments of the present invention are applied to specific products or technologies.
[0108] As Figure 4 shown, it is a flowchart of another embodiment of the category-level 6D pose estimation method provided by the present invention. Figure 4 In Figure 4 the specific implementation process is shown in Table 1.
[0109] Table 1. Example Table of the Implementation Process of the Category-Level 6D Pose Estimation Method
[0110]
[0111]
[0112] In Table 1, Neural Network 1 is the first mapping model, Neural Network 2 is the second mapping model, Neural Network 3 is the preset neural network model; Neural Network 4 is the pose estimation model.
[0113] The position offset Y, the position offset X, and the depth Z respectively correspond to the horizontal offset xc, the vertical offset yc, and the depth zc in the foregoing text. In addition, the component scale refers to the size of the length, width, and height of the circumscribed cube of the actual object component in the three-dimensional space. Although 6D pose estimation refers to estimating the position of an object in space and its own rotation angle, in this task, the estimation of the actual size of the object is generally also involved.
[0114] See Figure 5 , which is a schematic structural diagram of an embodiment of the category-level 6D pose estimation system provided by the present invention.
[0115] In a second aspect embodiment of the present invention, a category-level 6D pose estimation system is provided for implementing the category-level 6D pose estimation method described in any embodiment of the first aspect above. The system includes:
[0116] A data acquisition module 11 for obtaining the true circumscribed rectangle and true contour points of each component according to the image of the target object;
[0117] A standard mapping module 12 for respectively inputting the rectangle parameters of the true circumscribed rectangle and the true contour points into corresponding common mapping models for branch mapping to obtain standard rectangle parameters and contour latent features; wherein, the common mapping model is used to map the target object to a standard object model to obtain common geometric information;
[0118] A pose estimation module 13 for converting the standard rectangle parameters into circumscribed rectangle features and inputting the circumscribed rectangle features and the contour latent features into a pose estimation model to obtain the central rotation angle and depth of the target object;
[0119] An offset calculation module 14 for calculating the horizontal offset, vertical offset and offset rotation matrix of the target object based on the camera internal parameter matrix, the depth and the position of the center point of the target object on the image;
[0120] A rotation angle calculation module 15 for calculating the true rotation angle of the target object based on the offset rotation matrix and the central rotation angle.
[0121] Optionally, the standard mapping module 12 includes:
[0122] A circumscribed rectangle frame mapping unit 121 for inputting the rectangle parameters of all the true circumscribed rectangles into a first mapping model to obtain the standard rectangle parameters of all the components;
[0123] A sparse edge point encoding and decoding unit 122 for inputting all the true contour points into a second mapping model to obtain the standard contour points and contour latent features of all the components; wherein, the second mapping model includes: an encoder, a latent feature mean embedding layer and a decoder; the contour latent feature is the output of the latent feature mean embedding layer.
[0124] It should be noted that the main purpose of the circumscribed rectangle mapping unit 121 is to obtain the standard rectangle parameters after mapping; the main purpose of the sparse edge point encoding and decoding unit 122 is to perform encoding feature extraction on the sparse contour information of the object and then decode and restore it to the standardized level. The sparse edge point encoding and decoding unit 122 includes: a sparse edge point representation subunit 1221 and a feature encoding and decoding subunit 1222; wherein, the feature encoding and decoding subunit 1222 is used to obtain the contour latent features of all components.
[0125] In specific implementation, as Figure 6 shown, it is a schematic structural diagram of another embodiment of the category-level 6D pose estimation system provided by the present invention. The data acquisition module 11 includes: an object detection unit 111 and an object segmentation unit 112; wherein, the object detection unit 111 is used to perform target detection on the object to obtain the true circumscribed rectangle of each component in the object; the object segmentation unit 112 is used to segment the object to obtain the true contour points of each component in the object. The pose estimation module 13 includes: a feature extraction unit 131, a feature fusion unit 132, and a pose calculation unit 133; wherein, the feature extraction unit 131 is used to convert the standard rectangle parameters into circumscribed rectangle features; the feature fusion unit 132 is used to combine the circumscribed rectangle features and the contour latent features to form the input features of the pose estimation model; the pose calculation unit 133 is used to obtain the central rotation angle and depth of the target object.
[0126] It should be noted that the category-level 6D pose estimation system provided in the second aspect embodiment of the present invention can implement all the processes of the category-level 6D pose estimation method described in any of the above first aspect embodiments. The functions and achieved technical effects of each module in the system are respectively the same as those of the category-level 6D pose estimation method described in any of the above first aspect embodiments, and will not be elaborated here.
[0127] The third aspect embodiment of the present invention provides a computer-readable storage medium, and the computer-readable storage medium includes a stored computer program; wherein, the computer program controls the device where the computer-readable storage medium is located to execute the category-level 6D pose estimation method described in any of the above first aspect embodiments when running.
[0128] The fourth aspect embodiment of the present invention provides a computer program product, including computer instructions, and the computer instructions implement the category-level 6D pose estimation method described in any of the above first aspect embodiments when executed by a processor.
[0129] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A category-level 6D pose estimation method, characterized in that: include: According to the image of the target object, the real bounding rectangle and the real contour points of each component are obtained; Inputting the rectangle parameters of the true circumscribed rectangle and the true contour points into the corresponding commonality mapping model for branch mapping, respectively, to obtain standard rectangle parameters and contour potential features; wherein the commonality mapping model is used to map the target object to a standard object model to obtain common geometric information; Converting the standard rectangle parameters into circumscribed rectangle features, and inputting the circumscribed rectangle features and the contour potential features into a pose estimation model to obtain the central rotation angle and depth of the target object; Based on the camera intrinsic parameter matrix, the depth and the center point position of the target object on the image, a horizontal offset, a vertical offset and an offset rotation matrix of the target object are calculated; Based on the offset rotation matrix and the center rotation angle, the real rotation angle of the target object is calculated.
2. The class-level 6D pose estimation method according to claim 1, characterized in that: The rectangle parameters include: the horizontal coordinate of the center point of the circumscribed rectangle, the vertical coordinate of the center point of the circumscribed rectangle, the length of the long side of the circumscribed rectangle, the length of the short side of the circumscribed rectangle, and the rotation angle of the long side.
3. The class-level 6D pose estimation method according to claim 1, characterized in that: The step of inputting the rectangle parameters of the true circumscribed rectangle and the true contour points into the corresponding commonality mapping model for branch mapping to obtain standard rectangle parameters and contour potential features includes: Inputting the rectangle parameters of all the real circumscribed rectangles into the first mapping model to obtain the standard rectangle parameters of all the components; All the true contour points are input into the second mapping model to obtain the standard contour points and contour latent features of all the components; wherein the second mapping model comprises: an encoder, a latent feature mean embedding layer and a decoder; the contour latent feature is the output of the latent feature mean embedding layer.
4. The class-level 6D pose estimation method according to claim 1, characterized in that: The converting the standard rectangle parameters into the circumscribed rectangle features comprises: Input all the standard rectangle parameters into the preset neural network model to obtain the hidden layer output of the last layer; The hidden layer output is used as the circumscribed rectangle feature.
5. The class-level 6D pose estimation method according to claim 1, characterized in that: The horizontal offset and vertical offset of the target object are calculated by the following formula: Among them, zc is the depth; (u, v) is the coordinate corresponding to the center point position of the target object on the image; fx and fy are the focal lengths of the camera; (cx, cy) are the principal point coordinates of the camera; xc is the horizontal offset; and yc is the vertical offset.
6. The class-level 6D pose estimation method according to claim 1, characterized in that: The offset rotation matrix is obtained by the following steps: Get the position vector [xc, yc, zc] of the target object T ; Wherein, xc is the horizontal offset, yc is the vertical offset, and zc is the depth; Calculate the position vector [xc, yc, zc] T and the axis vector [0,0,zc] T The angle between The angle is transformed into a three-dimensional rotation matrix to obtain the offset rotation matrix.
7. The class-level 6D pose estimation method according to claim 1, characterized in that: The calculating the real rotation angle of the target object based on the offset rotation matrix and the center rotation angle includes: Performing a three-dimensional rotation matrix conversion on the central rotation angle to obtain a central rotation matrix; Multiplying the offset rotation matrix by the center rotation matrix to obtain a true rotation matrix; The true rotation matrix is converted into Euler angles to obtain the true rotation angle.
8. A category-level 6D pose estimation system, characterized in that: include: A data acquisition module, used to obtain the true circumscribed rectangle and true contour points of each component according to the image of the target object; A standard mapping module, used to input the rectangle parameters of the real circumscribed rectangle and the real contour points into the corresponding common mapping model for branch mapping, so as to obtain standard rectangle parameters and contour potential features; wherein the common mapping model is used to map the target object to a standard object model to obtain common geometric information; A pose estimation module, used for converting the standard rectangle parameters into circumscribed rectangle features, and inputting the circumscribed rectangle features and the contour potential features into a pose estimation model to obtain the central rotation angle and depth of the target object; An offset calculation module, used to calculate the horizontal offset, vertical offset and offset rotation matrix of the target object based on the camera intrinsic parameter matrix, the depth and the center point position of the target object on the image; The rotation angle calculation module is used to calculate the real rotation angle of the target object based on the offset rotation matrix and the center rotation angle.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program; wherein, when the computer program is run, it controls the device where the computer-readable storage medium is located to execute the category-level 6D pose estimation method according to any one of claims 1 to 7.
10. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the category-level 6D pose estimation method according to any one of claims 1 to 7.
Citation Information
Cited By
Method and device for quickly estimating 6D attitude of target object
CN120976317A