6D pose estimation method and device, electronic equipment and storage medium

Through the 2D object detection and image matching method, combined with the neural network model to update the pose residual, the problem of strong dependence on the 3D model in the prior art is solved, and efficient and accurate 6D pose estimation is achieved to adapt to the changes of different objects and individual objects.

CN120279089APending Publication Date: 2025-07-08江淮前沿技术协同创新中心
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510146446.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art has strong dependence on the 3D model in 6D pose estimation and has poor adaptability to scene changes, making it difficult to efficiently and accurately estimate the three-dimensional position and orientation of an object.

Method used

By acquiring the camera image and reference image set of the object to be measured, the initial translation vector and rotation vector are determined using the 2D object detection and image matching method, and the pose residual update is performed in combination with the preset neural network model to realize the 6D pose estimation.

Benefits of technology

There is no need to collect 6D pose data sets, and only retraining the 2D object detection model and a small amount of image matching data can be performed to perform 6D pose estimation, which improves the accuracy and adaptability of the estimation and can distinguish different objects from different individuals of the same object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279089A_ABST
    Figure CN120279089A_ABST
Patent Text Reader

Abstract

The invention discloses a 6D pose estimation method and device, electronic equipment and a storage medium, and relates to the technical field of computer vision. The method comprises the steps that a camera image of a measured object and a reference image set of the measured object are acquired, and the reference image set of the measured object comprises 6D poses of the measured object under different poses; determining an initial translation vector of the interested object according to the camera image of the detected object, a preset target detection model and the camera parameters; determining an initial rotation vector of the measured object according to the camera image of the measured object and the reference image set of the measured object; determining an initial 6D pose of the measured object according to the camera coordinate of the measured object, the initial translation vector of the measured object and the initial rotation vector of the measured object; and determining a pose residual error according to the initial 6D pose of the measured object and the reference image set of the measured object, and updating the initial 6D pose of the measured object according to the pose residual error to determine a target 6D pose of the measured object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular, to a 6D pose estimation method, a 6D pose estimation device, a computer-readable storage medium, and an electronic device. Background Art

[0002] Traditional computer vision tasks mainly focus on two-dimensional information such as object classification, detection, and segmentation. However, in many practical application scenarios, relying solely on two-dimensional information is far from sufficient. For example, in robot grasping and assembly tasks, the robot needs to know the three-dimensional position and orientation of the object in order to perform precise operations. This makes inferring the 6D pose of an object from a two-dimensional image a challenging problem.

[0003] 6D pose estimation is an important research direction in the fields of computer vision and robotics. Its core lies in accurately estimating the position and pose of an object in three-dimensional space (i.e., three position coordinates and three rotation angles, 6 degrees of freedom). With the rapid development of artificial intelligence technology, 6D pose estimation has broad prospects and significance in many practical applications. In intelligent manufacturing, the automated operation and precise assembly of robots are core technologies, and 6D pose estimation provides fundamental support for this. By accurately perceiving the pose of an object, the robot can complete precise grasping, assembly, and handling tasks, thereby significantly improving production efficiency and quality. In an actual industrial environment, the shape, size, material, and position of objects often vary, which poses extremely high requirements for the accuracy and robustness of pose estimation. In VR (Virtual Reality) and AR (Augmented Reality) scenarios, 6D pose estimation makes it possible to integrate virtual objects with the real world. Through accurate pose estimation, virtual objects can seamlessly interact and integrate with the real scene, enhancing the user's sense of immersion. Especially in AR applications, 6D pose estimation can help virtual objects be accurately projected into the real scene, and users can obtain a more realistic experience when viewing or interacting with virtual objects.

[0004] 3D vision has witnessed a huge development from robots, games to VR / AR. These applications have put forward higher requirements for 6D pose estimation methods. In the early research of 6D pose estimation, model matching methods were mainly used. Such methods usually need to obtain the 3D model of the object first, generate the feature points of the object using point cloud data or CAD (Computer Aided Design) models, and match these feature points with the feature points in the image to estimate the pose of the object. However, they are highly dependent on the 3D model and have poor adaptability to scene changes. With the development of deep learning technology, in recent years, 6D pose estimation methods based on deep learning have gradually become the mainstream. Deep learning methods construct large-scale datasets, extract image features using convolutional neural networks, and directly predict the 6D pose of the object through regression or classification methods. However, it is difficult to collect 6D pose datasets and the model training is troublesome. Summary of the Invention

[0005] This application aims to solve at least one of the technical problems in the related art to some extent. For this reason, the first objective of this application is to propose a 6D pose estimation method, which includes: obtaining the camera image of the object to be measured and the set of reference images of the object to be measured, where the set of reference images of the object to be measured includes the 6D poses of the object to be measured in different postures; determining the initial translation vector of the object of interest according to the camera image of the object to be measured, the preset target detection model and the camera parameters; determining the initial rotation vector of the object to be measured according to the camera image of the object to be measured and the set of reference images of the object to be measured; determining the initial 6D pose of the object to be measured according to the camera coordinates of the object to be measured, the initial translation vector of the object to be measured and the initial rotation vector of the object to be measured; determining the pose residual according to the initial 6D pose of the object to be measured and the set of reference images of the object to be measured, and updating the initial 6D pose of the object to be measured according to the pose residual to determine the target 6D pose of the object to be measured. In this way, the pose estimation method of this application can determine the 6D pose of the object to be measured by using 2D target detection and image matching methods, without collecting 6D pose datasets; for different objects, only the 2D target detection model and a small amount of image matching data need to be retrained to estimate the 6D pose, without a complex model training process, and at the same time, the advantages of 2D target detection can also be exerted, which can not only distinguish different objects, but also distinguish different individuals of the same object.

[0006] The second objective of this application is to propose a 6D pose estimation device.

[0007] The third objective of this application is to propose a computer-readable storage medium.

[0008] The fourth objective of this application is to propose an electronic device.

[0009] To achieve the above object, an embodiment of the first aspect of the present application proposes a 6D pose estimation method, which includes: obtaining a camera image of an object to be measured and a set of reference images of the object to be measured, where the set of reference images of the object to be measured includes the 6D poses of the object to be measured in different poses; determining an initial translation vector of an object of interest according to the camera image of the object to be measured, a preset target detection model, and camera parameters; determining an initial rotation vector of the object to be measured according to the camera image of the object to be measured and the set of reference images of the object to be measured; determining an initial 6D pose of the object to be measured according to the camera coordinates of the object to be measured, the initial translation vector of the object to be measured, and the initial rotation vector of the object to be measured; determining a pose residual according to the initial 6D pose of the object to be measured and the set of reference images of the object to be measured, and updating the initial 6D pose of the object to be measured according to the pose residual to determine the target 6D pose of the object to be measured.

[0010] According to an embodiment of the present application, determining a pose residual according to the initial 6D pose of the object to be measured and the set of reference images of the object to be measured, and updating the initial 6D pose of the object to be measured according to the pose residual to determine the target 6D pose of the object to be measured includes: obtaining the cosine similarity distance between the camera image of the object to be measured and each reference image in the set of reference images of the object to be measured; sorting all the cosine similarity distances in ascending order, and obtaining the reference images corresponding to the top preset number of cosine similarity distances as the target reference image set; determining a first feature information set based on the target reference image set and a preset 2D neural network model, and projecting each first feature information in the first feature information set to a 3D feature layer according to its pose to obtain the 3D feature of each first feature information, and calculating the mean and variance of the first feature information set based on the 3D features of all the first feature information in the first feature information set; determining a second feature information based on the camera image of the object to be measured and the preset 2D neural network model, and projecting the second feature information to the 3D feature layer according to the initial 6D pose of the object to be measured to obtain the 3D feature of the second feature information; determining a pose residual according to the mean and variance of the first feature information set, the 3D feature of the second feature information, and a preset 3D neural network model.

[0011] According to an embodiment of the present application, determining the initial rotation vector of the object to be measured according to the camera image of the object to be measured and the set of reference images of the object to be measured includes: obtaining the cosine similarity distance between the camera image of the object to be measured and each reference image in the set of reference images of the object to be measured; using the rotation vector of the reference image corresponding to the minimum value in the cosine similarity distances as the initial rotation vector of the object to be measured.

[0012] According to an embodiment of the present application, obtaining the cosine similarity distance between the camera image of the object to be measured and each reference image in the reference image set of the object to be measured includes: inputting the camera image of the object to be measured into a preset neural network model to determine a first feature vector; inputting the reference image set of the object to be measured into the preset neural network model to determine a second feature vector set; calculating the cosine similarity distance between the first feature vector and each second feature vector in the second feature vector set as the cosine similarity distance between the camera image of the object to be measured and each reference image in the reference image set of the object to be measured.

[0013] According to an embodiment of the present application, determining an initial translation vector of the object of interest based on the camera image of the object to be measured, a preset target detection model, and camera parameters includes: inputting the camera image of the object to be measured into the preset target detection model to determine the circumscribed rectangle of the object to be measured; obtaining the center point coordinates of the circumscribed rectangle of the object to be measured, the depth of the center point, and the camera parameters; determining the initial translation vector of the object to be measured based on the center point coordinates of the circumscribed rectangle of the object to be measured, the depth of the center point, and the camera parameters.

[0014] According to an embodiment of the present application, the center point coordinates of the circumscribed rectangle of the object to be measured include the abscissa of the center point of the circumscribed rectangle of the object to be measured in the pixel coordinate system and the ordinate of the center point of the circumscribed rectangle of the object to be measured in the pixel coordinate system, and the camera parameters include the focal length of the camera in the horizontal direction, the focal length of the camera in the vertical direction, the abscissa of the camera optical center in the pixel coordinate system, and the ordinate of the camera optical center in the pixel coordinate system. Determining the initial translation vector of the object to be measured based on the center point coordinates of the circumscribed rectangle of the object to be measured, the depth of the center point, and the camera parameters includes: determining a first difference based on the difference between the abscissa of the center point of the circumscribed rectangle of the object to be measured in the pixel coordinate system and the abscissa of the camera optical center in the pixel coordinate system; determining a second difference based on the difference between the ordinate of the camera optical center in the pixel coordinate system and the ordinate of the center point of the circumscribed rectangle of the object to be measured in the pixel coordinate system; determining a first ratio based on the ratio of the first difference to the focal length of the camera in the horizontal direction; determining a second ratio based on the ratio of the second difference to the focal length of the camera in the vertical direction; determining the initial translation vector of the object to be measured based on the product of the first ratio and the depth of the center point, the product of the second ratio and the depth of the center point, and the depth of the center point.

[0015] According to an embodiment of the present application, the preset target detection model is the YOLOv8 target detection model, and the preset neural network model is a siamese neural network model.

[0016] To achieve the above object, an embodiment of the second aspect of the present application provides a 6D pose estimation device, which includes: an acquisition module, configured to acquire a camera image of an object to be measured and a set of reference images of the object to be measured, where the set of reference images of the object to be measured includes the 6D poses of the object to be measured in different poses; a first determination module, configured to determine an initial translation vector of an object of interest according to the camera image of the object to be measured, a preset target detection model, and camera parameters; a second determination module, configured to determine an initial rotation vector of the object to be measured according to the camera image of the object to be measured and the set of reference images of the object to be measured; a third determination module, configured to determine an initial 6D pose of the object to be measured according to the camera coordinates of the object to be measured, the initial translation vector of the object to be measured, and the initial rotation vector of the object to be measured; a fourth determination module, configured to determine a pose residual according to the initial 6D pose of the object to be measured and the set of reference images of the object to be measured, and update the initial 6D pose of the object to be measured according to the pose residual to determine the target 6D pose of the object to be measured.

[0017] To achieve the above object, an embodiment of the third aspect of the present application provides a computer-readable storage medium, on which a 6D pose estimation program is stored. When the 6D pose estimation program is executed by a processor, the foregoing 6D pose estimation method is implemented.

[0018] To achieve the above object, an embodiment of the fourth aspect of the present application provides an electronic device, including a memory, a processor, and a 6D pose estimation program stored on the memory and executable on the processor. When the processor executes the 6D pose estimation program, the foregoing 6D pose estimation method is implemented.

[0019] 6D pose estimation method, device, electronic device and storage medium according to an embodiment of the present application. Obtain a camera image of an object to be measured and a set of reference images of the object to be measured. The set of reference images of the object to be measured includes the 6D poses of the object to be measured in different poses. Determine an initial translation vector of an object of interest according to the camera image of the object to be measured, a preset target detection model and camera parameters. Determine an initial rotation vector of the object to be measured according to the camera image of the object to be measured and the set of reference images of the object to be measured. Determine an initial 6D pose of the object to be measured according to the camera coordinates of the object to be measured, the initial translation vector of the object to be measured and the initial rotation vector of the object to be measured. Determine a pose residual according to the initial 6D pose of the object to be measured and the set of reference images of the object to be measured, and update the initial 6D pose of the object to be measured according to the pose residual to determine the target 6D pose of the object to be measured. Thus, the pose estimation method of the present application can determine the 6D pose of the object to be measured by using the methods of 2D object detection and image matching, without collecting a 6D pose data set. For different objects, only need to retrain the 2D object detection model and a small amount of image matching data to estimate the 6D pose, without a complex model training process, and can also give play to the advantages of 2D object detection, which can not only distinguish different objects, but also distinguish different individuals of the same object. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a flowchart of a 6D pose estimation method according to some embodiments of the present application;

[0021] Figure 2 is a schematic diagram of data acquisition according to some embodiments of the present application;

[0022] Figure 3 is a schematic diagram of pose regression according to some embodiments of the present application;

[0023] Figure 4 is a block schematic diagram of a 6D pose estimation device according to some embodiments of the present application;

[0024] Figure 5 is a block schematic diagram of an electronic device according to some embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, but should not be construed as limiting the present application.

[0026] The 6D pose estimation method, device, electronic device and storage medium according to the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0027] Figure 1 It is a flowchart of a 6D pose estimation method according to some embodiments of the present application. Referring to Figure 1 , the 6D pose estimation method of the embodiments of the present application may include the following steps:

[0028] S110, obtain a camera image of the object to be measured and a set of reference images of the object to be measured, and the set of reference images of the object to be measured includes the 6D poses of the object to be measured in different poses.

[0029] Specifically, the camera image of the object to be measured, such as the RGB (Red Green Blue) image of the object to be measured, can be obtained by camera shooting. Referring to Figure 2 , place the object to be measured statically at the shooting position, ensure that the object itself or the object background has sufficient texture, and use the camera to shoot a reference video of the object to be measured. Then, segment the captured reference video into an image sequence to form a set of reference images of the object to be measured (RGB image set of the object to be measured). Further, input the ordered image set into COLMAP (Color and Mosaic Photo scanner), through consistent search and incremental reconstruction, use SFM (Structure from Motion) to generate a 3D point cloud scene. Manually specify the object area by cropping the point cloud of the object, and at the same time manually specify the x+ direction and z+ direction of the point cloud object to complete the establishment of the coordinate system. Adopt the right-hand coordinate system principle to generate a txt file marked with the orientation of xyz, that is, generate the 6D poses of the object to be measured in different poses.

[0030] In this way, by simply using the camera to take a set of RGB pictures of the object to be measured, the three-dimensional point cloud data of the object and the set of reference images required for pose estimation can be generated through the algorithm, without obtaining an accurate CAD model of the object to be measured.

[0031] S120, determine the initial translation vector of the object of interest according to the camera image of the object to be measured, a preset target detection model, and camera parameters.

[0032] Specifically, the camera image of the object to be measured is usually very large, and the object to be measured only occupies a small part of the camera image. Therefore, in order to pay more attention to the object to be measured, input the camera image of the object to be measured into a preset target detection model for target detection, and segment the circumscribed rectangle of the object to be measured. Then, determine the initial translation vector of the object of interest according to the center point coordinates of the circumscribed rectangle, the depth of the center point of the circumscribed rectangle, and the camera parameters.

[0033] S130. Determine the initial rotation vector of the object to be measured based on the camera image of the object to be measured and the set of reference images of the object to be measured.

[0034] Specifically, perform image similarity matching on the camera image of the object to be measured and each reference image of the object to be measured in the set of reference images of the object to be measured. For example, calculate the similarity value between the camera image of the object to be measured and each reference image of the object to be measured in the set of reference images of the object to be measured, obtain the reference image that is most similar to the camera image of the object to be measured, and use the rotation vector of the most similar reference image as the initial rotation vector of the object to be measured.

[0035] S140. Determine the initial 6D pose of the object to be measured based on the camera coordinates of the object to be measured, the initial translation vector of the object to be measured, and the initial rotation vector of the object to be measured.

[0036] Specifically, after obtaining the camera image of the object to be measured, the feature points of the camera image can be determined by performing feature point detection on the camera image of the object to be measured, and the pixel coordinates of these feature points in the image can be directly read as the pixel coordinates of the object to be measured. Convert the pixel coordinates of the object to be measured into camera coordinates, and combine the initial translation vector of the object to be measured and the initial rotation vector of the object to be measured to determine the initial 6D pose of the object to be measured.

[0037] Exemplarily, assuming that the pixel coordinates of the object to be measured are (u, v), the pixel coordinates (u, v) are converted into image coordinates (x, y) through the following formula:

[0038]

[0039] where W represents the width of the image and H represents the height of the image.

[0040] The image coordinates (x, y) need to be converted into the camera coordinates (X, Y, Z) of the object to be measured through the following formula under the action of depth:

[0041]

[0042] where d represents the depth.

[0043] Convert the camera coordinates (X, Y, Z) of the object to be measured into the world coordinates (Xw, Yw, Zw) of the object to be measured through the following formula based on the camera coordinates of the object to be measured, the initial translation vector of the object to be measured, and the initial rotation vector of the object to be measured, so as to determine the initial 6D pose of the object to be measured:

[0044]

[0045] where R represents the initial rotation matrix of the object to be measured and t represents the initial translation vector of the object to be measured.

[0046] It should be noted that the initial rotation matrix can be calculated from the initial rotation vector.

[0047] S150. Determine the pose residual according to the initial 6D pose of the object to be measured and the set of reference images of the object to be measured, and update the initial 6D pose of the object to be measured according to the pose residual to determine the target 6D pose of the object to be measured.

[0048] Specifically, extract the feature information of the camera image of the object to be measured, and project the feature information of the camera image onto the 3D feature layer according to the initial 6D pose to form the 3D feature of the feature information of the camera image.

[0049] Extract the feature information of each reference image from the set of reference images of the object to be measured, and project the feature information of each reference image onto the 3D feature layer according to its pose to form the 3D feature of the feature information of the reference image.

[0050] Input the 3D feature of the feature information of the camera image and the 3D feature of the feature information of the reference image into a preset 3D neural network model, such as 3D CNN (3D Convolutional Neural Network), analyze the inconsistencies therein, regress to obtain the pose residual, so as to obtain a more accurate translation vector and rotation vector, and use the more accurate translation vector and rotation vector to update the initial 6D pose of the object to be measured to determine the target 6D pose of the object to be measured.

[0051] In this way, the pose estimation method of the present application can determine the 6D pose of the object to be measured by using the methods of 2D object detection and image matching, without collecting a 6D pose data set; for different objects, only need to retrain the 2D object detection model and a small amount of image matching data to estimate the 6D pose, without a complex model training process, and at the same time can also give play to the advantages of 2D object detection, not only can distinguish different objects, but also can distinguish different individuals of the same object.

[0052] In some embodiments, a pose residual is determined based on the initial 6D pose of the object to be measured and a set of reference images of the object to be measured, and the initial 6D pose of the object to be measured is updated based on the pose residual to determine the target 6D pose of the object to be measured, including: obtaining the cosine similarity distance between the camera image of the object to be measured and each reference image in the set of reference images of the object to be measured; sorting all the cosine similarity distances in ascending order, and obtaining the reference images corresponding to the top preset number of cosine similarity distances as the target reference image set; determining a first feature information set based on the target reference image set and a preset 2D neural network model, and projecting each first feature information in the first feature information set onto a 3D feature layer according to its pose to obtain the 3D feature of each first feature information, and calculating the mean and variance of the first feature information set based on the 3D features of all the first feature information in the first feature information set; determining a second feature information based on the camera image of the object to be measured and the preset 2D neural network model, and projecting the second feature information onto the 3D feature layer according to the initial 6D pose of the object to be measured to obtain the 3D feature of the second feature information; determining the pose residual according to the mean and variance of the first feature information set, the 3D feature of the second feature information, and the preset 3D neural network model.

[0053] Specifically, referring to Figure 3 , after obtaining the cosine similarity distance between the camera image of the object to be measured and each reference image in the set of reference images of the object to be measured, all the cosine similarity distances are sorted in ascending order. The smaller the distance, the more similar the images are. The reference images corresponding to the top preset number (for example, the top 6) of cosine similarity distances are obtained as the target reference image set.

[0054] Each reference image in the target reference image set is input into a preset 2D neural network model (such as 2DCNN) for feature extraction to form a first feature information set, and each first feature information in the first feature information set is projected onto a 3D feature layer according to its pose to obtain the 3D feature of each first feature information. After all the first feature information in the first feature information set is projected, the corresponding 3D features of the first feature information are spliced together, and statistical analysis is performed on all the 3D features in the first feature information set to calculate the mean and variance of the first feature information set.

[0055] The camera image of the object to be measured is input into the same preset 2D neural network model, such as 2DCNN (2D Convolutional Neural Network), for feature extraction to determine the second feature information, and the second feature information is projected onto the 3D feature layer according to the initial 6D pose of the object to be measured to obtain the 3D feature of the second feature information.

[0056] The mean and variance of the first feature information set and the 3D features of the second feature information are input into a preset 3D neural network model (such as 3D CNN) to analyze the inconsistency between the mean and variance of the first feature information set and the 3D features of the second feature information. After multiple iterative calculations, the pose residual is regressed.

[0057] In some embodiments, determining an initial rotation vector of the object to be measured according to the camera image of the object to be measured and a set of reference images of the object to be measured includes: obtaining the cosine similarity distance between the camera image of the object to be measured and each reference image in the set of reference images of the object to be measured; using the rotation vector of the reference image corresponding to the minimum value in the cosine similarity distances as the initial rotation vector of the object to be measured.

[0058] Specifically, after obtaining the cosine similarity distance between the camera image of the object to be measured and each reference image in the set of reference images of the object to be measured, all the cosine similarity distances are sorted in ascending order, the reference image corresponding to the minimum value in the cosine similarity distances is obtained, and the rotation vector of the reference image corresponding to the minimum value in the cosine similarity distances is used as the initial rotation vector of the object to be measured.

[0059] In some embodiments, obtaining the cosine similarity distance between the camera image of the object to be measured and each reference image in the set of reference images of the object to be measured includes: inputting the camera image of the object to be measured into a preset neural network model to determine a first feature vector; inputting the set of reference images of the object to be measured into the preset neural network model to determine a set of second feature vectors; calculating the cosine similarity distance between the first feature vector and each second feature vector in the set of second feature vectors as the cosine similarity distance between the camera image of the object to be measured and each reference image in the set of reference images of the object to be measured.

[0060] Specifically, the camera image of the object to be measured is input into a preset neural network model (such as a siamese neural network model) for feature extraction to determine a first feature vector. Each reference image in the set of reference images of the object to be measured is input into the preset neural network model (such as a siamese neural network model) for feature extraction to determine a second feature vector. After all the reference images in the set of reference images are subjected to feature extraction, a set of second feature vectors is formed. The cosine similarity distance between the first feature vector and each second feature vector in the set of second feature vectors is calculated respectively, for example, the element product of the first feature vector and the second feature vector, to obtain the cosine similarity distance between the camera image of the object to be measured and each reference image in the set of reference images of the object to be measured. It should be noted that the set of second feature vectors can be extracted and saved in a matrix in advance, which is beneficial to quickly predicting the rotation of the camera image during actual prediction.

[0061] In some embodiments, determining an initial translation vector of an object of interest based on a camera image of an object to be measured, a preset target detection model, and camera parameters includes: inputting the camera image of the object to be measured into the preset target detection model to determine a circumscribed rectangle of the object to be measured; obtaining the center point coordinates of the circumscribed rectangle of the object to be measured, the depth of the center point, and the camera parameters; and determining the initial translation vector of the object to be measured according to the center point coordinates of the circumscribed rectangle of the object to be measured, the depth of the center point, and the camera parameters.

[0062] Specifically, inputting the camera image of the object to be measured into the preset target detection model to determine the circumscribed rectangle of the object to be measured, connecting the two diagonals of the circumscribed rectangle, and the intersection point of the two diagonals is the center point of the circumscribed rectangle of the object to be measured, and the coordinates of this center point are (u c , v c ), where u c is the abscissa of the center point in the pixel coordinate system, and where v c is the ordinate of the center point in the pixel coordinate system. The depth of this center point can be obtained by a depth camera. The camera parameters include the focal length f x of the camera in the horizontal direction, the focal length f y of the camera in the vertical direction, the abscissa c x of the camera optical center in the pixel coordinate system, and the ordinate c y of the camera optical center in the pixel coordinate system. Among them, the camera parameters can be obtained through camera calibration. According to the abscissa u c of the center point in the pixel coordinate system, the ordinate v c of the center point in the pixel coordinate system, the focal length f x of the camera in the horizontal direction, the focal length f y of the camera in the vertical direction, the abscissa c x of the camera optical center in the pixel coordinate system, the ordinate c y of the camera optical center in the pixel coordinate system, and the depth of the center point, the initial translation vector of the object to be measured is determined.

[0063] In some embodiments, the center point coordinates of the circumscribed rectangle of the object to be measured include the abscissa of the center point of the circumscribed rectangle of the object to be measured in the pixel coordinate system and the ordinate of the center point of the circumscribed rectangle of the object to be measured in the pixel coordinate system. The camera parameters include the focal length of the camera in the horizontal direction, the focal length of the camera in the vertical direction, the abscissa of the camera optical center in the pixel coordinate system, and the ordinate of the camera optical center in the pixel coordinate system. Determining the initial translation vector of the object to be measured based on the center point coordinates of the circumscribed rectangle of the object to be measured, the depth of the center point, and the camera parameters includes: determining a first difference based on the difference between the abscissa of the center point of the circumscribed rectangle of the object to be measured in the pixel coordinate system and the abscissa of the camera optical center in the pixel coordinate system; determining a second difference based on the difference between the ordinate of the camera optical center in the pixel coordinate system and the ordinate of the center point of the circumscribed rectangle of the object to be measured in the pixel coordinate system; determining a first ratio based on the ratio of the first difference to the focal length of the camera in the horizontal direction; determining a second ratio based on the ratio of the second difference to the focal length of the camera in the vertical direction; and determining the initial translation vector of the object to be measured based on the product of the first ratio and the depth of the center point, the product of the second ratio and the depth of the center point, and the depth of the center point.

[0064] Exemplarily, the calculation formula of the initial translation vector is as follows:

[0065]

[0066] Z c = d

[0067]

[0068] where, u c represents the abscissa of the center point in the pixel coordinate system, v c represents the ordinate of the center point in the pixel coordinate system, c x represents the abscissa of the camera optical center in the pixel coordinate system, c y represents the ordinate of the camera optical center in the pixel coordinate system, f x represents the focal length of the camera in the horizontal direction, f y represents the focal length of the camera in the vertical direction, d represents the depth of the center point, and t represents the initial translation vector of the object to be measured.

[0069] In some embodiments, the preset object detection model is the YOLOv8 object detection model, and the preset neural network model is the Siamese neural network model.

[0070] Specifically, referring to Figure 2, after the set of reference images of the object to be measured, it is also necessary to manually annotate the 2D image frames (bounding rectangles) for the ordered images and create a VOC (a standard data set) data set. Use the VOC data set to train a preset object detection model, and the preset object detection model can be the YOLOv8 object detection model. By using the YOLOv8 object detection model, not only can the 6D poses of different types of objects be distinguished, but also the 6D poses of multiple objects of the same type can be distinguished, greatly expanding the scope of 6D pose estimation.

[0071] The preset neural network model can be a Siamese neural network model. The Siamese neural network consists of two or more sub-networks sharing weights and can well compare the similarity between the input camera image and the reference image of the object to be measured.

[0072] It should be noted that during the actual grasping process, after determining the target 6D pose of the object to be measured, that is, determining the 3D bounding rectangle of the object to be measured, the first set of 3D point coordinates (8 3D point coordinates) can be determined according to the vertex coordinates of the 3D bounding rectangle of the object to be measured. According to the actual size of the object to be measured, the second set of 3D point coordinates (8 3D point coordinates) can also be defined. Modify the first set of 3D point coordinates according to the second set of 3D point coordinates to obtain the target set of 3D point coordinates. Use EPNP to calculate the 6D pose according to the target set of 3D point coordinates, the corresponding 2D pixel coordinates, and the camera internal parameters. EPNP packages this 6D pose in the OpenCV library and can be directly called during the actual grasping process.

[0073] In summary, the pose estimation method of the present application can determine the 6D pose of the object to be measured by adopting the methods of 2D object detection and image matching, without collecting a 6D pose data set; for different objects, only need to retrain the 2D object detection model to estimate the 6D pose, without a complex model training process, and at the same time can also give play to the advantages of 2D object detection, not only can distinguish different objects, but also can distinguish different individuals of the same object.

[0074] Corresponding to the above embodiment, the present application also proposes a 6D pose estimation device.

[0075] Referring to Figure 4 , the 6D pose estimation device 400 includes: an acquisition module 410, a first determination module 420, a second determination module 430, a third determination module 440, and a fourth determination module 450.

[0076] Among them, the acquisition module 210 is configured to acquire a camera image of the object to be measured and a set of reference images of the object to be measured, and the set of reference images of the object to be measured includes the 6D poses of the object to be measured in different poses. The first determination module 220 is configured to determine an initial translation vector of the object of interest according to the camera image of the object to be measured, a preset target detection model, and camera parameters. The second determination module 230 is configured to determine an initial rotation vector of the object to be measured according to the camera image of the object to be measured and the set of reference images of the object to be measured. The third determination module 240 is configured to determine an initial 6D pose of the object to be measured according to the camera coordinates of the object to be measured, the initial translation vector of the object to be measured, and the initial rotation vector of the object to be measured. The fourth determination module 250 is configured to determine a pose residual according to the initial 6D pose of the object to be measured and the set of reference images of the object to be measured, and update the initial 6D pose of the object to be measured according to the pose residual to determine the target 6D pose of the object to be measured.

[0077] According to an embodiment of the present application, the fourth determination module 250 is specifically configured to obtain the cosine similarity distance between the camera image of the object to be measured and each reference image in the set of reference images of the object to be measured; sort all the cosine similarity distances in ascending order, and obtain the reference images corresponding to the first preset number of cosine similarity distances as the target reference image set; determine a first feature information set based on the target reference image set and a preset 2D neural network model, project each first feature information in the first feature information set to a 3D feature layer according to its pose, so as to obtain the 3D feature of each first feature information, and calculate the mean and variance of the first feature information set based on the 3D features of all the first feature information in the first feature information set; determine a second feature information based on the camera image of the object to be measured and a preset 2D neural network model, and project the second feature information to the 3D feature layer according to the initial 6D pose of the object to be measured, so as to obtain the 3D feature of the second feature information; determine the pose residual according to the mean and variance of the first feature information set, the 3D feature of the second feature information, and a preset 3D neural network model.

[0078] According to an embodiment of the present application, the second determination module 230 is specifically configured to obtain the cosine similarity distance between the camera image of the object to be measured and each reference image in the set of reference images of the object to be measured; use the rotation vector of the reference image corresponding to the minimum value in the cosine similarity distance as the initial rotation vector of the object to be measured.

[0079] According to an embodiment of the present application, the camera image of the object to be measured is input into a preset neural network model to determine a first feature vector; the reference image set of the object to be measured is input into the preset neural network model to determine a second feature vector set; the cosine similarity distance between the first feature vector and each second feature vector in the second feature vector set is calculated as the cosine similarity distance between the camera image of the object to be measured and each reference image in the reference image set of the object to be measured.

[0080] According to an embodiment of the present application, the first determination module 220 is specifically configured to input the camera image of the object to be measured into a preset object detection model to determine the circumscribed rectangle of the object to be measured; obtain the center point coordinates of the circumscribed rectangle of the object to be measured, the depth of the center point, and the camera parameters; and determine the initial translation vector of the object to be measured according to the center point coordinates of the circumscribed rectangle of the object to be measured, the depth of the center point, and the camera parameters.

[0081] According to an embodiment of the present application, the center point coordinates of the circumscribed rectangle of the object to be measured include the abscissa of the center point of the circumscribed rectangle of the object to be measured in the pixel coordinate system and the ordinate of the center point of the circumscribed rectangle of the object to be measured in the pixel coordinate system, and the camera parameters include the focal length of the camera in the horizontal direction, the focal length of the camera in the vertical direction, the abscissa of the camera optical center in the pixel coordinate system, and the ordinate of the camera optical center in the pixel coordinate system. The first determination module 220 is specifically configured to determine a first difference based on the difference between the abscissa of the center point of the circumscribed rectangle of the object to be measured in the pixel coordinate system and the abscissa of the camera optical center in the pixel coordinate system; determine a second difference based on the difference between the ordinate of the camera optical center in the pixel coordinate system and the ordinate of the center point of the circumscribed rectangle of the object to be measured in the pixel coordinate system; determine a first ratio based on the ratio of the first difference to the focal length of the camera in the horizontal direction; determine a second ratio based on the ratio of the second difference to the focal length of the camera in the vertical direction; and determine the initial translation vector of the object to be measured according to the product of the first ratio and the depth of the center point, the product of the second ratio and the depth of the center point, and the depth of the center point.

[0082] According to an embodiment of the present application, the preset object detection model is the YOLOv8 object detection model, and the preset neural network model is a siamese neural network model.

[0083] It should be noted that the above explanations of the embodiments and beneficial effects of the 6D pose estimation method are also applicable to the 6D pose estimation device of the embodiments of the present application. To avoid redundancy, no detailed elaboration will be made here.

[0084] Corresponding to the above embodiments, the present application also proposes a computer-readable storage medium.

[0085] The computer-readable storage medium of the present application stores a 6D pose estimation program thereon. When the 6D pose estimation program is executed by a processor, the foregoing 6D pose estimation method is implemented.

[0086] It should be noted that the above explanations of the embodiments and beneficial effects of the 6D pose estimation method are also applicable to the computer-readable storage medium of the embodiments of the present application. To avoid redundancy, no detailed elaboration will be made here.

[0087] Corresponding to the above embodiments, the present application also proposes an electronic device.

[0088] See Figure 5 As shown, the electronic device 500 of the present application includes a memory 510, a processor 520, and a 6D pose estimation program stored on the memory 510 and executable on the processor 520. When the processor executes the 6D pose estimation program, the foregoing 6D pose estimation method is implemented.

[0089] It should be noted that the above explanations of the embodiments and beneficial effects of the 6D pose estimation method are also applicable to the electronic device of the embodiments of the present application. To avoid redundancy, no detailed elaboration will be made here.

[0090] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.

[0091] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0092] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0093] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined.

[0094] In the present application, unless otherwise clearly specified and limited, the terms "mounted", "connected", "connected to", "fixed", etc. should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements or the interaction relationship between two elements, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0095] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A 6D pose estimation method, characterized in that, The method includes: Obtaining a camera image of the object to be measured and a set of reference images of the object to be measured, where the set of reference images of the object to be measured includes the 6D poses of the object to be measured in different poses; Determining an initial translation vector of the object of interest according to the camera image of the object to be measured, a preset target detection model, and camera parameters; Determining an initial rotation vector of the object to be measured according to the camera image of the object to be measured and the set of reference images of the object to be measured; Determining an initial 6D pose of the object to be measured according to the camera coordinates of the object to be measured, the initial translation vector of the object to be measured, and the initial rotation vector of the object to be measured; Determining a pose residual according to the initial 6D pose of the object to be measured and the set of reference images of the object to be measured, and updating the initial 6D pose of the object to be measured according to the pose residual to determine the target 6D pose of the object to be measured.

2. The 6D pose estimation method according to claim 1, wherein Determining a pose residual according to the initial 6D pose of the object to be measured and the set of reference images of the object to be measured, and updating the initial 6D pose of the object to be measured according to the pose residual to determine the target 6D pose of the object to be measured, including: Obtaining the cosine similarity distance between the camera image of the object to be measured and each reference image in the set of reference images of the object to be measured; Sorting all the cosine similarity distances in ascending order, and obtaining the reference images corresponding to the first preset number of the cosine similarity distances as the target reference image set; Determining a first set of feature information based on the target reference image set and a preset 2D neural network model, projecting each first feature information in the first set of feature information to a 3D feature layer according to its pose to obtain the 3D feature of each first feature information, and calculating the mean and variance of the first set of feature information based on the 3D features of all the first feature information in the first set of feature information; Determining a second feature information based on the camera image of the object to be measured and the preset 2D neural network model, and projecting the second feature information to the 3D feature layer according to the initial 6D pose of the object to be measured to obtain the 3D feature of the second feature information; Determining the pose residual according to the mean and variance of the first set of feature information, the 3D feature of the second feature information, and a preset 3D neural network model.

3. The 6D pose estimation method according to claim 1, wherein Determining the initial rotation vector of the object to be measured according to the camera image of the object to be measured and the set of reference images of the object to be measured, including: Obtaining the cosine similarity distance between the camera image of the object to be measured and each reference image in the set of reference images of the object to be measured; Taking the rotation vector of the reference image corresponding to the minimum value in the cosine similarity distance as the initial rotation vector of the object to be measured.

4. The 6D pose estimation method according to claim 2 or 3, characterized in that, Obtaining the cosine similarity distance between the camera image of the object to be measured and each reference image in the set of reference images of the object to be measured, including: Inputting the camera image of the object to be measured into a preset neural network model to determine a first feature vector; Input the reference image set of the object to be measured into the preset neural network model to determine the second feature vector set; Calculate the cosine similarity distance between the first feature vector and each second feature vector in the second feature vector set as the cosine similarity distance between the camera image of the object to be measured and each reference image in the reference image set of the object to be measured.

5. The 6D pose estimation method according to claim 1, characterized in that Determine the initial translation vector of the object of interest according to the camera image of the object to be measured, the preset object detection model and the camera parameters, including: Input the camera image of the object to be measured into the preset object detection model to determine the circumscribed rectangle of the object to be measured; Obtain the center point coordinates of the circumscribed rectangle of the object to be measured, the depth of the center point and the camera parameters; Determine the initial translation vector of the object to be measured according to the center point coordinates of the circumscribed rectangle of the object to be measured, the depth of the center point and the camera parameters.

6. The 6D pose estimation method according to claim 5, wherein The center point coordinates of the circumscribed rectangle of the object to be measured include the abscissa of the center point of the circumscribed rectangle of the object to be measured in the pixel coordinate system and the ordinate of the center point of the circumscribed rectangle of the object to be measured in the pixel coordinate system. The camera parameters include the focal length of the camera in the horizontal direction, the focal length of the camera in the vertical direction, the abscissa of the camera optical center in the pixel coordinate system and the ordinate of the camera optical center in the pixel coordinate system. Determining the initial translation vector of the object to be measured according to the center point coordinates of the circumscribed rectangle of the object to be measured, the depth of the center point and the camera parameters includes: Determine the first difference based on the difference between the abscissa of the center point of the circumscribed rectangle of the object to be measured in the pixel coordinate system and the abscissa of the camera optical center in the pixel coordinate system; Determine the second difference based on the difference between the ordinate of the camera optical center in the pixel coordinate system and the ordinate of the center point of the circumscribed rectangle of the object to be measured in the pixel coordinate system; Determine the first ratio based on the ratio of the first difference to the focal length of the camera in the horizontal direction; Determine the second ratio based on the ratio of the second difference to the focal length of the camera in the vertical direction; Determine the initial translation vector of the object to be measured according to the product of the first ratio and the depth of the center point, the product of the second ratio and the depth of the center point and the depth of the center point.

7. The 6D pose estimation method according to claim 4, wherein The preset object detection model is the YOLOv8 object detection model, and the preset neural network model is the Siamese neural network model.

8. A 6D pose estimation device, characterized in that, The device includes: An acquisition module for acquiring the camera image of the object to be measured and the reference image set of the object to be measured, where the reference image set of the object to be measured includes the 6D poses of the object to be measured in different poses; A first determination module for determining the initial translation vector of the object of interest according to the camera image of the object to be measured, the preset object detection model and the camera parameters; A second determination module for determining the initial rotation vector of the object to be measured according to the camera image of the object to be measured and the reference image set of the object to be measured; A third determination module, configured to determine an initial 6D pose of the object to be measured according to the camera coordinates of the object to be measured, the initial translation vector of the object to be measured, and the initial rotation vector of the object to be measured; A fourth determination module, configured to determine a pose residual according to the initial 6D pose of the object to be measured and the reference image set of the object to be measured, and update the initial 6D pose of the object to be measured according to the pose residual, so as to determine the target 6D pose of the object to be measured.

9. A computer-readable storage medium, characterized in that, A 6D pose estimation program is stored thereon, and when the 6D pose estimation program is executed by a processor, the 6D pose estimation method according to any one of claims 1-7 is implemented.

10. An electronic device, characterized in that, It includes a memory, a processor, and a 6D pose estimation program stored on the memory and operable on the processor. When the processor executes the 6D pose estimation program, the 6D pose estimation method according to any one of claims 1-7 is implemented.