Object six-degree-of-freedom pose estimation method based on three-dimensional geometric information registration

Through the method based on three-dimensional geometric information registration, three-dimensional point clouds under the object coordinate system are generated and registered, which solves the accuracy and robustness problems of the traditional two-dimensional image feature method in complex scenarios, and achieves high-precision and high-rootability object position estimation.

CN120107347AActive Publication Date: 2025-06-06SHANDONG UNIV +1

Patent Information

Application Number
CN202510591680.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-06
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

The traditional two-dimensional image features and the six-degree-of-freedom pose estimation method of objects dependent on low-precision sensors are susceptible to noise, occlusion and lighting changes in complex scenarios, resulting in insufficient positioning accuracy and poor robustness.

Method used

The object six-degree-of-freedom pose estimation method based on three-dimensional geometric information registration is adopted. The object detection module, the depth ensemble feature extraction network and the spatial 3D point cloud decoding network are used to generate the three-dimensional point cloud under the object coordinate system, and the point cloud is registered using the DCP geometric information registration model, and the rotation and translation transformation are calculated to achieve pose estimation.

Benefits of technology

It improves the accuracy and robustness of the six-degree-of-freedom pose estimation of the object, reduces the problems of noise interference and error accumulation in traditional methods, significantly improves processing speed and adaptability, and is suitable for industrial automation, robot grasping and other fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107347A_ABST
    Figure CN120107347A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and particularly relates to an object six-degree-of-freedom pose estimation method based on three-dimensional geometric information registration, which comprises a target detection module, a three-dimensional point cloud generation model and a geometric information registration pose estimation module. According to the method, geometric reconstruction, a geometric information high-precision registration technology and an object pose estimation task are combined, a geometric reconstruction module introduces object geometric feature extraction and feature decoding to realize deep understanding and modeling of an algorithm on object geometric features, the whole process accords with geometric disorder, and the accuracy of object pose estimation is improved. And the accuracy and scene adaptability of the algorithm are further improved. According to the method, accurate estimation of the six-degree-of-freedom pose of the object can be achieved, and powerful support is provided for the fields of industrial automation, robot vision, AR / VR, automatic driving, medical robots and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of image processing technology, and specifically relates to a method for estimating the six-degree-of-freedom pose of an object based on three-dimensional geometric information registration. Background Art

[0002] In order to meet the growing demand for object positioning and posture accuracy in industrial automation, robot precision grasping, intelligent manufacturing, and augmented reality, related industries are constantly seeking more efficient and accurate six-degree-of-freedom posture estimation technology. Traditional posture estimation methods often rely on two-dimensional image features or low-precision sensors. These methods are easily affected by factors such as noise, occlusion, and lighting changes in complex scenes, resulting in insufficient positioning accuracy and poor robustness.

[0003] Relying on traditional algorithms and manually designed feature matching methods, when faced with high-precision requirements, problems such as high computational complexity, slow convergence speed, and easy to fall into local optimality often occur, making it difficult to achieve real-time monitoring and accurate detection at all angles and all time periods. At the same time, although the existing point cloud processing technology has made up for the lack of two-dimensional information to a certain extent, there are bottlenecks such as low computational efficiency and error accumulation in the process of point cloud data acquisition, preprocessing and registration, which makes it difficult to meet the requirements of modern high-precision industrial applications.

[0004] At present, with the rapid development of high-precision three-dimensional data acquisition technologies such as laser scanning and depth cameras, as well as the widespread application of new algorithms such as deep learning and global optimization in the field of computer vision, high-precision geometric information registration has gradually become an effective way to solve the problem of six-degree-of-freedom pose estimation of objects. Using advanced geometric information registration technology, the three-dimensional structural information of objects can be fully extracted and utilized to achieve all-round, multi-level accurate estimation of the position and posture of objects, thereby greatly improving the accuracy and stability of pose detection, and providing strong technical support for the refined operation of industrial automation and intelligent control systems. Summary of the invention

[0005] In order to solve the above technical problems, the present invention provides a method for estimating the six-degree-of-freedom pose of an object based on three-dimensional geometric information registration, so as to improve the accuracy and robustness of the six-degree-of-freedom pose estimation of the object, and provide strong support for industrial automation, robot vision, AR / VR, autonomous driving, medical robots and other fields.

[0006] To achieve the above object, the technical solution of the present invention is as follows: A method for estimating the six-degree-of-freedom pose of an object based on three-dimensional geometric information registration, comprising a target detection module, a three-dimensional point cloud generation model, and a geometric information registration pose estimation module; Object detection module for object detection and bounding box prediction; The 3D point cloud generation model includes a deep set feature extraction network and a spatial 3D point cloud decoding network; Deep ensemble feature extraction network: It combines RGB and depth information and converts RGBD images into global feature representations through an encoder; Spatial 3D point cloud decoding network: Generates disordered point cloud through decoder, completes the reconstruction of object camera coordinate system and 3D point cloud generation based on depth information, and realizes high-quality geometric information reconstruction of multiple coordinate systems driven by depth image; The geometric information registration pose estimation module registers the decoded three-dimensional point cloud in the object coordinate system and the calculated point cloud in the real camera coordinate system. The former is used as the source point cloud and the latter as the target point cloud. The rotation and translation transformation from the source point cloud to the target point cloud is calculated to realize the six-degree-of-freedom pose estimation of the corresponding object.

[0007] Preferably, the specific steps of the target detection module training phase are as follows: (1) Based on the RGBD data in the pose estimation dataset and the bounding box information of each object in the image, the YOLO format training dataset is constructed, screened and enhanced in the pose estimation scenario; (2) By introducing the PGI mechanism, the model can learn different types of targets in a more accurate and efficient way through an adjustable gradient update strategy during training. The gradient weighting module dynamically adjusts the gradient weight according to the object category and detection difficulty. (3) Optimize the feature extraction and information fusion process of the YOLOv9 network. In the multi-level feature aggregation stage, information is aggregated between the convolutional layers of the YOLOv9 network, combining low-level detail features with high-level semantic features to capture object information of different scales and complex backgrounds. At the same time, the adaptive feature fusion module dynamically adjusts the fusion strategy of different levels of features according to the complexity of the object and background interference, so that the model can more accurately identify and locate the target; (4) The loss is calculated based on the cross entropy, and the gradient is calculated by back propagation to update the parameters to complete the model training.

[0008] Preferably, the steps of training the 3D point cloud generation model are as follows: (1) Based on the trained YOLOv9 target detection model, the pose estimation dataset is preprocessed. For the input RGB image , corresponding to the depth map , use the RGB part of the image for target detection, and get the bounding box of each object in the image as well as the category and confidence information ,in Indicates the object category, Indicates confidence level; (2) Setting the confidence threshold And filter out all qualified high-confidence objects to obtain a set of valid bounding boxes , based on the bounding box, the RGB image and depth map are intercepted to obtain an image set containing the RGBD information of each object , so far the data set is completed; (3) For each RGBD image of the object in each batch during the training phase, calculate the point cloud in the real camera coordinate system , based on the real pose label of each object, calculate the point cloud in the real object coordinate system ; (4) For the RGBD image of the object in each batch during the training phase, the spatial 3D point cloud decoding network predicts the point cloud in the object coordinate system ,calculate and The Chamfer Distance is used as the loss and the back-propagation training model is performed.

[0009] Preferably, the training steps of the geometric information registration pose estimation module are as follows: (1) Based on the object RGBD dataset, for each batch of object RGBD images in the training phase, calculate the point cloud in the real camera coordinate system ; (2) For the RGBD images of the objects in each batch during the training phase, the point cloud in the object coordinate system is predicted based on the trained 3D point cloud generation model. ; (3) Based on the DCP geometric information registration network, the point cloud under the object coordinates is predicted As the source point cloud, the point cloud in the real camera coordinate system As the target point cloud, predict the object's pose, including the rotation matrix with translation vector ; (4) Calculate the true rotation matrix and The cross entropy loss , calculate the true translation vector and The cross entropy loss , + Back-propagation gradient calculation and model parameter update are performed as loss.

[0010] Preferably, the target detection module realizes the recognition and positioning of objects in RGB images based on YOLOv9, and segments RGBD images of different objects. The specific steps of the reasoning stage are as follows: (1) For the input RGB image , the corresponding depth map is , use the RGB part of the image for target detection, and get the bounding box of each object in the image as well as the category and confidence information ,in Indicates the object category, Indicates confidence; (2) Setting the confidence threshold And filter out all qualified high-confidence objects to obtain a set of valid bounding boxes , based on the bounding box, the RGB image and depth map are intercepted to obtain an image set containing the RGBD information of each object .

[0011] Preferably, the specific steps of the three-dimensional point cloud generation model reasoning stage are as follows: (1) For RGBD image input, the input channels need to be expanded from 3 to 4 to accommodate the input data. The modified convolution kernel size is , after the first layer of convolution, we get the first feature map The entire feature extraction part is composed of multiple residual stacks. From the input to the final feature map after passing through all the residual blocks, the pooling layer is added to compress the spatial dimension of the feature map into a feature vector ; (2) The spatial 3D point cloud decoding network further decodes the features extracted from the image and defines the tree structure of the decoding network , where the tree structure satisfies The number of point cloud points N, the initial input is the initial feature vector , after the 0th layer, you get high vit , the dimension changes to obtain the first layer feature set ; (3) The next layers 1 to L-2 cannot directly use the fully connected layer to achieve the spatial mapping of the two-dimensional matrix. One-dimensional convolution is used for feature mapping. In order to maintain the feature direction of the decoding process, the initial feature vector is concatenated before each calculation to obtain According to the tree structure described above, the first layer can be expanded times the number of features, that is, at this time, the features are further mapped to At this time, the feature dimension transformation is performed to further expand the number of feature sets to obtain the first layer output feature set ; (4) Similarly, from the 2nd layer to the L-2nd layer, feature decoding and feature set accumulation are continuously performed, and the number of features in the set is continuously approached to the number of points in the point cloud to be generated, until the L-2nd layer, at which time the feature set is obtained Similarly, after concatenating the initial feature vectors, the last layer maps each feature into a 3D space to obtain the generated point cloud .

[0012] Preferably, the geometric information registration pose estimation module realizes the registration of the source point cloud and the target point cloud. The specific steps of the reasoning stage are as follows: (1) Based on the RGBD image of the object, the depth information and the object segmentation information, the point cloud in the camera coordinate system is calculated, and the farthest point sampling is further performed to reduce the number of point cloud points to N, and the result is , , based on the 3D point cloud generation model, the point cloud in the object coordinate system is obtained , ; The point cloud in the object coordinate system is regarded as the source point cloud, and the point cloud in the camera coordinate system is regarded as the target point cloud. That is, the transformation from the source point cloud to the target point cloud is the pose of the object in the picture at this time; (2) Based on the feature mapping part of the PointNet network as the DCP feature extraction network, each point is processed independently by a multi-layer perceptron with shared parameters, and the high-dimensional features of each point are extracted to obtain the feature set corresponding to the target point cloud and the source point cloud. and , and further apply maximum pooling to the two feature sets to obtain the global feature vectors and , and then concatenated with the feature set output by the last layer to obtain and , use MLP to implement equal-dimensional mapping for each feature in the set at this time, and obtain the final high-dimensional feature set and ; (3) Two point cloud high-dimensional feature sets are used for self-attention calculation. The input point cloud high-dimensional features go through , , Generate a query (Query, Q), key (Key, K) and value (Value, V), , , , map the features to the feature_dim dimension, calculate the attention weight matrix A through the dot product operation, and then model the global information learned by the attention mechanism into the value vector by clicking A and the value vector, so as to realize dynamic attention to different features in the entire feature set. , the two point cloud feature sets are self-attention calculated separately, and we get and , , , They represent the mapping matrices of query vector, key vector, and value vector in self-attention calculation respectively; (4) Perform cross-attention calculation to learn the correspondence between the two, and use the query matrix to map the feature vector of the first point cloud to obtain , a key and value mapping matrix to map the feature vectors of the second point cloud, , , after cross attention calculation, we can get and ; (5) Calculate the similarity matrix of two point cloud feature sets ,in Indicates the similarity between the i-th point of the source point cloud and the j-th point of the target point cloud. At this time, the position with the highest similarity becomes the corresponding relationship between the two point clouds, and a two-way consistency check is adopted, that is, the best match from the source point cloud point to the target point cloud point. The best match of the target point cloud to the source point cloud Must be consistent, , (6) Extracting source point clouds that meet the corresponding relationship With the target point cloud , calculate the covariance matrix Represent the geometric relationship between point clouds and perform singular value decomposition on the covariance matrix , calculate the pose rotation matrix , and further calculate the translation vector based on the center of mass position ,in and They represent the centroids of the source point cloud and target point cloud calculated according to the correspondence.

[0013] Preferably, the target detection module implements the training of the YOLOv9 target detection model based on the information of the bounding box in the pose estimation data set; the target detection module performs object target detection on the pose estimation data set, and cuts out the RGBD images of all objects to construct a subsequent training data set; the three-dimensional point cloud generation model implements the training of the three-dimensional space geometric reconstruction model composed of the image feature extractor and the TopNet point cloud decoder based on the constructed object image RGBD data set and the pose labels of all objects; the geometric information registration module completes the training of the geometric information registration model DCP based on the point cloud in the camera coordinate system calculated based on the object RGBD image and the point cloud in the object coordinate system generated by the three-dimensional space geometric reconstruction model, and combines the object's real pose label.

[0014] Preferably, an RGBD image is input, the image contains multiple objects, the target detection model detects the category information and bounding box information of each object in the image, and cuts out each object according to the bounding box information; for the RGBD image of each object, the high-dimensional feature extraction of the image and the decoding and generation of the three-dimensional point cloud are completed based on the three-dimensional point cloud generation model to obtain the point cloud in the object coordinate system, and at the same time, the real camera coordinate system geometric information registration model is calculated according to the image depth information to realize the estimation of the object pose.

[0015] Compared with the prior art, the present invention has the following beneficial effects: The present invention combines deep learning technology with object pose estimation, fully utilizes the category information and bounding box information of the object extracted by the target detection model, and combines the depth data in the RGBD image to achieve high-precision object cutting and three-dimensional point cloud generation. In this way, the high-dimensional features of each object can be accurately extracted, and point cloud data in the object coordinate system can be generated. This method not only improves the accuracy of object pose estimation, but also avoids the excessive reliance on image features in traditional methods during the processing process, so that the estimation results can maintain high stability in complex environments and different viewing angles.

[0016] By using the DCP geometric information registration model, the present invention can efficiently and accurately align the source point cloud and the target point cloud, thereby accurately estimating the object's pose. This technology not only solves the noise interference and error accumulation problems existing in traditional image-based pose estimation methods, but also significantly improves processing speed and robustness. It is particularly suitable for industrial automation, robot grasping and other fields that require high-precision pose estimation.

[0017] The six-degree-of-freedom pose estimation method of the present invention does not require complex calibration of the object surface or excessive manual design, has a high degree of automation, strong adaptability, can be quickly deployed and run in real time in a variety of scenarios, and greatly reduces the operation and maintenance costs of the system. In general, the present invention has significant advantages in improving object recognition and positioning accuracy, shortening processing time, and reducing hardware dependence, and has broad application prospects and market value. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art are briefly introduced below.

[0019] Figure 1 A schematic diagram disclosed in an embodiment of the present invention; Figure 2 This is a schematic diagram of a process disclosed in an embodiment of the present invention. DETAILED DESCRIPTION

[0020] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0021] A method for estimating the six-degree-of-freedom pose of an object based on three-dimensional geometric information registration, comprising a target detection module, a three-dimensional point cloud generation model, and a geometric information registration pose estimation module; The target detection module completes the target detection task in the object pose estimation scenario based on the YOLOv9 network, making full use of the network's programmable gradient information mechanism (PGI) and generalized efficient layer aggregation network mechanism (GELAN), ensuring that the local details of small targets can be preserved during the feature transfer process, significantly improving the recognition robustness of partially visible targets. At the same time, GELAN supports flexible multi-scale feature fusion and reduces the possible mismatching problem when objects are densely arranged through a hierarchical interaction mechanism; The 3D point cloud generation model includes a deep set feature extraction network and a spatial 3D point cloud decoding network. Based on image encoding and 3D geometric information decoding, the generation of 3D point clouds and reconstruction of object geometric information in the object coordinate system are realized. The deep set feature extraction network integrates RGB and depth information, and converts RGBD images into global feature representations through encoders. The spatial 3D point cloud decoding network generates unordered point clouds through decoders such as TopNet, which greatly improves the quality of 3D spatial geometric reconstruction. At the same time, based on depth information, the reconstruction of set information and the generation of 3D point clouds in the object camera coordinate system are completed, realizing high-quality geometric information reconstruction in multiple coordinate systems driven by deep images.

[0022] The geometric information registration pose estimation module is responsible for registering the decoded three-dimensional point cloud in the object coordinate system and the calculated point cloud in the real camera coordinate system. The former is used as the source point cloud and the latter is used as the target point cloud. The rotation and translation transformation from the source point cloud to the target point cloud is calculated to achieve the six-degree-of-freedom pose estimation of the corresponding object. At the same time, the registration process of this module complies with the disorder of the point cloud, ensuring that the registration result is not affected by the order of the point cloud input, thereby improving the robustness and adaptability of the pose estimation method. The specific method flow is as follows: Figure 2 As shown, the following steps are included: Training phase: Different from the traditional pose estimation method that uses end-to-end regression for the center point of the object, the target detection module in this method implements the training of the YOLOv9 target detection model based on the bounding box information in the pose estimation dataset; the target detection module performs object target detection on the pose estimation dataset, and cuts out the RGBD images of all objects to construct the subsequent training dataset; the 3D point cloud generation model implements the training of the 3D space geometric reconstruction model composed of the image feature extractor and the TopNet point cloud decoder based on the constructed object image RGBD dataset and the pose labels of all objects; the geometric information registration module completes the training of the geometric information registration model DCP (Deep Closest Point) based on the point cloud in the camera coordinate system calculated based on the object RGBD image and the point cloud in the object coordinate system generated by the 3D space geometric reconstruction model, and combines the object's real pose label.

[0023] Inference stage: Input RGBD image, the image contains multiple objects, the target detection model detects the category information and bounding box information of each object in the image, and cuts out each object according to the bounding box information; for each object's RGBD image, the high-dimensional feature extraction and 3D point cloud decoding generation of the image are completed based on the 3D point cloud generation model to obtain the point cloud in the object coordinate system, and at the same time, the point cloud in the real camera coordinate system is calculated according to the image depth information; the predicted point cloud in the object coordinate system is used as the source point cloud, and the point cloud in the real camera coordinate system is used as the target point cloud, and the object pose is estimated through the DCP geometric information registration model.

[0024] 1. Target detection model training: (1) Based on the RGBD data in the pose estimation dataset and the bounding box information of each object in the image, the YOLO format training dataset is constructed, screened and enhanced in the pose estimation scenario; (2) Adopting the programmable gradient information mechanism (PGI), by introducing the PGI mechanism, the gradient optimization capability of the YOLOv9 network in the target detection task is enhanced. The PGI mechanism allows the model to learn different types of targets in a more accurate and efficient manner through an adjustable gradient update strategy during the training process. The gradient weighting module dynamically adjusts the gradient weight according to the object category and detection difficulty. For objects that are more difficult to detect (such as small objects or occluded objects), PGI increases their gradient weights so that the model pays more attention to these difficult-to-detect targets during training. On the other hand, in the gradient clipping part, gradient clipping is used during the gradient calculation process to avoid the gradient explosion problem and ensure stable convergence of the network.

[0025] (3) The generalized efficient layer aggregation network mechanism (GELAN) is used to optimize the feature extraction and information fusion process of the YOLOv9 network. The GELAN mechanism improves the model's ability to detect multi-scale targets by effectively aggregating features at different levels in the network. In the multi-level feature aggregation stage, information aggregation is performed between the convolutional layers of the YOLOv9 network, combining low-level detail features with high-level semantic features to capture object information at different scales and in complex backgrounds. At the same time, the adaptive feature fusion module dynamically adjusts the fusion strategy of features at different levels according to the complexity of the object and background interference, allowing the model to more accurately identify and locate targets.

[0026] (4) The loss is calculated based on the cross entropy, and the gradient is calculated by back propagation to update the parameters to complete the model training.

[0027] The reasoning process of the target detection module is as follows: (1) For the input RGB image , the corresponding depth map is . Use the RGB part of the image for target detection to obtain the bounding box (BouningBox) of each object in the image as well as the category and confidence information ,in Indicates the object category, Indicates confidence.

[0028] (2) Setting the confidence threshold And filter out all qualified high-confidence objects to obtain a set of valid bounding boxes Based on the bounding box, the RGB image and depth map are intercepted to obtain an image set containing the RGBD information of each object. .

[0029] 2. The training steps of the 3D point cloud generation model are as follows: (1) Based on the trained YOLOv9 target detection model, the pose estimation dataset is preprocessed. For the input RGB image , corresponding to the depth map , use the RGB part of the image for target detection, and get the bounding box of each object in the image as well as the category and confidence information ,in Indicates the object category, Indicates confidence level; (2) Setting the confidence threshold And filter out all qualified high-confidence objects to obtain a set of valid bounding boxes , based on the bounding box, the RGB image and depth map are intercepted to obtain an image set containing the RGBD information of each object , so far the data set is completed; (3) For each RGBD image of the object in each batch during the training phase, calculate the point cloud in the real camera coordinate system , based on the real pose label of each object, calculate the point cloud in the real object coordinate system ; (4) For the RGBD image of the object in each batch during the training phase, the spatial 3D point cloud decoding network predicts the point cloud in the object coordinate system ,calculate and The Chamfer Distance is used as the loss and the back-propagation training model is performed.

[0030] The reasoning steps of the 3D point cloud generation model are as follows: (1) For RGBD image input, the input channels need to be expanded from 3 to 4 to accommodate the input data. The modified convolution kernel size is , after the first layer of convolution, we get the first feature map The entire feature extraction part is composed of multiple residual stacks. From the input to the final feature map after passing through all the residual blocks, the pooling layer is added to compress the spatial dimension of the feature map into a feature vector ; (2) The spatial 3D point cloud decoding network further decodes the features extracted from the image and defines the tree structure of the decoding network , where the tree structure satisfies The number of point cloud points N, when the number of point cloud points is set to 2048, the common decoding tree structure is {2: [32, 64], 4: [4, 8, 8, 8], 6: [2, 4, 4, 4, 4, 4], 8: [2, 2, 2, 2, 2, 4, 4, 4]}. The initial input is the initial feature vector , after the 0th layer, high vit , the dimension changes to obtain the first layer feature set ; (3) The next layers 1 to L-2 cannot directly use the fully connected layer to achieve the spatial mapping of the two-dimensional matrix. One-dimensional convolution is used for feature mapping. In order to maintain the feature direction of the decoding process, the initial feature vector is concatenated before each calculation to obtain According to the tree structure described above, the first layer can be expanded times the number of features, that is, at this time, the features are further mapped to At this time, the feature dimension transformation is performed to further expand the number of feature sets to obtain the first layer output feature set ; (4) Similarly, from the 2nd layer to the L-2nd layer, feature decoding and feature set accumulation are continuously performed, and the number of features in the set is continuously approached to the number of points in the point cloud to be generated, until the L-2nd layer, at which time the feature set is obtained Similarly, after concatenating the initial feature vectors, the last layer maps each feature into a 3D space to obtain the generated point cloud .

[0031] 3. The training steps of the geometric information registration pose estimation module are as follows: (1) Based on the object RGBD dataset, for each batch of object RGBD images in the training phase, calculate the point cloud in the real camera coordinate system ; (2) For the RGBD images of the objects in each batch during the training phase, the point cloud in the object coordinate system is predicted based on the trained 3D point cloud generation model. ; (3) Based on the DCP geometric information registration network, the point cloud under the object coordinates is predicted As the source point cloud, the point cloud in the real camera coordinate system As the target point cloud, predict the object's pose, including the rotation matrix with translation vector ; (4) Calculate the true rotation matrix and The cross entropy loss , calculate the true translation vector and The cross entropy loss , + Back-propagation gradient calculation and model parameter update are performed as loss.

[0032] The geometric information registration pose estimation module realizes the registration of the source point cloud and the target point cloud. The specific steps of the inference stage are as follows: (1) Based on the RGBD image of the object, the depth information and the object segmentation information, the point cloud in the camera coordinate system is calculated, and the farthest point sampling is further performed to reduce the number of point cloud points to N, and the result is , , based on the 3D point cloud generation model, the point cloud in the object coordinate system is obtained , ; The point cloud in the object coordinate system is regarded as the source point cloud, and the point cloud in the camera coordinate system is regarded as the target point cloud. That is, the transformation from the source point cloud to the target point cloud is the pose of the object in the picture at this time; (2) Based on the feature mapping part of the PointNet network as the DCP feature extraction network, each point is processed independently by a multi-layer perceptron with shared parameters, and the high-dimensional features of each point are extracted to obtain the feature set corresponding to the target point cloud and the source point cloud. and , and further apply maximum pooling to the two feature sets to obtain the global feature vectors and , and then concatenated with the feature set output by the last layer to obtain and , use MLP to implement equal-dimensional mapping for each feature in the set at this time, and obtain the final high-dimensional feature set and ; (3) Two point cloud high-dimensional feature sets are used for self-attention calculation. The input point cloud high-dimensional features go through , , Generate a query (Query, Q), key (Key, K) and value (Value, V), , , , map the features to the feature_dim dimension, calculate the attention weight matrix A through the dot product operation, and then model the global information learned by the attention mechanism into the value vector by clicking A and the value vector, so as to realize dynamic attention to different features in the entire feature set. , the two point cloud feature sets are self-attention calculated separately, and we get and , , , They represent the mapping matrices of query vector, key vector, and value vector in self-attention calculation respectively; (4) Perform cross-attention calculation to learn the correspondence between the two, and use the query matrix to map the feature vector of the first point cloud to obtain , a key and value mapping matrix to map the feature vectors of the second point cloud, , , after cross attention calculation, we can get and ; (5) Calculate the similarity matrix of two point cloud feature sets ,in Indicates the similarity between the i-th point of the source point cloud and the j-th point of the target point cloud. At this time, the position with the highest similarity becomes the corresponding relationship between the two point clouds, and a two-way consistency check is adopted, that is, the best match from the source point cloud point to the target point cloud point. The best match of the target point cloud to the source point cloud Must be consistent, , (6) Extracting source point clouds that meet the corresponding relationship With the target point cloud , calculate the covariance matrix Represent the geometric relationship between point clouds and perform singular value decomposition on the covariance matrix , calculate the pose rotation matrix , and further calculate the translation vector based on the center of mass position ,in and Respectively represent the centroids of the source point cloud and the target point cloud calculated according to the corresponding relationship. So far, the pose estimation from the source point cloud to the target point cloud (including the rotation matrix R and the translation vector t) is completed.

[0033] The present application also provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes Figure 2 The methods provided in the various optional methods are therefore not described in detail here.

[0034] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0035] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in this description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0036] The method and related apparatus provided by the embodiment of the present application are described with reference to the method flow chart and / or structural diagram provided by the embodiment of the present application. Specifically, each process and / or box in the method flow chart and / or structural diagram, as well as the combination of the processes and / or boxes in the flow chart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable device to generate a machine, so that the instructions executed by the processor of the computer or other programmable device generate instructions for implementing the process in the process. Figure 1 A process or multiple processes and / or structures Figure 1 The computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including the instruction device, or are transmitted through a computer-readable storage medium. Computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The instruction device is implemented in the process Figure 1 A process or multiple processes and / or structures Figure 1 These computer program instructions can also be loaded onto a computer or other programmable device so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide the functions for implementing the process. Figure 1 A flow or multiple flows and / or structures illustrate the steps of the functions specified in one block or multiple blocks.

Claims

1. A method for estimating the six-degree-of-freedom pose of an object based on three-dimensional geometric information registration, characterized in that: It includes target detection module, 3D point cloud generation model, and geometric information registration pose estimation module; Object detection module for object detection and bounding box prediction; The 3D point cloud generation model includes a deep set feature extraction network and a spatial 3D point cloud decoding network; Deep ensemble feature extraction network: It combines RGB and depth information and converts RGBD images into global feature representations through an encoder; Spatial 3D point cloud decoding network: Generates disordered point cloud through decoder, completes the reconstruction of object camera coordinate system and 3D point cloud generation based on depth information, and realizes high-quality geometric information reconstruction of multiple coordinate systems driven by depth image; The geometric information registration pose estimation module registers the decoded three-dimensional point cloud in the object coordinate system and the calculated point cloud in the real camera coordinate system. The former is used as the source point cloud and the latter as the target point cloud. The rotation and translation transformation from the source point cloud to the target point cloud is calculated to realize the six-degree-of-freedom pose estimation of the corresponding object.

2. The method for estimating the six-degree-of-freedom pose of an object based on three-dimensional geometric information registration according to claim 1, characterized in that: The specific steps of the target detection module training phase are as follows: (1) Based on the RGBD data in the pose estimation dataset and the bounding box information of each object in the image, the YOLO format training dataset is constructed, screened and enhanced in the pose estimation scenario; (2) By introducing the PGI mechanism, the model can learn different types of targets in a more accurate and efficient way through an adjustable gradient update strategy during training. The gradient weighting module dynamically adjusts the gradient weight according to the object category and detection difficulty. (3) Optimize the feature extraction and information fusion process of the YOLOv9 network. In the multi-level feature aggregation stage, information is aggregated between the convolutional layers of the YOLOv9 network, combining low-level detail features with high-level semantic features to capture object information of different scales and complex backgrounds. At the same time, the adaptive feature fusion module dynamically adjusts the fusion strategy of different levels of features according to the complexity of the object and background interference, so that the model can more accurately identify and locate the target; (4) The loss is calculated based on the cross entropy, and the gradient is calculated by back propagation to update the parameters to complete the model training.

3. The method for estimating the six-degree-of-freedom pose of an object based on three-dimensional geometric information registration according to claim 1, characterized in that: The training steps for the 3D point cloud generation model are as follows: (1) Based on the trained YOLOv9 target detection model, the pose estimation dataset is preprocessed. For the input RGB image , corresponding to the depth map , use the RGB part of the image for target detection, and get the bounding box of each object in the image as well as the category and confidence information ,in Indicates the object category, Indicates confidence; (2) Setting the confidence threshold And filter out all qualified high-confidence objects to obtain a set of valid bounding boxes , based on the bounding box, the RGB image and depth map are intercepted to obtain an image set containing the RGBD information of each object , so far the data set is completed; (3) For each RGBD image of the object in each batch during the training phase, calculate the point cloud in the real camera coordinate system , based on the real pose label of each object, calculate the point cloud in the real object coordinate system ; (4) For the RGBD image of the object in each batch during the training phase, the spatial 3D point cloud decoding network predicts the point cloud in the object coordinate system ,calculate and The Chamfer Distance is used as the loss and the back-propagation training model is performed.

4. The method for estimating the six-degree-of-freedom pose of an object based on three-dimensional geometric information registration according to claim 1, characterized in that: The training steps of the geometric information registration pose estimation module are as follows: (1) Based on the object RGBD dataset, for each batch of object RGBD images in the training phase, calculate the point cloud in the real camera coordinate system ; (2) For the RGBD images of the objects in each batch during the training phase, the point cloud in the object coordinate system is predicted based on the trained 3D point cloud generation model. ; (3) Based on the DCP geometric information registration network, the point cloud under the object coordinates is predicted As the source point cloud, the point cloud in the real camera coordinate system As the target point cloud, predict the object's pose, including the rotation matrix with translation vector ; (4) Calculate the true rotation matrix and The cross entropy loss , calculate the true translation vector and The cross entropy loss , + Back-propagation gradient calculation and model parameter update are performed as loss.

5. The method for estimating the six-degree-of-freedom pose of an object based on three-dimensional geometric information registration according to claim 1, characterized in that: The target detection module realizes the recognition and positioning of objects in RGB images based on YOLOv9, and segments RGBD images of different objects. The specific steps of the inference stage are as follows: (1) For the input RGB image , the corresponding depth map is , use the RGB part of the image for target detection, and get the bounding box of each object in the image as well as the category and confidence information ,in Indicates the object category, Indicates confidence; (2) Setting the confidence threshold And filter out all qualified high-confidence objects to obtain a set of valid bounding boxes , based on the bounding box, the RGB image and depth map are intercepted to obtain an image set containing the RGBD information of each object .

6. The method for estimating the six-degree-of-freedom pose of an object based on three-dimensional geometric information registration according to claim 1, characterized in that: The specific steps of the 3D point cloud generation model reasoning stage are as follows: (1) For RGBD image input, the input channels need to be expanded from 3 to 4 to accommodate the input data. The modified convolution kernel size is , after the first layer of convolution, we get the first feature map The entire feature extraction part is composed of multiple residual stacks. From the input to the final feature map after passing through all the residual blocks, the pooling layer is added to compress the spatial dimension of the feature map into a feature vector ; (2) The spatial 3D point cloud decoding network further decodes the features extracted from the image and defines the tree structure of the decoding network , where the tree structure satisfies The number of point cloud points N, the initial input is the initial feature vector , after the 0th layer, you get high vit , the dimension changes to obtain the first layer feature set ; (3) The next layers 1 to L-2 cannot directly use the fully connected layer to achieve the spatial mapping of the two-dimensional matrix. One-dimensional convolution is used for feature mapping. In order to maintain the feature direction of the decoding process, the initial feature vector is concatenated before each calculation to obtain According to the tree structure described above, the first layer can be expanded times the number of features, that is, at this time, the features are further mapped to At this time, the feature dimension transformation is performed to further expand the number of feature sets to obtain the first layer output feature set ; (4) Similarly, from the 2nd layer to the L-2nd layer, feature decoding and feature set accumulation are continuously performed, and the number of features in the set is continuously approached to the number of points in the point cloud to be generated, until the L-2nd layer, at which time the feature set is obtained Similarly, after concatenating the initial feature vectors, the last layer maps each feature into a 3D space to obtain the generated point cloud .

7. The method for estimating the six-degree-of-freedom pose of an object based on three-dimensional geometric information registration according to claim 1, characterized in that: The geometric information registration pose estimation module realizes the registration of the source point cloud and the target point cloud. The specific steps of the inference stage are as follows: (1) Based on the RGBD image of the object, the depth information and the object segmentation information, the point cloud in the camera coordinate system is calculated, and the farthest point sampling is further performed to reduce the number of point cloud points to N, and the result is , , based on the 3D point cloud generation model, the point cloud in the object coordinate system is obtained , ; The point cloud in the object coordinate system is regarded as the source point cloud, and the point cloud in the camera coordinate system is regarded as the target point cloud. That is, the transformation from the source point cloud to the target point cloud is the pose of the object in the picture at this time; (2) Based on the feature mapping part of the PointNet network as the DCP feature extraction network, each point is processed independently by a multi-layer perceptron with shared parameters, and the high-dimensional features of each point are extracted to obtain the feature set corresponding to the target point cloud and the source point cloud. and , and further apply maximum pooling to the two feature sets to obtain the global feature vectors and , and then concatenated with the feature set output by the last layer to obtain and , use MLP to implement equal-dimensional mapping for each feature in the set at this time, and obtain the final high-dimensional feature set and ; (3) Two point cloud high-dimensional feature sets are used for self-attention calculation. The input point cloud high-dimensional features go through , , Generate a query (Query, Q), key (Key, K) and value (Value, V), , , , map the features to the feature_dim dimension, calculate the attention weight matrix A through the dot product operation, and then model the global information learned by the attention mechanism into the value vector by clicking A and the value vector, so as to realize dynamic attention to different features in the entire feature set. , the two point cloud feature sets are self-attention calculated separately, and we get and , , , They represent the mapping matrices of query vector, key vector, and value vector in self-attention calculation respectively; (4) Perform cross-attention calculation to learn the correspondence between the two, and use the query matrix to map the feature vector of the first point cloud to obtain , a key and value mapping matrix to map the feature vectors of the second point cloud, , , after cross attention calculation, we can get and ; (5) Calculate the similarity matrix of two point cloud feature sets ,in Indicates the similarity between the i-th point of the source point cloud and the j-th point of the target point cloud. At this time, the position with the highest similarity becomes the corresponding relationship between the two point clouds, and a two-way consistency check is adopted, that is, the best match from the source point cloud point to the target point cloud point. The best match of the target point cloud to the source point cloud Must be consistent, , (6) Extracting source point clouds that meet the corresponding relationship With the target point cloud , calculate the covariance matrix Represent the geometric relationship between point clouds and perform singular value decomposition on the covariance matrix , calculate the pose rotation matrix , and further calculate the translation vector based on the center of mass position ,in and They represent the centroids of the source point cloud and target point cloud calculated according to the correspondence.

8. The method for estimating the six-degree-of-freedom pose of an object based on three-dimensional geometric information registration according to claim 1, characterized in that: The object detection module implements YOLOv9 object detection model training based on the bounding box information in the pose estimation dataset; The target detection module performs object target detection on the pose estimation dataset, and cuts out RGBD images of all objects to construct subsequent training datasets; the 3D point cloud generation model implements the training of the 3D space geometric reconstruction model composed of the image feature extractor and the TopNet point cloud decoder based on the constructed object image RGBD dataset and the pose labels of all objects; the geometric information registration module completes the training of the geometric information registration model DCP based on the point cloud in the camera coordinate system calculated based on the object RGBD image and the point cloud in the object coordinate system generated by the 3D space geometric reconstruction model, and combines the object's true pose label.

9. The method for estimating the six-degree-of-freedom pose of an object based on three-dimensional geometric information registration according to claim 1, characterized in that: Input RGBD image, the image contains multiple objects, the object detection model detects the category information and bounding box information of each object in the image, and cuts out each object according to the bounding box information; For each object's RGBD image, the high-dimensional feature extraction and 3D point cloud decoding generation of the image are completed based on the 3D point cloud generation model to obtain the point cloud in the object coordinate system. At the same time, the real camera coordinate system geometric information registration model is calculated according to the image depth information to realize the estimation of the object's pose.

Citation Information

Patent Citations

  • Robot grasp pose estimation method based on object recognition depth learning model

    CN109102547A

  • Three-dimensional reconstruction method and device based on RGB camera and laser sensor, and server

    CN113379815A

  • 6dof object pose estimation method based on single RGB image

    CN116958262A

  • Geometric information enhancement-based category-level 6D attitude estimation method

    CN118261979A

  • Object six-degree-of-freedom pose estimation method based on deep learning

    CN119515980A

Cited By

  • Image three-dimensional point cloud generation method based on low-deviation feature extraction and enhancement

    CN120411381A

  • Image three-dimensional point cloud generation method based on low bias feature extraction and enhancement

    CN120411381B

  • Part manufacturing-oriented lattice structure mechanical property prediction method

    CN120597564A

  • Part surface point cloud acquisition method based on 6D pose estimation

    CN120876738A

  • Roadside vision three-dimensional target sensing method, device, equipment and medium

    CN120953945A