Panel 6D pose estimation method in complex scene

By iteratively optimizing the ConvGRU network, combining ResNet and bidirectional RNN to extract spatiotemporal features, and using feature pyramids and PnP modules to correct errors, the accuracy and robustness issues of plate pose estimation in complex scenarios are solved, achieving efficient and accurate pose estimation, which is suitable for industrial inspection and robot navigation.

CN120599031APending Publication Date: 2025-09-05BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510638681.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

In complex scenarios, sheet metal pose estimation faces problems such as occlusion, weak texture features, and poor lighting, resulting in poor robustness of traditional methods, insufficient pose estimation accuracy and robustness, and difficulty meeting the requirements of industrial robot grasping and sorting.

Method used

The ConvGRU iterative optimization network is adopted, combined with ResNet and bidirectional RNN networks to extract spatiotemporal features, feature pyramid and PnP module are used to correct geometric reprojection errors, the pose is optimized by Gauss-Newton method, and Lie algebraic pose incremental update is introduced to achieve high-precision pose estimation.

Benefits of technology

It improves the accuracy and robustness of pose estimation in complex scenarios, is suitable for fields such as industrial inspection and robot navigation, meets real-time requirements, and enhances adaptability in changing environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599031A_ABST
    Figure CN120599031A_ABST
Patent Text Reader

Abstract

The invention provides a method for iteratively optimizing a 6D pose of an object through ConvGRU in a complex scene. In a complex scene with serious shielding and insufficient texture information, image information of an object is collected through an industrial camera, and a corresponding rendering image is generated by using PyTorch3D. An observation image and rendered images under multiple postures are used as input, and cross-view geometric consistency image features are constructed by using a weight-shared Resnet network and a bidirectional RNN network. Then, correlation features between the images are coded through correlation volume, and the correlation features and semantic features are input into ConvGRU together for iterative updating; a confidence coefficient weight output by the network is used for dynamically suppressing interference of a low-quality area, an output corresponding field correction amount constructs a re-projection error through a differentiable Perspective-n-Point (PnP) module, a Gaussian-Newton method is introduced into a Lie group space to carry out nonlinear least square solution, a pose increment is calculated in a corresponding Lie algebra space, and the pose of the low-quality area is calculated. And completing iterative updating of the attitude.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical direction of three-dimensional plate recognition in the field of computer vision, and in particular relates to a 6D pose estimation method for plates based on ConvGRU iterative optimization, which is suitable for complex scenes. Background Art

[0002] The cutting and sorting of sheet metal is widely used in furniture, shipbuilding, automotive, and other fields. Industrial robots combined with vision technology facilitate automated sheet metal sorting. Sheet metal pose estimation is a fundamental research area and a key technology for enabling precise grasping, sorting, and assembly by industrial robots. 6D pose estimation encompasses not only the sheet metal's 3D position (X, Y, Z) but also its 3D pose (Roll, Pitch, Yaw) in space. One of the key aspects of this technology is capturing as clear and extractable image information as possible. However, complex scenes often present several challenges. First, sheet metal types are numerous and diverse, with varying incoming postures and often overlapped or stacked, resulting in a chaotic and disorganized image. Second, the sheet metal sorting environment is extremely complex. Conveyor areas on shipbuilding production lines are often surrounded by numerous interfering factors, such as oil, dust, and rust. Furthermore, lighting conditions are often poor, resulting in low overall image quality and limited useful information for feature extraction. Furthermore, sheet metal surfaces generally lack significant texture features, exhibiting a typical weak texture pattern. Traditional detection and matching mechanisms based on sparse keypoints still have certain limitations. On the one hand, occlusion can prevent some keypoints from being correctly predicted, leading to significant deviations in the pose results of the Perspective-n-Point (PnP) solution. On the other hand, the traditional single-shot inference process lacks an effective mechanism to handle such outliers, resulting in poor robustness.

[0003] Therefore, how to solve the problem that pose estimation based on key point correspondence is sensitive to outliers is of great significance for improving the accuracy of pose estimation and achieving a coordinated improvement in the efficiency and accuracy of manipulator grasping. Summary of the Invention

[0004] To address the shortcomings of existing pose estimation methods, this paper proposes an iterative pose estimation algorithm that integrates deep learning and geometric optimization. It uses a ConvGRU iterative optimization network to update pixel-level correspondence fields and confidence weights, introduces a PnP module to construct geometric reprojection errors, and uses the Gauss-Newton method to solve pose increments, thereby achieving efficient and accurate pose updates.

[0005] The technical solution of the present invention is: a method for iteratively optimizing the 6D pose of a plate through ConvGRU in complex scenes. It includes the following steps:

[0006] a. Rendering image generation and spatiotemporal feature extraction;

[0007] To build an iterative pose optimization system, we first generate a rendered image corresponding to the observed image using PyTorch3D. We then feed the observed and rendered images into a weight-sharing ResNet network to extract spatial features. Next, we use a weight-sharing bidirectional RNN network to perform temporal modeling of these spatial features. We fuse the features from the forward loop (from observed image to rendered image) and the backward loop (from rendered image to observed image) to construct geometrically consistent features for the image across viewpoints.

[0008] b. Correlation feature construction and matching area extraction;

[0009] To encode the similarity between the observed and rendered images, a feature pyramid is first used to process multi-scale features, adapting to different image sizes and ensuring the integrity of feature extraction. Image correlation volumes are calculated from different viewpoints, defining dynamic matching regions guided by initial poses to extract correlation features between the images. Multi-scale analysis ensures accurate matching of feature points under different conditions, minimizing mismatches caused by occlusion or viewpoint changes. An algorithm then identifies the best matching region, optimizing feature point alignment and subsequent processing steps.

[0010] c. Lightweight optimization of ConvGRU network parameters;

[0011] To address parameter optimization issues, we employed depthwise separable convolution techniques to improve the ConvGRU network structure. This splits the convolution operation into two dimensions: spatial and channel processing, reducing redundant computations and the number of parameters. We also experimentally adjusted the network architecture to ensure pose estimation accuracy remains unchanged, thereby improving system resource efficiency.

[0012] d. Pose optimization framework based on reprojection error;

[0013] To address the problem of translating pixel-level geometric information into actual pose updates, an optimization framework was established, using reprojection error as the primary optimization metric and error correction via a PnP module. In each optimization iteration, the plate pose is updated and solved using error feedback using the Gauss-Newton method, refining the scene reconstruction accuracy. Furthermore, a Lie algebraic pose incremental update mechanism was incorporated to ensure accurate pose updates during geometric transformations. The ultimate goal was to minimize reprojection error and enhance pose resolution capabilities in complex scenes. This process involved optimizing the objective function and combining it with real-world scene testing to continuously adjust various parameter settings to ensure system stability and efficient motion correction capabilities.

[0014] Beneficial effects:

[0015] The present invention proposes a method for estimating the 6D pose of a plate in complex scenes by iteratively optimizing the ConvGRU. By generating high-precision rendered images and extracting spatiotemporal features and consistency features, the accuracy and robustness problems of plate pose estimation in complex scenes are solved. This method uses feature pyramids and dynamic matching region extraction technology to effectively overcome feature matching deviations caused by factors such as perspective changes and occlusions, thereby improving the accuracy of scene reconstruction. At the same time, through the lightweight optimized ConvGRU network, the accuracy of pose estimation is maintained while reducing computing resource consumption, making it suitable for application scenarios with high real-time requirements. The pose optimization framework based on reprojection error ensures stable output of high-precision pose estimation under complex lighting and geometric conditions through explicit geometric constraints. This method significantly improves the adaptability of the pose estimation system in changing environments, and is suitable for precise positioning tasks in fields such as industrial inspection, robot navigation, and augmented reality. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 Schematic diagram of the ResNet network structure provided by the present invention;

[0017] Figure 2 A schematic diagram of the bidirectional RNN network structure provided by the present invention;

[0018] Figure 3 A schematic diagram of feature correspondence between pixel pairs of the observed image and the rendered image provided by the present invention;

[0019] Figure 4 Schematic diagram of standard convolution and depth-wise separable convolution provided by the present invention; Figure 5 Schematic diagram of the overall structure of the ConvGRU network provided by the present invention; DETAILED DESCRIPTION

[0020] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0021] An embodiment of the present invention discloses a method for iteratively optimizing the 6D pose of a plate in a complex scene by using ConvGRU, which includes the following steps:

[0022] Step 1: Rendering image generation and spatiotemporal feature extraction;

[0023] After the camera collects the image information of the plate to be captured, first use PyTorch3D to generate corresponding rendering images at different perspectives to establish a clear geometric reference image. Next, define the camera parameters of the rendering view, including camera position, direction, viewing angle range, and projection method. The intrinsic parameters of the camera are determined by the RealSense camera parameters, while the extrinsic parameters are based on the initial pose estimation. The program increases or decreases the pitch angle, yaw angle, or roll angle by a certain angle to automatically generate multi-perspective rendering images. At the same time, the rendering material is uniformly set to PMMA. After the generation of the rendered image is completed, the ResNet network is used as the backbone for spatial feature extraction, and convolution processing is performed on the observed image and images under multiple rendering perspectives respectively. For the specific network structure of ResNet, please refer to Figure 1 As shown. In the temporal feature extraction stage, a bidirectional RNN network is introduced to treat the observed image and the multi-view rendered image as a pseudo time series, and gradually unify their feature expressions through bidirectional temporal propagation. Bidirectional RNN network structure reference Figure 2 As shown in Figure 2, cross-perspective semantic modeling has been achieved, and the geometric consistency and context adaptability of features have been effectively improved.

[0024] Step 2: Correlation feature construction and matching area extraction;

[0025] Based on the previously extracted features, correlation features are constructed to encode the spatial geometric similarity between the observed image and the rendered image. First, given a set of observed images and rendered images to be matched, the similarity between any pixels in the two feature maps is calculated. The dot product between the feature vectors of the same point in the two images is used as the similarity measure to construct a single-layer correlation volume. This is used to traverse every pixel in the image and record the feature similarity between all pixel pairs. Figure 3 As shown. Next, a feature pyramid structure is introduced so that effective correspondence can be established between images at different scales. The features of the rendered image are downsampled layer by layer, keeping the original resolution of the feature map of the observed image unchanged, and performing correlation calculations. Subsequently, based on the initial estimated plate pose, a depth-enhanced pinhole projection model is introduced to reversely project each pixel in the observed image into a three-dimensional coordinate point, and then project it to the rendering perspective through geometric transformation, thereby establishing a pixel alignment relationship across perspectives. Finally, based on pyramid sampling, a method of dynamically adjusting the spatial coverage of the extraction window at different scales is used. A smaller sampling window is used at high levels to avoid matching confusion caused by low resolution, and a larger window is used at low levels to enhance the perception of local details, thereby achieving adaptive adjustment of the matching strategy at different scales.

[0026] Step 3: Lightweight optimization of ConvGRU network parameters;

[0027] After constructing the correlation features and semantic features for pose estimation, we introduce a convolutional gated recurrent unit (ConvGRU) temporal network structure with spatial modeling capabilities to establish a gradually convergent pixel-level dense correspondence field between the observed image and the rendered image. Next, we perform a lightweight optimization on the ConvGRU structure, replacing all standard convolution operations with depthwise separable convolutions. Figure 4 As shown. The traditional input feature map size is H*W*C0, the number of output channels is C1, the traditional standard convolution uses a four-dimensional convolution kernel of K*K*C0*C1, and the corresponding total computational cost is H*W*K*K*C0*C1. The depthwise separable convolution is divided into two steps, with computational costs of H*W*C0*C1 and H*W*(K2*C0+C0*C1) respectively. The computational cost is the original At this point, the lightweight optimization of the ConvGRU network is completed.

[0028] Step 4: Pose optimization framework based on reprojection error;

[0029] After building the ConvGRU network and performing lightweight optimization, we obtain two types of intermediate results that are highly relevant to pose estimation: pixel-by-pixel depth enhancement corresponding field correction, and spatial confidence. The input and output structure of the overall network is as follows: Figure 5 As shown. In the network structure, a differentiable PnP module is introduced. The depth map, camera intrinsic parameters and the current estimated pose are used to reversely project the pixels in the observation image into three-dimensional space, and then mapped to the rendering view through pose transformation to construct the cross-view pixel position correspondence. By comparing the cross-view pixel correspondence results using the observed image pose with the dense correspondence field output by the iterative loop, an optimization objective function can be established to further update the current observed image pose estimate and perform geometric reprojection modeling on the ConvGRU output. Finally, the Gauss-Newton method is used for optimization to minimize a nonlinear projection error function about the pose G0∈SE(3). The pose update cannot be processed by simple matrix addition operations, but must be processed by multiplication operations within the group. The concept of Lie algebra se(3) is introduced, and the Lie algebra pose increment is selected. (including 3 rotation parameters and 3 translation parameters) as the updated result of Newton iteration, thereby maintaining the orthogonality of the rotation matrix and the geometric consistency of the translation vector.

[0030] At this point, according to the method proposed in this embodiment for estimating the 6D pose of the plate through iterative optimization of ConvGRU in a complex scenario, efficient and fast pose estimation of the plate can be achieved in fields such as robot control, so as to facilitate further operations such as grasping and sorting.

[0031] Of course, the present invention may have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art may of course make various corresponding changes and modifications based on the present invention, but these corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present invention.

Claims

1. A method for estimating the 6D pose of an object in a complex scene by iteratively optimizing a ConvGRU, the method comprising: S1: Generate rendered images and extract spatiotemporal features. Use PyTorch3D to generate rendered images of objects from different perspectives. Utilize a weight-sharing ResNet network and a bidirectional RNN network to construct geometrically consistent image features across perspectives. S2: Correlation feature construction and matching region extraction. Correlation volume and feature pyramid are used to calculate the multi-scale correlation features of the observed image and the rendered image. The depth-enhanced pinhole projection model is combined to establish pixel alignment relationships and adaptively adjust the matching strategy at different scales. S3: ConvGRU network parameters are lightly optimized, and the ConvGRU network structure is improved by using depthwise separable convolution to reduce the amount of computation while maintaining pose estimation accuracy, and to construct a pixel-level dense correspondence field; S4: A pose optimization framework based on reprojection error, which uses the Perspective-n-Point (PnP) module and the Gauss-Newton method to optimize pose estimation, and combines the Lie algebra incremental update mechanism to minimize the reprojection error to achieve high-precision 6D pose estimation.

2. The method for estimating the 6D pose of an object in a complex scene by iteratively optimizing ConvGRU according to claim 1, wherein: When using the ResNet network to extract image semantic features, a bidirectional RNN network is introduced to optimize the temporal consistency of the observed image and rendered image features.

3. The method for estimating the 6D pose of an object in a complex scene by iteratively optimizing ConvGRU according to claim 1, wherein: When calculating the correlation features between the observed image and the rendered image, the feature pyramid is used to achieve multi-scale matching. Small windows are used in the high layers to reduce matching confusion, while large windows are used in the low layers to enhance detail perception.

4. The method for estimating the 6D pose of an object in a complex scene by iteratively optimizing ConvGRU according to claim 1, wherein: When iteratively optimizing the estimated 6D pose of the object, the pixels are back-projected into 3D space through the PnP module to establish cross-view correspondences. The Gauss-Newton method is also used to minimize the reprojection error and update the pose estimate. Finally, the Lie algebra se(3) is introduced to parameterize the pose increment, ensuring the orthogonality of the rotation matrix and the geometric consistency of the translation vector.