Pose optimization method based on SCFlow2

By introducing 3D scene flow and depth information regularization methods in SCFlow2, the existing methods are solved inadequate generalization ability of new object estimation, and efficient pose optimization for plug and play is achieved, improving the accuracy and robustness of pose estimation.

CN120355745APending Publication Date: 2025-07-22XIDIAN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510432982.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing 6D pose estimation method of object based on RGB images needs to be retrained or fine-tuned when processing new objects not involved in the training process, and the depth information utilization is low efficiency, so it is unable to effectively fuse RGB images and depth information, resulting in insufficient generalization ability.

Method used

Using the pose optimization method based on SCFlow2, a plug-and-play end-to-end pose optimization network is constructed by embedding rigid body motion in 3D scene flow and introducing the target's 3D shape prior to the estimation network, and using depth information as regularization in iteration, an end-to-end pose optimization network is constructed, and a 4D-related pyramid is constructed using RGB-D fusion features, and an iterative update is performed in combination with scene flow prediction pose residuals.

Benefits of technology

It improves the accuracy and robustness of pose estimation, and can achieve multi-pose assumption optimization performance under a single initial pose assumption, realizes plug-and-play, improves pose optimization speed and accuracy, and significantly improves the BOP AR index on multi-reference datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355745A_ABST
    Figure CN120355745A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of pose optimization, and particularly relates to an SCFlow2-based pose optimization method, which constructs a plug-and-play end-to-end pose optimization network by embedding rigid body motion in a 3D scene flow and introducing 3D shape priori of a target into an estimation network and taking depth information as regularization in iteration at the same time. Only synthetic data is used for training, and a 6D pose estimation task of a target which is not seen in a training process in a real scene can be generalized without any fine adjustment for a specific object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pose optimization, and particularly relates to a pose optimization method based on SCFlow2. Background Art

[0002] Currently, the 6D pose estimation methods for objects based on RGB images can be classified into: (1) methods based on 2D-3D correspondence. First, a deep neural network is used to establish the 2D-3D correspondence from the pixel coordinate system to the object coordinate system, and then the PnP algorithm is used to solve the 6D pose of the target object from the 2D-3D correspondence; (2) using a neural network to directly predict the 6D pose of the target object in the image end-to-end. Compared with the two-stage methods based on 2D-3D correspondence, although the end-to-end method is more conducive to deployment in practical application scenarios, its accuracy is often not as good as that of the methods based on 2D-3D correspondence. To accurately estimate the 6D pose of an object, various different methods have been proposed, including methods based on RGB-D images, methods based on depth images, methods based on point clouds, and methods based on RGB images, etc.; however, most of the current existing methods assume that there are many real images of the target, and the network model is trained based on these real images. Such methods usually show accurate pose estimation effects for the trained targets, but cannot handle new objects not involved in the training process. Retraining or fine-tuning can make the network adapt to new objects, but this is usually troublesome and time-consuming, limiting the application of the algorithm in practical scenarios. To achieve the 6D pose estimation of new objects not involved in the training process, a pose estimation algorithm with stronger generalization ability needs to be designed.

[0003] Methods that solely rely on RGB images usually face challenges such as cluttered backgrounds, lighting changes, and texture differences. When depth information is known, combining RGB images with depth information can enhance the ability to extract target geometric data. The main challenge of RGB-D based methods lies in how to make full use of the appearance information from RGB images and the geometric information from depth images, and at the same time fuse the two-modal information to improve the generalization ability of the algorithm.

[0004] Although these methods have shown good results in the pose estimation of new objects, there are still some problems in the algorithm design itself. For example, the current methods do not consider the prior information of the target object's shape, making them less efficient in the feature matching process of rendering-comparison. Secondly, the rendering-comparison strategy of these methods is only a 2D matching process. When the depth information is known, methods such as ICP or Kabsch are used to utilize the depth information, but the depth information is not used in the network to further narrow the matching space, so it is a suboptimal strategy. In addition, for the problem of large initial pose errors, these methods usually use the pose optimization strategy of multi-pose hypotheses and an additional scoring network to select the final output result to improve the accuracy of pose estimation, which requires additional computational costs and the results are unstable. Summary of the Invention

[0005] The object of the present invention is to: aiming at the above existing problems, the present invention provides a pose optimization method based on SCFlow2. By embedding the rigid body motion in the 3D scene flow and introducing the 3D shape prior of the target into the estimation network, and at the same time using the depth information as regularization in the iteration, a plug-and-play end-to-end pose optimization network is constructed, which can be generalized to the 6D pose estimation task of unseen targets in the real scene training process only using synthetic data training without any fine-tuning for specific objects.

[0006] The technical solution adopted by the present invention is as follows:

[0007] A pose optimization method based on SCFlow2, the method includes:

[0008] Obtain data, the data includes the target object image and the initial pose rendering map of the target object;

[0009] Construct a 4D correlation pyramid, extract and fuse the RGB and depth features of the target object image and the initial pose rendering map of the target object, construct a 4D correlation volume through the fusion feature dot product, and obtain correlation features;

[0010] The intermediate flow predictor uses the correlation features to update the dense transformation field, and the dense transformation field is used as the input of the pose predictor based on scene optical flow to predict a pose residual, update the pose, and complete the refined update of the pose through several iterative processes.

[0011] Further, the construction of the 4D correlation pyramid is specifically:

[0012] For the input target object RGB image I s and the image I of the initial pose rendering r , first obtain the dense feature maps F s and F r, and the depth maps D corresponding to the two images s and D r , lift the two depth maps to point clouds, denoted as C s and C r , then extract the three-dimensional structure information of the target, then extract the local feature information for each region to generate global information. For the low-resolution features i and p, connect the features at the same index and input them into the network to generate a global feature vector; for the original target image feature s and the initial pose rendering image feature r, perform dense fusion respectively to generate global context, and finally construct a correlation feature 4D correlation pyramid using the dot product of the fused features.

[0013] Furthermore, the dense feature maps F s and F r are obtained at 1 / 8 resolution using the encoder DINOv2 with shared parameters. For the depth maps D s and D r corresponding to the two images, lift the two depth maps to point clouds according to the camera intrinsic matrix, denoted as C s and C r .

[0014] Furthermore, the acquisition of the three-dimensional information specifically includes:

[0015] Downsample the original point cloud to 1 / 8 resolution, denoted as C′ s and C′ r , centered on C′ s and C′ r perform grouping and query operations on the complete point clouds C s and C r , extract local feature information from each region using a multi-layer perceptron MLP, and use max pooling to generate global information and

[0016] Furthermore, the correlation feature 4D correlation pyramid is specifically as follows:

[0017]

[0018] Furthermore, the update of the dense transformation field is specifically as follows:

[0019] Given I r and I s and the corresponding point clouds C r and C s , predict a dense transformation field T ∈ SE(3) H×W representing the three-dimensional motion between the two point clouds, mapping points from I r to I s, for C r The point X at index i in i , whose projection pixel coordinates are x i =(u i , v i , z i ), perform the transformation T i ·X i to obtain the corresponding point X' in C s , whose projection pixel coordinates x' i =(u' i , v' i , z' i ); The scene flow vector is defined as the difference f = x' i - x i , where the first two components represent the standard optical flow and the last component describes the depth difference between two frames. An iterative update of the scene flow is performed using a recurrent model with GRU as the intermediate scene flow predictor. In the k-th iteration, the update of the hidden state is expressed as: i

[0020] h k = GRU(L C (F k-1 ), T k-1 , h k-1 ; Θ) (2)

[0021] In the formula, L C (F k-1 ) represents the relevant feature map indexed by the pose-induced flow in the (k - 1) th -th iteration, T k-1 is the dense transformation field from the previous iteration, Θ represents the GRU network parameters, and h k is the hidden state feature, which will be used to predict the correction of the current scene flow (r x , r y , r z ). For mapping correction. The updated scene flow is expressed as f k = f k-1 + (r x , r y , r z ), and the mapping correction will be used as the input of the Dense-SE3 layer to update the SE(3) motion and generate a new transformation field T' k .

[0022] Furthermore, the pose residual is obtained as follows:

[0023] Use the dense transformation field as the input to predict the global pose residual; Take the dense transformation field T' k ​Expressed as a 4×4 transformation matrix, it is encoded using a three-layer 2D convolutional network, and the updated global pose residual ΔP k The six-dimensional representation of the rotation matrix and the scaled representation of the translation vector in k will be output by two fully connected layers with different dimensions. In the k-th iteration, the pose residual ΔP k is used for pose update: The updated pose P k is used to calculate the pose-induced flow, which is used to impose a shape constraint on the lookup operation in the next iteration. Based on the updated pose ΔP k a new local dense transformation field T is generated k , where each pixel-level 3D motion describes the same updated pose residual. Calculate the initial pose P0 = [R0|t0] and P k = [R k |t k , and the pose residual between them is as follows:

[0024] ΔR k = R k ·R0, Δt k = t k - ΔR k ·t0(3)

[0025] Finally, the residual pose [ΔR k |Δt k is copied H×W times to generate the updated three-dimensional motion field T k ∈ SE(3) H×W , which is used as the input for the next iteration.

[0026] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are:

[0027] A pose optimization method based on SCFlow2 of the present invention can further improve the accuracy of the estimation result on the basis of the pose estimated by other pose estimation algorithms. Due to the addition of an extra regularization term provided by 3D geometric features and the integration of rigid body motion representation and 3D shape prior during the iteration process, it has higher robustness for the pose estimation of occluded objects, and can achieve the performance of pose optimization with multiple pose hypotheses only on the basis of a single initial pose hypothesis, improving the algorithm efficiency. Finally, the present invention realizes the plug-and-play pose optimization effect without any additional training for specific objects; the pose optimization speed of 0.18 s / object is achieved on an NVIDIA RTX-3090 GPU. Compared with the benchmark model SCFlow, under the same initial pose, the BOP AR index of the pose optimization effect is increased by 4%. On the seven benchmark datasets of BOP, the BOP AR under the unseen condition can reach 75.2, and the BOP AR index under the seen condition can reach 86.0, ranking among the top in terms of accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is the architecture schematic diagram of a pose optimization method based on SCFlow2 of the present invention;

[0029] Figure 2 is the workflow diagram of the benchmark algorithm SCFlow adopted by the present invention;

[0030] Figure 3 is the workflow diagram of SCFlow2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The present invention will be described in detail below with reference to the drawings.

[0032] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0033] Embodiment

[0034] This embodiment provides a pose optimization method based on SCFlow2, and the specific architecture design is as Figure 1As shown. Given an RGB-D image and a 3D model of the target object, the goal of this embodiment is to estimate the 6D pose of the visible target object in the image. Assume that the intrinsic matrix of the camera is known, and the initial pose of the target object is obtained by other existing algorithms; given the initial pose of the target object, the pose optimization network optimizes the pose using the matching relationship between the rendered RGB-D image and the input RGB-D; first, a fusion encoder with shared weights is used to extract and fuse these two sets of RGB and depth features respectively, and a 4D correlation volume is constructed through the dot product of the fused features, that is, a 4D correlation pyramid is constructed. The intermediate flow predictor will update the dense 3D transformation field using the correlation features as the input of the pose predictor based on scene optical flow, predict a residual pose once, and finally update the pose; through several iterative processes, the refinement update of the pose is completed.

[0035] In this embodiment, the construction of the 4D correlation pyramid is specifically as follows:

[0036] Based on the basic framework of the RAFT optical flow estimation algorithm, this embodiment first constructs a 4D correlation pyramid that describes the visual similarity between two images for two images; this feature pyramid is usually composed of the dot products of RGB image features at different resolutions; conventional RAFT only uses RGB image features to construct the pyramid. However, in fact, the depth data corresponding to the RGB image describes the 3D geometric features of the object and can provide additional correlation information. Therefore, the goal of this embodiment is to use the fused RGB-D features in the two sets to construct the correlation pyramid; the DenseFusion method realizes the fusion of RGB and depth information by extracting dense pixel-level feature embeddings. It processes the two data sources separately and generates multiple global features to capture rich global context. Inspired by this, this embodiment develops a similar global dense fusion network F θ to integrate information from two modalities; however, the key difference from this method is that this embodiment performs dense fusion on low-resolution feature maps rather than visible pixel levels to maintain the continuity of the pyramid features.

[0037] For the input RGB image I s and the corresponding image I rendered using the initial pose r , first use the encoder DINOv2 with shared parameters to obtain their respective dense feature maps F s and F r at a resolution of 1 / 8. For the depths D s and D r corresponding to the two RGB images, first lift the two depth maps to point clouds according to the camera intrinsic matrix, denoted as C s and C r; Then, a shared-parameter architecture similar to PointNet++ is used to extract their 3D structural information. However, it is usually difficult and time-consuming to extract local features from all points in the point cloud and aggregate global information, especially when the resolution of the depth map is very high. In addition, downsampling the point cloud before feature extraction may lead to the loss of detailed information. Therefore, in this embodiment, some modifications are made to the PointNet++ paradigm. First, the original point cloud is downsampled to 1 / 8 of the resolution, denoted as C′ s and C′ r . To retain the information of the entire point cloud, grouping and query operations are performed on the complete point clouds C s and C′ r centered on C′ s and C r ; Then, a multi-layer perceptron MLP is used to extract local feature information from each region, and max pooling is used to generate global information and

[0038] For the low-resolution features i and p, the features at the same index are concatenated and fed into the network to generate a global feature vector. For s and r, dense fusion is performed respectively to generate global context while retaining the spatial continuity between the features; Finally, a correlation feature 4D Volume is constructed using the dot product of the fused features:

[0039]

[0040] Introduction of 3D scene flow:

[0041] The correlation features are usually further used to predict the optical flow f(u, v). The optical flow can map the pixel position x r in I r =(u, v) to the corresponding position in I s : x s =(u, v)+f(u, v). However, this optical flow only contains 2D correspondence information, which may lead to the loss of depth information during the forward pass of the network. To solve this problem, this embodiment introduces 3D scene flow as an intermediate representation of the pose residual. Given I r and I s and their corresponding point clouds C r and C s , the core of the scene flow is to predict a dense transformation field T∈SE(3) H×W to represent the 3D motion between the two point clouds, mapping points from I r to I s . Specifically, for the point X r with index i in C i , its projected pixel coordinates are Xi =(u i , v i , z i ), a transformation T i ·X i can be performed to obtain the corresponding point X' s in C i , whose projected pixel coordinates x' i =(u' i , v' i , z' i ); the scene flow vector is defined as the difference f = x' i - x i , where the first two components represent the standard optical flow and the last component describes the depth difference between two frames. To achieve iterative update of the scene flow, this embodiment uses GRU as a recurrent model for the intermediate scene flow predictor. In the k-th iteration, the update of the hidden state is expressed as:

[0042] h k = GRU(L C (F k-1 ), T k-1 , h k-1 ; Θ) (2)

[0043] In the formula, L C (F k-1 ) represents the relevant feature map indexed by the pose-induced flow in the (k - 1) th -th iteration, T k-1 is the dense transformation field from the previous iteration, which will be discussed later, Θ represents the GRU network parameters, and h k is the hidden state feature, which will be used to predict the correction of the current scene flow (r x , r y , r z ), which can be called "mapping correction". The updated scene flow can be expressed as f k = f k-1 +(r x , r y , r z ). The mapping correction will be used as the input of the Dense-SE3 layer to update the SE(3) motion and generate a new transformation field T' k .

[0044] Pose residual acquisition:

[0045] The goal of this embodiment is to predict the pose residual of a single object described in I r and I s . Considering the rigidity of the object in the 6D pose estimation scenario, T' kAll pixel-level 3D motions in k are theoretically described by the same pose residuals. However, local pose residual prediction is vulnerable to object occlusion and environmental noise. Therefore, in this embodiment, a small network is designed to predict the global pose residuals using the dense transformation field as the input. First, the dense transformation field T′ k is represented as a 4×4 transformation matrix, and then it is encoded using a three-layer 2D convolutional network. The updated global pose residuals ΔP k in the six-dimensional representation of the rotation matrix and the scaled representation of the translation vector will be output by two fully connected layers with different dimensions. At the k-th iteration, the pose residuals ΔP will be used for pose update: k The updated pose P k will be used to calculate the pose-induced flow, which is used to impose a shape constraint on the lookup operation in the next iteration. In addition, a new local dense transformation field T k will be generated based on the updated pose ΔP k , where each pixel-level 3D motion describes the same updated pose residuals. Specifically, the pose residuals between the initial pose P0 = [R0|t0] and P k = [R k |t k are calculated as follows:

[0046] ΔR k = R k ·R0, Δt k = t k - ΔR k ·t0 (3)

[0047] Then, the residual pose [ΔR k |Δt k is copied H×W times to generate the updated three-dimensional motion field T H×W ∈ SE(3)

[0048] For comparison:

[0049] Given the 3D model of the target object, the images I1 and depth maps D1 are rendered according to the initial pose, and then the network is used to compare the rendered outputs with the real inputs I2 and D2 for pose optimization. As Figure 2 shown, the working process of the benchmark algorithm SCFlow is presented. Although this method adds 3D shape constraints in the optimization loop, it regards the matching process as a pure 2D problem, which is less efficient in capturing 3D motions and cannot be used for RGBD inputs. For the additional depth information, a common practice is to use the RANSAC-Kabsch algorithm for pose optimization in the second stage. However, this method is locally optimal at each stage.

[0050] As shown Figure 3 in the figure, the working process of the SCFlow2 of the present invention is shown. Aiming at the problems existing in SCFlow, the present invention introduces an intermediate representation based on scene flow to capture 3D motion in network optimization. In addition, the present invention embeds depth into the loop by formulating depth as an additional regularization to iteratively guide the correlation search. SCFlow2 is end-to-end trainable and can be well generalized to new objects.

[0051] In summary, the present invention designs a plug-and-play pose optimizer for the problem that conventional pose optimization algorithms need to be retrained or fine-tuned for specific objects to work properly. It can be generalized to the pose optimization tasks of objects not seen in the training process only by training with a large-scale virtual dataset. In the process of establishing the 4D Volume correlation matrix, the present invention first proposes a method for establishing the correlation matrix based on RGB-D fusion features, introduces 3D geometric features, and provides an additional regularization term for the correlation indexing process. In the iterative process, the present invention uses 3D scene flow representing pixel-level rigid body motion to replace 2D optical flow as the intermediate representation, and combines the rigid body motion representation in the 3D scene flow with the 3D shape prior of the target to predict the global pose residual.

[0052] Specific embodiments are applied in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A pose optimization method based on SCFlow2, characterized in that, The method includes: Obtain data, where the data includes an image of a target object and an initial pose rendering of the target object; Construct a 4D correlation pyramid, extract and fuse the RGB and depth features of the target object image and the initial pose rendering of the target object, construct a 4D correlation volume through the dot product of the fused feature points, and obtain correlation features; The intermediate flow predictor updates the dense transformation field using the correlation features. The dense transformation field serves as the input to the pose predictor based on scene optical flow, predicts a pose residual once, updates the pose, and completes the refined update of the pose through several iterative processes.

2. The pose optimization method based on SCFlow2 according to claim 1, wherein Specifically, the construction of the 4D correlation pyramid is as follows: For the input RGB image I of the target object s and the image I rendered with the initial pose r , first obtain the dense feature maps F s and F r of the two images respectively, as well as the depth maps D s and D r corresponding to the two images. Lift the two depth maps to point clouds, denoted as C s and C r respectively. Then extract the three-dimensional structure information of the target, and then extract the local feature information of each region to generate global information. For the low-resolution features i and p, connect the features at the same index and input them into the network to generate a global feature vector; for the original target image feature s and the initial pose rendered image feature r, perform dense fusion respectively to generate global context. Finally, construct a correlation feature 4D correlation pyramid using the dot product of the fused features.

3. A pose optimization method based on SCFlow2 according to claim 2, characterized in that, The dense feature map F s and F r are obtained at 1 / 8 resolution using the encoder DINOv2 with shared parameters. For the depth maps D s and D r corresponding to the two images, the two depth maps are lifted to point clouds according to the camera intrinsic matrix, denoted as C s and C r .

4. A pose optimization method based on SCFlow2 according to claim 2, characterized in that, Specifically, the acquisition of the three-dimensional information includes: Downsample the original point cloud to 1 / 8 resolution, denoted as C′ s and C′ r , centered on C′ s and C′ r perform grouping and query operations on the complete point clouds C s and C r Extract local feature information from each region using a multi-layer perceptron MLP, and generate global information using max pooling and 5. A pose optimization method based on SCFlow2 according to claim 2, characterized in that The 4D correlation pyramid of the correlation features is specifically as follows:

6. A pose optimization method based on SCFlow2 according to claim 2, characterized in that Specifically, the update of the dense transformation field is as follows: Given I r and I s and the corresponding point cloud C r and C s , predict a dense transformation field T ∈ SE(3) H×W representing the 3D motion between two point clouds, mapping points from I r to I s . For a point X r with index i in C i , whose projected pixel coordinates are x i =(u i , v i , z i ), perform the transformation T i ·X i to obtain the corresponding point X′ s in C i , whose projected pixel coordinates x′ i =(u′ i , v′ i , z′ i ); the scene flow vector is defined as the difference f = x′ i - x i , where the first two components represent the standard optical flow and the last component describes the depth difference between two frames. An iterative update of the scene flow is performed using a recurrent model with GRU as the intermediate scene flow predictor. In the k-th iteration, the update of the hidden state is expressed as: h k = GRU(L C (F k-1 ), T k-1 , h k-1 ; Θ) (2) where L C (F k-1 ) represents the relevant feature map of the pose-induced flow index in the (k - 1) th -th iteration, T k-1 is the dense transformation field from the previous iteration, Θ represents the GRU network parameters, h k is the hidden state feature, which will be used to predict the correction of the current scene flow (r x , r y , r z ), for mapping correction; the updated scene flow is represented as f k = f k-1 +(r x , r y , r z ), the mapping correction will be used as the input of the Dense-SE3 layer to update the SE(3) motion and generate a new transformation field T′ k .

7. A pose optimization method based on SCFlow2 according to claim 6, characterized in that, Specifically, the acquisition of the pose residual is as follows: Use the dense transformation field as input to predict the global pose residual; the dense transformation field T′ k is represented as a 4×4 transformation matrix and encoded using a three-layer 2D convolutional network. The updated global pose residual ΔP k The six-dimensional representation of the rotation matrix and the scaled representation of the translation vector in will be output by two fully connected layers with different dimensions. In the k-th iteration, the pose residual ΔP l is used for pose update: The updated pose P k is used to calculate the pose-induced flow, which is used to impose a shape constraint on the lookup operation in the next iteration. A new local dense transformation field T k will be generated based on the updated pose ΔP k , where each pixel-level 3D motion describes the same updated pose residual. Calculate the pose residual between the initial pose P0 = [R0|t0] and P k = [R k |t k as follows: ΔR k = R k ·R0, Δt k = t k -ΔR k ·t0(3) Finally, copy the residual pose [ΔR k |Δt k and generate the updated three-dimensional motion field T H×W times k ∈ SE(3) H×W , which is used as the input for the next iteration.