Camera pose estimation algorithm in dynamic scene

The proposed method uses dense optical flow, motion segmentation, and Transformer-based feature extraction to enhance camera pose estimation in dynamic scenes, addressing inaccuracies by refining model parameters and leveraging global context, achieving superior performance in dynamic environments.

CN120318320APending Publication Date: 2025-07-15GUIZHOU UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510373753.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The prior art has low accuracy in camera position estimation and insufficient utilization of image context information in dynamic scenarios, resulting in inaccurate fusion of positioning errors and virtual information in robot navigation and augmented reality applications.

Method used

The optical flow estimation neural network is used to obtain dense optical flow, combine the motion segmentation network and the Transformer module to extract context information, extract feature through the residual neural network, and calculate the camera position using the fully connected neural network to construct a loss function for model training.

Benefits of technology

It significantly improves the accuracy and stability of camera pose estimation in dynamic scenarios, and can accurately estimate the six-degree-of-freedom pose without the help of auxiliary equipment, improving the accuracy of robot navigation and augmented reality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318320A_ABST
    Figure CN120318320A_ABST
Patent Text Reader

Abstract

The invention discloses a camera pose estimation algorithm in a dynamic scene, and the algorithm comprises the following steps: S1, inputting a video frame into an optical flow estimation neural network (CNN), and obtaining a dense optical flow of adjacent image frames; s2, inputting the dense optical flow and the video frame into a motion segmentation network, obtaining a binary mask, and obtaining a motion segmentation result by using the mask thresholding optical flow; s3, inputting a motion segmentation result into a Transform module, and fully extracting context information; s4, inputting the context information extracted in the step S3 into a residual neural network, and performing feature extraction; and S5, calculating the content in the step S4 by using a full-connection neural network to obtain a camera pose. Through the algorithm, static and dynamic areas can be accurately segmented in a complex dynamic environment, and the contextual information of the image is fully utilized to realize the accurate estimation of the six-degree-of-freedom pose of the camera.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multi - technology cross - integration, and particularly relates to a camera pose estimation algorithm in a dynamic scene. Background Art

[0002] Visual odometry is a core technology in the fields of computer vision and robotics. Its main task is to use a continuous image sequence to estimate the motion of a camera, so as to accurately determine its position and orientation in three - dimensional space. Accurate camera pose estimation is the key to ensuring the stable and efficient operation of many applications such as three - dimensional reconstruction, robot navigation, and augmented reality in complex and changing scenarios.

[0003] Traditional pose estimation methods are based on the assumption of a static environment, that is, it is considered that the lighting conditions in the scene remain constant and all objects are in a static state. However, in actual application scenarios, there are a large number of dynamic objects, and these dynamic objects will introduce significant visual interference, resulting in obvious feature changes between consecutive image frames, and causing large errors in camera pose estimation. For example, in the application of robot navigation, if there are fast - moving pedestrians and vehicles around, the image information obtained by the camera carried by the robot will be interfered, resulting in deviation of the robot's positioning and inability to accurately execute the navigation task; in the application of augmented reality, inaccurate camera poses caused by moving objects will make the virtual information unable to be accurately fused with the real scene, seriously affecting the user experience.

[0004] When faced with a dynamic scene, the mainstream strategy of algorithms based on structural features is to rely on manually - designed feature extraction and motion estimation models to filter out dynamic feature points in the image, and only retain static feature points for camera pose estimation. However, such algorithms are prone to false detection problems in dynamic object recognition and motion judgment, and are likely to produce large errors when there are a large number of dynamic objects in the scene. Deep - learning - based methods usually use semantic segmentation networks to distinguish static objects and dynamic objects, but they only identify potential moving objects based on object categories and cannot accurately judge the motion state of objects, which may reduce the stability of the algorithm. At the same time, after filtering out the dynamic regions, most end - to - end solutions use convolutional neural networks (CNNs) to extract image features, but these features are often local, ignoring the spatial position relationship between adjacent pixels in the image. These methods that directly map camera poses using two - dimensional image features do not make full use of the rich context information in the image and lack an understanding of the global context. Summary of the Invention

[0005] The present invention aims to provide a camera pose estimation algorithm in a dynamic scene to solve the problems of low accuracy of camera pose estimation and insufficient utilization of image context information in the prior art in a dynamic scene.

[0006] A camera pose estimation algorithm in a dynamic scenario in this solution includes the following steps: S1. Input the video frame into the optical flow estimation neural network (CNN) to obtain the dense optical flow of adjacent image frames; S2. Input the dense optical flow and the video frame into the motion segmentation network to obtain a binary mask, and use the mask to threshold the optical flow to obtain the motion segmentation result; S3. Input the motion segmentation result into the Transformer module to fully extract the context information; S4. Input the tensor output in step S3 into the residual neural network for feature extraction; S5. Calculate the content extracted in step S4 using the fully connected neural network to obtain the camera pose.

[0007] Furthermore, in step S1, the convolutional neural network (CNN) is used to obtain the dense optical flow of the image frame: input adjacent video frames , , where C is the number of input channels, and respectively for After multiple convolutions: , and then calculate the feature correlation: , and finally output the optical flow result through optical flow refinement .

[0008] Furthermore, in step S2, the motion segmentation result is obtained: use the original input video frame , , where C is the number of input channels, and first pass through convolution: (1) where represents the convolution operation, W1 is the weight of the convolution, and b1 is the bias term; pass through the max pooling operation: (2) s is the stride; pass through the transposed convolution operation: (3) where represents the transposed convolution operation, is the transposed convolution kernel, is the bias of the transposed convolution kernel; perform upsampling: (4) Further process the feature map using convolution: (5) Obtain the pixel-level probability map, and perform thresholding processing on the probability of each pixel: (6) Set the optical flow within the mask to zero to obtain the segmentation result of dynamic and static pixels.

[0009] Further, in step S3, context information is extracted: the motion segmentation result is input into the attention network. For the input frame sequence , for , an embedding transformation is performed for subsequent feature fusion. and the tensor after feature embedding are concatenated to obtain for subsequent feature processing: (7) (8) Denote as having a dimension of , and perform a further feature transformation on to obtain and adjust the dimension to adapt to subsequent convolution operations.

[0010] (9) (10) Fuse the features of and , and increase the dimensions of and to adapt to the concatenation operation: (11) (12) (13) Concatenate and tensors and further compress the features to obtain a tensor representing the global features for subsequent attention calculation: (14) (15) Perform an attention operation on the tensor after global average pooling and mean calculation to generate attention weights: (16) (17) Normalize the attention weights: (18) Multiply the tensor and the normalized attention weight tensor element-wise, and then sum over dimension 2 to obtain the final output tensor , the weighted fusion of features is achieved: (19) Furthermore, in step S4, a residual neural network is used for feature extraction: Input tensor into the convolution, where represents the convolution operation, W1 is the weight of the convolution, b1 is the bias term, and the output size is B1 ∈ 56×80×32: (20) The output size of the first - layer residual is B2 ∈ 28×40×64: (21) The output size of the second - layer residual is B3 ∈ 14×20×128: (22) The output size of the third - layer residual is B4 ∈ 7×10×128: (23) The output size of the fourth - layer residual is B5 ∈ 4×5×256: (24) The output size of the fifth - layer residual is B6 ∈ 2×3×256: (25) Furthermore, in step S5, the pose information is obtained through a fully - connected layer, and the feature map is unfolded into a one - dimensional vector: (26) Obtain the translational part of the pose: (27) (28) (29) Obtain the rotational part of the pose: (30) (31) (32) Finally, the translational and rotational results are concatenated and input into a six - dimensional pose: (33) Furthermore, for a continuous input frame in one of the , the loss function is constructed by minimizing the difference between the network-predicted pose and the corresponding ground truth, and the model is trained and validated using a pose estimation dataset.

[0011] (34) In the formula , represents the estimated value, which consists of the translation and rotation loss between the predicted pose of frame B and the ground truth, and the translation amount is normalized to constrain the pose estimation neural network. represents the rotation estimation value. represents the translation estimation value, and R and T represent the ground truth.

[0012] The working principle and beneficial effects of this solution are as follows: (1) The loss function provides a quantitative index for the model to measure the gap between the prediction result and the real data. Through continuous iterative processes, the model continuously adjusts its parameters according to the feedback of the loss function, gradually narrowing the difference between the predicted value and the target value, thus achieving accurate fitting of the training data; (2) This algorithm integrates advanced deep learning technologies such as optical flow matching and attention mechanism, aiming to accurately extract and process image features, thereby significantly improving the performance and efficiency of learning; (3) Experimental results show that this algorithm can estimate the six-degree-of-freedom pose of the camera in a dynamic scene relatively accurately without the aid of auxiliary equipment. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is the schematic diagram of an algorithm for estimating the pose of a camera in a dynamic scene according to the present invention; Figure 2 is the visualization comparison diagram of the ATE trajectory of part of the sequences in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] The following is a more detailed description through specific embodiments: The embodiment is basically as shown in the attached Figure 1 figure: The present invention provides an algorithm for estimating the pose of a camera in a dynamic scene. It includes the following steps: S1. Input adjacent video frames , , where C is the number of input channels, and after passing through multiple layers of convolution respectively for : , then calculate the feature correlation: , and finally refine the output optical flow result through optical flow.

[0015] S2. Use the original input video frame , first pass through convolution as shown in formula (1), where Conv(∙)Denotes a convolution operation, W 1 is the weight of the convolution, b 1 is the bias term; After the max pooling operation, as shown in Equation (2), s is the stride; After the transposed convolution operation, as shown in Equation (3), where Denotes the transposed convolution operation, is the transposed convolution kernel, is the bias of the transposed convolution kernel; The upsampling operation is as shown in Equation (4); Further process the feature map using convolution, as shown in Equation (5), to obtain a pixel-level probability map. Threshold each pixel in the optical flow map obtained in step S1 according to the probability map, as shown in Equation 6, to obtain the segmentation result of dynamic and static pixels.

[0016] (1) (2) (3) (4) (5) (6) S3. Input the motion segmentation result obtained in step S2 into the attention network. For an input of size [B, C, H, W] x , perform feature embedding and splicing operations using Equations (7), (8), (9), and (10). Denote the dimension of as . The local convolution and global average pooling operations are as shown in Equations (11), (12), (13), (14), and (15). The attention output is as shown in Equations (16), (17), (18), and (19).

[0017] (7) (8) (9) (10) (11) (12) (13) (14) (15) (16) (17) (18) (19) S4. Input the tensor obtained in step S3 into the convolution, where the convolution is as shown in formula (20), and represents the convolution operation, Conv(∙) 1 is the weight of the convolution, W 1 is the bias term, and the output size is b 1 B ; the first-layer residual is as shown in formula (21), and the output size is ∈50×80×32 2 B ; the second-layer residual is as shown in formula (22), and the output size is ∈9×9×64 3 B ; the third-layer residual is as shown in formula (23), and the output size is ∈14×20×128 4 B ; the fourth-layer residual is as shown in formula (24), and the output size is ∈7×10×128 5 B ; the fifth-layer residual is as shown in formula (25), and the output size is ∈4×5×256 6 B 6 ∈2 ×3×256 .

[0018] (20) (21) (22) (23) (24) (25) S5. Expand the feature map in step S4 into a one-dimensional vector, as shown in formula (23). For the translation part, through formulas (24), (25), and (26), and for the rotation part, through formulas (27), (28), and (29). Finally, concatenate the translation and rotation results to output a six-dimensional pose, as shown in formula (30).

[0019] (26) (27) (28) (29) (30) (31) (32) (33) Example 1: The data in the AirDOS-Shibuya dataset was used as the initial data for Example 1. The experimental results of using the present invention on the AirDOS-Shibuya dataset were compared with the experimental results of methods such as DROID-SLAM, AirDOS, ORB-SLAM, VDO-SLAM, and DynaSLAM. The results are shown in Table 1 below. Figure 2 A visual comparison chart of the ATE trajectories on some sequences was provided. The black trajectory represents the reference trajectory generated from the true pose data, while the red trajectory is the estimated trajectory drawn based on the pose data predicted by the model. The higher the prediction accuracy, the higher the coincidence degree between the estimated trajectory and the reference trajectory. Table 1 ATE (m) results on the dynamic sequences of the AirDOS-Shibuya dataset

[0020] Example 2: The data in the KITTI dataset was used as the initial data for Example 2. The experimental results of using the present invention on the dynamic sequences of the KITTI dataset were compared with the experimental results of methods such as DROID-SLAM, AirDOS, ORB-SLAM, VDO-SLAM, and DynaSLAM. The results are shown in Table 2 below.

[0021] Table 2 ATE (m) results on the dynamic sequences of the KITTI dataset

[0022] Table 1 and Table 2 respectively represent the ATE results on the dynamic sequences of the AirDOS-Shibuya and KITTI datasets. Among them, the smaller the value, the higher the accuracy. The best and the second-best VO performances are marked in bold and underlined respectively. For example, the data "0.0271" corresponding to the present method in Sequence I of Table 1 is the best, and the data "0.0327" corresponding to the DytanVO method in Sequence I of Table 1 is the second-best. "-" indicates that some SLAM methods failed to complete the test due to initialization failure. In the AirDOS-Shibuya dataset, the present method showed the best ATE results on 6 out of 7 sequences (Sequence I, II, III, IV, V, and Sequence VII), and generally improved the accuracy by 20% compared with DytanVO. In challenging scenarios (Sequence VI and Sequence VII), the results are very close to DytanVO. From Figure 2It can be seen that the camera motion trajectory predicted by the method in this paper is highly consistent with the true reference trajectory. In the KTTTI dataset, this method outperforms other VO methods in 4 out of 8 dynamic sequences (sequences 00, 02, 03, and sequence 10). This result shows that the proposed method can still maintain stable and reliable performance when facing challenging scenarios. The method of the present invention achieves the best results in most sequences and can accurately estimate the six-degree-of-freedom pose of the camera in a dynamic scene without the aid of auxiliary devices. The above are only embodiments of the present invention, and common knowledge such as specific structures and characteristics well known in the art are not described in detail herein. It should be noted that for those skilled in the art, without departing from the structure of the present invention, several deformations and improvements can still be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicability of the patent. The protection scope required by this application shall be subject to the content of its claims, and the specific implementation manners and the like recorded in the specification can be used to interpret the content of the claims.

Claims

1. A camera pose estimation algorithm in a dynamic scenario, characterized in that: It includes the following steps: S1. Input the video frame into the optical flow estimation neural network (CNN) to obtain the dense optical flow of adjacent image frames; S2. Input the dense optical flow and the video frame into the motion segmentation network to obtain a binary mask, and use the mask to threshold the optical flow to obtain the motion segmentation result; S3. Input the motion segmentation result into the Transformer module to fully extract context information; S4. Input the tensor output in step S3 into the residual neural network for feature extraction; S5. Calculate the content extracted in step S4 using a fully connected neural network to obtain the camera pose.

2. The camera pose estimation algorithm in a dynamic scenario according to claim 1, wherein: The step S1 uses a Convolutional Neural Network (CNN) to obtain the dense optical flow of the image frame: input adjacent video frames , , where C is the number of input channels, respectively for After multiple layers of convolution: , then calculate the feature correlation: , and finally output the optical flow result through optical flow refinement .

3. A camera pose estimation algorithm in a dynamic scenario according to claim 2, characterized in that: The motion segmentation result obtained in step S2: using the original input video frame , , where C is the number of input channels. First, it passes through convolution as shown in formula (1), where Conv(∙) represents the convolution operation, W1 is the weight of the convolution, and b1 is the bias term; then through the max pooling operation, as shown in formula (2), s is the stride; then through the transposed convolution operation, as shown in formula (3), where represents the transposed convolution operation, is the transposed convolution kernel, is the bias of the transposed convolution kernel; the upsampling operation is as shown in formula (4); further use convolution to process the feature map, as shown in formula (5), W2 is the weight of the convolution, b2 is the bias term, to obtain the pixel-level probability map, and threshold the probability of each pixel, as shown in formula 6, to obtain the segmentation result of dynamic and static pixels; (1) (2) (3) (4) (5) (6)。 4. The camera pose estimation algorithm in a dynamic scenario according to claim 3, wherein: The step S3 extracts context information: The motion segmentation result is input into the attention network. For an input of size [B, C, H, W] x , feature embedding and splicing operations are performed using the formulas shown in (7), (8), (9), and (10). Denote The dimension of is ; Local convolution and global average pooling operations are shown in formulas (11), (12), (13), (14), and (15), and the attention output is shown in formulas (16), (17), (18), and (19); (7) (8) (9) (10) (11) (12) (13) (14) (15) (16) (17) (18) (19)。 5. The camera pose estimation algorithm in a dynamic scenario according to claim 4, wherein: The step S4 uses a residual neural network for feature extraction: the input tensor is fed into the convolution. The convolution is shown in formula (20), where Conv(∙) represents the convolution operation, W1 is the weight of the convolution, b1 is the bias term, and the output size is B1 ∈ 50×80×32; the first-layer residual is shown in formula (21), and the output size is B2 ∈ 9×9×64; The second layer of residual is shown in formula (22), and the output size is B3 ∈ 14×20×128; the third layer of residual is shown in formula (23), and the output size is B4 ∈ 7×10×128; the fourth layer of residual is shown in formula (24), and the output size is B5 ∈ 4×5×256; the fifth layer of residual is shown in formula (25), and the output size is B6 ∈ 2×3×256; (20) (21) (22) (23) (24) (25)。 6. The camera pose estimation algorithm in a dynamic scenario according to claim 5, characterized in that: In step S5, the pose information is obtained through a fully connected layer: the feature map is unfolded into a one-dimensional vector, as shown in formula (26). For the translation part, through formulas (27), (28), and (29), and for the rotation part, through formulas (30), (31), and (32). Finally, the translation and rotation results are concatenated to output a six-dimensional pose, as shown in formula (33); (26) (27) (28) (29) (30) (31) (32) (33)。 7. The camera pose estimation algorithm in a dynamic scenario according to claim 6, characterized in that: For a continuous input frame in the sequence, a loss function is constructed by minimizing the difference between the pose estimation value of the network and the corresponding ground truth ; (34) In the formula, , represents the rotation estimation value, represents the translation estimation value, R and T represent the true values, and the translation amount is normalized to constrain the pose estimation neural network.

Citation Information

Cited By

  • Camera motion estimation method and system under foreground mask based on optical flow guidance

    CN120602799A