Optical flow pose estimation method based on context semantics and pixel uncertainty
By constructing an optical flow position estimation method based on context semantics and pixel uncertainty, the problem of unstable position estimation of monocular visual odometer in complex environments is solved, and stable robustness and efficient calculations are achieved in dynamic scenes and low-texture areas.
Patent Information
- Application Number
- CN202510765568.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The existing monocular visual odometry method is difficult to take into account dynamic robustness, texture adaptability and structural simplicity in complex environments, resulting in unstable position estimation, especially in dynamic scenes and low-textured areas.
Using an optical flow pose estimation method based on context semantics and pixel uncertainty, a dual-channel ResNetFPN and cyclic RNN optical flow estimation network with shared weights is used to construct a tightly coupled optical flow estimation and pose regression joint architecture to realize end-to-end deep network design.
It significantly improves the stability and robustness of the model in complex environments, enhances the response ability to static areas and texture-rich areas, reduces the computational complexity, and is suitable for real-time application scenarios.
Smart Images

Figure CN120279101A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of robot navigation and positioning, computer vision and three-dimensional reconstruction, and in particular to an optical flow pose estimation method based on context semantics and pixel uncertainty. Background Art
[0002] Monocular visual odometry is one of the key technologies to achieve autonomous navigation and precise positioning of robots. Its basic principle is to estimate the position and posture of the camera by analyzing the pixels between frames of continuous images and calculating the movement of the feature points in them, thereby inferring the translation and rotation of the camera. Currently, there are two main methods for mainstream monocular visual odometry: one is to use traditional feature extraction and optical flow calculation to estimate the posture, and the other is to use deep learning to train the model through a neural network to estimate the camera posture.
[0003] In traditional methods, feature extraction is to extract key feature points from images and estimate camera motion by matching key point pairs in continuous images. It is robust to changes in illumination and scale, but performs poorly in low-texture scenes. The optical flow method, on the other hand, does not extract discrete feature points, but optimizes the camera pose by constructing a photometric error function to minimize the pixel grayscale of the current frame after projection of the previous frame, based on the assumption of photometric consistency. It also performs poorly in scenes with large photometric changes. Deep learning-based methods learn features and estimate poses directly from image sequences, without manually designing feature extraction and matching processes, and directly output the relative pose changes of the camera. A classic example is DeepVO, which uses a combination of CNN and RNN to learn the pose changes of the camera from a monocular image sequence end-to-end, which improves accuracy and robustness compared to traditional methods, but still faces challenges in complex scenes.
[0004] The typical representative of the traditional geometry-based monocular visual odometer method is ORB-SLAM2, which relies on manually designed feature extraction and matching processes, usually including: first, detecting image key points and generating descriptors through the ORB feature extractor; then using the Hamming distance between descriptors for feature matching; then using the PnP algorithm to estimate the camera pose; and finally optimizing the pose and map points through bundle adjustment. However, this type of method has the following defects: it is not robust to dynamic scenes, and moving objects will destroy the feature matching relationship and affect the accuracy of pose estimation; in low-texture areas such as the sky and walls, the estimation is unstable due to the lack of effective feature points; it is highly dependent on manual parameter adjustment and lacks adaptability; it uses a modular non-end-to-end structure, which is prone to overall performance degradation due to error accumulation.
[0005] For deep learning-based monocular visual odometry methods, they mainly rely on deep neural networks to directly regress the camera pose from images, including two methods: direct regression and optical flow-assisted regression. Among them, the direct regression method uses a CNN network to extract features from two frames of images and outputs translation and rotation parameters through a fully connected layer (K. Konda and R.Memisevic, “Learning visual odometry with a convolutional network,” inInternational Conference on Computer Vision Theory and Applications, vol. 2,pp. 486–490, SciTePress, 2015.) (S. Wang, R. Clark, H. Wen, and N. Trigoni,“Deepvo: Towards end-to-end visual odometry with deep recurrent convolutionalneural networks,” in 2017 IEEE international conference on robotics andautomation (ICRA), pp. 2043–2050, IEEE, 2017.); while for optical flow-assisted methods such as TartanVO and DROID-VO, they first use an optical flow network to estimate the motion between pixels, then combine the camera intrinsics to form intermediate features as the input to the pose network, and finally regress the camera pose (Wang, Wenshan, Yaoyu Hu, and Sebastian Scherer. "Tartanvo: A generalizable learning-based vo." Conference on Robot Learning .PMLR, 2021.) (Teed, Zachary, and Jia Deng. "Droid-slam: Deep visual slam formonocular, stereo, and rgb-d cameras." Advances in neural information processing systems34 (2021): 16558-16569.). Although the learning-based method has certain end-to-end characteristics and shows good generalization ability, it also has deficiencies: it lacks an attention mechanism for dynamic domains or low-texture regions, which affects the stability and robustness of pose estimation; it fails to fully exploit the semantic information in the image, restricting the model's understanding ability of complex scenes; the overall network structure is complex, consuming a large amount of computing resources and being unfavorable for real-time deployment on resource-constrained devices.
[0006] Therefore, the existing technology has not been able to balance dynamic robustness, texture adaptability, and structural simplicity while ensuring real-time performance, and there is still much room for improvement. Summary of the Invention
[0007] Aiming at the deficiencies of the existing technology, the present invention provides an optical flow pose estimation method based on context semantics and pixel uncertainty, aiming to effectively solve various challenges faced by existing visual odometry methods when estimating the camera pose in complex environments.
[0008] The present invention is realized through the following technical solutions: An optical flow pose estimation method based on context semantics and pixel uncertainty, comprising the following steps: (1) Use a monocular camera to collect a continuous video frame sequence, and preprocess the input image. (2) Input two consecutive frames of images into a pre-trained dual-path ResNetFPN with shared weights and a recurrent RNN optical flow estimation network to obtain an optical flow map; then extract context semantic feature information according to the current frame input into the ResNetFPN network; and obtain an initial optical flow uncertainty map through an internal convolutional layer, and then obtain the final optical flow map and uncertainty map through RNN network cyclic estimation. (3) Convert the camera internal parameters into an internal parameter layer; and combine the optical flow output by the optical flow network Figure 1 to splice into an optical flow internal parameter layer, and crop and scale it to 1 / 4 resolution. (4) Use the scaled optical flow internal parameter layer and uncertainty map as the input of the pose prediction network to enter a U-shaped network architecture, and input the context semantic feature information and the features extracted by the internal parameter layer in the second layer of the encoder into the context semantic attention module of the encoder placed in the U-shaped network architecture for attention fusion operation. (5) Predict the pixel-level camera pose through the decoder in the U-shaped network architecture, use the uncertainty map to screen the local pose, and finally adopt a hierarchical weighted method to fuse the global pose and the local pose to obtain the final camera pose.
[0009] Further, the preprocessing in step (1) is specifically as follows: Adjust the normalized pixel value range of the input image to [0, 1] as the input of the optical flow network, where the resolution of each frame of the image is H×W, the number of channels is 3, and the frame rate remains unchanged.
[0010] Specifically, the context semantic feature information in step (2) is extracted by the ResNet network, which contains image semantics and texture information and is used to enhance the semantic expression ability of the subsequent pose network.
[0011] Specifically, for the uncertainty map in step (2), its uncertainty information is applied to the screening in pixel-level pose to suppress noise.
[0012] Further, converting the camera intrinsics to the intrinsics layer in step (3) is specifically as follows: The camera intrinsics include the focal length and the optical center. At the same time, an intrinsics layer of the same size as the optical flow map is generated according to the focal length and the optical center parameters, and it is concatenated with the optical flow map to obtain the optical flow intrinsics layer.
[0013] Further, step (4) specifically includes the following sub-steps: (4.1) Extract features from the optical flow intrinsics layer through 3 layers of convolution, that is, the first layer downsamples with a stride of 2 to 60×80, and the second and third layers have a stride of 1, and the output number of channels is 32; (4.2) It includes five residual modules from layer1 to layer5. Among them, a context semantic attention CAPA module is added to the second layer of the encoder. There are two inputs, the features extracted from the second layer of the optical flow intrinsics and the semantic features. The complementary fusion of the context semantic feature information and the motion information is realized through channel attention and pixel attention; then the fused features continue to be used by the following three residual modules to extract features; the output number of channels is 64, 128, 128, 256, 256 in sequence, and the resolution of the feature map gradually decreases to 1×2.
[0014] Further, the specific steps of step (4.2) are as follows: The context semantic attention module includes two sub-modules, namely, channel attention (CA) and pixel attention (PA). First, the context semantic features are subjected to feature extraction and integration through a convolutional layer and a ReLU function to maintain dimensional consistency. Subsequently, an element-wise addition operation is performed with the features extracted in the second layer of the optical flow internal parameters to achieve preliminary complementarity between semantic information and motion information. The added features are input into the context semantic attention module. This module adjusts the weights of different channels through global pooling and attention weight calculation, multiplies the added features element-wise with the calculated channel weights to generate enhanced channel features. Subsequently, the channel features are used as the input of the pixel attention (PA) sub-module to generate a pixel feature attention map, which is multiplied element-wise with the channel features to obtain pixel-enhanced features. Finally, the pixel-enhanced features are added to the features extracted in the second layer to obtain the final attention features.
[0015] Further, the U-shaped network architecture in step (4) is specifically as follows: A U-shaped encoding-decoding network design is adopted, and its encoder backbone is ResNet32, which is used to receive the spliced optical flow map, internal parameter layer, and context semantic features. Its decoder restores the spatial resolution through skip connections and outputs a pixel-level pose prediction map.
[0016] Further, step (5) specifically includes the following sub-steps: (5.1) Since it is uncertain whether the figure has been reduced to the same size as the spliced features, it is first necessary to use two learnable convolutional layers to increase its dimension to a pixel translation uncertainty map and a rotation uncertainty map, respectively representing the uncertainty distributions of each component of translation and rotation, so as to provide independent uncertainty estimates for each dimension of camera translation and rotation. (5.2) The right half of the U-shaped network architecture is responsible for decoding and predicting the camera translation and rotation at the pixel level. Subsequently, it is fused with the uncertainty map output by the optical flow network to screen the most robust local translation amount and rotation amount. At the same time, a global camera pose estimation network parallel to the local camera pose estimation network is added at the encoder output end to directly predict the translation amount and rotation amount of the global pose based on high-level features. Finally, the camera pose is obtained through the weighted combination of the fused local pose and global pose. (5.3) First, divide the camera translation and rotation, as well as the translation uncertainty map and rotation uncertainty map of pixels, into small pixel blocks. Then, index the translation amount and rotation amount of the pixel with the lowest uncertainty from each small pixel block. Based on the current uncertainty, use the Softmax function to calculate the weights of each pixel pose. Finally, obtain a more robust local pose translation amount and rotation amount through weighted summation. The global pose prediction module takes the high-level features of the encoder as input, and then integrates the global features through two independent fully connected layers respectively and estimates the translation amount and rotation amount of the global pose, and fuses and weighted sums them with the translation amount and rotation amount of the local pose to obtain the final camera pose. Among them, the weights of the global pose prediction value and the local pose prediction value each take 0.5.
[0017] The beneficial effects of the present invention are as follows: 1. The tightly coupled joint architecture of optical flow estimation and pose regression constructed by the present invention. This architecture adopts an end-to-end deep network design, integrating multiple processes such as optical flow prediction, context semantic feature information extraction, attention mechanism, and pose decoding into one, avoiding the problem of error accumulation in the transfer between modules in traditional methods. At the same time, a context semantic attention module (CAPA) and an uncertainty screening module are introduced to achieve adaptive coupling optimization of the optical flow network information and enhancement of scene semantic understanding. At the same time, the optimal local pixel-level pose information of uncertainty screening greatly increases the anti-interference ability of the network to data outside the training domain. Finally, while simplifying the structure, the system stability and integration are significantly improved. The present invention designs an explicit context semantic feature guidance mechanism. By constructing a semantic weight map and fusing it with the encoded feature map, the model is guided to focus on key regions strongly related to pose estimation (such as static landmarks, fixed structures), and dynamically suppress the attention to low-related regions. This mechanism significantly improves the understanding ability and feature utilization efficiency of the model in complex environments. Especially in dynamic scenes, weak texture regions, and blurred images, it can still maintain stable and robust estimation performance.
[0018] 2. By introducing a context semantic attention module and an uncertainty screening mechanism, the present invention significantly improves the response ability of the model to static regions and texture-rich regions, and effectively suppresses the influence of dynamic interference and image blur. At the same time, the overall network is constructed by lightweight residual modules, combined with the input strategy of reducing the size of the optical flow map, reducing the model complexity and computational load, enabling the system to have higher operating efficiency while maintaining high precision, and being more suitable for real-time application scenarios.
[0019] 3. The present invention uses the uncertainty information of the optical flow network to screen the effective area, ensuring the accuracy and generalization ability of the model. At the same time, through structural lightweighting and module optimization, it can still operate efficiently on resource-constrained platforms such as embedded devices, with good real-time performance and generalization ability, and is applicable to visual positioning tasks in various complex scenarios such as autonomous driving, robot navigation, and augmented reality. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0021] Figure 1 It is the overall framework diagram of the present invention; Figure 2 It is the context semantic attention module diagram of the present invention; Figure 3 It is the local pose network architecture diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application.
[0023] This example describes an optical flow pose coupling estimation network based on parallel context semantics and pixel uncertainty. As Figure 1 , this system deeply couples the optical flow network and the pose network through an end-to-end deep learning network, uses the optical flow, context semantic features, and uncertainty features output by the optical flow network as the input information of the pose network, and deeply couples and jointly trains and optimizes each module of the two to solve the problems of accuracy and robustness of traditional monocular visual odometry in dynamic scenes, lighting changes, and blurry conditions. The following details its implementation manner from aspects such as system composition, method steps, and technical effects.
[0024] Scene description: Any monocular camera with known internal parameters acquires image video frames through the monocular camera, calculates the relative pose of the camera only through two frames of images, and finally infers the motion trajectory of the camera under the entire video stream.
[0025] 1. System Composition The product structure of this system consists of the following core components, and each component is tightly connected through data flow and calculation logic: (1) Monocular Camera Function: It can continuously capture video frames for shooting and input them into the system as a data sequence.
[0026] Position relationship: It can be installed at the front end or the top end of any autonomously movable device (such as a robot or an autonomous vehicle), and the lens is oriented without obstacle blockage.
[0027] Parameters: The internal parameters of the camera are known, the frame rate is fixed, and it can output RGB three-channel images with unlimited resolution, and the RGB three channels are output.
[0028] (2) Optical flow network module Function: Responsible for predicting the optical flow map between adjacent video frames, characterizing the movement of each pixel point.
[0029] Connection relationship: The optical flow network receives consecutive frames output by the monocular camera and outputs the optical flow map, context semantic feature information, and uncertainty information to the pose estimation network.
[0030] Implementation method: Based on a pre-trained optical flow estimation network, such as the Sea-RAFT network (Wang Y, Lipson L, Deng J. Sea-raft: Simple, efficient, accurate raft for optical flow[C] / / European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2024: 36-54). For example Figure 1 after the monocular camera collects consecutive image frames, the image data is taken as input and sequentially fed into the optical flow network to extract context semantic information and uncertainty features, and then the optical flow internal parameters are spliced. At the same time, it is jointly input into the pose network together with the context semantic feature information and uncertainty information for joint training and inference. Finally, the computing device outputs the camera motion pose corresponding to the current frame.
[0031] (3) Pose network module Function: Corely estimate global and local pose information, responsible for optical flow map and internal parameter layer feature extraction, and at the same time fuse the context semantic feature information and uncertainty map from the optical flow network, and jointly optimize and regress the pose with different modules.
[0032] Composition: It includes a ResNet32 encoder, a context semantic attention module, a global camera motion pose prediction branch containing a fully connected layer, and an uncertainty screening local camera motion pose prediction branch Connection relationship: Such as Figure 1As shown in the pose network, the pose network receives the feature map after the fusion of optical flow and camera intrinsics, the context semantic feature information extracted by the optical flow network, and the uncertainty map. It completes the context semantic feature fusion through the context semantic attention module, and uses the uncertainty filtering module to complete the local optimal pose filtering. Finally, the local pose is combined with the global pose weighted to obtain the output camera six-degree-of-freedom pose estimation result.
[0033] 2. Method steps The camera pose estimation of this system is realized through the following specific steps, and the steps are clear and operable: Step 1: Data acquisition Use a monocular camera to collect a continuous video frame sequence , where represents the th frame image, and the frame ordinal number increases from 1 to . The resolution of each frame image is H×W (H is the height, H is the width, in pixels), and the number of channels is 3 (RGB format). The frame rate needs to be kept stable (such as 30 frames per second) to ensure temporal continuity. The input image needs to be preprocessed, and the normalized pixel value range is adjusted to [0,1] as the input of the optical flow network.
[0034] Step 2: Prediction of optical flow map, context semantic feature information and uncertainty map As Figure 1 shown in the optical flow network, the specific process is to input two consecutive frame images into the pre-trained dual-path ResNetFPN with shared weights and the recurrent RNN optical flow estimation network to obtain the optical flow map , and the two dimensions represent the motion components in the horizontal and vertical directions. At the same time, according to the current frame input into the ResNetFPN network to extract the deep context semantic feature information at time , this feature can better reflect the high-level semantic information in the image, such as the category of objects, the context relationship of the scene, etc. At the same time, the uncertainty map
[0035] will also be obtained through the RNN network. The uncertainty map is upsampled by convolution to obtain the pixel uncertainty map, and this uncertainty information is used to filter the pixel-level pose to suppress the influence of noise.
[0036] Step 3: Intrinsic layer generation and feature splicing As Figure 1 shown, according to the camera intrinsics (focal length and optical center Generate the internal parameter layer (For subsequent training of different camera data to improve generalization ability), concatenate the optical flow map and the internal parameter layer in the channel dimension (i.e., optical flow internal parameters) to generate input features , and finally align and perform interpolation to scale to a resolution of 1 / 4 size to improve training efficiency
[0037] Step 4: Fusion of the optical flow internal parameter encoder and the context semantic feature information As Figure 1 shown, the main process of the encoder of the pose network is as follows. The optical flow internal parameters and the uncertainty map after scaling in step 3 are used as the inputs of the pose prediction network and enter the U-shaped network architecture. The input optical flow internal parameter class The features extracted in the second layer of the encoder of the U-shaped network architecture and the context semantic features are input into the context semantic attention module for attention fusion operation. Specifically: (4.1) The optical flow internal parameter features are first used to extract features through an initialized 3-layer convolution (the first layer has a stride of 2 for downsampling to 60×80, and the second and third layers have a stride of 1), and the output number of channels is 32
[0038] (4.2) Then, it passes through a residual network containing five residual modules (layer1 to layer5). A context semantic attention CAPA module is added to the second layer of the encoder. There are two input optical flow internal parameter second layer features and semantic features , and complementary fusion of the context semantic feature information and the motion information is realized through channel attention and pixel attention to enhance the semantic expression ability of the optical flow features. Subsequently, the fused features continue to be used by the following three residual modules to extract features, and the output number of channels is 64, 128, 128, 256, 256 in sequence, and the resolution of the final feature map is gradually reduced to 1×2
[0039] The context semantic attention module (CAPA module) in step (4.2) is specifically as follows The context semantic attention module: As Figure 2 shown, it includes two sub-modules: channel attention (Context attention, CA) and pixel attention (Pixel attention, PA)
[0040] First, the context semantic feature information realizes feature extraction and integration through a convolutional layer and the ReLU function to maintain dimensional consistency, and then combines with Perform the element-wise addition operation to achieve the initial complementarity of semantic information and motion information. The added features are input into the context semantic attention module. This module adjusts the weights of different channels through global pooling and attention weight calculation, and multiplies the added features element-wise with the calculated channel weights to generate enhanced channel features , and its calculation formula is as follows: .
[0041] Subsequently, input into the pixel attention sub-module to generate a pixel feature attention map, and multiply it element-wise with to obtain the pixel-enhanced feature .
[0042] Finally, add and to obtain the final attention feature , and its output formula is as follows: .
[0043] The core idea of the context semantic attention module is to refine the input feature map in two stages, achieve the complementary fusion of context semantic feature information and motion information, enhance the semantic expression ability of the optical flow features. Especially in the scenarios of dynamic objects or complex backgrounds, this module can effectively distinguish the foreground and background and reduce the interference of irrelevant regions. At the same time, the element-wise addition operation promotes the gradient flow and improves the training stability of the model
[0044] Step 5: Optical flow intrinsic parameter decoder and uncertainty screening module In the optical flow pose estimation method, as Figure 3 shown, the local pose estimation network uses the optical flow intrinsic parameter decoder uncertainty screening module to finely predict the camera pose by using the deep features of the encoder and the uncertainty information, improving the estimation robustness and accuracy
[0045] (5.1) The decoder takes the last layer feature of the encoder as the input, fuses multi-stage features through a U-shaped network and skip connections, and restores the pixel-level translation and rotation at the original spatial resolution. The uncertainty map is upsampled to a pixel translation uncertainty map and a pixel rotation uncertainty map through two layers of convolution to provide confidence estimates for translation and rotation
[0046] (5.2) The uncertainty screening module divides , , and into pieces The translation amount with the lowest index uncertainty for small pixel blocks and the rotation amount , calculate the weights using the Softmax function , and obtain the translation amount of the local pose through weighted summation and the rotation amount .
[0047] Step 6: Camera pose calculation First, use the global pose prediction module of the pose network as shown Figure 1 . The global pose prediction module contains two fully connected branches (each branch structure is 512→128→32→3) that respectively regress the translation parameters (3D vector) and the rotation parameters , and its input is the flattened form of the output features of the last layer of the encoder . Finally, the translation amount of the global pose is output as and the rotation amount .
[0048] Finally, fuse and weight the local pose with and the global pose and to obtain the final camera pose , , where .
[0049] 7. Network training (a) Data preparation Load the training dataset and the validation dataset, which contain the following data: The image data contains RGB images of two consecutive frames. The pose ground truth: The true translation and rotation parameters of the camera, which are converted into the Lie algebra form through transformation and . The camera internal parameter information includes the focal length and the optical center parameters , which are used to generate the internal parameter layer .
[0050] (b) Input data preprocessing: Use random cropping for the images and the internal parameter layer to increase the generalization during the training process. Subsequently, adjust the images to a unified size of 240×320, and ensure that the cropped area matches the field of view in combination with the camera internal parameters. If there is optical flow data, it is necessary to downsample the optical flow data and scale its flow amount so that its resolution and value are consistent with the cropped images
[0051] (c) Model training parameters: The training process is divided into two stages. In the first stage, the present invention first trains the real optical flow and pose for 200 epochs using the analogy enhancement strategy combined with a separate pose network, with a batch size of 256 and a learning rate of 1e-4. In the second stage, the present invention loads the pose network trained in the first stage as pre-trained weights into the global pose estimation network of CPVO, and jointly optimizes and trains the optical flow pose coupling network in combination with the image experiment analogy enhancement strategy. The number of training iterations is 100, the batch size is 128, and the learning rate is 1e-5. The input image size of this optical flow network is 240×320, and the input image size of the pose network is 120×160, reducing the GPU memory occupancy rate under the condition of extremely low loss accuracy. The present invention uses the AdamW optimizer, = 0.9, = 0.999, and introduces a multi-step learning rate scheduler (MultiStepLR). The training in the first stage decays at the 60th and 100th epochs, and the second stage decays at the 50th time, with a decay rate of 0.2 for both. The models of the present invention are all based on PyTorch and are trained on a single NVIDIA RTX 4090 GPU.
[0052] (d) Loss function: During the training process of the model, the setting of the loss function directly affects the optimization effect of the network model parameters. It is inherently unreasonable to recover the absolute motion scale of the camera from a single monocular camera. Although the absolute scale can be obtained through model overfitting, its generalization performance is extremely poor. Therefore, the present invention normalizes the camera translation amount between two frames, and the expression is as follows: ; where, is to avoid errors caused by dividing by zero.
[0053] The final loss function of the present invention is as follows: ; where, represents the normalized predicted translation amount, represents the normalized true translation amount, is the predicted rotation amount, is the true rotation amount, is the predicted pose transformation vector, is the true pose transformation vector. The L1 loss is calculated pairwise between the translation amount and the rotation amount and added together to obtain the final loss function.
[0054] Since the scale of the monocular camera movement itself cannot be observed from the monocular image sequence, the loss function of the camera pose , scale ambiguity only affects translation, and the loss of rotation remains unchanged. Therefore, the loss function defines the camera pose loss function : ; Among them, and are the predicted values of the camera motion, and are the true values of the camera motion, to avoid incorrect calculations of dividing by zero in the formula.
[0055] The protection scope of the present invention is not limited to the above specific implementation structures, but also covers equivalent replacements and extended solutions based on the substantial content of the technical solution, including but not limited to the following situations: (a) Expansion and lightweight adjustment of the network architecture: such as changes in the number of layers of the residual module and changes in the position of the attention context semantic attention module; (b) Replacement of the uncertainty modeling strategy: such as using methods like Bayesian estimation and information entropy to replace the current prediction mechanism; (c) Changes in the context semantic fusion method: such as introducing Transformer, Cross-Attention or other context encoding modules to achieve semantic guidance; (d) Replacement of the optimizer and training strategy: such as replacing the optimizer (such as SGD, Adam) and changing the loss function design.
[0056] Therefore, as long as the adjustment does not deviate from the basic concept and technical key points of the present invention regarding optical flow-pose coupling, context semantic fusion, and pixel-level uncertainty fusion, it should be regarded as within the protection scope defined by the claims of the present invention.
Claims
1. An optical flow pose estimation method based on context semantics and pixel uncertainty, characterized in that The steps include the following: (1) Use a monocular camera to collect a sequence of consecutive video frames, and preprocess the input images; (2) Input two consecutive frames of images into a pre-trained dual-path ResNetFPN with shared weights and a recurrent RNN optical flow estimation network to obtain an optical flow map; then extract context semantic feature information according to the current frame input into the ResNetFPN network; And obtain an initial optical flow uncertainty map through an internal convolutional layer, and then obtain the final optical flow map and uncertainty map through RNN network loop estimation; (3) Convert the camera internal parameters into an internal parameter layer; and splice the optical flow map output by the optical flow network together to form an optical flow internal parameter layer, and crop and scale it to 1 / 4 resolution; (4) Use the scaled optical flow internal parameter layer and uncertainty map as the input of the pose prediction network to enter a U-shaped network architecture, and input the context semantic feature information and the features extracted by the internal parameter layer in the second layer of the encoder into the context semantic attention module of the encoder placed in the U-shaped network architecture for attention fusion operation; (5) Predict the pixel-level camera pose through the decoder in the U-shaped network architecture, use the uncertainty map to screen the local pose, and finally adopt a hierarchical weighted method to fuse the global pose and the local pose to obtain the final camera pose.
2. The optical flow pose estimation method according to claim 1, wherein The preprocessing in step (1) is specifically: adjust the normalized pixel value range of the input image to [0,1] as the input of the optical flow network, where the resolution of each frame of image is H×W, the number of channels is 3, and the frame rate remains unchanged.
3. The optical flow pose estimation method according to claim 1, wherein The context semantic feature information in step (2) is extracted by the ResNet network, which contains image semantics and texture information and is used to enhance the semantic expression ability of the subsequent pose network.
4. The optical flow pose estimation method according to claim 1, wherein For the uncertainty map in step (2), its uncertainty information is applied to the screening in pixel-level pose to suppress noise.
5. The optical flow pose estimation method according to claim 1, wherein The conversion of the camera internal parameters into an internal parameter layer in step (3) is specifically that the camera internal parameters include the focal length and the optical center. At the same time, an internal parameter layer of the same size as the optical flow map is generated according to the focal length and optical center parameters, and it is spliced with the optical flow map to obtain the optical flow internal parameter layer.
6. The optical flow pose estimation method according to claim 1, wherein Step (4) specifically includes the following sub-steps: (4.1) Extract features from the optical flow internal parameter layer through 3 layers of convolution, that is, the first layer downsamples with a stride of 2 to 60×80, and the second and third layers have a stride of 1, and the output number of channels is 32; (4.2) It includes five residual modules from layer1 to layer5. Among them, a context semantic attention CAPA module is added to the second layer of the encoder. There are two inputs, the features extracted from the second layer of the optical flow internal parameter and the semantic features, and the complementary fusion of the context semantic feature information and the motion information is realized through channel attention and pixel attention; then the fused features are continued to be used by the following three residual modules to extract features; The output number of channels is 64, 128, 128, 256, 256 in turn, and the resolution of the feature map decreases layer by layer to 1×2.
7. The optical flow pose estimation method according to claim 1, wherein The specific steps of step (4.2) are as follows: The context semantic attention module includes two sub-modules, namely channel attention CA and pixel attention PA. First, the context semantic features are subjected to feature extraction and integration through a convolutional layer and the ReLU function to maintain dimensional consistency. Subsequently, an element-wise addition operation is performed with the features extracted in the second layer of the optical flow internal parameters to achieve preliminary complementarity between semantic information and motion information. The added features are input into the context semantic attention module. This module adjusts the weights of different channels through global pooling and attention weight calculation, multiplies the added features element-wise with the calculated channel weights to generate enhanced channel features. Subsequently, the channel features are used as the input of the pixel attention PA sub-module to generate a pixel feature attention map, and it is multiplied element-wise with the channel features to obtain pixel-enhanced features. Finally, the pixel-enhanced features are added to the features extracted in the second layer to obtain the final attention features.
8. The optical flow pose estimation method according to claim 1, characterized in that The specific U-shaped network architecture in step (4) is as follows: A U-shaped encoding-decoding network design is adopted, and its encoder backbone is ResNet32, which is used to receive the spliced optical flow map, internal parameter layer, and context semantic features. Its decoder restores the spatial resolution through skip connections and outputs a pixel-level pose prediction map.
9. The optical flow pose estimation method according to claim 1, characterized in that The specific steps of step (5) include the following sub-steps: (5.1) Since the uncertainty map has been reduced to the same size as the spliced features, it is first necessary to up-dimension it to a translation uncertainty map and a rotation uncertainty map of pixels through two learnable convolutional layers, which respectively represent the uncertainty distributions of each component of translation and rotation, so as to provide independent uncertainty estimates for each dimension of camera translation and rotation. (5.2) The right half of the U-shaped network architecture is responsible for decoding and predicting the camera translation and rotation at the pixel level. Subsequently, it is fused with the uncertainty map output by the optical flow network to screen the most robust local translation amount and rotation amount. At the same time, a global camera pose estimation network parallel to the local camera pose estimation network is added at the encoder output end, which directly predicts the translation amount and rotation amount of the global pose based on high-level features. Finally, the camera pose is obtained through the weighted combination of the fused local pose and global pose. (5.3) First, the camera translation and rotation, as well as the translation uncertainty map and rotation uncertainty map of pixels, are divided into small pixel blocks of. Then, the translation amount and rotation amount of the pixel with the lowest uncertainty are indexed from each small pixel block, and based on the current uncertainty, the weights of each pixel pose are calculated using the Softmax function. Finally, a more robust local pose translation amount and rotation amount are obtained through weighted summation. The global pose prediction module takes the high-level features of the encoder as input, and then integrates the global features and estimates the translation amount and rotation amount of the global pose through two independent fully connected layers, and fuses and weights them with the translation amount and rotation amount of the local pose to obtain the final camera pose. Among them, the weights of the global pose prediction value and the local pose prediction value each take 0.5.
Citation Information
Patent Citations
Adaptive speedometer based on RGBD camera and optimization key frame selection method
CN115272460A
Method for combining end-to-end unsupervised laser odometer and semantic segmentation
CN116342883A
Monocular image depth estimation method and system
CN117437274A
Vehicle determination system
WO2023047809A1
Cited By
Video image stabilization method based on online path smoothing network
CN120583314A