End-to-end monocular visual odometer method fusing space-time semantic information

By introducing the optical flow-semantic-depth three-modal cross-attention mechanism and the Transformer model into the monocular visual odometer, the robustness of the monocular visual odometer under dynamic environment and lighting changes is solved, and higher generalization ability and positioning accuracy are achieved.

CN120088332AActive Publication Date: 2025-06-03ZHEJIANG UNIV

Patent Information

Application Number
CN202510578595.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-03
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

The existing monocular visual odometry method fails in dynamic environments and is not robust to illumination changes, resulting in limited accuracy and stability of pose estimation.

Method used

The end-to-end monocular visual odometry method is adopted to integrate spatiotemporal semantic information, and the dynamic fusion of geometric and semantic features is achieved through the optical flow-semantic-depth cross-attention mechanism, and the monocular depth estimation information and large data volume training Transformer model is introduced to enhance the dynamic scene adaptability and cross-domain generalization capabilities of the model.

Benefits of technology

It significantly improves the generalization ability and robustness of the model, can accurately estimate the camera position under dynamic scenes and complex lighting conditions, reduces matching errors caused by brightness changes, and improves positioning accuracy and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088332A_ABST
    Figure CN120088332A_ABST
Patent Text Reader

Abstract

The invention discloses an end-to-end monocular visual odometer method fusing space-time semantic information. According to the method, continuous image sequence frames are collected through a color monocular camera, and a multi-information fusion end-to-end deep learning framework is constructed; a heterogeneous training domain is adopted to set various data set course sharing parameter fusion training, continuous image sequences are input, and the end-to-end deep learning framework is dynamically coupled with hidden state feature vectors output historically, so that a feature mapping relation of time sequence perception is formed; and interpretable feature decoupling of the static background elements and the dynamic entity objects in the scene is realized. After iterative feature fusion, the system outputs sparse depth and camera motion poses which conform to scene geometric constraints, so that a camera trajectory estimation model with high robustness and strong generalization ability in a complex environment is constructed. According to the method, the positioning precision and stability of the monocular vision odometer are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention application relates to fields such as robot autonomous positioning, UAV navigation, and autonomous driving, and specifically provides an end-to-end monocular visual odometry method that fuses spatio-temporal semantic information. Background Art

[0002] Visual Simultaneous Localization and Mapping (V-SLAM), as the core technology for environmental perception in modern unmanned autonomous systems, has important application values in fields such as robot navigation and augmented reality. Visual Odometry (VO), as the front-end pose estimation module of the V-SLAM system, initializes the projection structure through the geometric correlation between consecutive image sequences captured by a monocular camera, then estimates the internal parameters of the camera through self-calibration, restores the movement of the camera, and calculates its six-degree-of-freedom (6-DoF) motion trajectory. However, existing VO methods face dual challenges in practical applications: firstly, the interference of dynamic objects causes the failure of the static scene assumption; secondly, the monocular scale ambiguity and complex lighting changes affect the robustness of pose estimation. Therefore, constructing a monocular visual odometry framework that can adapt to dynamic environments and has strong generalization ability has become the key technical bottleneck for improving the reliability of environmental perception in autonomous systems.

[0003] Traditional visual odometry methods are mainly divided into two categories: feature point-based geometric methods and direct methods, both of which have significant limitations: Feature point-based methods construct a reprojection error function through feature extraction and matching such as SIFT and ORB, and use Bundle Adjustment for pose optimization. However, the feature matching process has a high failure rate in dynamic scenes and repetitive texture environments, and the highly non-convex optimization problem is prone to falling into local optima. The error accumulation caused by the initial pose deviation makes the trajectory drift increase quadratically; for direct methods, they directly construct pixel intensity consistency constraints by bypassing feature extraction, which improves the calculation efficiency but is extremely sensitive to lighting changes. When the signal-to-noise ratio of the photometric error drops significantly, the pose estimation will fail. Although the advantage of this method is that it saves the time for feature extraction, the cost is that it needs to use all the information in the image, resulting in a much larger scale of the full-image optimization problem than using feature points. Therefore, the computational complexity of the visual odometry (VO) based on the direct method requires GPU acceleration and it is difficult to achieve low-power real-time processing on embedded devices. Both types of methods have essential defects in terms of dynamic adaptability, algorithm robustness, and hardware deployment efficiency, restricting the application reliability of visual odometry in actual complex scenes.

[0004] On the other hand, with the development of deep learning, learning-based methods have shown excellent performance in many vision tasks, including image classification (ResNet), object detection (YOLO), and depth estimation (MiDaS), etc. However, in the field of learning-based visual odometry, existing learning-based VO models exhibit significant performance degradation in cross-domain tests, which is a generalization problem. Currently, most VO models adopt the "closed-set training" paradigm, that is, training and testing on the same dataset. At the same time, some multi-task learning frameworks only test their generalization ability on depth prediction, rather than on camera pose estimation. Therefore, there is still a need for a more effective strategy to estimate camera poses and improve the generalization ability of large-parameter models. Summary of the Invention

[0005] The present invention proposes an end-to-end monocular visual odometry method that fuses spatio-temporal semantic information, and realizes the dynamic fusion of geometric and semantic features through an optical flow-semantic-depth three-modal cross-attention mechanism, aiming to solve problems such as no object-level matching in traditional methods, inability to initialize camera poses in low-texture and high-exposure environments, and increasing cumulative errors. At the same time, monocular depth estimation information and a large amount of data are introduced to train the Transformer model, increasing its adaptability to dynamic scenes and cross-domain generalization ability.

[0006] To achieve the above invention objectives, the technical solutions proposed by the present invention are as follows: An end-to-end monocular visual odometry method that fuses spatio-temporal semantic information, comprising the following steps: (1) Collect a continuous image sequence through a color monocular camera to obtain an image frame sequence , , … , where represents the frame ordinal number of the current frame; (2) Construct an end-to-end deep learning localization Transformer framework, which includes: an optical flow estimation network, a multi-scale optical flow feature extraction network, an optical flow intrinsic parameter encoder, a depth semantic feature encoder, a multi-modal spatio-temporal feature fusion decoder, and a camera pose prediction layer; (3) Use a multi-source heterogeneous dataset and a self-constructed dynamic blur dataset for curriculum sharing parameter learning, where an adaptive moment estimation optimizer with weight decay is used in combination with a learning rate decay strategy based on the Lambda function to learn and train the network weights; (4) Input the continuous image sequence and collected by the color monocular camera into this constructed end-to-end deep learning localization Transformer framework to obtain sparse depth and camera relative pose transformation information between image frames.

[0007] Specifically, the optical flow estimation network in step (2) adopts the SEA-RAFT network architecture, and realizes the representation of motion information by constructing a dense optical flow field estimated between adjacent frame images.

[0008] Specifically, the multi-scale optical flow feature extraction network in step (2) uses the architecture of ResNet-50 as the feature backbone, and constructs a core feature extraction module by introducing a residual skip connection and a cross-level feature fusion mechanism; the multi-scale optical flow feature extraction network includes a five-level downsampling architecture, constructs a learnable parameter embedding layer using the camera intrinsic matrix, fuses it with the optical flow features at the channel level as the input, and compresses the spatial dimension of the fused information of the optical flow field and the intrinsic parameters by convolution. Finally, a composite optical flow feature map containing shallow geometric information and deep semantic information is generated.

[0009] Specifically, the optical flow intrinsic parameter encoder in step (2) uses a deformable attention mechanism to splice the composite optical flow feature map of shallow geometric information and deep semantic information generated by the multi-scale optical flow feature extraction network and the camera intrinsic parameter layer, and uses a 3-layer deformable attention mechanism to aggregate its feature encoding through sparse sampling of dynamic features to generate a multi-scale optical flow intrinsic parameter encoded feature vector.

[0010] Furthermore, the depth semantic feature encoder in step (2) is implemented as follows: (5.1) Use the encoder of the Depth Anything v2 model to perform multi-level semantic feature extraction on the input image frame; (5.2) Input the extracted semantic features into a multi-layer perceptron MLP for feature dimension encoding; (5.3) Synchronously use the depth decoder branch of Depth Anything v2 to estimate the scene depth; (5.4) Fuse the semantic features encoded by the MLP in step (5.2) with the position encoding to generate the key-value pair features required by the attention mechanism, and input them into the cross-attention module of the feature fusion decoder.

[0011] Furthermore, the feature fusion decoder in step (2) realizes temporal information fusion as follows: (1) Adopt a three-level cascaded Transformer decoder architecture, which can realize iterative three-time estimation of pose and depth. Each decoder layer contains a composite attention mechanism module, which is integrated in order of processing: a temporal self-attention module based on the BEVFormer architecture, a deformable attention module, a standard self-attention module, and a cross-attention module; (2) Deploy a reference point adjustment module composed of a fully connected neural network between adjacent decoder layers. The execution of this module is as follows: Based on the feature output of the previous layer, dynamically update the reference point coordinates input to the next layer, and simultaneously predict the depth value and confidence weight of the reference point corresponding to the current frame; (3) During the update process of the reference point coordinates, perform a gradient separation operation on the adjusted reference point coordinate tensor to block the gradient transmission of the coordinate parameters during the backpropagation process.

[0012] Further, the camera pose prediction layer in step (2) is used to predict the relative motion of the camera between two frames of pictures. This motion is a relative pose transformation matrix with six degrees of freedom components , and its architecture includes a 1D convolution layer and a parallel two-branch fully connected network structure. Among them, the translation prediction branch: consists of three fully connected network layers, and outputs a three-dimensional translation vector to predict the relative amount of camera translation , where is the 3D translation amount, and the other rotation prediction branch: consists of three fully connected network layers, and outputs the Lie algebra parameter representation of the rotation matrix to predict the 3D rotation amount of the camera.

[0013] Further, step (3) uses a multi-source heterogeneous dataset and a self-constructed dynamic blur dataset for curriculum sharing parameter learning, and progressively trains the model in the following way: Gradually divide the heterogeneous training data by source from low to high in terms of difficulty, construct a multi-course learning sequence based on the scene complexity, and sort the training batches in ascending order of data difficulty; In each training cycle, synchronously extract four types of data to form a mixed batch at a sampling rate proportional to the scale of the multi-source heterogeneous dataset and the self-constructed dynamic blur dataset, and input it into the Transformer framework with shared parameters for joint training; Through the curriculum learning mechanism with increasing difficulty, enable the model to gradually adapt from simple geometric scenes to complex dynamic blur environments, and simultaneously improve the generalization ability and training convergence efficiency of motion estimation.

[0014] Further, the sparse depth and camera relative pose transformation information between image frames obtained in step (4) are used to estimate the camera motion trajectory through the following iterative spatio-temporal feature fusion mechanism: (9.1) Optical flow feature generation stage: Input the consecutive input frames and into the SEA-RAFT optical flow estimation network, and output a pixel-level optical flow field ; (9.2) Multi-modal feature encoding stage: Concatenate the optical flow field and the camera intrinsic parameter layer and input them into the optical flow intrinsic parameter encoder to generate a multi-scale optical flow intrinsic parameter encoded feature vector ; Subsequently, the current frame is input into the DepthAnything v2 model to synchronously extract multi-level semantic features and depth feature maps , and through a multi-layer perceptron MLP, and are weighted and summed after consistent dimensionality conversion to generate enhanced depth semantic features ; (9.3) Cross-modal attention fusion stage: In the feature fusion decoder, , and the historical hidden features are input. Temporal information is fused through temporal self-attention, the deformable attention module realizes adaptive optical flow spatial feature aggregation, and the cross-attention module realizes depth semantic information fusion; (9.4) Temporal recursive optimization stage: The output feature of the last layer of the current decoder is input into the camera pose prediction layer to output the six-degree-of-freedom relative pose of the camera between the current frames; meanwhile, a depth map of sparse reference points is generated through a fully connected neural network filter; and when processing and in the next time step, is input as part of the temporal self-attention module to construct cross-frame feature associations. In this way, over time, each time frame incorporates the latent features of the previous frame, and continuously optimizes the camera's motion trajectory in the time dimension similar to a recurrent neural network RNN; (9.5) Through a progressive curriculum learning mechanism, the recursive feature transfer process is kept stably convergent in complex scenarios, and finally a trajectory optimization result with temporal chain dependence is formed.

[0015] Compared with the prior art, the present invention has the following beneficial effects: The present invention constructs an end-to-end monocular visual odometry model based on the Transformer encoder-decoder architecture. Through the multi-course joint training strategy of multi-source heterogeneous datasets, the generalization ability and robustness of the model are significantly improved. The present invention uses a Deformable Transformer encoder to extract multi-scale features of optical flow, forming a pyramidal feature representation, and inputs it into the decoder to fuse coarse and fine-grained information, thereby constructing an efficient model for predicting the relative motion of the camera. At the same time, to improve the adaptability and generalization ability in low-texture and high-exposure scenarios, the decoder incorporates semantic features and depth features provided by the depth semantic encoder through the Cross Attention mechanism, expanding feature matching from single-pixel differences to association matching based on object-level semantics and relative depth, effectively reducing the matching error caused by brightness changes. In addition to spatial information, the present invention introduces a temporal self-attention module, which dynamically represents the motion trend of objects in the environment by fusing historical temporal features, thereby improving the perception ability in dynamic scenarios. The present invention uses a network decoder containing three decoder blocks. Each decoder block outputs an intermediate prediction result and uses it as the input of the next layer, gradually realizing the coarse-to-fine prediction of the relative pose of the camera through three-level iteration. Compared with traditional single prediction methods, this structure significantly improves the accuracy and stability of pose estimation. Description of the Drawings

[0016] Figure 1 is the overall architecture diagram of the Transformer of the present invention; Figure 2 is the flow chart of the application of the present invention in temporal information; Figure 3 is the specific structure diagram of the camera pose prediction layer of the present invention. Detailed Embodiments

[0017] To more clearly elaborate the technical objectives, scheme compositions, and beneficial effects of the present invention, the following will detail the preferred embodiments of the present invention with reference to the accompanying drawings of the specification, where: The embodiment of the present invention provides an end-to-end monocular visual odometry method that fuses spatio-temporal semantic information. The specific implementation process is as follows: (1) Collect a continuous image sequence through a color monocular camera to obtain an image frame sequence , , … , where represents the frame ordinal number of the current frame.

[0018] (2)Construct an end-to-end deep learning localization Transformer framework, which includes: an optical flow estimation network, a multi-scale optical flow feature extraction network, an optical flow intrinsic parameter encoder, a depth semantic feature encoder, a multi-modal spatio-temporal feature fusion decoder, and a camera pose prediction layer (as Figure 3 shown).

[0019] Optical flow is a fundamental task in computer vision, which is used to estimate the two-dimensional motion information of each pixel between video frames and is of great significance for downstream tasks such as camera motion prediction, 3D reconstruction, and action recognition. As Figure 1 shown, the present invention first inputs two frames of images and the input optical flow into the optical flow estimation network to generate a pixel-level optical flow field. The optical flow estimation network adopts the SEA-RAFT architecture, and its advantage lies in balancing efficient inference and high-precision results.

[0020] Subsequently, considering the differences in the intrinsic parameters of cameras in different datasets, the camera intrinsic parameter layer is concatenated with the optical flow field and jointly input into the optical flow intrinsic parameter encoder. The optical flow intrinsic parameter encoding process is as follows: multi-scale features of the optical flow intrinsic parameter fusion are extracted through a multi-scale feature extraction network and corresponding position encodings are generated; after concatenating the multi-scale features and adding them to the position encodings, they are input into the three-level attention module of the encoder. To optimize GPU memory usage, the present invention introduces a deformable attention mechanism (Deformable Attention) in the cross-attention layers of the optical flow intrinsic parameter encoder and the subsequent feature fusion decoder. Each attention module contains a Deformable Self-Attention layer and a feed-forward network (FFN) layer, and finally outputs multi-scale optical flow intrinsic parameter encoded features. At the same time, the image is input into the depth semantic feature encoder to extract depth information and semantic information. The backbone framework of the depth semantic feature encoding uses the Depth anything v2 depth estimation model architecture (Yang L, Kang B, Huang Z, et al. Depth anything v2[J]. Advances in Neural Information Processing Systems, 2024, 37: 21875-21911.), in which semantic information based on the large vision model (DINOv2) and depth information based on the DPT depth estimation head are obtained. After unifying the dimensions of the semantic and depth information through a linear layer, the depth information is used as the position encoding and the semantic information is used as the features to be fused with each other to generate depth semantic encoded features.

[0021] After obtaining the multi-scale optical flow intrinsic encoding features and depth semantic encoding features, the present invention uses a feature fusion decoder to fuse the two features. The decoder consists of three levels of attention modules. Each level of the module sequentially includes temporal self-attention, deformable attention, self-attention, and cross-attention, which respectively implement the functions of cross-frame temporal feature fusion, adaptive aggregation of optical flow intrinsic features, improvement of feature quality, and incorporation of depth semantic features to assist in camera relative pose and depth estimation. Figure 2 As shown in and the output of the model at two frames, and and the randomly initialized inputs of the inference models at two frames are concatenated and input into the temporal self-attention of the feature fusion decoder together; then, its output is input into the next deformable attention module. The deformable attention serves as an optical flow intrinsic feature fusion module. Another input of this module is the multi-scale optical flow intrinsic encoding features, and its purpose is to incorporate the optical flow into the decoder; subsequently, its own features are weighted by the self-attention module; finally, the depth semantic encoding features are incorporated into the decoder using cross-attention to assist in supervising the decoder's calculation of the camera relative pose and depth estimation. Between each level of the module, there is a reference point filtering module, which is used to dynamically update the reference point coordinates and predict the depth and confidence, and at the same time adjust the gradient transmission of the post-reference points through gradient blocking operations to optimize the training process.

[0022] Finally, the output features of the feature fusion decoder are input into the camera pose prediction layer for processing to predict the relative pose transformation matrix of the camera between two frames of images. The prediction layer includes a 1D convolution layer and a parallel double-branch fully connected network (each branch consists of three fully connected layers), which respectively output the three-dimensional translation and the three-dimensional rotation of the camera, jointly constituting the relative pose transformation matrix with six degrees of freedom components.

[0023] During the training process of the model, the setting of the loss function directly affects the optimization effect of the network model parameters. The present invention designs corresponding loss functions for the camera pose prediction and depth estimation tasks respectively. Among them, the camera pose loss is used to optimize the camera translation and rotation parameters, and the depth loss is used to measure the accuracy of depth estimation. The total loss function of the model is defined as (where , ): ; Since the scale of the monocular camera movement itself cannot be observed from the monocular image sequence, the loss function of the camera pose , scale ambiguity only affects translation , rotation The loss remains unchanged. Therefore, the camera pose loss function is defined according to the TartanVO loss function : ; where and are the predicted values of camera motion, and are the true values of camera motion, Avoid incorrect calculations of dividing by zero in the formula.

[0024] For depth prediction, the present invention uses relative depth for supervision, which is to consider unknown or variable scales and use scale-invariant loss in the logarithmic depth space. At the same time, the multi-scale, scale-invariant gradient matching term is adapted to the disparity space. The purpose of this gradient matching term is to keep the discontinuities sharp and consistent with the discontinuities in the ground truth. For the overall depth prediction loss function of a frame of picture , where is 0.5, ; ; where ; where ; At the same time, combined with the confidence value evaluation of the region where the specific reference point is located, the heteroscedastic arbitrary uncertainty setting is used for the depth prediction loss function of the specific region of the reference point : ; where and are the depth prediction value and the true value respectively; is the set of reference points of the th frame of picture, is the pixel position of a single reference point, is the confidence of the corresponding reference point, where is predicted by the fully connected layer of the filter.

[0025] The final total depth loss function is : .

[0026] (3) Use multi-source heterogeneous datasets for curriculum sharing parameter learning, where the AdamW (Adaptive Moment Estimation Optimizer with Weight Decay) combined with the LambdaLR (Learning Rate Adjuster based on Lambda function) learning rate piecewise decay strategy is used to learn and train the network weights.

[0027] In view of the significant domain differences among different datasets, the present invention adopts a curriculum learning method to construct training batches to enhance the generalization ability of the model. Specifically, first, the dataset is divided into multiple parts according to its source. Then, based on the multi-curriculum learning strategy, the training data is graded according to the scene complexity (such as static scene vs. dynamic scene, texture richness), and the training batches are sampled in ascending order. In each batch, all datasets are sampled synchronously according to their scale ratios to ensure data diversity. Subsequently, the sampled data is input into the model with shared parameters for training. Through this progressive learning method, the model can first master the simple scene features and gradually adapt to complex scenes, thereby further improving the generalization ability and convergence effect of the model.

[0028] For the hyperparameters in the training process, the present invention uses 8 heads for all attention modules and sets the number of queries to 100, and the queries are learnable embeddings with predicted 2D reference points. The channel and the latent feature dimensions of all MLPs in the invention are set to 256. This model can be optionally trained on a single RTX 4090 GPU. The present invention trains for 200 epochs, with a batch size of 24 and a learning rate of 2×10 4 . The optimizer uses AdamW with a weight decay of 10 4 , and the learning rate is reduced by 0.1 times at 100 and 150 epochs.

[0029] (4) The continuous image sequence collected by the color monocular camera and is input into this Transformer framework. Finally, the sparse depth and the relative pose transformation information of the camera between image frames are obtained.

[0030] The overall process of the present invention is as follows: First, the continuous images collected by the color monocular camera are used as the initialization sequence and are input into the optical flow estimation network (SEA-RAFT) to estimate the optical flow and generate pixel-level optical flow . Subsequently, the optical flow and the camera intrinsics are concatenated and input into the optical flow intrinsics encoder to generate multi-scale optical flow intrinsics encoded feature vectors . At the same time, is input into the depth estimation network (Depth Anything v2) to synchronously extract multi-level semantic features and the depth feature map , and through the multi-layer perceptron (MLP) for and Perform dimension alignment and weighted summation to generate enhanced deep semantic features .

[0031] In the feature fusion decoder, for the initial time frame, since there is no output feature of the previous frame in the time dimension, a randomly initialized feature vector is used as the decoder input query. Subsequently, the optical flow intrinsic parameter encoding features are aggregated through the Deformable Attention mechanism. And integrate deep semantic features through the Cross Attention mechanism The queries output by each layer of the decoder are input into the camera pose prediction layer and the fully connected neural network filter, respectively, to output the camera relative pose and sparse depth information. At the same time, the output features of the last layer of the current decoder are retained as the time series input of the next time frame.

[0032] In subsequent time frames (such as processing and When the last frame is read, the output features retained from the previous time frame are concatenated with the randomly initialized query as the input of the current decoder. The temporal information is fused through the Temporal Self-Attention module, and then the feature fusion and prediction steps are repeated. Through this temporal recursive mechanism, the model continuously incorporates the potential features of historical frames, similar to a recurrent neural network, and continuously optimizes the estimation accuracy of the camera motion trajectory in the time dimension.

[0033] In summary, the present invention proposes a monocular visual odometry method based on deep learning, and innovatively applies the Transformer network codec architecture to the camera positioning task. The cross-attention module effectively integrates depth and semantic information, thereby enhancing the robustness of the model to environmental changes; at the same time, the temporal self-attention mechanism is used to dynamically couple historical features to achieve feature decoupling of static background and dynamic objects in the scene, significantly improving the ability to simulate long-term dependencies. The model predicts the relative pose and sparse depth of the camera respectively through parallel modules, further optimizing the positioning accuracy. In the training stage, a progressive course learning mechanism is adopted, combined with multi-source heterogeneous data sets, so that the model gradually adapts from simple scenes to complex environments, greatly improving the generalization ability and convergence effect. Compared with traditional monocular visual odometry, the positioning performance of the present invention in dynamic and complex scenes is better.

[0034] The above embodiments fully exemplify the core solution and innovative benefits of the present invention through specific technical implementation paths and comparative experimental data. It should be noted that the above embodiments are only optional practice examples of the technical solution of the present invention and do not constitute a limitation on the protection scope of the present invention. Any adaptive modifications based on the essence of the technical solution of the present invention (such as alternative implementations of the same algorithm), functional expansions (including but not limited to adding new training constraint terms), and technical adjustments with equivalent replacement effects to the features of the prior art (such as hyperparameter modification or using other optimizers to replace the AdamW optimizer, etc.) shall fall within the protection scope defined by the claims of the present invention. When those skilled in the art implement the present invention, the adjustment of non-core parameters such as data ratio parameters and the specific functional form of the learning rate decay curve is also regarded as a reasonable application scope of the technical solution of the present invention.

Claims

1. An end-to-end monocular visual odometry method integrating spatiotemporal semantic information, characterized in that: The following steps are involved: (1) Collect continuous image sequences through a color monocular camera to obtain image frame sequences , , … ,in Indicates the frame number of the current frame; (2) Construct an end-to-end deep learning positioning Transformer framework, which includes: an optical flow estimation network, a multi-scale optical flow feature extraction network, an optical flow intrinsic parameter encoder, a deep semantic feature encoder, a multimodal spatiotemporal feature fusion decoder, and a camera pose prediction layer; (3) Utilize multi-source heterogeneous datasets and self-constructed dynamic fuzzy datasets to learn course shared parameters, using an adaptive moment estimation optimizer with weight decay combined with a learning rate adjuster based on Lambda function and a learning rate segmented decay strategy to learn and train network weights; (4) Continuous image sequence acquired by a color monocular camera and Input into the constructed end-to-end deep learning positioning Transformer framework to obtain sparse depth and camera relative pose transformation information between image frames.

2. The end-to-end monocular visual odometer method integrating spatiotemporal semantic information according to claim 1, characterized in that: The optical flow estimation network in step (2) adopts the SEA-RAFT network architecture, and realizes the representation of motion information by constructing a dense optical flow field estimated between adjacent frame images.

3. The end-to-end monocular visual odometer method integrating spatiotemporal semantic information according to claim 1, characterized in that: The multi-scale optical flow feature extraction network of step (2) adopts the architecture of ResNet-50 as the feature backbone, and constructs a core feature extraction module by introducing residual jump connection and cross-level feature fusion mechanism; the multi-scale optical flow feature extraction network includes a five-level downsampling architecture, uses the camera intrinsic parameter matrix to construct a learnable parameter embedding layer, and fuses it with the optical flow feature at the channel level as input, and uses convolution to compress the spatial dimension of the optical flow field and the intrinsic parameter fusion information, and finally generates a composite optical flow feature map containing shallow geometric information and deep semantic information.

4. The end-to-end monocular visual odometer method integrating spatiotemporal semantic information according to claim 1, characterized in that: The optical flow intrinsic parameter encoder in step (2) uses a deformable attention mechanism to aggregate the shallow geometric information, deep semantic information, and camera intrinsic parameter layer of the composite optical flow feature map generated by the multi-scale optical flow feature extraction network, and encodes its features through sparsely sampled dynamic features using a three-layer deformable attention mechanism. Generate multi-scale optical flow intrinsic encoding feature vectors.

5. The end-to-end monocular visual odometer method integrating spatiotemporal semantic information according to claim 1, characterized in that: The deep semantic feature encoder in step (2) is implemented in the following way: (5.1) Use the encoder of the Depth Anything v2 model to extract multi-level semantic features from the input image frame; (5.2) Input the extracted semantic features into the multi-layer perceptron MLP for feature dimension encoding; (5.3) Synchronously use the depth decoder branch of Depth Anything v2 to estimate scene depth; (5.4) The semantic features encoded by MLP in step (5.2) are fused with the position encoding to generate the key-value pair features required by the attention mechanism and input into the cross-attention module of the feature fusion decoder.

6. The end-to-end monocular visual odometer method integrating spatiotemporal semantic information according to claim 1, characterized in that: The feature fusion decoder of step (2) realizes the fusion of time series information in the following way: (1) A three-level cascaded Transformer decoder architecture is used to achieve iterative three-time estimation of pose and depth. Each decoder layer contains a composite attention mechanism module, which is integrated in the following order: temporal self-attention module based on the BEVFormer architecture, deformable attention module, standard self-attention module, and cross-attention module; (2) A reference point adjustment module consisting of a fully connected neural network is deployed between adjacent decoder layers. The module is specifically implemented as follows: based on the feature output of the previous layer, the reference point coordinates input to the next layer are dynamically updated, and the depth value and confidence weight of the corresponding reference point of the current frame are simultaneously predicted; (3) During the reference point coordinate update process, a gradient separation operation is performed on the adjusted reference point coordinate tensor to block the gradient transfer of the coordinate parameters during the back propagation process.

7. The end-to-end monocular visual odometer method integrating spatiotemporal semantic information according to claim 1, characterized in that: The camera pose prediction layer of step (2) is used to predict the relative motion of the camera between two frames of images. This motion is a relative pose transformation matrix with six degrees of freedom components: Its architecture consists of a layer of 1D convolution and a parallel two-branch fully connected network structure, in which the translation prediction branch: consists of a three-layer fully connected network, outputting a three-dimensional translation vector to predict the relative amount of camera translation ,in is the 3D translation, and another rotation prediction branch: it consists of a three-layer fully connected network and outputs a rotation matrix The Lie algebra parameters represent the predicted 3D rotation of the camera.

8. The end-to-end monocular visual odometer method integrating spatiotemporal semantic information according to claim 1, characterized in that: The step (3) uses a multi-source heterogeneous data set and a self-constructed dynamic fuzzy data set to learn course shared parameters, and progressively trains the model in the following manner: progressively divides the heterogeneous training data from low to high difficulty according to the source, constructs a multi-course learning sequence based on the complexity of the scene, and sorts the training batches in order from low to high data difficulty; In each training cycle, four types of data are synchronously extracted at a sampling rate proportional to the size of the multi-source heterogeneous datasets and the independently constructed dynamic fuzzy dataset to form a mixed batch, and input into the Transformer framework with shared parameters for joint training; through a curriculum learning mechanism with increasing difficulty, the model gradually adapts from simple geometric scenes to complex dynamic fuzzy environments, and simultaneously improves the generalization ability of motion estimation and the training convergence efficiency.

9. The end-to-end monocular visual odometer method integrating spatiotemporal semantic information according to claim 1, characterized in that: The sparse depth and camera relative pose transformation information between image frames is obtained in step (4), and the camera motion trajectory is estimated through the following iterative spatiotemporal feature fusion mechanism: (9.1) Optical flow feature generation stage: continuous input frames and Input SEA-RAFT optical flow estimation network, output pixel-level optical flow field ; (9.2) Multimodal feature encoding stage: Optical flow field and the camera intrinsic layer Concatenate the input optical flow intrinsic encoder to generate a multi-scale optical flow intrinsic encoding feature vector ; Then, the current frame Input the Depth Anythingv2 model to extract multi-level semantic features simultaneously and deep feature maps , and through the multi-layer perceptron MLP and Dimension transformation is performed consistently with weighted summation to generate enhanced deep semantic features ; (9.3) Cross-modal attention fusion stage: In the feature fusion decoder, , Hidden features with history Input, temporal information is fused through temporal self-attention, the deformable attention module realizes adaptive optical flow spatial feature aggregation, and the cross attention module realizes deep semantic information fusion; (9.4) Timing recursive optimization stage: Output features of the last layer of the current decoder Input the camera pose prediction layer and output the six-degree-of-freedom relative pose of the camera between the current frames; at the same time, a depth map of sparse reference points is generated through a fully connected neural network filter; and in the next timing step, and When As part of the input of the temporal self-attention module, it builds cross-frame feature associations, so that as time goes by, each time frame incorporates the potential features of the previous frame, similar to the recurrent neural network RNN, which continuously optimizes the camera's motion trajectory in the time dimension; (9.5) Through the progressive curriculum learning mechanism, the recursive feature transfer process can maintain stable convergence in complex scenarios, and finally form a trajectory optimization result with temporal chain dependence.

Citation Information

Patent Citations

  • Visual odometer method based on semantic prior

    CN112819853A

  • Deep privileged visual odometer method based on cross-modal knowledge distillation

    CN114743105A

  • Method for constructing visual odometer through image robustness based on deep learning

    CN117079072A

  • Transform-based end-to-end multi-frame joint pose estimation method and device

    CN118505808A

  • Fuzzy robust visual mileage calculation method based on deblurring network

    CN119006721A

Cited By

  • Visual inertial odometer based on deep learning and use method thereof

    CN120252705A

  • Monocular visual odometer positioning method based on end-to-end deep learning

    CN120259619A

  • End-to-end monocular visual odometer method for adaptively adjusting attention domain

    CN120298500A

  • End-to-end underwater three-dimensional reconstruction method and system based on underwater imaging model

    CN120635333A

  • Monocular vision and sparse IMU-based rehabilitation action whole body attitude estimation method and system

    CN120673471A