An end-to-end monocular visual odometry method with adaptive adjustment of attention domain

Through adaptive adjustment of the end-to-end monocular visual odometry method of the domain of concern, combined with optical flow prediction, feature extraction and semantic feature fusion, the problem of insufficient accuracy of traditional methods in dynamic scenes and low-texture areas is solved, and lightweight and real-time pose estimation is achieved, suitable for autonomous driving and robot navigation.

CN120298500BActive Publication Date: 2025-08-26ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510779933.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-08-26
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

The traditional monocular visual odometry method performs poorly in dynamic scenes and low-texture areas, lacks fine processing and optimization mechanisms for features, and the learning-based method is insufficiently adaptable to real-time and complex scenes, and does not fully integrate semantic information.

Method used

The end-to-end monocular visual odometry method of adaptively adjusting the domain of attention is adopted, combining optical flow prediction, feature extraction, attention optimization and pose regression, and dual attention mechanisms and semantic features fusion are introduced to seamlessly integrate optical flow estimation and pose estimation through deep learning design.

Benefits of technology

It improves the estimation accuracy and robustness in dynamic scenes and low-texture areas, reduces system complexity, enhances scene comprehension capabilities, and realizes lightweight and real-time pose estimation, suitable for autonomous driving and robot navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298500B_ABST
    Figure CN120298500B_ABST
Patent Text Reader

Abstract

The present invention discloses an end-to-end monocular visual odometry method with adaptive adjustment of the focus domain. This method constructs a compact connection architecture between an optical flow estimation network and a pose estimation network, and seamlessly integrates optical flow prediction, feature extraction, attention optimization, and pose regression based on an end-to-end deep learning design. A dual attention mechanism (CBAM module) and semantic feature fusion are introduced to filter out interference from dynamic objects, low-texture, and blurred areas. Combined with semantic feature fusion, the model's ability to understand scenes is enhanced, thereby optimizing pose estimation accuracy and network structure. A lightweight CBAM module and efficient residual block design are used to reduce computational complexity, enabling adaptive adjustment of feature extraction and enhanced scene understanding. The present invention reduces system complexity, eliminating the need for manual design and optimization of multiple modules. The end-to-end design avoids error accumulation between modules, resulting in more stable overall performance and enabling efficient pose estimation on resource-constrained embedded devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of robot navigation and positioning, autonomous driving, and computer vision technology, and in particular relates to an end-to-end monocular visual odometry method for adaptively adjusting a focus domain. Background Art

[0002] Traditional methods (such as ORB-SLAM2) use a modular design, encompassing multiple independent modules such as feature extraction, matching, pose estimation, and optimization. These modules require individual tuning, resulting in high system complexity and difficult maintenance. Learning-based methods (such as DeepVO) employ end-to-end networks, but their structures are relatively simple and lack refined feature processing and optimization mechanisms.

[0003] Furthermore, traditional methods perform poorly in dynamic scenes and low-texture areas. For example, ORB-SLAM2 performs only corner detection and uses bundle adjustment for pose optimization. However, its feature point matching process has a high failure rate in dynamic scenes and environments with repetitive textures. Learning-based methods (such as DeepVO and TartanVO) have shown some improvement, but their structures are not compact. Even compact networks, such as DROID-SLAM, are complex and bulky, and cannot meet real-time requirements.

[0004] Existing traditional methods rely entirely on geometric features and ignore semantic information. Learning-based methods (such as DeepVO and TartanVO) only use optical flow or image features to directly infer pose, failing to incorporate deep, rich semantic information as a guide. This results in limited scene understanding capabilities. Furthermore, existing technologies perform well in static, texture-rich scenes but struggle in dynamic or low-texture scenes. While learning-based methods have shown improvements, they lack real-time performance and adaptability to complex scenes. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this paper presents an end-to-end monocular visual odometry method with adaptive focus adjustment. Based on an end-to-end deep learning design, this method seamlessly integrates optical flow prediction, feature extraction, attention optimization, and pose regression. It also incorporates a dual attention mechanism (CBAM module) and semantic feature fusion to achieve adaptive adjustment of feature extraction and enhance scene understanding.

[0006] The present invention is achieved through the following technical solutions: an end-to-end monocular visual odometry method for adaptively adjusting the focus domain, comprising the following steps:

[0007] (1) Use a monocular camera to capture a continuous sequence of video frames; then preprocess the input image and adjust the normalized pixel value range;

[0008] (2) Input two consecutive frames of images into the pre-trained optical flow estimation network, and combine the hidden layer features and cost bodies obtained by the two ResNetFPN networks with the initialized confidence map and initialized optical flow obtained by the confidence upsampling module and the optical flow upsampling module respectively. Figure 1 After the same input RNN network goes through the cycle, the final predicted optical flow map, that is, the optical flow information, is obtained;

[0009] (3) Generate an intrinsic reference layer based on the camera intrinsic parameters; crop the optical flow information and the intrinsic reference layer, scale them by a quarter, and then concatenate them in the channel dimension to generate the input features of the pose network;

[0010] (4) Using the semantic information fusion mechanism to shift attention to the feature level, by calculating the weight of each channel or region in the feature map, the original feature map is weighted fused according to the calculated semantic weight to achieve the fusion of semantic information and optical flow information;

[0011] (5) The input features of the pose network generated in step (3) are input into the pose estimation network for processing, namely, feature extraction of optical flow intrinsic parameters and dual attention optimization, and finally the pose of the camera is obtained.

[0012] Specifically, the resolution of each frame of the input image in step (1) is , the format is RGB three channels; its frame rate remains stable at 30 frames per second; and the preprocessed normalized pixel value range is adjusted to [0,1].

[0013] Specifically, in step (2), the optical flow estimation network is estimated using a pre-trained Sea-RAFT network, and its input is two consecutive frames of images, which are passed through two ResNetFPN networks, one of which extracts the hidden layer features and semantic features of the image required by the subsequent RNN network through a layer of convolution; the image features extracted by the other ResNetFPN network are used to calculate the relevant cost body between the two image features, and the confidence and optical flow are initialized by obtaining the initialization graphs through the confidence upsampling module and the optical flow upsampling module respectively; then, the initialized confidence graph, the initialized optical flow graph, the hidden layer features and the cost body are input into the RNN network together and passed through N cycles to obtain the final optical flow information.

[0014] Specifically, the camera intrinsic parameters in step (3) need to be calibrated in advance and kept constant or dynamically updated with each frame. The camera intrinsic parameters are converted into the intrinsic parameter layer as input to increase the generalization ability of this model on different cameras.

[0015] Furthermore, the step (4) specifically includes the following sub-steps:

[0016] (4.1) The semantic features are adjusted to the same resolution as the feature map of layer 2 in the middle of the ResNet backbone network through bilinear interpolation, and then the semantic weights are generated through convolution and Sigmoid function;

[0017] (4.2) Finally, the semantic weight is multiplied element-wise with the layer2 output feature map to achieve the fusion of semantic information and optical flow information.

[0018] Furthermore, the step (5) specifically includes the following sub-steps:

[0019] (5.1) Concatenate the input features generated by the pose network Input the initial convolution module, which is a 3-layer 2D convolution operation, i.e., one downsampling with a stride of 2 and two convolutions with a stride of 1; the number of output channels is 32, the feature map size is halved, and the output dimension is ;

[0020] (5.2) Extracting optical flow intrinsic parameter features through a ResNet-based backbone network; the ResNet-based backbone network is a five-layer residual module, namely layer 1 to layer 5; the number of output channels of each layer is 64, 128, 128, 256, and 256, and the feature map resolution is reduced to 1×2 layer by layer;

[0021] (5.3) A dual attention mechanism is used to refine the input feature map in two stages. First, the channel attention module focuses on the channel part, and then the spatial attention module focuses on the information part. This is used for channel optimization and spatial optimization in end-to-end VO. The channel optimization is used to suppress noise or irrelevant information; the spatial optimization focuses on static background areas and reduces the weight of dynamic objects or low-texture areas. Finally, the CBAM module is introduced in each residual module layer and the final output layer.

[0022] (5.4) The pose inference network includes a camera translation prediction network and a camera rotation prediction network; the camera translation prediction network is composed of a three-layer fully connected network, which is used to output a three-dimensional translation vector, that is, the relative amount of camera translation between two frames of images; the camera rotation prediction network is composed of a three-layer fully connected network, which is used to output the Lie algebraic parameters of the rotation matrix, that is, the three-dimensional rotation amount of the camera between two frames of images, that is, the final pose of the camera is obtained by splicing.

[0023] Furthermore, the dual attention mechanism in step (5.3) includes two modules, as follows:

[0024] (a) Channel attention module: perform global average pooling and maximum pooling on the input features to generate average pooling features and max pooling features The channel description vectors capture global statistical information and significant features respectively; the two channel description vectors are respectively passed through two 2D convolutions with a convolution kernel of 1×1, where the first layer of 1×1 convolution reduces the number of channels from Reduce to , which is the reduction rate, take 16, as the hidden layer; the second layer 1×1 convolution, the number of channels from Restore to , as the output layer, and then add the two convolution results; generate the channel attention weight through the Sigmoid function, and finally output the channel attention feature With the input feature map Multiply element by element to obtain the channel-optimized feature map;

[0025] (b) Spatial attention module: Pooling and concatenating the channel-optimized feature maps, that is, performing average pooling and maximum pooling on the input features along the channel axis to generate two average pooling features and max pooling features The spatial description map is then spliced ​​along the channel axis to form features ; Then use 7×7 2D convolution to process the splicing results and output the spatial attention map of a single channel; finally, pass the Sigmoid function Generate spatial attention weights, and the final output is to multiply the spatial attention weights by the input feature map element by element to obtain the spatially optimized feature map.

[0026] The present invention includes the following beneficial effects:

[0027] (1) This paper proposes a compact architecture that connects the optical flow estimation network and the pose estimation network. Based on an end-to-end deep learning design, it seamlessly integrates optical flow prediction, feature extraction, attention optimization, and pose regression. This reduces system complexity, eliminates the need for manual design and tuning of multiple modules, and the end-to-end design avoids error accumulation between modules, resulting in a more compact architecture and more stable performance.

[0028] (2) The present invention adaptively focuses on key features such as static backgrounds through a dual attention mechanism, filters out interference from dynamic objects, low textures, and blurred areas, and combines semantic feature fusion to enhance the model's ability to understand the scene, thereby optimizing pose estimation accuracy. At the same time, the network structure is optimized, and a lightweight CBAM module and efficient residual block design are used to reduce computational complexity. This makes the present invention lightweight and compact.

[0029] (3) By fusing semantic information, the present invention enables the model to better understand the scene structure and object relationships in the image. For example, in complex environments, the model can distinguish foreground objects from the background, improve the perception of the overall context, and thus enhance the capture of task-related information. The semantic weights enable the model to adaptively adjust the focus on different features according to task requirements. In the visual odometry task, the model can focus on the features that are most useful for pose estimation (such as static landmarks) rather than irrelevant dynamic interference, thereby optimizing feature utilization efficiency. In dynamic scenes, low-texture areas, or blurred images, the semantic information fusion module filters out noise and interference features through semantic guidance, allowing the model to maintain stable performance and accuracy.

[0030] (4) With its comprehensive advantages of accuracy, robustness, and real-time performance, this invention is applicable to fields such as autonomous driving, robot navigation, and augmented reality. It can achieve efficient pose estimation on resource-constrained embedded devices, thus broadening the scope of applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0032] Figure 1 This is a diagram of the overall end-to-end network architecture of the present invention;

[0033] Figure 2 This is a diagram of the optical flow network architecture of the present invention;

[0034] Figure 3 This is a diagram of the posture network architecture of the present invention;

[0035] Figure 4 Schematic diagram of the dual attention (CBAM) module of the present invention. DETAILED DESCRIPTION

[0036] Exemplary embodiments are described in detail herein, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numerals in different drawings represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with this application.

[0037] This example describes an end-to-end monocular visual odometry system that adaptively adjusts the focus area based on semantic features combined with a dual attention mechanism. This system utilizes an end-to-end deep learning network, combined with optical flow network prediction, a dual attention mechanism (Convolutional Block Attention Module, or CBAM), and semantic feature fusion, to address the accuracy and robustness issues faced by traditional monocular visual odometry in dynamic scenes, changing lighting, and ambiguous conditions. The following details its implementation, including system architecture, method steps, process requirements, and technical results.

[0038] Scenario description: Any monocular camera with known intrinsic parameters can acquire image and video frames through the monocular camera, calculate the relative position of the camera using only two frames of images, and finally infer the motion trajectory of the camera in the entire video stream.

[0039] 1. System Configuration

[0040] The product structure of this system consists of the following core components, such as Figure 1 As shown, the components are closely connected through data flow and calculation logic:

[0041] (1) Monocular camera

[0042] Function: Collect continuous video frame sequences as the input data source of the system.

[0043] Position relationship: Fixedly installed on the front end of a mobile device (such as a robot or unmanned vehicle), with the lens facing forward to capture environmental images.

[0044] Parameters: frame rate of about 30 frames per second, unlimited resolution, RGB three-channel output, and known parameters.

[0045] (2) Optical flow network module

[0046] Function: Predict the optical flow information between adjacent video frames and represent the motion of pixels.

[0047] Connection relationship: such as Figure 2 , receive two consecutive frames output by the monocular camera , output optical flow information and semantic features To the pose estimation network.

[0048] Implementation: Based on pre-trained optical flow estimation network, such as Figure 2As shown in the figure, the Sea-RAFT network was selected in this study (Wang Y, Lipson L, Deng J. Sea-raft: Simple, efficient, accurate raft foroptical flow[C] / / European Conference on Computer Vision. Cham: SpringerNature Switzerland, 2024: 36-54.).

[0049] (4) Pose network module

[0050] Function: The core pose estimation module is responsible for the extraction, optimization and pose regression of optical flow intrinsic parameter fusion features.

[0051] Composition: Contains the initial convolution module, residual network, dual attention (CBAM) module, semantic feature fusion module and fully connected camera motion prediction branch.

[0052] Connection relationship: Receives optical flow intrinsic fusion information and semantic features, fuses them through modules, and outputs the final camera pose.

[0053] The action relationship between each component is as follows: after the monocular camera captures the image, the data flows through the optical flow network, the intrinsic parameter layer generation and the optical flow splicing module in sequence, and then enters the pose estimation network for fusion processing, and finally the pose result is output by the computing device.

[0054] 2. Methods and Steps

[0055] This system estimates the camera pose through the following specific steps, which are clear and operational:

[0056] Step 1: Data Collection

[0057] Capture continuous video frame sequences using a monocular camera .

[0058] The resolution of each frame of image input for processing is 480×640, the format is RGB three channels, and the frame rate is kept stable at 30 frames / second; and the pre-processed normalized pixel value range is adjusted to [0,1] to ensure the standardization of the input data.

[0059] Step 2: Optical flow and semantic feature prediction Figure 2 As shown, two consecutive frames Input optical flow estimation network, which uses pre-trained Sea-RAFT network for estimation, and its input is two frames , after two ResNetFPN networks that have been loaded with pre-trained weights, one of them extracts the hidden layer features required by the subsequent RNN network through a layer of convolution. At the same time, using the current frame Input the semantic extraction network together to extract the Semantic feature information of the frame ; Another extracted image feature is used to calculate the correlation cost body between the two image features; the confidence upsampling module and the optical flow upsampling module are used to obtain the initialized confidence map and initialized optical flow map respectively. Subsequently, the initialized confidence map, initialized optical flow map, hidden layer features and cost body are input into the RNN network together through the cycle to obtain the final predicted optical flow information. (B is the batch size, 2 is the x and y direction components of the optical flow, is the height of the channel, is the width of the channel).

[0060] Step 3: Internal reference layer generation and optical flow splicing

[0061] According to the camera internal parameters (focal length ) Generate internal reference layer , which is used for subsequent training with different camera data to improve generalization ability.

[0062] Concatenate the optical flow information and the intrinsic reference layer in the channel dimension to generate input features , and finally align and perform interpolation scaling to improve training efficiency.

[0063] Step 4: Semantic feature fusion

[0064] like Figure 3 As shown in the figure, the attention is turned to the feature level by using the mechanism of semantic information fusion. By calculating the weight of each channel or region in the feature map, the original feature map is weighted fused according to the calculated weight; that is, the semantic feature fusion module is to combine the semantic feature information The resolution is adjusted by bilinear interpolation and aligned with the output feature of layer2 (i.e. the second layer) (B×128×15×20). The semantic weight is generated by 1×1 convolution and Sigmoid function. , multiplied element-by-element with the layer2 feature, realizes the fusion of semantic information and optical flow information, and also provides attention to the key areas of the scene. The calculation expression of the Sigmoid function is: .

[0065] Step 5: Feature extraction and dual attention optimization of optical flow intrinsic parameters

[0066] The input features generated after splicing Input pose estimation network for processing:

[0067] (5.1) Concatenate the generated input features The initial convolution module of the input pose network is a 3-layer 2D convolution operation (the first layer has a stride of 2 downsampled to 60×80, and the second and third layers have a stride of 1). The number of output channels is 32, the feature map size is halved, and the output dimension is .

[0068] (5.2) Extracting optical flow intrinsic parameter features through a ResNet-based backbone network; the ResNet-based backbone network is a residual network including five layers of residual modules (layer1 to layer5), with the number of output channels being 64, 128, 128, 256, and 256, respectively, and the feature map resolution being reduced to 1×2 layer by layer.

[0069] (5.3) If Figure 3 As shown in the figure, a dual attention mechanism is used to refine the input feature map in two stages. First, the channel attention module adjusts the attention of the channel part, and then the spatial attention module adjusts the attention of the image texture information part. It is used for channel optimization and spatial optimization in end-to-end VO. The channel optimization is used to suppress noise or irrelevant information; the spatial optimization focuses on the static background area and reduces the weight of dynamic objects or low-texture areas.

[0070] The dual attention mechanism (CBAM module) includes a channel attention module and a spatial attention module, such as Figure 4 As shown, the channel attention module performs global average pooling and maximum pooling on the input features to generate average pooling features. and max pooling features The channel description vectors capture global statistical information and significant features respectively; the two channel description vectors are respectively passed through two 2D convolutions with a convolution kernel of 1×1, where the first layer of 1×1 convolution reduces the number of channels from Reduce to , which is the reduction rate, take 16, as the hidden layer; the second layer 1×1 convolution, the number of channels from Restore to , as the output layer, and then add the two convolution results; generate the channel attention weight through the Sigmoid function, and finally output the channel attention feature With the input feature map Multiply element by element to get the channel-optimized feature map. Spatial attention module: Pool and concatenate the channel-optimized feature map, that is, perform average pooling and maximum pooling on the input features along the channel axis to generate two average pooling features. and max pooling features The spatial description map is then spliced ​​along the channel axis to form features ; Then use 7×7 2D convolution to process the splicing results and output the spatial attention map of a single channel; finally, generate the spatial attention weight through the Sigmoid function, and the final output is to multiply the spatial attention weight by the input feature map element by element to obtain the spatially optimized feature map.

[0071] Regarding the use of the CBAM module, this invention embeds the CBAM module (blue) into each residual module (green) and layer5 output. It then optimizes feature representation through channel and spatial attention mechanisms. Furthermore, a CBAM module is introduced into the final output layer to adjust the focus of the final output features. The cascaded design of the CBAM modules adaptively focuses on information-rich areas, improving the model's ability to focus on dynamic regions and enhancing the robustness and accuracy of feature representation. Embedding the CBAM module into the residual module and layer5 output optimizes features at each layer, significantly improving the accuracy of camera pose estimation.

[0072] Step 6: Pose Inference

[0073] The pose inference network includes a camera translation prediction network and a camera rotation prediction network. The camera translation prediction network is composed of a three-layer fully connected network and is used to output a three-dimensional translation vector, that is, the relative amount of camera translation between two frames of images. The camera rotation prediction network is composed of a three-layer fully connected network and is used to output the Lie algebra parameters of the rotation matrix, that is, the three-dimensional rotation amount of the camera between two frames of images, that is, the final pose of the camera is obtained by splicing. The details are as follows:

[0074] (6.1) Flatten the features output by layer 5 combined with the CBAM module into a B×512 vector and input it into the pose inference network to predict the camera pose.

[0075] (6.2) The pose inference network consists of two fully connected branches (each branch has a structure of 512→128→32→3) that regress the translation parameters (3D vector) and rotation parameters Finally, the two are spliced ​​into the final pose .

[0076] 3. Training conditions

[0077] To ensure the feasibility of the technical solution, the following are the specific training parameters and conditions:

[0078] (a) Network training

[0079] Training epochs: 150.

[0080] Initial learning rate: 0.0001, learning rate adjustment step: [60,120] (respectively reduce the learning rate in the 60th and 120th rounds

[0081] Rate).

[0082] Weight decay: 0.0001.

[0083] Data preprocessing: Crop the original image to 480×640, generate the internal reference layer, and perform the optical flow and internal reference data

[0084] Samples are scaled to match the network input.

[0085] Optimizer: AdamW, loss function: L1 loss (measures the difference between the predicted pose and the actual pose).

[0086] (b) Reasoning Environment

[0087] Hardware: NVIDIA Jetson TX2, Orin Nano, or similar embedded device with GPU acceleration support.

[0088] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the contents disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed in this application.

[0089] It will be understood that the present application is not limited to the exact construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof.

Claims

1. An end-to-end monocular visual odometry method with adaptive adjustment of the focus region, characterized by: The following steps are involved: (1) Use a monocular camera to capture a continuous sequence of video frames; then preprocess the input image and adjust the normalized pixel value range; (2) Two consecutive frames of images are input into the pre-trained optical flow estimation network, and the hidden layer features and cost bodies obtained by the two ResNetFPN networks are input into the RNN network together with the initialized confidence map and initialized optical flow map obtained by the confidence upsampling module and the optical flow upsampling module respectively. After N cycles, the final predicted optical flow map, i.e., optical flow information, is obtained, and the optical flow estimation network outputs semantic features; (3) Generate an intrinsic reference layer based on the camera intrinsic parameters; crop the optical flow information and the intrinsic reference layer, scale them by a quarter, and then concatenate them in the channel dimension to generate the input features of the pose network; (4) Using the semantic information fusion mechanism to shift attention to the feature level, by calculating the weight of each channel or region in the feature map, the original feature map is weighted fused according to the calculated semantic weight to achieve the fusion of semantic information and optical flow information; specifically, it includes the following sub-steps: (4.1) The semantic features are adjusted to the same size as the feature map resolution of the second layer layer 2 of the ResNet backbone network through bilinear interpolation, and then the semantic weights are generated through 1×1 convolution and Sigmoid function; the ResNet backbone network is the backbone network of the pose estimation network; (4.2) Finally, the semantic weight is multiplied element-by-element with the layer2 output feature map to achieve the fusion of semantic information and optical flow information; (5) The input features of the pose network generated in step (3) are input into the pose estimation network for processing, namely, feature extraction of optical flow intrinsic parameters and dual attention optimization, and finally the pose of the camera is obtained.

2. The end-to-end monocular visual odometry method with adaptive adjustment of the focus area according to claim 1 is characterized in that: In step (1), the resolution of each frame of the input image is H×W, and the format is RGB three-channel; its frame rate is kept stable at 30 frames per second; and the preprocessed normalized pixel value range is adjusted to [0, 1].

3. The end-to-end monocular visual odometry method with adaptive adjustment of the focus area according to claim 1 is characterized in that: In the step (2), the optical flow estimation network is estimated using a pre-trained Sea-RAFT network, and its input is two consecutive frames of images, which are passed through two ResNetFPN networks, one of which extracts the hidden layer features and semantic features of the image required by the subsequent RNN network through a layer of convolution; the image features extracted by the other ResNetFPN network are used to calculate the relevant cost body between the two image features, and the confidence and optical flow are initialized by obtaining initialization maps through the confidence upsampling module and the optical flow upsampling module respectively; then, the initialized confidence map, the initialized optical flow map, the hidden layer features and the cost body are input into the RNN network together and cycled N times to obtain the final optical flow information.

4. The end-to-end monocular visual odometry method with adaptive adjustment of the focus area according to claim 1, characterized in that: The camera intrinsic parameters in step (3) need to be calibrated in advance and kept constant or dynamically updated with each frame. Converting the camera intrinsic parameters into an intrinsic parameter layer as input increases the generalization ability of this model on different cameras.

5. The end-to-end monocular visual odometry method with adaptive adjustment of the focus area according to claim 1, characterized in that: The step (5) specifically includes the following sub-steps: (5.1) Input features generated by splicing pose networks Input the initial convolution module, which is a 3-layer 2D convolution operation, i.e., one downsampling with a stride of 2 and two convolutions with a stride of 1; the number of output channels is 32, the feature map size is halved, and the output dimension is (5.2) Extracting optical flow intrinsic features through a ResNet-based backbone network; The ResNet-based backbone network is a five-layer residual module, layer 1 to layer 5; the number of output channels of each layer is 64, 128, 128, 256, 256, and the feature map resolution is reduced to 1×2 layer by layer; (5.3) A dual attention mechanism is used to refine the input feature map in two stages. First, the channel attention module focuses on the channel part, and then the spatial attention module focuses on the information part. This is used for channel optimization and spatial optimization in end-to-end VO. The channel optimization is used to suppress noise or irrelevant information. The spatial optimization focuses on the static background area and reduces the weight of dynamic objects or low-texture areas; finally, the CBAM module is introduced in each layer of residual module and the final output layer; (5.4) The pose inference network includes a camera translation prediction network and a camera rotation prediction network; the camera translation prediction network is composed of a three-layer fully connected network, which is used to output a three-dimensional translation vector, that is, the relative amount of camera translation between two frames of images; the camera rotation prediction network is composed of a three-layer fully connected network, which is used to output the Lie algebraic parameters of the rotation matrix, that is, the three-dimensional rotation amount of the camera between two frames of images, that is, the final pose of the camera is obtained by splicing.

6. The end-to-end monocular visual odometry method with adaptive adjustment of the focus area according to claim 5, characterized in that: The dual attention mechanism in step (5.3) includes two modules, as follows: (a) Channel attention module: perform global average pooling and maximum pooling on the input features to generate average pooling features and max pooling features The channel description vectors capture global statistical information and salient features respectively; The two channel description vectors are respectively passed through two 2D convolutions with a convolution kernel of 1×1. The first layer of 1×1 convolution reduces the number of channels from C to C / r, where r is the reduction rate, which is 16, as the hidden layer; the second layer of 1×1 convolution restores the number of channels from C / r to C, as the output layer, and then the two convolution results are added; the channel attention weight is generated by the Sigmoid function, and finally the channel attention output feature M is obtained. c (F) is multiplied element-by-element with the input feature map F to obtain the channel-optimized feature map; (b) Spatial attention module: Pooling and splicing the channel-optimized feature maps, that is, performing average pooling and maximum pooling on the input features along the channel axis to generate two average pooling features and max pooling features The spatial description map is then spliced ​​along the channel axis to form features Then use 7×7 2D convolution to process the splicing results and output the spatial attention map of a single channel; finally, the spatial attention weight is generated by the Sigmoid function σ. The final output is to multiply the spatial attention weight by the input feature map element by element to obtain the spatially optimized feature map.

Citation Information

Patent Citations

  • Three-dimensional reconstruction method of arbitrary point tracking network based on dynamic perception

    CN118212364A

  • Scene perception and reconstruction system based on geometric semantic information fusion

    CN119810317A