Visual inertia odometer method based on lightweight design
By adopting lightweight neural networks and adaptive optimization strategies in the visual inertial odometer method, the resource restricted and low-cost IMU noise problems in the prior art are solved, and efficient visual inertial data fusion and positioning accuracy are achieved.
Patent Information
- Application Number
- CN202510297030.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
AI Technical Summary
The existing visual inertial odometry method is difficult to operate in real time on resource-constrained hardware platforms, and lacks adaptability to noise and drift problems of low-cost IMUs, resulting in insufficient positioning accuracy and robustness.
Lightweight neural network and adaptive optimization strategies are adopted to extract visual features through lightweight convolutional neural networks, lightweight deep neural networks perform pre-integration modeling of IMU data, and a tightly coupled optimizer is used to build an adaptive weighted loss function, combining graph neural networks to realize scene semantic recognition and loopback detection.
It realizes efficient visual inertial data fusion on resource-constrained embedded devices, improves positioning accuracy and robustness, and reduces computing resource occupancy.
Smart Images

Figure CN120141528A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision, inertial navigation, and embedded system optimization, and more specifically, to a visual-inertial odometry method based on a lightweight neural network and a tightly coupled sensor fusion algorithm. Background Art
[0002] In fields such as autonomous driving and robot navigation, accurate real-time positioning and environmental perception are core requirements. Traditional single-sensor positioning technologies (such as pure visual odometry or inertial navigation) often suffer from reduced accuracy or tracking loss due to factors such as environmental light changes, rapid movement, or sensor noise. Visual-inertial odometry (VIO) can theoretically combine the complementary advantages of both by fusing camera and inertial measurement unit (IMU) data - vision provides high-precision scene constraints, and the IMU provides high-frequency motion estimation, thus enhancing robustness in complex scenarios.
[0003] In recent years, with the popularity of embedded devices and mobile applications, the demand for low-power, high-real-time lightweight positioning technologies has surged. However, existing visual-inertial odometry methods are mostly based on tightly coupled optimization frameworks. Although they can achieve high accuracy, they rely on non-linear optimization algorithms (such as sliding window optimization, factor graph optimization), resulting in high computational complexity and large memory occupancy, making it difficult to run in real-time on resource-constrained hardware platforms. In addition, traditional methods usually assume the use of high-precision IMU sensors and lack adaptability to the noise and drift problems of low-cost IMUs, further limiting their universality in practical scenarios.
[0004] In summary, in the prior art, the visual-inertial odometry method still needs to optimize the efficient fusion of visual and inertial data. At the same time, it is necessary to consider how to reduce the computational resource occupancy while improving the positioning accuracy, making the model more lightweight to meet the requirements of resource-constrained embedded devices. Summary of the Invention
[0005] The present invention aims to provide a visual-inertial odometry method based on lightweight design to solve the above deficiencies of the prior art, more efficiently fuse visual and inertial data, improve positioning accuracy and robustness, and reduce computational resource occupancy through a lightweight neural network and an adaptive optimization strategy, making it more suitable for resource-constrained embedded devices. To this end, the present invention adopts the following technical solutions.
[0006] A visual-inertial odometry method based on lightweight design, the method comprising the following steps:
[0007] S1: Synchronize the RGB-format visual images collected by the camera and the accelerometer and gyroscope data of the inertial measurement unit (IMU), and preprocess the visual images and IMU data.
[0008] S2: Extract image features through a lightweight Convolutional Neural Network (CNN), detect feature points and generate descriptors for key frames, and perform feature matching through a Transformer.
[0009] S3: Use a lightweight deep neural network to perform pre-integration modeling on IMU data, and predict the relative motion information between adjacent key frames through end-to-end learning, including displacement increment, velocity change, and attitude rotation amount, to solve the problem that traditional pre-integration is sensitive to noise and reduce the accumulation of integration errors.
[0010] S4: Input the visual features and the IMU motion estimation output by the neural network into a lightweight tightly-coupled optimizer, construct an adaptive weighted loss function for visual reprojection error and IMU motion residuals, and jointly optimize pose, velocity, and sensor bias parameters using sliding window non-linear optimization technology.
[0011] S5: Implement scene semantic recognition through a lightweight Graph Neural Network (GNN), perform fast loop detection in combination with the Bag of Words model (BoW), and use pose graph optimization to sparsely adjust the global trajectory and map to suppress the cumulative error during long-term operation.
[0012] S6: Output the 6-degree-of-freedom pose (position and attitude quaternion), real-time velocity vector, and sparse feature map of the environment of the system.
[0013] The embodiments of the present invention bring the following beneficial effects: The present application provides a visual-inertial odometry method based on lightweight design. On the one hand, it reduces the occupation of computing resources through lightweight neural networks and model compression technology, making the model run more efficiently on embedded devices; on the other hand, through an adaptive sensor fusion strategy and a tightly-coupled optimization framework, it makes full use of the complementary characteristics of visual and inertial data to improve positioning accuracy and system robustness.
[0014] Other features and advantages of the present invention will be described in the subsequent specification, and part of them will become obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the specific embodiments of the present invention, the drawings required for the specific embodiments will be briefly introduced below. The drawings in the following description only relate to some embodiments of the present invention and do not limit the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1: Specific operation flowchart of the method provided by the embodiment of the present invention;
[0017] Figure 2 : Structural diagram between the steps of the method provided by the embodiment of the present invention. Detailed implementation manners
[0018] To make the objectives, technical solutions, and advantages of the present invention more understandable, the present invention will be further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0019] The flowchart and structural diagram of the present invention are respectively as Figure 1 、 Figure 2 shown. A visual-inertial odometry method based on lightweight design uses a multi-modal data synchronization and preprocessing module to preprocess video images and IMU data, etc. Then, the video images are input into the Visual Backbone of the visual feature extraction module for feature extraction, and the preprocessed IMU data is input into the IMU motion estimation module. Then, a tightly coupled optimization and adaptive fusion module is used to optimize the pose, and then loop detection and correction are performed through the loop detection and correction module. Finally, the pose, velocity, and environmental three-dimensional map are output in real time. The method includes the following steps:
[0020] S1: The visual image data mainly comes from a D435-i camera, where the video format is RGB video or RGB-D video, the resolution is 640×480, the frame rate is 30Hz, and the size of a single frame image does not exceed 1MB. The inertial measurement unit (IMU) data includes a three-axis accelerometer and a gyroscope, and the sampling frequency is 200Hz. The data streams of the camera and the IMU are synchronized through a hardware timestamp, and the message_filters toolkit of the ROS framework is called for time alignment to ensure that the maximum time deviation between the visual frame and the IMU data is less than 1ms.
[0021] The visual image and IMU data are preprocessed. First, the visual image is dedistorted. The undistort function of OpenCV is used to eliminate lens distortion based on the camera calibration parameters, and the pixel value is normalized from [0, 255] to [-1, 1]; then the Adaptive Motion Thresholding method is used to filter key frames. When the average displacement of the optical flow feature points between adjacent frames exceeds 5 pixels, it is determined to be a key frame. For the IMU data, the high-frequency noise is first suppressed by sliding window mean filtering (window size 10ms), and the gravity direction is estimated based on the acceleration data of the initial static segment, and the IMU coordinate system is rotated and aligned; finally, the IMU data is linearly interpolated between the visual key frames to generate an inertial data stream that is time-aligned with the key frames. The final output is a time-aligned visual-inertial data packet, including visual keyframes (640×480 RGB images, tensor dimensions are [Batch×3×H×W]), interpolated IMU data blocks (each keyframe corresponds to 10 sets of inertial data, dimensions are [Batch×10×6], 6 channels include accelerometer xyz, gyroscope xyz) and timestamp synchronization error (recorded as millisecond offset for error compensation in subsequent optimization modules). The specific implementation formulas for keyframe screening and preprocessing are as follows:
[0022]
[0023] Among them, v i is the i-th visual-inertial data stream, f j is the jth key frame, f′ j is the jth key frame after preprocessing, K f is the filtered keyframe set, AdaptiveMotionThreshold(·) is the adaptive motion threshold filtering method, and Preprocess(·) is the data preprocessing operation.
[0024] S2: After the key frame data is passed to the visual module, the visual features are extracted through the lightweight visual backbone network (Visual Backbone). Traditional methods usually use convolutional neural networks (CNN) or 3D convolutional networks (such as I3D) for feature extraction, but these methods are difficult to efficiently capture the spatiotemporal relationship between different regions of the image. To this end, this module introduces a lightweight visual backbone network based on EfficientNet-B0, and combines it with transfer learning technology, using the weight initialization model pre-trained on ImageNet to improve feature extraction capabilities while reducing training parameters.
[0025] EfficientNet-B0 balances the depth, width, and resolution of the network through the compound scaling strategy, significantly reducing the computational complexity while ensuring performance. Its output is a low-dimensional sparse visual feature representation Fv, with dimensions [Batch×128×H / 16×W / 16], where 128 is the number of feature channels, and H / 16 and W / 16 are the spatial resolutions of the feature maps.
[0026] The visual feature extraction process can be expressed by the following formula:
[0027] Fv = EfficientNet-B0(fj′) (3)
[0028] where fj is the preprocessed key frame, Fv is the extracted visual feature representation, and EfficientNet-B0(·) is the lightweight visual backbone network.
[0029] S3: Input the IMU data into a lightweight deep neural network (DNN), and predict the relative motion information between adjacent key frames through end-to-end learning. This network consists of two fully connected layers (FC) and one long short-term memory network (LSTM). The input is an IMU data block, and the output is the relative displacement increment Δp, the velocity change Δv, and the attitude rotation amount Δq. The IMU motion estimation process can be expressed by the following formula:
[0030] Δp, Δv, Δq = DNN-IMU(I k ) (4)
[0031] where I k is the k-th IMU data block, Δp is the relative displacement increment, Δv is the velocity change; Δq is the attitude rotation amount (represented by quaternion), and DNN-IMU(·) is the lightweight IMU motion estimation network.
[0032] S4: Input the visual feature Fv and the IMU motion estimation Δp, Δv, Δq into a tightly coupled optimizer, construct an adaptive weighted loss function for the visual reprojection error and the IMU motion residual, and jointly optimize the pose, velocity, and sensor bias parameters using the sliding window nonlinear optimization technique. The tightly coupled optimization process can be expressed by the following formula:
[0033] Loss = λ 1 ·ReprojectionError(F v ) + λ 2 ·IMUResidual(Δp, Δv, Δq) (5)
[0034] where λ 1 、λ 2is the adaptive weight coefficient, ReprojectionError(·) is the visual reprojection error, and IMUResidual(·) is the IMU motion residual.
[0035] S5: Implement scene semantic recognition through a lightweight graph neural network (GNN), perform fast loop closure detection in combination with the bag-of-words model (BoW), and use pose graph optimization to sparsely adjust the global trajectory and map to suppress the cumulative error during long-term operation. The loop closure detection and global correction process can be expressed by the following formula:
[0036] T corrected = PoseGraphOptimization(T initial , LoopClosure) (6)
[0037] where T initial is the initial pose estimate, and LoopClosure is the loop closure detection result; T corrected is the corrected global pose.
[0038] S6: Output the 6-degree-of-freedom pose (position and attitude quaternion), real-time velocity vector, and environmental sparse feature map of the system to support the low-latency and high-robustness positioning requirements in scenarios such as autonomous driving and drone navigation.
[0039] As described above, the above are only the specific implementation manners of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A visual inertial odometry method based on lightweight design, characterized in that: The method comprises the following steps: S1: Synchronize visual images and inertial measurement unit data. Read RGB images and IMU data from the D435-i camera and perform preprocessing. S2: Extract visual features. Image features are extracted through a lightweight convolutional neural network (CNN), feature point detection and descriptor generation are performed on key frames, and feature matching is performed through Transformer. S3: IMU pre-integration and motion estimation. By using a deep neural network to pre-integrate the IMU data, the relative motion information between adjacent key frames and the precise motion estimation in a short time can be obtained. S4: Tightly coupled optimization fusion. The Bundle Adjustment optimization algorithm is used to process the visual feature information and IMU pre-integration results to estimate the system pose and velocity. S5: Global optimization and loop detection: The loop detection module identifies visited scenes and combines global optimization methods to correct accumulated errors, further improving the accuracy and consistency of positioning and mapping. S6: Output position and map information. Output the system's real-time position estimation, velocity information, and a three-dimensional map of the environment in real time.
2. According to claim 1, a visual inertial odometry method based on lightweight design is characterized in that: In step S1, the visual image collected by the camera D435-i and the accelerometer and gyroscope data collected by the IMU sensor are synchronized, and the visual image is stored in RGB format. At the same time, the IMU data is preprocessed to obtain a time-aligned visual-inertial data stream.
3. According to claim 1, a visual inertial odometer method based on lightweight design is characterized in that: In step S2, the CNN convolution module is used to extract visual features, and the NMS method is used to reduce the influence of noise to improve the extraction of feature points. The Lin-Transformer method is used to match the feature points. The specific implementation formula of the Lin-Transformer is as follows: F(Z)=ReLU(ZW1+b1)W2+b2 (1) Among them, Z is the input feature matrix, W1 and W2 are input feature matrices, and b1 and b2 are bias terms.
4. According to claim 1, a visual inertial odometer method based on lightweight design is characterized in that: In step S3, a lightweight deep neural network is used to perform pre-integration modeling on the IMU data. The relative motion information between adjacent key frames is predicted through end-to-end learning, including displacement increment, velocity change, and attitude rotation. This solves the problem of traditional pre-integration being sensitive to noise and reduces the accumulation of integral errors.
5. According to claim 1, a visual inertial odometer method based on lightweight design is characterized in that: In step S4, the visual features and the IMU motion estimation output by the neural network are input into a lightweight tightly coupled optimizer to construct an adaptive weighted loss function of the visual reprojection error and the IMU motion residual, and the sliding window nonlinear optimization technology is used to jointly optimize the pose, velocity and sensor bias parameters.
6. The visual inertial odometer method based on lightweight design according to claim 1, characterized in that: In step S5, scene semantic recognition is realized through a lightweight graph neural network (GNN), and fast loop detection is performed in combination with the bag-of-words model (BoW). The pose graph optimization is used to perform sparse adjustment on the global trajectory and map to suppress the accumulated error of long-term operation.
Citation Information
Cited By
Pose information correction method, head-mounted display device and medium
CN120821388A
Vehicle-mounted AR-HUD real-time alignment system and method based on multi-modal fusion
CN121632183A