Robust visual inertial navigation method for dynamic and static factor decoupling in high dynamic environment
By generating stable confidence factors and inertial factors to distinguish between steady-state and dynamic features, and combining dynamic compensation constraints and global optimization, the trajectory drift problem of visual inertial navigation systems in dynamic environments is solved, achieving efficient robustness and accuracy improvement.
Patent Information
- Application Number
- CN202511626211.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-03-03
AI Technical Summary
Visual inertial navigation systems face interference from dynamic elements in highly dynamic environments, leading to trajectory drift and inaccurate pose estimation. Existing methods have limitations in computational efficiency and generalization ability.
A two-factor constraint module is used to generate stable confidence factors and stable inertia factors. A dynamic compensation constraint module is used to identify dynamic features. A global multi-constraint optimization module is used to decouple dynamic and static residuals to improve robustness.
It improves the robustness and trajectory accuracy of visual inertial navigation systems in complex environments, eliminates the need for scene-specific training, and adapts to dynamic changes in different scenarios.
Smart Images

Figure CN121594859A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this application relate to the field of robot navigation technology, and in particular to a robust visual-inertial navigation method for decoupling dynamic and static factors in a highly dynamic environment. Background Technology
[0002] Visual-inertial navigation systems (VINS) play a fundamental role in state estimation across various applications, including robot navigation, autonomous driving, virtual reality, and augmented reality. The use of inertial measurement units (IMUs) to assist camera-based vision-based methods has garnered significant attention in visual simultaneous localization and mapping (VSLAM) and visual inertial odometry research, thanks to their small size, low cost, complementary advantages, and independence from external signals. However, VINS was initially based on the theoretical assumption that landmarks are implicitly static, leading to significant trajectory drift in dynamic environments. Improving the accuracy and robustness of state estimation in highly dynamic environments remains a key research challenge to enhance the practical application of VINS.
[0003] In highly dynamic real-world scenarios, visual-inertial navigation systems (VIS) face three main challenges: scenes containing numerous dynamic elements, such as pedestrian traffic in a shopping mall or vehicular traffic on city streets; pixel occupancy of dynamic objects, which can lead to tracking failures when they occupy a large area of the image's field of view (pixels); and traditional problems such as regions with weak texture and variations in lighting. These challenges cause traditional multi-view geometric constraints to introduce significant errors when handling dynamic features. The fundamental reason is that dynamic features from moving objects violate the implicit assumption of static landmarks—because they cannot always be re-observed at the same spatial location in different frames. Therefore, using such dynamic features for inter-frame matching leads to inaccurate pose estimation.
[0004] In this context, accurately separating and processing dynamic elements is crucial to overcoming the challenges. Traditional methods primarily employ the RANSAC algorithm to detect and filter dynamic visual features, but this approach is only effective in scenes with small-scale dynamic features or when features follow specific degenerate motion patterns. Some researchers have explored dynamic object detection through deep clustering, feature reprojection, or deep learning, while others have utilized dynamic features for joint optimization to improve the accuracy of pose estimation. However, these deep learning-based dynamic feature detection methods are limited to predefined dynamic objects, lack generalization ability across different scenes, and are generally unsuitable for computationally constrained or highly dynamic environments. Therefore, they suffer from limitations such as low computational efficiency, poor generalization ability, and reliance on prior models. Summary of the Invention
[0005] To address the aforementioned technical issues, embodiments of this application propose a robust visual inertial navigation method for decoupling dynamic and static factors in highly dynamic environments. This method aims to improve the robustness of visual inertial navigation systems in complex environments through decoupling of dynamic and static factors and dynamic compensation constraints, without requiring scene-specific training.
[0006] To achieve the above objectives, embodiments of this application propose a robust visual-inertial navigation method for decoupling dynamic and static factors in highly dynamic environments, the method comprising the following steps: Obtain prior pose information and visual feature data of the inertial measurement unit from the navigation system; Using a two-factor constraint module, based on the prior pose information and visual feature data of the inertial measurement unit, a stable confidence factor and a stable inertial factor of the visual features are generated, and a two-factor residual is generated at the same time; the two-factor residual is used to distinguish between steady-state features and non-steady-state features. Using the dynamic compensation constraint module, based on the prior pose information and visual feature data of the inertial measurement unit, effective dynamic features are identified and dynamic effect residuals are generated; By utilizing the global multi-constraint optimization module, and combining the basic residuals, two-factor residuals, and dynamic effect residuals of the navigation system, the dynamic and static residuals are decoupled and jointly optimized to obtain the optimized navigation trajectory.
[0007] To achieve the above objectives, embodiments of this application also propose a robust visual-inertial navigation system with decoupling of dynamic and static factors in a highly dynamic environment, the system comprising: The data acquisition module is used to acquire prior pose information and visual feature data of the inertial measurement unit from the navigation system; The two-factor constraint module is used to generate stable confidence factors and stable inertial factors of visual features based on the prior pose information and visual feature data of the inertial measurement unit, and at the same time generate two-factor residuals; the two-factor residuals are used to distinguish between steady-state features and non-steady-state features. The dynamic compensation constraint module is used to identify effective dynamic features and generate dynamic effect residuals based on the prior pose information and visual feature data of the inertial measurement unit. The global multi-constraint optimization module is used to combine the basic residuals, two-factor residuals and dynamic effect residuals of the navigation system, decouple the dynamic and static residuals and perform joint optimization to obtain the optimized navigation trajectory.
[0008] To achieve the above objectives, embodiments of this application also propose an electronic device, including: a processor and a memory, wherein the memory stores instructions executable by the processor, and the processor is configured to execute the instructions such that the electronic device can implement a robust visual-inertial navigation method for decoupling dynamic and static factors in a high-dynamic environment as described above.
[0009] To achieve the above objectives, embodiments of this application also propose a computer-readable storage medium storing a computer program that, when executed by a processor, enables a robust visual-inertial navigation method for decoupling dynamic and static factors in a highly dynamic environment, as described above.
[0010] This application proposes a robust visual-inertial navigation method for decoupling dynamic and static factors in highly dynamic environments. First, it acquires prior pose information and visual feature data from the inertial measurement unit (IMU) of the navigation system. Then, using a two-factor constraint module, based on the IMU's prior pose information and visual feature data, it generates stable confidence factors and stable inertial factors for the visual features, and simultaneously generates two-factor residuals, thus initially distinguishing stable and dynamic features in the visual feature data. Next, using a dynamic compensation constraint module, based on the IMU's prior pose information and visual feature data, it identifies effective dynamic features and generates dynamic effect residuals. Finally, using a global multi-constraint optimization module, it combines the navigation system's basic residuals, two-factor residuals, and dynamic effect residuals to decouple the dynamic and static residuals and perform joint optimization, obtaining the optimized navigation trajectory. This achieves improved robustness of the visual-inertial navigation system in complex environments through dynamic and static factor decoupling and dynamic compensation constraints, without requiring scene-specific training.
[0011] Optionally, the prior pose information and visual feature data of the inertial measurement unit are obtained from the navigation system, including: preprocessing and initializing the image data collected by the visual sensor in the navigation system and the data collected by the inertial measurement unit to obtain the pre-integration result and visual feature data of the inertial measurement unit; and obtaining the prior pose information of the inertial measurement unit based on the pre-integration result of the inertial measurement unit and the initial state of the navigation system.
[0012] Optionally, based on the prior pose information of the inertial measurement unit and the visual feature data, a stable confidence factor and a stable inertial factor for the visual features are generated, and a two-factor residual is generated simultaneously. This includes: calculating the distance between the predicted position value of the visual feature in the visual feature data and the actual observation position based on the prior pose information of the inertial measurement unit; determining the stable confidence factor of the visual feature based on the distance between the predicted position value of each visual feature and the actual observation position; obtaining the stable inertial factor of the visual feature based on the stable confidence factor of the visual feature by using the number of consecutive tracking times of the visual feature in the visual feature data within the sliding window; and obtaining the two-factor residual based on the stable confidence factor and the stable inertial factor.
[0013] Optionally, the stable confidence factor of the visual feature is determined based on the distance between the predicted location value of each visual feature and the actual observation location, and is expressed by the following formulas (1) to (4): (1); (2); (3); (4); in, Indicates the stable confidence factor. This represents the visual reprojection residual term. Represent each visual feature steady-state confidence level, Indicates hyperparameters, It represents the steady-state confidence level. A supplementary characterization term for the stability confidence level. Representing visual feature data, This represents the current optimized state vector. Represents the loss function. Indicates the first The first frame of the image Information matrix of visual feature observations; The stability confidence factor based on visual features, and the stability inertia factor of visual features obtained by using the number of consecutive tracking times of visual features within the sliding window in the visual feature data, are expressed by the following formula (5): (5); in, The stability inertia factor representing visual characteristics. Representing visual features The number of times it is continuously tracked in the current sliding window. Indicates the confidence level of historical stability; The two-factor residuals, obtained based on the stable confidence factor and the stable inertia factor, are expressed by the following formula (6): (6); in, It is the regularization strength of the stability confidence factor, used to affect the convexity of the convergence loss and the magnitude of the gradient. It is the momentum intensity of the stable inertial factor.
[0014] Optionally, using a dynamic compensation constraint module, based on the prior pose information of the inertial measurement unit and visual feature data, effective dynamic features are identified and dynamic effect residuals are generated. This includes: calculating the observation position and static projection point of the visual features in the visual feature data in consecutive frames based on the prior pose information of the inertial measurement unit; calculating the motion angle and motion distance of the visual features on the normalized plane based on the observation position and static projection point of the visual feature points in consecutive frames; and identifying visual features that meet the threshold judgment as effective dynamic features based on the motion angle and motion distance of the visual features on the normalized plane, and generating dynamic effect residuals.
[0015] Optionally, based on the prior pose information of the inertial measurement unit, the observation position and static projection point of the visual features in the visual feature data in consecutive frames are calculated, expressed by the following formulas (7) and (8): (7); (8); Among them, visual feature points In image frame The observation location is In the image frame The actual two-dimensional observation location is Static projection points are ; The motion angle and distance of visual features on the normalized plane are used to identify effective dynamic features that meet the threshold judgment, which are expressed by the following formulas (9) to (10): (9); (10); in, Indicates the angle of motion. Indicates the distance traveled. Indicates the minimum angle of motion. Indicates the minimum torque. Indicates the maximum torque; The generated dynamic effect residuals are expressed by the following formulas (11) to (14): (11); (12); in, This represents the pre-integrated pose of the inertial measurement unit. Represents motion vector exist The projection of motion on the surface; use compensate get The formula is as follows: (13); The dynamic effect residuals are generated using static reprojection, as shown in the following formula: (14); in, Represents the dynamic effect residuals. This represents the visual reprojection function.
[0016] Optionally, let the vector of the state to be optimized in the sliding window be denoted as . , , Indicates the sliding window size is The The state of each frame, This represents the external parameters from the navigation system's camera to the inertial measurement unit's coordinate system. It is the first Motion compensation vectors for each motion effect, and the state vector to be optimized. By location ,speed Rotation Accelerometer bias and gyroscope bias The basic residuals of a navigation system consist of marginal residuals and inertial measurement unit residuals. The method utilizes a global multi-constraint optimization module, combining the basic residuals, two-factor residuals, and dynamic effect residuals of the navigation system, to decouple the dynamic and static residuals and perform joint optimization, resulting in an optimized navigation trajectory, including: The least squares nonlinear optimization function of the sliding window is constructed and expressed by formula (15): (15); in, Indicates marginalized residuals, Represents the residual of the inertial measurement unit. Indicates two-factor residuals, Represents the dynamic effect residual; By iteratively adjusting the state vector to be optimized, the least squares nonlinear optimization function is brought to converge, thereby obtaining the optimized navigation trajectory. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies of this application will be briefly introduced below. Obviously, the following drawings are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. The drawings described herein are only used to explain this application and are not intended to limit this application.
[0018] Figure 1 This is a flowchart of a robust visual-inertial navigation method for decoupling dynamic and static factors in a highly dynamic environment, provided in one embodiment of this application. Figure 2 This is a schematic diagram of the operation flow of a robust visual-inertial navigation method for decoupling dynamic and static factors provided in one embodiment of this application; Figure 3 This is an overall framework diagram of the two-factor constraint module, motion compensation constraint module, and global multi-constraint optimization module provided in one embodiment of this application; Figure 4 This is an effect decomposition model diagram of multi-frame dynamic features observed in VINS provided in one embodiment of this application; Figure 5 This is a schematic diagram of the structure of a robust visual-inertial navigation system for decoupling dynamic and static factors in a highly dynamic environment, provided in another embodiment of this application; Figure 6 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. Those skilled in the art will understand that many technical details have been presented in the embodiments of this application to facilitate better understanding. However, the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments. The division of the following embodiments is for ease of description and should not constitute any limitation on the specific implementation of this application. The following embodiments can be combined with and referenced by each other without contradiction.
[0020] VINS (Virtual Inertial Navigation) plays a fundamental role in state estimation across various applications, including robot navigation, autonomous driving, virtual reality, and augmented reality. IMU-assisted camera-based vision methods have garnered significant attention in visual simultaneous localization and mapping (VSM) and visual inertial odometry (VIO), thanks to their small size, low cost, complementary advantages, and independence from external signals. However, VINS was initially based on the theoretical assumption that landmarks are implicitly static, leading to significant trajectory drift in dynamic environments. Improving the accuracy and robustness of state estimation in highly dynamic environments remains a key research challenge to enhance the practical application of VINS.
[0021] In highly dynamic real-world scenarios, visual-inertial navigation systems (VIS) face three main challenges: scenes containing numerous dynamic elements, such as pedestrian traffic in a shopping mall or vehicular traffic on city streets; pixel occupancy of dynamic objects, which can lead to tracking failures when they occupy a large area of the image's field of view (pixels); and traditional problems such as regions with weak texture and variations in lighting. These challenges cause traditional multi-view geometric constraints to introduce significant errors when handling dynamic features. The fundamental reason is that dynamic features from moving objects violate the implicit assumption of static landmarks—because they cannot always be re-observed at the same spatial location in different frames. Therefore, using such dynamic features for inter-frame matching leads to inaccurate pose estimation.
[0022] In this context, accurately separating and processing dynamic elements is crucial to overcoming the challenges. Traditional methods primarily employ the RANSAC algorithm to detect and filter dynamic visual features, but this approach is only effective in scenes with small-scale dynamic features or when features follow specific degenerate motion patterns. Some researchers have explored dynamic object detection through deep clustering, feature reprojection, or deep learning, while others have utilized dynamic features for joint optimization to improve the accuracy of pose estimation. However, these deep learning-based dynamic feature detection methods are limited to predefined dynamic objects, lack generalization ability across different scenes, and are generally unsuitable for computationally constrained or highly dynamic environments. Therefore, they suffer from limitations such as low computational efficiency, poor generalization ability, and reliance on prior models.
[0023] In view of this, the embodiments of this application focus on highly dynamic environments with large-area occlusion. To address the fundamental limitations of using deep learning or semantic prior models to identify dynamic features in terms of computational efficiency and adaptability to highly dynamic scenes, a robust visual inertial navigation method for decoupling dynamic and static factors in highly dynamic environments is proposed. The aim is to improve the robustness of the visual inertial navigation system in complex environments by decoupling dynamic and static factors and dynamic compensation constraints, without requiring scene-specific training.
[0024] One embodiment of this application proposes a robust visual-inertial navigation method for decoupling dynamic and static factors in a highly dynamic environment, applied to an electronic device. The electronic device can be a terminal or a server; this embodiment and subsequent embodiments will use a server as an example. The implementation details of the robust visual-inertial navigation method for decoupling dynamic and static factors in a highly dynamic environment proposed in this embodiment are described below. The following implementation details are provided for ease of understanding and are not essential for implementing this solution.
[0025] The specific process of the robust visual-inertial navigation method for decoupling dynamic and static factors in a high-dynamic environment proposed in this embodiment can be described as follows: Figure 1 As shown, it includes: Step 101: Obtain the prior pose information and visual feature data of the inertial measurement unit from the navigation system.
[0026] In one possible embodiment, step 101 includes: preprocessing and initializing the image data collected by the visual sensor and the data collected by the inertial measurement unit in the navigation system to obtain the pre-integration result and visual feature data of the inertial measurement unit; and obtaining the prior pose information of the inertial measurement unit based on the pre-integration result of the inertial measurement unit and the initial state of the navigation system.
[0027] For example, the image data acquired by the vision sensor can be: feature tracking and extraction using the KLT sparse optical flow algorithm to ensure that each frame of the image maintains 100–300 uniformly distributed feature points. After distortion correction and normalization, the feature points are projected onto a unit spherical coordinate system, and mismatched points are removed using the RANSAC algorithm to obtain the visual feature data.
[0028] Specifically, in acquiring visual feature data, feature extraction and tracking can be performed. For each feature tracked in a frame, distortion correction and normalization are performed using camera intrinsics to obtain its coordinates on the normalized plane. On the other hand, IMU measurements are aligned with timestamp-based image frames, and vectors related to system position, velocity, and rotation are obtained through IMU pre-integration. For each new frame output by the camera, existing feature points are tracked using the KLT sparse optical flow algorithm while detecting new corner features to ensure that the number of feature points in each image remains between 100 and 300. Two-dimensional feature points are first distorted and then projected onto a unit sphere after outlier removal. Subsequently, the detector achieves a uniform distribution of feature points by setting the minimum pixel interval between adjacent feature points, and finally, outlier removal is completed using the RANSAC algorithm based on the fundamental matrix model.
[0029] For example, the data collected by the inertial measurement unit can be time-stamped and calibrated, the relative motion increments (including position, velocity, and rotation) between adjacent keyframes can be calculated through pre-integration, and an error covariance matrix can be constructed.
[0030] Specifically, in acquiring the pre-integration results of the IMU, the relative motion increment (including relative attitude change, relative velocity change, and relative position change) is obtained by integrating the IMU data within a time window between two adjacent visual keyframes. Simultaneously, the error covariance matrix of this increment is constructed to quantify the uncertainty introduced by noise (such as Gaussian noise and random walk) during the integration process. An error correction mechanism is introduced concurrently during processing: on the one hand, by modeling the IMU noise characteristics (e.g., incorporating the statistical parameters of Gaussian noise and random walk into the covariance calculation), the error propagation pattern is predicted in advance; on the other hand, it is often combined with zero-bias compensation, i.e., by estimating and correcting the IMU zero-bias error in real time, reducing the impact of zero-bias drift on integration and avoiding excessive accumulation of error over integration time. The final output "IMU pre-integration results (e.g., relative motion increment + error covariance)" can be directly used for subsequent visual IMU fusion, providing accurate and low-drift inertial motion constraints for navigation system initialization (e.g., initial attitude and velocity estimation).
[0031] Understandably, VINS can provide reliable initial values for all state variables. Without accurate initial values, directly fusing IMU measurements is prone to mismatches because a monocular camera cannot directly provide scale. The situation becomes even more complex when IMU measurements are significantly affected by bias, requiring a robust initialization process to ensure the method's applicability. For VINS, the states that need to be estimated during the initialization phase include: accelerometer bias. and gyroscope bias And the position, velocity, and rotation at a real scale within the sliding window. Among these, the accelerometer bias... With gravity Related. Due to the gravity vector, the accelerometer bias term Difficult to initialize. However, accelerometer bias. The impact on the initialization results is minimal, so no estimation is performed at this stage.
[0032] It is understandable that by aligning visual SfM (structure for motion recovery) with IMU pre-integration results, and jointly estimating the scale factor, gravity vector, velocity state, and IMU zero bias, the problems of monocular visual scale blur and IMU zero drift can be solved.
[0033] Step 102: Using the two-factor constraint module, based on the prior pose information of the inertial measurement unit and the visual feature data, generate the stable confidence factor and stable inertial factor of the visual features, and at the same time generate the two-factor residuals.
[0034] Among them, the two-factor residuals are used to distinguish between steady-state and non-steady-state characteristics.
[0035] Understandably, in the two-factor constraint module, feature data (e.g., visual features) can be extracted into stable confidence factors and stable inertia factors, thereby enabling the differentiation between steady-state and non-steady-state features. Compared with related techniques, the two-factor constraint module replaces the Huber loss function of visual residuals in traditional bundle adjustment optimization, assigning small weights and continuity to unstable visual residuals, preventing abnormally large residual values from causing deviations in state variables during local optimization.
[0036] In one possible embodiment, step 102 includes: calculating the distance between the predicted position value of the visual feature in the visual feature data and the actual observation position based on the prior pose information of the inertial measurement unit; determining the stability confidence factor of the visual feature based on the distance between the predicted position value of each visual feature and the actual observation position; obtaining the stability inertial factor of the visual feature based on the stability confidence factor of the visual feature and the number of consecutive tracking times of the visual feature in the visual feature data within the sliding window; and obtaining the two-factor residual based on the stability confidence factor and the stability inertial factor.
[0037] In one possible embodiment, the stable confidence factor of a visual feature is determined based on the distance between the predicted location value of each visual feature and the actual observation location, expressed by the following formulas (1) to (4): (1); (2); (3); (4); in, Indicates the stable confidence factor. This represents the visual reprojection residual term. Represent each visual feature steady-state confidence level, Indicates hyperparameters, It represents the steady-state confidence level. A supplementary characterization term for the stability confidence level. Representing visual feature data, This represents the current optimized state vector. Represents the loss function. Indicates the first The first frame of the image Information matrix of visual feature observations.
[0038] For example, for the stability confidence factor of visual features, the current optimization state can first be estimated using IMU prior pose information and previous optimization states. Then, using the current optimization state and visual feature data, and based on the visual reprojection residual technique, the visual reprojection residual term can be calculated. The stability confidence is then updated while keeping the current optimization state fixed.
[0039] For example, each visual feature Stability confidence The closer the value is to 1, the greater the probability that the visual feature is a steady-state feature; conversely, the closer the value is to 1, the greater the probability that the visual feature is a steady-state feature. The closer it is to 0, the more likely the visual feature is to be a non-stationary feature, making it more prone to mismatches and dynamic changes.
[0040] In one possible embodiment, the stability confidence factor based on visual features, and the stability inertia factor of visual features obtained by utilizing the number of consecutive tracking times of visual features within the sliding window in the visual feature data, is expressed by the following formula (5): (5); in, The stability inertia factor representing visual characteristics. Representing visual features The number of times it is continuously tracked in the current sliding window. This indicates the confidence level of historical stability.
[0041] For example, regarding the stability inertia factor, the stability confidence of the forward-facing state based on visual features is first incorporated to address situations where IMU pre-integration is temporarily inaccurate. In this case, the stability confidence becomes less reliable due to the unreliability of the input IMU prior pose information. Furthermore, an accurate value cannot be estimated during optimization. The proposed stabilization inertia factor, in the absence of reliable IMU prior pose, will cause the current stability confidence to tend towards the previous state. Furthermore, this inertial trend will continue with the number of optimization iterations. It increases with the increase of.
[0042] In one possible embodiment, the two-factor residuals, based on the stability confidence factor and the stability inertia factor, are obtained and expressed by the following formula (6): (6); in, It is the regularization strength of the stability confidence factor, used to affect the convexity of the convergence loss and the magnitude of the gradient. It is the momentum intensity of the stable inertial factor.
[0043] For example, the momentum strength of regularization strength and stabilization inertia factor is selected on a validation set (such as the VIODE dataset) through ablation experiments, depending on experimental parameter tuning, and is typically tried between 0.1 and 10. The convexity of the convergence loss and the magnitude of the gradient affect the convergence loss; for If the system is operating in a high-dynamic or high-IMU-noise scenario (such as a drone making a sharp turn), the value should be increased appropriately to enhance stability; if the environment is relatively static, the value can be decreased to improve the response speed to real dynamic objects.
[0044] Step 103: Using the dynamic compensation constraint module, based on the prior pose information and visual feature data of the inertial measurement unit, identify effective dynamic features and generate dynamic effect residuals.
[0045] In one possible embodiment, step 103 above includes: using a dynamic compensation constraint module, based on the prior pose information of the inertial measurement unit and visual feature data, identifying effective dynamic features and generating dynamic effect residuals, including: based on the prior pose information of the inertial measurement unit, calculating the observation position and static projection point of the visual features in the visual feature data in consecutive frames; based on the observation position and static projection point of the visual feature points in consecutive frames, calculating the motion angle and motion distance of the visual features on the normalized plane; based on the motion angle and motion distance of the visual features on the normalized plane, identifying visual features that meet the threshold judgment as effective dynamic features, and generating dynamic effect residuals.
[0046] In one possible embodiment, based on the prior pose information of the inertial measurement unit, the observation position and static projection point of the visual features in the visual feature data in consecutive frames are calculated, expressed by the following formulas (7) and (8): (7); (8); Among them, visual feature points In image frame The observation location is In the image frame The actual two-dimensional observation location is Static projection points are ; The motion angle and distance of visual features on the normalized plane are used to identify effective dynamic features that meet the threshold judgment, which are expressed by the following formulas (9) to (10): (9); (10); in, Indicates the angle of motion. Indicates the distance traveled. Indicates the minimum angle of motion. Indicates the minimum torque. Indicates the maximum torque; For example, the first formula above is used to verify the consistency of motion. If it holds true, it can be inferred that the visual feature did not move in the expected direction observed in the coordinate system as a steady-state feature, so the steady-state feature is likely a non-steady-state feature. Since the first formula cannot distinguish between dynamic features and mismatched features in non-steady-state features, the second formula above further evaluates the degree of motion, describing the visual observation. Compared with IMU predicted values The distance between them. A lower bound on the degree of motion is set, meaning only distances reaching a certain threshold represent valid motion in the feature; smaller differences are likely steady-state features; while anomalous features caused by prior pose errors or mismatches are detected through... Perform filtering. , , The parameters are obtained by tuning on a dataset through ablation experiments, and are generally set to... , , Combining these two methods allows for the identification of the required motion characteristics from both steady-state and unsteady-state features, laying the foundation for subsequent calculations of dynamic effect residuals. Understandably, considering dynamic feature points Its own movement , Compared with actual observed values The differences are distinct, and these differences can be represented by motion vectors. It is known that the dynamic characteristic residuals cannot be calculated using classical static reprojection theory formulas; otherwise… It will incorrectly tend towards zero and fail to optimize the system state. This leads to motion. After For reference, let's assume yes In the previous frame The projection before motion. The dynamic effect residual of dynamic characteristics can be defined as the projection before motion. With post-motion projection The reprojection residuals between them. Prior poses are pre-integrated using IMU. Calculate motion vectors In image frame Pre-motion projection on .
[0047] The generated dynamic effect residuals are expressed by the following formulas (11) to (14): (11); (12); in, This represents the pre-integrated pose of the inertial measurement unit. Represents motion vector exist The projection of motion on the surface; use compensate get The formula is as follows: (13); The dynamic effect residuals are generated using static reprojection, as shown in the following formula: (14); in, Represents the dynamic effect residuals. This represents the visual reprojection function.
[0048] Step 104: Using the global multi-constraint optimization module, the basic residuals, two-factor residuals, and dynamic effect residuals of the navigation system are combined to decouple the static and dynamic residuals and perform joint optimization to obtain the optimized navigation trajectory.
[0049] Understandably, by using the dynamic effect residuals of tightly coupled dynamic features, the pose estimation problem of visual inertia in high dynamic scenes can be transformed into an improved version of a nonlinear optimization problem based on a sliding window.
[0050] In one possible embodiment, the state vector to be optimized in the sliding window is denoted as . , , Indicates the sliding window size is The The state of each frame, This represents the external parameters from the navigation system's camera to the inertial measurement unit's coordinate system. It is the first Motion compensation vectors for each motion effect, and the state vector to be optimized. By location ,speed Rotation Accelerometer bias and gyroscope bias The basic residuals of a navigation system consist of marginal residuals and inertial measurement unit residuals. In one possible embodiment, step 104 above includes: The least squares nonlinear optimization function of the sliding window is constructed and expressed by formula (15): (15); in, Indicates marginalized residuals, Represents the residual of the inertial measurement unit. Indicates two-factor residuals, Represents the dynamic effect residual; By iteratively adjusting the state vector to be optimized, the least squares nonlinear optimization function is brought to converge, thereby obtaining the optimized navigation trajectory.
[0051] For example, the above formula (15) consists of the marginal residual, the IMU residual, the two-factor residual from the two-factor constraint module, and the dynamic effect residual from the dynamic compensation constraint module. State vector It is the core variable of the entire optimization problem, and its final output is the optimal global pose solution (i.e., the optimized navigation trajectory) that the system needs to estimate within the sliding window. It is continuously updated through nonlinear optimization (such as bundle adjustment) to obtain the robot trajectory and environmental structure estimate that best matches the sensor observations (IMU data and visual features).
[0052] This application proposes a robust visual-inertial navigation method for decoupling dynamic and static factors in highly dynamic environments. First, it acquires prior pose information and visual feature data from the inertial measurement unit (IMU) of the navigation system. Then, using a two-factor constraint module, based on the IMU's prior pose information and visual feature data, it generates stable confidence factors and stable inertial factors for the visual features, and simultaneously generates two-factor residuals, thus initially distinguishing stable and dynamic features in the visual feature data. Next, using a dynamic compensation constraint module, based on the IMU's prior pose information and visual feature data, it identifies effective dynamic features and generates dynamic effect residuals. Finally, using a global multi-constraint optimization module, it combines the navigation system's basic residuals, two-factor residuals, and dynamic effect residuals to decouple the dynamic and static residuals and perform joint optimization, obtaining the optimized navigation trajectory. This achieves improved robustness of the visual-inertial navigation system in complex environments through dynamic and static factor decoupling and dynamic compensation constraints, without requiring scene-specific training.
[0053] To further describe steps 101 to 104 in the embodiments of this application, the following can be combined with... Figures 2 to 4 An exemplary description is provided. For example... Figure 2 As shown, Figure 2This is a schematic diagram illustrating the operation flow of a robust visual-inertial navigation method with decoupling of dynamic and static factors provided in one embodiment of this application. First, in the tightly coupled optimization module, a new frame (a new image frame and corresponding IMU data) is input, and then a state limit is defined within a sliding window. Combined with multi-constraint residual fusion, the final optimized state limit is output. The multi-constraint residuals include IMU residuals, two-factor visual residuals, dynamic effect residuals, and marginalization prior residuals. Next, in the dynamic feature differentiation module, visual features and preliminary depth estimation are extracted from the image. Using the relative motion prior information between two frames provided by the IMU, the angle between the actual motion direction of the feature point and the motion direction of the static background estimated based on the IMU prior is calculated. The actual distance moved by the feature point is calculated, and then it is checked whether the feature point has a valid depth estimate. Finally, after being determined to be a valid depth feature, it is identified as a dynamic feature.
[0054] Combined Figure 3 As shown, Figure 3 This is an overall framework diagram of the two-factor constraint module, motion compensation constraint module, and global multi-constraint optimization module provided in one embodiment of this application. First, a sliding window is used for optimization, followed by IMU pre-integration prior pose estimation. Then, in the two-factor constraint and dynamic feature recognition, weights are initialized, and stable confidence factors and stable inertia factors are determined to generate the final weights for the current optimization iteration. Finally, in the multi-constraint joint optimization, the optimized navigation trajectory is output through the input of various constraint residuals combined with robust bundle adjustment.
[0055] Combined Figure 4 As shown, Figure 4 This is an effect decomposition model diagram of multi-frame dynamic features observed in VINS provided in one embodiment of this application; Figure 4 This reflects how the embodiments of this application decompose the motion effects of dynamic features and transform them into optimizable residual terms, thereby improving the robustness of the system in highly dynamic environments.
[0056] Understandably, by integrating multivariate constraints—dynamic and static residual factors and dynamic motion compensation constraints—the robustness of visual inertial navigation methods can be improved without relying on prior models or scene-specific training. This method models dynamic feature motion and treats it as an optimizable residual under dynamic compensation constraints. During bundle adjustment (BA), the method dynamically adjusts the feature stability confidence, effectively suppressing error propagation caused by instantaneous dynamics. Comprehensive evaluation using datasets shows that the method provided in this application demonstrates excellent performance in trajectory accuracy and robustness, especially in complex scenes containing mixed moving objects, severe occlusion, and poor lighting conditions, establishing a general optimization paradigm for the robustness of VINS in highly dynamic real-world environments.
[0057] The steps described above are for clarity only. In implementation, they can be combined into one step, or some steps can be broken down into multiple steps, as long as they involve the same logical relationship, they are all within the scope of protection of this application. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, without changing the core design of the algorithm and process, are also within the scope of protection of this application.
[0058] Another embodiment of this application proposes a robust visual-inertial navigation system with decoupling of dynamic and static factors in a highly dynamic environment. The details of this robust visual-inertial navigation system with decoupling of dynamic and static factors in a highly dynamic environment are described below. The following implementation details are provided for ease of understanding and are not essential for implementing this example. Figure 5 This is a schematic diagram of the structure of a robust visual-inertial navigation system with decoupling of dynamic and static factors in a high-dynamic environment, as proposed in this embodiment, including: The data acquisition module 510 is used to acquire prior pose information and visual feature data of the inertial measurement unit from the navigation system; The two-factor constraint module 520 is used to generate stable confidence factors and stable inertial factors of visual features based on the prior pose information and visual feature data of the inertial measurement unit, and at the same time generate two-factor residuals; wherein, the two-factor residuals are used to distinguish between steady-state features and non-steady-state features. The dynamic compensation constraint module 530 is used to identify effective dynamic features and generate dynamic effect residuals based on the prior pose information and visual feature data of the inertial measurement unit. The global multi-constraint optimization module 540 is used to combine the basic residuals, two-factor residuals and dynamic effect residuals of the navigation system, decouple the dynamic and static residuals and perform joint optimization to obtain the optimized navigation trajectory.
[0059] It is not difficult to see that this embodiment is a system embodiment corresponding to the above method embodiments, and this embodiment can be implemented in conjunction with the above method embodiments. The relevant technical details and technical effects mentioned in the above method embodiments are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above method embodiments.
[0060] It is worth mentioning that all modules and units involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed in this application; however, this does not mean that other units do not exist in this embodiment.
[0061] Another embodiment of this application provides an electronic device, such as Figure 6 As shown, it includes a processor 61 and a memory 62. The memory 62 stores instructions that the processor 61 can execute. When the processor 61 is configured to execute the instructions, the electronic device can realize a robust visual-inertial navigation method for decoupling dynamic and static factors in a high dynamic environment as described in the above method embodiment.
[0062] The memory and processor are connected via a bus, which includes any number of interconnecting buses and bridges, connecting various circuits of one or more processors and the memory. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single component or multiple components, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.
[0063] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.
[0064] Another embodiment of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, can implement a robust visual-inertial navigation method for decoupling dynamic and static factors in a highly dynamic environment, as described in the above method embodiments.
[0065] That is, those skilled in the art will understand that all or part of the steps in the above method embodiments can be implemented by a program instructing related hardware. The program is stored in a storage medium and includes several instructions to cause a device (such as a microcontroller, chip, etc.) or processor to execute all or part of the steps of the method in the method embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.
[0066] Those skilled in the art will understand that the above embodiments are specific implementations of this application, and in practical applications, various changes can be made in form and detail without departing from the spirit and scope of this application. For those skilled in the art, several improvements and modifications can be made without departing from the principles of this application, and these improvements and modifications are also considered to be within the scope of protection of this application.
Claims
1. A robust visual-inertial navigation method for decoupling dynamic and static factors in a highly dynamic environment, characterized in that, The method includes: Obtain prior pose information and visual feature data of the inertial measurement unit from the navigation system; Using a two-factor constraint module, based on the prior pose information and visual feature data of the inertial measurement unit, a stable confidence factor and a stable inertial factor of the visual features are generated, and a two-factor residual is generated at the same time; the two-factor residual is used to distinguish between steady-state features and non-steady-state features. Using the dynamic compensation constraint module, based on the prior pose information and visual feature data of the inertial measurement unit, effective dynamic features are identified and dynamic effect residuals are generated; By utilizing the global multi-constraint optimization module, and combining the basic residuals, two-factor residuals, and dynamic effect residuals of the navigation system, the dynamic and static residuals are decoupled and jointly optimized to obtain the optimized navigation trajectory.
2. The method according to claim 1, characterized in that, The acquisition of prior pose information and visual feature data of the inertial measurement unit from the navigation system includes: The image data acquired by the visual sensor and the data acquired by the inertial measurement unit in the navigation system are preprocessed and initialized to obtain the pre-integration results and visual feature data of the inertial measurement unit; Based on the pre-integration results of the inertial measurement unit and the initial state of the navigation system, the prior pose information of the inertial measurement unit is obtained.
3. The method according to claim 1, characterized in that, The prior pose information and visual feature data based on the inertial measurement unit generate stable confidence factors and stable inertial factors for the visual features, and simultaneously generate two-factor residuals, including: Based on the prior pose information of the inertial measurement unit, the distance between the predicted position of the visual feature in the visual feature data and the actual observation position is calculated. The stability confidence factor of each visual feature is determined based on the distance between the predicted location value and the actual observation location. The stability confidence factor based on visual features is obtained by utilizing the number of consecutive tracking times of visual features within a sliding window in the visual feature data. Based on the stability confidence factor and the stability inertia factor, the two-factor residuals are obtained.
4. The method according to claim 3, characterized in that, The stable confidence factor for visual features, determined by the distance between the predicted location value and the actual observation location of each visual feature, is expressed by the following formulas (1) to (4): (1); (2); (3); (4); in, Indicates the stable confidence factor. This represents the visual reprojection residual term. Represent each visual feature steady-state confidence level, Indicates hyperparameters, It represents the steady-state confidence level. A supplementary characterization term for the stability confidence level. Representing visual feature data, This represents the current optimized state vector. Represents the loss function. Indicates the first The first frame of the image Information matrix of visual feature observations; The stability confidence factor based on visual features, and the stability inertia factor of visual features obtained by using the number of consecutive tracking times of visual features within the sliding window in the visual feature data, are expressed by the following formula (5): (5); in, The stability inertia factor representing visual features. Representing visual features The number of times it is continuously tracked in the current sliding window. Indicates the confidence level of historical stability; The two-factor residuals, obtained based on the stable confidence factor and the stable inertia factor, are expressed by the following formula (6): (6); in, It is the regularization strength of the stability confidence factor, used to affect the convexity of the convergence loss and the magnitude of the gradient. It is the momentum intensity of the stable inertial factor.
5. The method according to claim 4, characterized in that, The method of utilizing the dynamic compensation constraint module, based on the prior pose information and visual feature data of the inertial measurement unit, to identify effective dynamic features and generate dynamic effect residuals includes: Based on the prior pose information of the inertial measurement unit, the observation position and static projection point of the visual features in the visual feature data in consecutive frames are calculated. Based on the observation position and static projection point of visual feature points in consecutive frames, calculate the motion angle and motion distance of visual features on the normalized plane. Based on the motion angle and distance of visual features on the normalized plane, effective dynamic features are identified by visual features that meet the threshold judgment, and dynamic effect residuals are generated.
6. The method according to claim 5, characterized in that, The prior pose information based on the inertial measurement unit is used to calculate the observation position and static projection point of the visual features in the visual feature data in consecutive frames, which is expressed by the following formulas (7) and (8): (7); (8); Among them, visual feature points In image frame The observation location is In the image frame The actual two-dimensional observation location is Static projection points are ; The motion angle and distance of visual features on the normalized plane are used to identify effective dynamic features that meet the threshold judgment, which are expressed by the following formulas (9) to (10): (9); (10); in, Indicates the angle of motion. Indicates the distance traveled. Indicates the minimum angle of motion. Indicates the minimum torque. Indicates the maximum torque; The generated dynamic effect residuals are expressed by the following formulas (11) to (14): (11); (12); in, This represents the pre-integrated pose of the inertial measurement unit. Represents motion vector exist The projection of motion on the surface; use compensate get The formula is as follows: (13); The dynamic effect residuals are generated using static reprojection, as shown in the following formula: (14); in, Represents the dynamic effect residuals. This represents the visual reprojection function.
7. The method according to claim 6, characterized in that, Let the vector of the state to be optimized in the sliding window be . , , Indicates the sliding window size is The The state of each frame, This represents the external parameters from the navigation system's camera to the inertial measurement unit's coordinate system. It is the first Motion compensation vectors for each motion effect, and the state vector to be optimized. By location ,speed Rotation Accelerometer bias and gyroscope bias The basic residuals of a navigation system consist of marginal residuals and inertial measurement unit residuals. The method utilizes a global multi-constraint optimization module, combining the basic residuals, two-factor residuals, and dynamic effect residuals of the navigation system, to decouple the dynamic and static residuals and perform joint optimization, resulting in an optimized navigation trajectory, including: The least squares nonlinear optimization function of the sliding window is constructed and expressed by formula (15): (15); in, Indicates marginalized residuals, Represents the residual of the inertial measurement unit. Indicates two-factor residuals, Represents the dynamic effect residual; By iteratively adjusting the state vector to be optimized, the least squares nonlinear optimization function is brought to converge, thereby obtaining the optimized navigation trajectory.
8. A robust visual-inertial navigation system with decoupling of dynamic and static factors in a highly dynamic environment, characterized in that, The system includes: The data acquisition module is used to acquire prior pose information and visual feature data of the inertial measurement unit from the navigation system; The two-factor constraint module is used to generate stable confidence factors and stable inertial factors of visual features based on the prior pose information and visual feature data of the inertial measurement unit, and at the same time generate two-factor residuals; the two-factor residuals are used to distinguish between steady-state features and non-steady-state features. The dynamic compensation constraint module is used to identify effective dynamic features and generate dynamic effect residuals based on the prior pose information and visual feature data of the inertial measurement unit. The global multi-constraint optimization module is used to combine the basic residuals, two-factor residuals and dynamic effect residuals of the navigation system, decouple the dynamic and static residuals and perform joint optimization to obtain the optimized navigation trajectory.
9. An electronic device, characterized in that, include: The processor and memory, wherein the memory stores instructions that the processor can execute, and the processor is configured to, when executing the instructions, enable the electronic device to implement a robust visual-inertial navigation method for decoupling dynamic and static factors in a high-dynamic environment as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it can implement a robust visual-inertial navigation method for decoupling dynamic and static factors in a high-dynamic environment as described in any one of claims 1 to 7.
Citation Information
Cited By
A dynamic object scene recognition method and system for compensating for camera self-motion
CN122244101A