Adaptive Covariance Method Based on Feature Observation Number and IMU Pre-Integration
Through the adaptive covariance method of feature observations and IMU pre-integration, the error matching problem caused by dynamic targets in the visual positioning system in the dynamic environment is solved, and the accuracy and robustness of the visual positioning system are improved.
Patent Information
- Application Number
- CN202211063510.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-01
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-09-01
AI Technical Summary
In the visual positioning system under a dynamic environment, dynamic targets lead to incorrect feature matching, affect positioning accuracy, and fail to effectively deal with the uncertainty of different feature points, resulting in excessive removal or too few visual measurements.
Adaptive covariance method based on feature observations and IMU pre-integration is adopted, and the adaptive covariance matrix is calculated through the number of feature point tracking times and IMU pre-integration modeling inverse depth uncertainty, and the adaptive covariance matrix is calculated, and the reprojection error weighted by uncertainty is fused, and dynamic objects are used as visual constraint information to solve the camera pose.
It effectively alleviates the influence of dynamic targets, avoids distorted feature distribution and too few visual measurements, and improves the positioning accuracy of the visual positioning system.
Smart Images

Figure CN115457127B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence and computer vision measurement, and particularly relates to an adaptive covariance method based on the number of feature observations and IMU pre-integration. Background Art
[0002] In the past few decades, computer vision has been widely studied, providing accurate and low-cost solutions for autonomous positioning systems. Traditional vision-based positioning schemes are all based on an ideal static environment, so that the motion calculated by feature point matching is only caused by the camera motion, and thus the pose estimation of the camera can be obtained. However, in the actual scene, there may be dynamic targets, and these dynamic targets will lead to incorrect feature matching and incorrect use of the information of dynamic feature points for camera motion estimation, which is fatal to the vision-based positioning system based on feature matching. Therefore, the research on vision-based positioning systems applicable to dynamic environments has always been a hot issue.
[0003] Currently, the SLAM problem in a dynamic environment focuses more on how to distinguish the dynamic area from the static area, remove the dynamic area, and only use the feature points in the static background to calculate the camera pose. By removing dynamic objects, the influence of dynamic objects on the camera pose solution can be reduced. However, similar to the scenario where RANSAC fails, when there are a large number of dynamic objects in the scene, directly removing the features on the dynamic objects will lead to excessive feature removal, which will not only seriously affect the geometric shape of the feature distribution, but also seriously reduce the number of feature point pairs participating in visual positioning, resulting in limited visual measurement and affecting the positioning accuracy. In addition, most VIO systems do not consider the uncertainty of the features used for relative pose estimation. After removing outliers from feature matching, it is assumed that all feature points have the same noise distribution, that is, each pair of matched feature points contributes equally to the pose. However, in the real environment, the noise distribution of each point will have different types due to different positions and environments. Therefore, different two-dimensional features should have different contribution degrees. Figure 4 Shows the situation where different feature points have different uncertainties, where Figure 4 a is an isotropic and independent and identically distributed covariance matrix diagram, Figure 4 b is an anisotropic and non-independent and identically distributed covariance matrix diagram. The higher the uncertainty of the feature point, the smaller the contribution to the optimization function, and vice versa. Summary of the Invention
[0004] The present invention precisely aims at the problems existing in the prior art and provides an adaptive covariance method based on the number of feature observations and IMU pre-integration. Feature points on a dynamic object are used as visual constraint information to solve the camera pose, so as to improve the performance of the visual positioning system. First, feature points are extracted and image feature data is aligned with IMU data; feature matching is achieved by using image pyramid optical flow tracking; then the number of times of feature point tracking is calculated, and the number of observations of each feature point in the current frame is obtained in the time dimension. By calculating the number of observations of the feature points, the mean and standard deviation of the number of observations of the feature points in the current frame are calculated, so as to determine the static observation weight of the feature points in the current frame; the inverse depth uncertainty is modeled through IMU pre-integration, and then the adaptive covariance is calculated through the number of feature observations and IMU pre-integration to obtain an adaptive covariance matrix. Finally, the covariance matrix of the feature points is fused with the error function of pose solution to construct a reprojection error with uncertainty weighting, realizing the adaptive covariance method based on the number of feature observations and IMU pre-integration, effectively alleviating the influence of dynamic targets, and avoiding the distortion of feature distribution caused by excessive rejection and too few visual measurements.
[0005] To achieve the above object, the technical solution adopted by the present invention is: an adaptive covariance method based on the number of feature observations and IMU pre-integration, including the following steps:
[0006] S1: Extract feature points and align the image feature data with the IMU data. The alignment strategy is: the timestamp of the first IMU data is less than the timestamp at the end of the previous frame image, and the timestamp of the last IMU data is the first one greater than the timestamp at the end of the current frame image;
[0007] S2: Achieve feature matching by using image pyramid optical flow tracking;
[0008] S3: Calculate the number of times of feature point tracking, obtain the number of observations of each feature point in the current frame in the time dimension, calculate the mean and standard deviation of the number of observations of the feature points in the current frame by calculating the number of observations of the feature points, so as to determine the static observation weight of the feature points in the current frame;
[0009] S4: Model the inverse depth uncertainty through IMU pre-integration. The inverse depth uncertainty of the feature point is inversely proportional to the cross product of the normalized coordinates of the feature point in the camera frame and the inter-frame translation vector, and is proportional to the feature matching uncertainty δx. The inverse depth uncertainty is represented by the pre-integration term of the IMU as follows:
[0010]
[0011] where, x i = K -1 p i ,pi is the homogeneous coordinate of the feature point on the image plane, represents x i skew-symmetric matrix of, is the translation vector obtained by IMU pre-integration between the i-th frame image and the j-th frame image, δx = f -1 δp is the radius of the circle formed by the uncertainty on the normalized plane, f is the focal length of the camera, δp is the error generated during feature matching of the feature point, and δd represents the uncertainty of the feature point in the j-th frame;
[0012] S5: Calculate the adaptive covariance using the number of feature point observations obtained in step S3 and the IMU pre-integration obtained in step S4 to obtain an adaptive covariance matrix. For the l-th feature point, its static observation weight is W(p l ), its inverse depth uncertainty is δd l , and its adaptive covariance matrix Q l is:
[0013] Q l = W c W(p l )δd l I 2×2
[0014] where, W c is the original covariance matrix of the feature point, and I 2×2 is the identity matrix;
[0015] S6: Fuse the covariance matrix of the feature point calculated above with the error function of pose solution to construct an uncertainty-weighted reprojection error. The uncertainty-weighted reprojection error is:
[0016]
[0017] p l is the two-dimensional coordinate of the l-th feature point on the current frame image, n is the total number of feature points. Decomposing the covariance matrix of the feature point can obtain Q -1 = U∑ -1 U T , where,
[0018]
[0019]
[0020] Compared with the prior art, the present invention has made some improvements in the processing of dynamic objects, and is no longer limited to eliminating dynamic objects. An adaptive covariance method based on feature observation number and IMU pre-integration is proposed, and the feature points on the dynamic objects are used as visual constraint information to solve the camera pose, so as to improve the performance of the visual positioning system. When solving the camera pose, the feature observation weight of each feature point is obtained by statistically modeling the number of feature point observations in the time dimension, and the uncertainty of the inverse depth of the feature point is obtained by IMU pre-integration. The covariance is improved by combining the above two factors, and the uncertainty of each feature point is modeled according to the feature observation number and the translation and rotation constraints generated by IMU pre-integration, and this uncertainty is propagated as relative pose uncertainty. This method effectively alleviates the impact of dynamic targets, and avoids excessive elimination leading to feature distribution distortion and too few visual measurements, and has higher accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a flow chart of an adaptive covariance method based on characteristic observation number and IMU pre-integration of the present invention;
[0022] Figure 2 This is the schematic diagram of the epipolar geometry constraint;
[0023] Figure 3 is the feature matching error map;
[0024] Figure 4 is a plot of the covariance matrices of different distributions, where
[0025] Figure 4 a is the covariance matrix diagram of isotropic and independent and identically distributed;
[0026] Figure 4 b is the covariance matrix diagram of anisotropic and non-independent and identically distributed;
[0027] Figure 5 is a comparison chart of the experimental results of the present invention, wherein
[0028] Figure 5 a is the original vins-fusion experiment result diagram;
[0029] Figure 5 b is the improved VINS-fusion experiment result diagram. DETAILED DESCRIPTION
[0030] The present invention will be further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.
[0031] Example 1
[0032] An adaptive covariance technique based on the number of feature observations and IMU pre-integration, as Figure 1 shown, the specific method is as follows:
[0033] Step S1: Extract feature points and align the image feature data with the IMU data;
[0034] In this method, Shi-Tomasi corner points are extracted, and a mask strategy is used to make the extracted feature points evenly distributed. The following alignment strategy is defined: the timestamp of the first IMU data is less than the timestamp at the end of the previous frame of the image, and the timestamp of the last IMU data is the first one greater than the timestamp at the end of the current frame of the image;
[0035] Step S2: Classify the point cloud and use image pyramid optical flow tracking to achieve feature matching;
[0036] Use image pyramid optical flow tracking to achieve feature point matching. An image pyramid is established for each frame of the image to achieve optical flow tracking in the scale space, which has spatial scale invariance and is fast. There is no need to extract the descriptors of the feature points for matching;
[0037] Step S3: Calculate the number of times the feature points are tracked;
[0038] If a feature point is a static feature point, then the number of times this feature point is judged as a static feature point will be very large. The number of times of feature tracking determines the quality of feature tracking. The more times a feature point is tracked between frames, the better the quality of the feature point. Therefore, we statistically count the number of observations of each feature point in the current frame in the time dimension. Assume that the i-th feature point in the k-th frame of the image is represented by denoted as,
[0039] From the first frame to the current frame, if the feature point p i is observed in the k-th frame, then the number of observations of p i ;
[0040] N k (p i ) = N k-1 (p i ) + 1
[0041] where, N k (p i ) is the number of observations of the feature point p i in the k-th frame. When the feature point is first tracked by optical flow after extraction, at this time the feature point has experienced two frames of images, so the initial value is set to 1;
[0042] If the feature point p i is not observed, then N k (p i) = N k-1 (p i ), after calculating the number of observations of the current frame features, clear the observation number information of this feature point, and set N k (p i ) = 0;
[0043] If N k (p i ) is greater than the observation number threshold T v , then N k (p i ) = Tv;
[0044] Calculate the mean μ k and standard deviation σ k of the number of observations of the current frame features, as follows:
[0045]
[0046]
[0047] where n k is the number of feature points in the k-th frame. Using the mean μ k and standard deviation σ k of the number of observations of the feature points, the static observation weight of each feature point in the current frame can be calculated, that is
[0048]
[0049] where α is a real number greater than zero.
[0050] The fewer the number of times a feature point is tracked, the greater the probability that it belongs to a dynamic feature point, and the higher the uncertainty of this feature point;
[0051] Step S4: Model the inverse depth uncertainty through IMU pre-integration:
[0052] For visual localization based on feature methods, the most widely used strategy for relative pose estimation is to use the fundamental matrix for solution. This method highly depends on the accurate correspondence between feature points. In many visual localization systems, the uncertainty of feature points is not considered. After removing outliers from feature matching, the remaining feature matching results have the same contribution to the objective function. However, different 2D feature points should have different error distributions, which is related to image quality and whether the feature points have stable depth information. The depth estimation of feature points is numerically unstable. At the same time, the uncertainty of pose estimation will also increase. Since the inverse depth information has better numerical stability than depth information, the present invention adopts an uncertainty estimation method based on IMU pre-integration to estimate this depth uncertainty;
[0053] Assume that the position information of the spatial point P in the world coordinate system is P w = [X Y Z] T , assume that the position of the projected feature point of the spatial point P in the i-th frame is p i , and the position of the projected feature point in the j-th frame is p j . According to the epipolar geometry constraint as shown in Figure 2 , the following equation can be obtained:
[0054]
[0055]
[0056]
[0057] where Z i is the depth of P w in the i-th frame, p i = [u i v i 1] T is the homogeneous coordinate of the feature point on the image plane, K is the camera internal parameter, is the rotation matrix and translation vector of the i-th frame in the world coordinate system.
[0058] According to the derivation, it can be obtained:
[0059]
[0060] Let
[0061] It can be obtained:
[0062]
[0063] Multiply both sides of the above equation by which represents the skew-symmetric matrix of x i .
[0064] Get:
[0065]
[0066]
[0067]
[0068] As can be seen from the above formula, the uncertainty of the inverse depth of the feature point is related to the error of feature matching and the pose transformation between the i-th frame and the j-th frame. Assuming that the uncertainty generated during feature matching is isotropic and uniformly distributed, that is, an error term with an uncertainty of δp will be generated during the feature matching of all feature points. On the normalized plane, the uncertainty will form a circle with a radius of δx = f -1 δp, where f is the focal length of the camera. The following formula gives the uncertainty of the inverse depth δd;
[0069]
[0070] According to the properties of the outer product, it can be obtained that:
[0071]
[0072] It is known that ||δx j || ≤ δx, RR T = I;
[0073] The relationship between the uncertainty of the inverse depth of the feature point and the normalized coordinates of the feature point in the camera frame and the inter-frame translation vector can be obtained, specifically:
[0074]
[0075] It can be concluded from the above formula that the uncertainty of the inverse depth of the feature point is inversely proportional to the cross product of the normalized coordinates of the feature point in the camera frame and the inter-frame translation vector, and is directly proportional to the feature matching uncertainty δx. The uncertainty of feature matching is as Figure 3 shown. For the inter-frame rotation matrix and translation vector, the present invention uses IMU pre-integration to obtain the inter-frame rotation matrix and translation vector. The uncertainty of feature matching depends on different feature extraction methods and feature matching methods. The present invention assumes that the uncertainty of feature matching at this time is a constant empirical value. Therefore, the inverse depth uncertainty can be expressed by the pre-integration term of the IMU;
[0076] Step S5: Calculate the adaptive covariance through the feature point tracking quality and IMU pre-integration;
[0077] By measuring the uncertainty of the feature from the above two aspects, the adaptive covariance matrix can be deduced as:
[0078]
[0079] Step S6: Pose estimation method based on adaptive covariance;
[0080] Fuse the two-dimensional feature point uncertainty calculated above with the error function of pose solution to construct a reprojection error with uncertainty weighting, so as to reflect that feature points with different uncertainties should contribute differently to the solution during the solution process;
[0081] Use Q to represent the adaptive covariance matrix. The covariance matrix of two-dimensional feature points is a positive semi-definite symmetric matrix. Singular value decomposition of the covariance matrix can obtain:
[0082] Q -1 = UΣ -1 U T
[0083] Among them,
[0084] According to the above formula, the affine matrix H 2×2 :
[0085]
[0086] Through the affine matrix H 2×2 Transform the coordinates of two-dimensional measurement points and reprojection points in the original data space to the coordinates in the uncertainty-weighted covariance data space. It is a measure of the contribution of feature uncertainty to the objective function. Feature points with smaller feature uncertainty have a greater contribution degree to the objective function. U T is a rotation matrix that rotates the various ellipses obtained from the covariance to obtain a positive ellipse. By combining Scale the ellipse to obtain a circle. In the original data space, the noise at the feature points is anisotropic and non-independent and identically distributed. Through the affine matrix H 2×2 Transform the data into the weighted covariance data space, and the noise becomes isotropic and independent and identically distributed. Construct the uncertainty-weighted reprojection error according to the transformed data as shown in the following formula:
[0087]
[0088] Through the transformation matrix H 2×2 Allocate the feature point uncertainty to the measurement points and reprojection points, so that the uncertainty of each feature point makes different contributions to the objective function, thus adapting to the pose solution under different feature uncertainties; complete the adaptive covariance technology of feature observation number and IMU pre-integration through the above steps.
[0089] The method of the present invention was tested in VINS-FUSION. Multiple tests were carried out under the total trajectory length of 3724.2 m in the KITTI00 dataset, and the GNSS positioning result was used as the ground truth to evaluate the trajectory accuracy. The experimental results are as Figure 5 shown.Figure 5 a is the experimental result graph of the original VINS-Fusion, Figure 5 b is the experimental result graph of the improved VINS-Fusion. From left to right, they are the three-dimensional trajectory comparison graph, the three-axis position comparison graph, and the three-axis attitude angle comparison graph. The results show that the root mean square error of the trajectory is reduced from 4.52 m to 3.46 m, greatly improving the performance of the visual positioning system and having higher accuracy.
[0090] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.
Claims
1. An adaptive covariance method based on the number of feature observations and IMU pre-integration, characterized in that , including the following steps: S1: Extract feature points and align the image feature data with the IMU data. The alignment strategy is as follows: The timestamp of the first IMU data is less than the timestamp at the end of the previous frame of the image, and the timestamp of the last IMU data is the first one greater than the timestamp at the end of the current frame of the image. S2: Achieve feature matching by using image pyramid optical flow tracking. S3: Calculate the number of times of feature point tracking, obtain the number of observations of each feature point in the current frame in the time dimension, calculate the mean and standard deviation of the number of observations of the feature points in the current frame by calculating the number of observations of the feature points, so as to determine the static observation weight W(p i ); S4: Model the inverse depth uncertainty through IMU pre-integration. The inverse depth uncertainty of the feature point is inversely proportional to the cross-product coordinates of the normalized vector of the feature point in the camera frame and the inter-frame translation, and is proportional to the feature matching uncertainty δx. The inverse depth uncertainty is represented by the pre-integration term of the IMU as follows: where x i = K -1 p i and p i is the homogeneous coordinate of the feature point on the image plane, denotes the skew-symmetric matrix of x i , is the translation vector obtained by IMU pre-integration between the i-th frame image and the j-th frame image, δx = f -1 δp is the radius of the circle formed by the uncertainty on the normalized plane, f is the focal length of the camera, δp is the error generated during feature matching of the feature point, and δd represents the uncertainty of the inverse depth of the feature point in the j-th frame; S5: Calculate the adaptive covariance using the number of feature point observations obtained in step S3 and the IMU pre-integration obtained in step S4 to obtain an adaptive covariance matrix. For the l-th feature point, its static observation weight is W(p l ), its inverse depth uncertainty is δd l , and its adaptive covariance matrix Q l is: Q l = W c W(p l )δd l I 2×2 Among them, W c is the original covariance matrix of the feature points, and I 2×2 is the identity matrix; S6: Fuse the covariance matrix of the calculated feature points with the error function of pose solution to construct an uncertainty-weighted reprojection error. The uncertainty-weighted reprojection error is: where p l is the two-dimensional coordinate of the l-th feature point on the current frame image, n is the total number of feature points, and decomposing the covariance matrix of the feature points gives Q -1 = U∑ -1 U T , where 2. The adaptive covariance method based on the number of feature observations and IMU pre-integration according to claim 1, characterized in that, In step S1, Shi-Tomasi corner points are extracted, and a masking strategy is used to make the distribution of the extracted feature points uniform. The masking strategy is: Taking the circle with the extracted feature point as the center and r as the radius as the mask, and the coordinates of the remaining extracted feature points do not fall within the mask area.
3. The adaptive covariance method based on the number of feature observations and IMU pre-integration according to claim 1 or 2, characterized in that, Step S2 further includes: S21: Obtain the number of observations of each feature point in the current frame in the time dimension; assume that the $i$-th feature point in the $k$-th frame image is represented by denote From the first frame to the current frame, if feature point p i is observed in the k-th frame, then the number of observations of p i ; N k (p i ) = N k-1 (p i ) + 1 where N k (p i ) is the number of observations of feature point p i at the k-th frame, which is first tracked by optical flow after feature point extraction; If N k (p i ) is greater than the observation number threshold T v , then N k (p i ) = Tv; S22: Calculate the mean μ of the number of feature point observations in the current frame k and the standard deviation σ k , specifically as follows: where n k is the number of feature points in the k-th frame; S23: According to the calculation result of step S22, calculate the static observation weight of each feature point in the current frame, specifically: where α is a real number greater than zero.
4. The adaptive covariance method based on the number of feature observations and IMU pre-integration according to claim 3, wherein: In the step S21, if the feature point p i is observed, the feature point p i is first tracked by optical flow after extraction. At this time, the feature point p i has experienced two frames of images, and the initial value is set to 1; If the feature point p i is not observed, then N k (p i ) = N k-1 (p i ). After the calculation of the number of feature observations in the current frame is completed, clear the observation number information of this feature point, and set N k (p i ) = 0.
5. The adaptive covariance method based on the number of feature observations and IMU pre-integration according to claim 4, wherein: In step S3, the number of times of feature tracking determines the quality of feature tracking. The more times the feature point is tracked between frames, the better the quality of the feature point.