Visual inertial positioning method based on dynamic target detection and semantic information constraint

Through the visual-inertial positioning method of dynamic target detection and semantic information constraints, dynamic target feature points are identified and eliminated, and semantic consistency constraints are constructed, which solves the robustness and accuracy problems of visual-inertial positioning in dynamic environments and achieves high-precision pose estimation.

CN120740573APending Publication Date: 2025-10-03BEIJING INST OF TECH
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202510588587.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing visual-inertial localization methods have poor robustness in dynamic environments, are easily disturbed by dynamic objects, and lack the modeling and utilization of high-level semantic information in images, resulting in feature matching errors and unstable state estimation.

Method used

A dynamic target detection mechanism is introduced to identify and eliminate dynamic target feature points through the target detection network, and semantic information consistency constraints are constructed on the back end to enhance the estimation stability of the system in areas with weak or repeated textures. A joint optimization framework is constructed to unify vision, IMU and semantic residuals.

Benefits of technology

Significantly reduce the mismatch rate of dynamic features, improve the positioning accuracy and robustness of the system in dynamic environments, and adapt to complex and changing real dynamic scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120740573A_ABST
    Figure CN120740573A_ABST
Patent Text Reader

Abstract

The invention discloses a visual inertial positioning method based on dynamic target detection and semantic information constraint, and belongs to the field of motion estimation and dynamic environment processing. According to the method, a dynamic target detection mechanism is introduced, the dynamic target is effectively detected based on the target detection network, inertial navigation information and geometric constraints, the dynamic target is effectively recognized in the image processing process, the corresponding dynamic feature points are screened out, the mismatching rate of the dynamic features is remarkably reduced, and high-quality observation input is provided for back-end optimization. In the back-end sliding window optimization stage, a semantic information consistency constraint method is constructed, and the estimation stability of the system in a weak texture area or a repeated texture area is enhanced by utilizing the consistency of feature points in the same semantic area on a geometric structure. Visual inertia pose estimation is realized based on dynamic target detection and semantic information constraint, and high-robustness and high-precision pose estimation can still be realized in a complex environment with dynamic interference of pedestrians, vehicles and the like and severe scene change.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of motion estimation and dynamic environment processing, and relates to a visual inertial positioning method based on dynamic target detection and semantic information constraint. Background Art

[0002] Visual-inertial positioning technology, which combines camera and inertial measurement unit (IMU) information, has been widely used in scenarios such as autonomous driving, drone navigation, and robotic mobility. By fusing visual information with acceleration and angular velocity data from an IMU, this technology enables continuous, real-time, and precise positioning of a device in three-dimensional space without relying on external positioning devices such as GPS or lidar.

[0003] Existing visual-inertial positioning methods can be roughly divided into two categories: one is the filtering method (such as MSCKF), which is mainly based on the extended Kalman filter. It achieves state estimation by linearizing the current state and observations, and has high computational efficiency; the other is the optimization method (such as VINS-Mono, VINS-Fusion, OKVIS), which jointly models multiple key frames and IMU measurements through a sliding window mechanism and uses a nonlinear optimization algorithm to iteratively solve the state, which performs better in terms of accuracy and robustness. In addition, there is a type of system based on direct methods (such as DSO and VI-DSO), which directly uses the grayscale values ​​of image pixels to optimize the photometric error. It is suitable for scenes with insufficient texture, but is more sensitive to changes in lighting and exposure.

[0004] While the aforementioned methods achieve good positioning accuracy and robustness in most static or semi-static environments, most existing technologies are still based on the "static environment assumption" and lack the ability to effectively identify and process dynamic interference in the scene, such as moving objects like pedestrians and vehicles. When large dynamic objects are present in an image, traditional methods have difficulty distinguishing the moving objects from the static background, resulting in feature matching errors, observation residual anomalies, and ultimately causing the system's state estimation to drift or even fail.

[0005] Furthermore, current mainstream methods generally rely on low-level visual features (such as corners and edges) for state estimation, lacking the modeling and utilization of high-level semantic information in images. This leads to insufficient system constraints and unstable estimation in scenes with repetitive or sparse textures, and also lacks robustness and versatility. Therefore, improving the system's ability to perceive interfering targets in dynamic environments and integrating semantic information to strengthen observation constraints have become key difficulties and important development directions in current visual-inertial positioning research.

[0006] The present invention aims to solve the problems of existing visual-inertial positioning methods, such as poor robustness in dynamic environments, susceptibility to interference from dynamic objects, and failure to consider semantic consistency in back-end optimization. This paper proposes a visual-inertial odometry method based on dynamic target detection and semantic information consistency constraints, which can achieve more accurate and stable navigation positioning in complex scenarios with large-scale dynamic objects. Summary of the Invention

[0007] In order to solve the problems of feature matching errors and state estimation drift caused by the presence of dynamic objects in the environment, the purpose of the present invention is to provide a visual-inertial pose estimation method with dynamic target detection and semantic information constraints, which can significantly reduce the dynamic feature mismatch rate; and utilize the consistency of the geometric structure of feature points in the same semantic area to enhance the estimation stability of the system in areas with weak texture or repeated texture. The present invention realizes visual-inertial pose estimation based on dynamic target detection and semantic information constraints. In complex environments with dynamic interference such as pedestrians and vehicles and drastic scene changes, it can still achieve highly robust and high-precision pose estimation, effectively improving the practicality and reliability of the visual-inertial navigation system.

[0008] The object of the present invention is achieved through the following technical solutions:

[0009] The present invention discloses a visual-inertial positioning method based on dynamic target detection and semantic information constraints. By introducing a dynamic target detection mechanism, it effectively detects dynamic targets based on a target detection network, inertial navigation information, and geometric constraints. During image processing, it effectively identifies dynamic targets and filters out corresponding dynamic feature points, significantly reducing the mismatch rate of dynamic features and providing high-quality observation input for back-end optimization. In the back-end sliding window optimization stage, a semantic information consistency constraint method is constructed, which utilizes the consistency of the geometric structure of feature points within the same semantic region to enhance the estimation stability of the system in areas with weak or repeated textures. Visual-inertial pose estimation is achieved based on dynamic target detection and semantic information constraints.

[0010] The visual inertial positioning method based on dynamic target detection and semantic information constraint disclosed in the present invention includes the following steps:

[0011] Step 1: Use the camera to capture a continuous image sequence; at the same time, obtain angular velocity and acceleration information through the IMU, which is called IMU information; align the image sequence and IMU information according to the timestamp; pre-integrate the IMU information between every two frames of images; and obtain the IMU pre-integration model;

[0012] Step 2: Target detection; prediction of detection box positions based on IMU pre-integration; dynamic echelon division based on DIoU; use the target detection network to identify and label the dynamic object categories in the image collected in step 1; remove feature points from the first echelon dynamic object area to reduce the interference of dynamic targets on pose estimation at the source;

[0013] Step 3: Based on the elimination of dynamic target features, the optical flow method is used for tracking. The second-tier potential dynamic targets are identified based on epipolar geometry constraints to determine the internal and external points, and the dynamic confidence level of the target is determined to determine the dynamic feature points.

[0014] Step 4: During the back-end nonlinear optimization process, based on the semantic information from the previous detection, consistency constraints are imposed on the feature points matched between frames. The semantic labels of the same feature points should be consistent, and a semantic consistency residual is constructed to enhance the global consistency of the system and mitigate error propagation caused by weak dynamic and pseudo-static targets.

[0015] Step 5: Construct a joint optimization problem, incorporate visual residuals, IMU residuals, and semantic consistency residuals into a unified optimization framework, and update the camera state variables in real time through a graph optimization method within a sliding window. The state variables include camera position and attitude.

[0016] Furthermore, the specific implementation method of step one is:

[0017] Use a monocular camera to capture continuous image sequences, and use the IMU to obtain high-frequency angular velocity and acceleration information; align data according to timestamps; obtain external parameters between the camera and IMU, including the rotation matrix and pan vector;

[0018] Inter-frame pre-integration: Given a deviation estimate, the relative changes in position, velocity, and posture between the i-th frame and the j-th frame are:

[0019]

[0020] Where, ΔR ij is the rotation matrix increment between the i-th frame and the j-th frame image, Δv ij is the velocity matrix increment between the i-th frame and the j-th frame image, Δp ij is the position transformation matrix between the i-th frame and the j-th frame image, ω I and are the gyroscope and accelerometer observation values ​​of the IMU, b g and b a is the deviation of the gyroscope and accelerometer, η g and η ais the white noise of the gyroscope and accelerometer, Δt is the time interval between the i-th and j-th image frames;

[0021] The IMU pre-integration measurement model is obtained by separating the noise and performing a first-order Taylor expansion, ignoring the near-zero high-order terms:

[0022]

[0023] in and are the pre-integrated measurement values ​​of the rotation, velocity change, and position change between the i-th and j-th frames, respectively. i is the rotation matrix from the IMU to the world coordinates of the i-th frame, ν i and p i are the velocity and position of the IMU in the i-th frame, δφ ij is the measurement noise, and g is the acceleration due to gravity.

[0024] Furthermore, the target detection method in step 2 is:

[0025] To effectively suppress the interference of dynamic objects on pose estimation, a lightweight target detection network YOLOv5s is introduced to quickly detect potential dynamic targets in the image. Each potential dynamic target is tracked and missed detections are compensated. The target state is dynamically judged based on the improved IoU method based on the IMU pre-integration results, and feature points in the first-tier dynamic target area are eliminated. The specific steps include the following:

[0026] Load and deploy the YOLOv5s network: Use the target detection network to identify common dynamic object categories in the image collected in step 1: Input the image in step 1 into the target detection network to obtain the target detection result in represents the bounding box coordinates of the k-th detection box in the i-th frame, For its corresponding category label; convert all detection results into pixel-level ROI areas;

[0027] Continuity is determined based on the target detection results, with two scenarios: 1) the target is detected in all consecutive image frames; 2) the target is successfully detected in the previous frame but fails in the next frame. For target detection failures, the previous frame's detection results are tracked to predict subsequent detection results and compensate for missed detections.

[0028] For the current detection frame I j and the previous detection frame I j-1 , defines the pixel velocity of the Kth dynamic object between frames:

[0029]

[0030] in denote the center pixel position of the K-th dynamic object in frame j and frame j-1 respectively;

[0031] Then the next detection frame I j+1 The weighted predicted velocity of the dynamic object for:

[0032]

[0033] If the object is not detected in the next frame, the bounding box of the object contains the pixel locations of the corner points The predicted detection bounding box in the new frame will be updated based on the weighted prediction speed The corresponding update formula is:

[0034]

[0035] After the missed detection compensation, the category label of the object remains unchanged, and its bounding box information and category information are placed in the target detection result collection to complete the missed detection compensation; when the missed detection time exceeds the threshold, the compensation of the dynamic object will be abandoned.

[0036] Furthermore, the method for predicting the position of the detection frame based on IMU pre-integration in step 2 is:

[0037] An object detection box in the i-th frame is B i , the image coordinates of the four corner points are expressed as At a known depth And the camera intrinsic parameter matrix K, combined with the camera-IMU calibration extrinsic parameter and IMU pre-integrated quantity (ΔR ij ,Δp ij ), directly propagate the entire detection frame forward to the j-th frame image plane, and the predicted position is recorded as The corner points of are given by the following uniform expressions:

[0038]

[0039] Where k = 1, 2, 3, 4 represents the four corner points of the box; π(·) is the camera perspective projection function, which maps a 3D point to a 2D image plane:

[0040]

[0041] Where p is the coordinates of the three points (x, y, z) and the four corner points Recombining to get the predicted bounding box

[0042] Furthermore, the method of dynamic echelon division based on DIoU in step 2 is:

[0043] Based on image detection results B t Get the position of the prediction box with IMU pre-integration information Get the detection box B obtained by the target detection network in the next frame image t+1 , and calculate B t+1 and DIoU indicator between:

[0044]

[0045] where b semantic and b IMU Represents the semantic detection box B of the new frame t+1 And obtain the predicted box based on IMU information The center point coordinates, ρ 2 (b semantic ,b IMU ) represents the Euclidean distance between the two center points, and C represents the diagonal length of the minimum circumscribed rectangle of the two image frames;

[0046] According to the DIoU value and the preset two-level thresholds τ1 and τ2, the targets are divided into the following three categories:

[0047]

[0048] Delete the dynamic feature points in the first echelon.

[0049] Furthermore, the specific implementation method of step three is:

[0050] ① Feature point tracking

[0051] Use Lucas-Kanade optical flow method to track the remaining feature points between frames; t and the t+1th frame image I t+1 , the pixel position of the feature point in the tth frame is (x, y), and in the t+1 frame it is (x+δx, y+δy), and the speed of the pixel point on the x-axis and y-axis is and The grayscale gradient is and The time derivative of pixel grayscale is Satisfy the basic constraints of optical flow:

[0052]

[0053] Solve equation (10) to obtain the optical flow velocities u and v of the feature points, and then obtain the position p2 of the feature point p1 on the t-th frame image in the t+1 frame image, and obtain the well-matched feature point pair (p1, p2) between the two frames;

[0054] ②Calculate the basic matrix

[0055] Based on the known inter-frame camera pose transformation rotation matrix R and translation vector t, the basic matrix F between the two frames is obtained:

[0056] F=K -T ·[t] × ·R·K -1 (11)

[0057] Where K is the camera intrinsic parameter matrix, [t] × is the antisymmetric matrix of the translation vector;

[0058] ③ Judgment of internal and external points based on epipolar distance

[0059] Based on the feature point pair (p1, p2) obtained in ①, we know that the point p1 in the t-th frame and its corresponding epipolar line l2 = Fp1. Then, in the t+1-th frame, based on the matching point p2, we can get the point-to-episode distance:

[0060]

[0061] If the epipolar distance d>τ, p2 is judged as an outlier, and τ is the set threshold;

[0062] ④ Target dynamic state determination

[0063] All feature points in the target detection frame are subjected to the above distance calculation and inside-outside point judgment, and the feature point set in the detection frame is recorded as P = {p1, p2, ..., p n}, where the external point set is P outlier , based on the proportion of external points, define the confidence C of the target as a dynamic target dynamic for:

[0064]

[0065] When C dynamic When ∈ is greater than the threshold, the detected target is considered to be a dynamic target; that is:

[0066]

[0067] For all dynamic targets in this stage, their corresponding feature points are removed; all the remaining feature points on the frame image are high-quality feature points u i .

[0068] Furthermore, the specific implementation method of step four is:

[0069] Select reference frame i and target frame j, and set the high-quality feature point u in reference frame i i The depth value is d i , the corresponding semantic label is in is a set of static semantic labels defined; according to the camera pose transformation (R, t) between the reference frame and the target frame, the high-quality feature points u in the reference frame are transformed into i Project it into the target frame to get the predicted position The specific calculation is as follows:

[0070]

[0071] Where, π(·) is the projection function of the camera model;

[0072] Based on the predicted location and its semantic label l, in the local neighborhood of the image Search for semantic features of the same category within the same category to obtain matching points for:

[0073]

[0074] according to and The semantic consistency residual term is obtained as:

[0075]

[0076] Furthermore, the specific implementation method of step five is:

[0077] Combine the semantic consistency residual e from step 4 sem Construct a joint optimization function with marginalization residual, IMU measurement residual, and visual reprojection residual:

[0078]

[0079] where {r p ,H p} represents the prior information passed down by the marginalization process, and denote the IMU measurement residual and visual reprojection residual, ρ(·) is the robust kernel function, and λ sem is the weight coefficient of the semantic error term, which is used to adjust its influence in the overall optimization; the set Represents all IMU measurements in the sliding window, the set Represents the set of feature points that have been observed at least twice in the current sliding window, set Ns All semantic feature points in the current sliding window; is the current sliding window, expressed as:

[0080]

[0081] in Indicates the IMU state at the time of image acquisition of the kth frame. The IMU state includes the position of the IMU in the world coordinate system. speed and posture And the acceleration bias b in the body coordinate system a and gyroscope bias b g ; Represents the parameters from IMU to camera, including the position of the camera in the IMU coordinate system and the rotation of the camera relative to the IMU n represents the total number of key frames, m represents the total number of feature points in the sliding window, and λ i represents the inverse depth of the i-th feature point since the first observation, l i Represents the semantic label category of the i-th feature point since the first observation.

[0082] Furthermore, in step three, ∈ is adjusted according to the characteristics of the dataset and set between 0.3 and 0.6.

[0083] Beneficial effects:

[0084] 1. The visual-inertial positioning method based on dynamic target detection and semantic information constraints disclosed in this invention constructs a robust visual-inertial positioning method for dynamic environments by integrating vision, IMU and semantic information. It introduces target position prediction and dynamic echelon division strategies based on IMU pre-integration at the front end to effectively identify and eliminate unstable feature points on dynamic objects, thereby improving the stability of front-end tracking.

[0085] 2. The visual-inertial positioning method based on dynamic target detection and semantic information constraints disclosed in the present invention introduces semantic consistency constraints in the back-end optimization process. By ensuring the semantic label consistency of matching features, it suppresses the error propagation caused by weak dynamic and pseudo-static targets, and enhances the constraint force of the global optimization of the state estimation process.

[0086] 3. The visual-inertial positioning method based on dynamic target detection and semantic information constraints disclosed in the present invention integrates vision, inertia and semantic residuals into a joint optimization framework, effectively improving the accuracy and robustness of state estimation and adapting to more complex and changeable real dynamic scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] Figure 1This is a flow chart of the visual-inertial positioning method based on dynamic target detection and semantic information constraints disclosed in the present invention;

[0088] Figure 2 Schematic diagram of sensor data alignment and IMU pre-integration;

[0089] Figure 3 DIoU diagram;

[0090] Figure 4 It is a dynamic target detection function module;

[0091] Figure 5 is the epipolar geometry diagram of the dynamic target;

[0092] Figure 6 This is the framework diagram of the visual inertial positioning system;

[0093] Figure 7 To compare the method effect trajectories;

[0094] Figure 8 APE comparison of method effects; DETAILED DESCRIPTION

[0095] The present invention is described in detail below with reference to the accompanying drawings and embodiments.

[0096] like Figure 1 As shown, the visual inertial positioning method based on dynamic target detection and semantic information constraint disclosed in this embodiment is specifically implemented as follows:

[0097] Step 1: Use a monocular camera to capture a continuous image sequence, and use the IMU to obtain high-frequency angular velocity and acceleration information; obtain the external parameters between the camera and IMU, including the rotation matrix and pan vector; data alignment is performed according to timestamps, such as Figure 2 As shown, IMU pre-integration is performed between adjacent frames of the aligned camera: given the deviation estimate, the relative changes in position, velocity, and attitude between the i-th frame and the j-th frame image are:

[0098]

[0099] Where, ΔR ij is the rotation matrix increment between the i-th frame and the j-th frame image, Δv ij is the velocity matrix increment between the i-th frame and the j-th frame image, Δp ij is the position transformation matrix between the i-th frame and the j-th frame image, ω I and are the gyroscope and accelerometer observation values ​​of the IMU, b g and b ais the deviation of the gyroscope and accelerometer, η g and η a is the white noise of the gyroscope and accelerometer, Δt is the time interval between the i-th and j-th image frames;

[0100] By separating the noise and performing a first-order Taylor expansion and ignoring the near-zero high-order terms, the IMU pre-integration measurement model can be obtained as follows:

[0101]

[0102] in and are the pre-integrated measurement values ​​of the rotation, velocity change, and position change between the i-th and j-th frames, respectively. i is the rotation matrix from the IMU to the world coordinates of the i-th frame, ν i and p i are the velocity and position of the IMU in the i-th frame, δφ ij is the measurement noise, and g is the acceleration due to gravity.

[0103] Step 2: To effectively suppress the interference of dynamic objects on pose estimation, a lightweight target detection network YOLOv5s is introduced to quickly detect potential dynamic targets in the image. Each potential dynamic target is tracked and compensated for missed detection. The target state is dynamically judged based on the improved IoU method based on the IMU pre-integration results, and feature points in the first-tier dynamic target area are eliminated. The specific steps include the following:

[0104] 1. Load and deploy the YOLOv5s network

[0105] Use a pre-trained YOLOv5s model or a model trained through transfer learning based on an independent dataset; load and deploy the YOLOv5s network: use the YOLOv5 target detection network to identify common dynamic object categories in the image collected in step 1, including "person", "car", "bicycle", "motorbike", and "bus": input the image from step 1 into the target detection network to obtain target detection results in represents the bounding box coordinates of the k-th detection box in the i-th frame, is its corresponding category label; all detection results are converted into pixel-level ROI areas.

[0106] 2. Missed detection compensation

[0107] Continuity is determined based on the target detection results, with two scenarios: 1) the target is detected in all consecutive image frames; 2) the target is successfully detected in the previous frame but fails in the next frame. For target detection failures, the previous frame's detection results are tracked to predict subsequent detection results and compensate for missed detections.

[0108] For the current detection frame I j and the previous detection frame I j-1 , defines the pixel velocity of the Kth dynamic object between frames:

[0109]

[0110] in Represent the center pixel position of the Kth dynamic object in frame j and frame j-1 respectively; then the next detection frame I j+1 The weighted predicted velocity of the dynamic object for:

[0111]

[0112] If the object is not detected in the next frame, the bounding box of the object contains the pixel locations of the corner points The predicted detection bounding box in the new frame will be updated based on the weighted prediction speed The corresponding update formula is:

[0113]

[0114] After the missed detection compensation, the category label of the object remains unchanged, and its bounding box information and category information are placed in the target detection result collection to complete the missed detection compensation; when the missed detection time exceeds the threshold, the compensation of the dynamic object will be abandoned.

[0115] 3. Detection frame position prediction based on IMU pre-integration

[0116] An object detection box in the i-th frame is B i , the image coordinates of the four corner points are expressed as At a known depth And the camera intrinsic parameter matrix K, combined with the camera-IMU calibration extrinsic parameter and IMU pre-integrated quantity (ΔR ij ,Δp ij ), directly propagate the entire detection frame forward to the j-th frame image plane, and the predicted position is recorded as The corner points of are given by the following uniform expressions:

[0117]

[0118] Where k = 1, 2, 3, 4 represents the four corner points of the box; π(·) is the camera perspective projection function, which maps a 3D point to a 2D image plane:

[0119]

[0120] Where p is the coordinates of the three points (x, y, z) and the four corner points Recombining to get the predicted bounding box

[0121] 4. Dynamic echelon division

[0122] Based on image detection results B t Get the position of the prediction box with IMU pre-integration information Get the detection box B obtained by the target detection network in the next frame image t+1 , and calculate B t+1 and DIoU indicators between Figure 3 shown):

[0123]

[0124] where b semantic and b IMU Represents the semantic detection box B of the new frame t+1 And obtain the predicted box based on IMU information The center point coordinates, ρ 2 (b semantic ,b IMU ) represents the Euclidean distance between the two center points, and C represents the diagonal length of the minimum circumscribed rectangle of the two image frames;

[0125] According to the DIoU value and the preset two-level thresholds τ1 and τ2, the targets are divided into the following three categories:

[0126]

[0127] Delete the dynamic feature points in the first echelon;

[0128] Step 3: Step 3 and step 2 work together to realize the dynamic target detection function. The algorithm flow is as follows: Figure 4 As shown in the figure, target detection and missed detection compensation are performed through camera images to obtain potential dynamic targets; IMU pre-integration is combined with potential dynamic targets to predict target detection frames based on IMU information; dynamic targets are initially screened based on DIoU, and feature points are extracted and tracked in the image after dynamic targets are screened out; the dynamic and static distribution of these feature points is judged based on epipolar geometry constraints, thereby further confirming dynamic targets. The specific implementation method of step three is,

[0129] 1. Feature point tracking

[0130] Use Lucas-Kanade optical flow method to track the remaining feature points between frames; t and the t+1th frame image I t+1 , the pixel position of the feature point in the tth frame is (x, y), and in the t+1 frame it is (x+δx, y+δy), and the speed of the pixel point on the x-axis and y-axis is and The grayscale gradient is and The time derivative of pixel grayscale is Satisfy the basic constraints of optical flow:

[0131]

[0132] Solve equation (10) to obtain the optical flow velocities u and v of the feature points, and then obtain the position p2 of the feature point p1 on the t-th frame image in the t+1 frame image, and obtain the well-matched feature point pair (p1, p2) between the two frames;

[0133] 2. Calculate the basic matrix

[0134] Based on the known inter-frame camera pose transformation rotation matrix R and translation vector t, the basic matrix F between the two frames is obtained:

[0135] F=K -T ·[t] × ·R·K -1 (30)

[0136] Where K is the camera intrinsic parameter matrix, [t] × is the antisymmetric matrix of the translation vector;

[0137] 3. Judgment of internal and external points based on epipolar distance

[0138] In dynamic cases, epipolar geometry transformation is as follows Figure 5 As shown in Figure 1, when the dynamic target changes position, the feature corresponding projection point no longer falls on the epipolar line but falls outside the epipolar line. Based on the feature point pair (p1, p2) obtained in step 1, it can be seen that the point p1 in the tth frame and its corresponding epipolar line l2 = Fp1, then in the t+1th frame based on the matching point p2, the point-to-episode distance can be obtained:

[0139]

[0140] If the epipolar distance d>τ, p2 is judged as an outlier, and τ is the set threshold;

[0141] 4. Target dynamic state determination

[0142] All feature points in the target detection frame are subjected to the above distance calculation and inside-outside point judgment, and the feature point set in the detection frame is recorded as P = {p1, p2, ..., p n}, where the external point set is P outlier , based on the proportion of external points, define the confidence C of the target as a dynamic target dynamic for:

[0143]

[0144] When C dynamic When ∈ is greater than the threshold, the detected target is considered to be a dynamic target; that is:

[0145]

[0146] Where ∈ can be adjusted according to the characteristics of the data set and is generally set between 0.3 and 0.6. For all dynamic targets in this stage, their corresponding feature points are removed; the remaining feature points on the frame image are high-quality feature points u i .

[0147] Step 4: Select reference frame i and target frame j, and set the high-quality feature point u in reference frame i i The depth value is d i , the corresponding semantic label is in is a set of static semantic labels defined; according to the camera pose transformation (R, t) between the reference frame and the target frame, the high-quality feature points u in the reference frame are transformed into i Project it into the target frame to get the predicted position The specific calculation is as follows:

[0148]

[0149] Where, π(·) is the projection function of the camera model;

[0150] Based on the predicted location and its semantic label l, in the local neighborhood of the image Search for semantic features of the same category within the same category to obtain matching points for:

[0151]

[0152] according to and The semantic consistency residual term is obtained as:

[0153]

[0154] Step 5: Combine the semantic consistency residual e from step 4 semConstruct a joint optimization function with marginalization residual, IMU measurement residual, and visual reprojection residual:

[0155]

[0156] where {r p ,H p} represents the prior information passed down by the marginalization process, and denote the IMU measurement residual and visual reprojection residual, ρ(·) is the robust kernel function, and λ sem is the weight coefficient of the semantic error term, which is used to adjust its influence in the overall optimization; the set Represents all IMU measurements in the sliding window, the set Represents the set of feature points that have been observed at least twice in the current sliding window, set N s All semantic feature points in the current sliding window; is the current sliding window, expressed as:

[0157]

[0158] in Indicates the IMU state at the time of image acquisition of the kth frame. The IMU state includes the position of the IMU in the world coordinate system. speed and posture And the acceleration bias b in the body coordinate system a and gyroscope bias b g ; Represents the parameters from IMU to camera, including the position of the camera in the IMU coordinate system and the rotation of the camera relative to the IMU n represents the total number of key frames, m represents the total number of feature points in the sliding window, and λ i represents the inverse depth of the i-th feature point since the first observation, l i Represents the semantic label category of the i-th feature point since the first observation. Figure 6 As shown in the figure, this method is built on the VINS-Mono architecture. After the system is initialized, the front-end inputs high-quality feature points to the VIO back-end for optimization. The BA strategy is used to obtain the maximum a posteriori estimate by minimizing the sum of the prior term and the Mahalanobis norm of all measurement residuals, and the pose information of the camera and IMU is solved. The effect of the method of the present invention is compared with other mainstream methods in a dynamic environment. The trajectory comparison diagram of the pose obtained by data solution is shown in the figure below. Figure 7As shown in the figure, the results show that the proposed method effectively improves the accuracy of visual inertial positioning and is more consistent with the ground-truth. The specific absolute position error (APE) data in the method comparison is as follows: Figure 8 As shown, the method of the present invention significantly improves the robustness by more than half under the same conditions, verifying the effectiveness of the present invention.

[0159] In response to the problems of poor robustness of existing visual-inertial positioning algorithms in dynamic environments, serious interference from dynamic targets, and insufficient use of semantic information in back-end optimization, the present invention proposes a visual-inertial positioning method based on dynamic target detection and semantic consistency constraints. This method fully combines the front-end dynamic target detection and feature elimination mechanism, inter-frame semantic information constraints, and the back-end sliding window joint optimization strategy, effectively improving the positioning accuracy and robustness of the system in complex dynamic scenes. By eliminating interfering feature points in target areas of different dynamic levels, mismatching and error propagation are suppressed; the introduction of semantic consistency residuals enhances the matching reliability of feature points and the stability of global optimization; the tightly coupled optimization framework finally constructed can achieve efficient fusion of multi-source information and real-time status updates. This embodiment realizes visual-inertial pose estimation based on dynamic target detection and semantic information constraints. In complex environments with dynamic interference such as pedestrians and vehicles and drastic scene changes, it can still achieve highly robust and high-precision pose estimation, effectively improving the practicality and reliability of the visual-inertial navigation system.

[0160] In summary, the above are only preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A visual-inertial positioning method based on dynamic target detection and semantic information constraints, characterized by: The visual-inertial positioning method based on dynamic target detection and semantic information constraints introduces a dynamic target detection mechanism. It effectively detects dynamic targets based on the target detection network, inertial navigation information, and geometric constraints. During the image processing process, it effectively identifies dynamic targets and filters out corresponding dynamic feature points, significantly reducing the dynamic feature mismatch rate and providing high-quality observation input for back-end optimization. In the back-end sliding window optimization stage, a semantic information consistency constraint method is constructed to utilize the geometric consistency of feature points within the same semantic region to enhance the estimation stability of the system in areas with weak or repeated textures. Visual-inertial pose estimation is achieved based on dynamic target detection and semantic information constraints.

2. The method according to claim 1, wherein The steps include: Step 1: Use the camera to capture a continuous image sequence; at the same time, use the IMU to obtain angular velocity and acceleration information, which is called IMU information; align the image sequence and IMU information according to the timestamp; Pre-integrate the IMU information between every two frames of images to obtain an IMU pre-integration model; Step 2: Target detection; prediction of detection box positions based on IMU pre-integration; dynamic echelon division based on DIoU; use the target detection network to identify and label the dynamic object categories in the image collected in step 1; remove feature points from the first echelon dynamic object area to reduce the interference of dynamic targets on pose estimation at the source; Step 3: Based on the elimination of dynamic target features, the optical flow method is used for tracking. The second-tier potential dynamic targets are identified based on epipolar geometry constraints to determine the internal and external points, and the dynamic confidence level of the target is determined to determine the dynamic feature points. Step 4: During the back-end nonlinear optimization process, based on the semantic information from the previous detection, consistency constraints are imposed on the feature points matched between frames. The semantic labels of the same feature points should be consistent, and a semantic consistency residual is constructed to enhance the global consistency of the system and mitigate error propagation caused by weak dynamic and pseudo-static targets. Step 5: Construct a joint optimization problem, incorporate visual residuals, IMU residuals, and semantic consistency residuals into a unified optimization framework, and update the camera state variables in real time through a graph optimization method within a sliding window. The state variables include camera position and attitude.

3. The method according to claim 2, wherein: The specific implementation method of step one is: Use a monocular camera to capture continuous image sequences, and use the IMU to obtain high-frequency angular velocity and acceleration information; align data according to timestamps; obtain external parameters between the camera and IMU, including the rotation matrix and pan vector; Inter-frame pre-integration: Given a deviation estimate, the relative changes in position, velocity, and posture between the i-th frame and the j-th frame are: Where, ΔR ij is the rotation matrix increment between the i-th frame and the j-th frame image, Δv ij is the velocity matrix increment between the i-th frame and the j-th frame image, Δp ij is the position transformation matrix between the i-th frame and the j-th frame image, ω I and are the gyroscope and accelerometer observation values ​​of the IMU, b g and b a is the deviation of the gyroscope and accelerometer, η g and η a is the white noise of the gyroscope and accelerometer, Δt is the time interval between the i-th and j-th image frames; The IMU pre-integration measurement model is obtained by separating the noise and performing a first-order Taylor expansion, ignoring the near-zero high-order terms: in and are the pre-integrated measurement values ​​of the rotation, velocity change, and position change between the i-th and j-th frames, respectively. i is the rotation matrix from the IMU to the world coordinates of the i-th frame, ν i and p i are the velocity and position of the IMU in the i-th frame, δφ ij is the measurement noise, and g is the acceleration due to gravity.

4. The method according to claim 3, wherein: The target detection method in step 2 is: To effectively suppress the interference of dynamic objects on pose estimation, a lightweight target detection network YOLOv5s is introduced to quickly detect potential dynamic targets in the image. Each potential dynamic target is tracked and missed detections are compensated. The target state is dynamically judged based on the improved IoU method based on the IMU pre-integration results, and feature points in the first-tier dynamic target area are eliminated. The specific steps include the following: Load and deploy the YOLOv5s network: Use the target detection network to identify common dynamic object categories in the image collected in step 1: Input the image in step 1 into the target detection network to obtain the target detection result in represents the bounding box coordinates of the k-th detection box in the i-th frame, For its corresponding category label; convert all detection results into pixel-level ROI areas; Continuity is determined based on the target detection results, with two scenarios: 1) the target is detected in all consecutive image frames; 2) the target is successfully detected in the previous frame but fails in the next frame. For target detection failures, the previous frame's detection results are tracked to predict subsequent detection results and compensate for missed detections. For the current detection frame I j and the previous detection frame I j-1 , defines the pixel velocity of the Kth dynamic object between frames: in denote the center pixel position of the K-th dynamic object in frame j and frame j-1 respectively; Then the next detection frame I j+1 The weighted predicted velocity of the dynamic object for: If the object is not detected in the next frame, the bounding box of the object contains the pixel locations of the corner points The predicted detection bounding box in the new frame will be updated based on the weighted prediction speed The corresponding update formula is: After the missed detection compensation, the category label of the object remains unchanged, and its bounding box information and category information are placed in the target detection result collection to complete the missed detection compensation; when the missed detection time exceeds the threshold, the compensation of the dynamic object will be abandoned.

5. The method according to claim 4, wherein: The method for predicting the detection frame position based on IMU pre-integration in step 2 is: An object detection box in the i-th frame is B i , the image coordinates of the four corner points are expressed as At a known depth And the camera intrinsic parameter matrix K, combined with the camera-IMU calibration extrinsic parameter and IMU pre-integrated quantity (ΔR ij ,Δp ij ), directly propagate the entire detection frame forward to the j-th frame image plane, and the predicted position is recorded as The corner points of are given by the following uniform expressions: Where k = 1, 2, 3, 4 represents the four corner points of the box; π(·) is the camera perspective projection function, which maps a 3D point to a 2D image plane: Where p is the coordinates of the three points (x, y, z) and the four corner points Recombining to get the predicted bounding box 6. The method according to claim 5, wherein: The method of dynamic echelon division based on DIoU in step 2 is: Based on image detection results B t Get the position of the prediction box with IMU pre-integration information Get the detection box B obtained by the target detection network in the next frame image t+1 , and calculate B t+1 and DIoU indicator between: where b semantic and b IMU Represents the semantic detection box B of the new frame t+1 And obtain the predicted box based on IMU information The center point coordinates, ρ 2 (b semantic ,b IMU ) represents the Euclidean distance between the two center points, and C represents the diagonal length of the minimum circumscribed rectangle of the two image frames; According to the DIoU value and the preset two-level thresholds τ1 and τ2, the targets are divided into the following three categories: Delete the dynamic feature points in the first echelon.

7. The method according to claim 6, wherein: The specific implementation method of step three is: ① Feature point tracking Use Lucas-Kanade optical flow method to track the remaining feature points between frames; t and the t+1th frame image I t+1 , the pixel position of the feature point in the tth frame is (x, y), and in the t+1 frame it is (x+δx, y+δy), and the speed of the pixel point on the x-axis and y-axis is and The grayscale gradient is and The time derivative of pixel grayscale is Satisfy the basic constraints of optical flow: Solve equation (10) to obtain the optical flow velocities u and v of the feature points, and then obtain the position p2 of the feature point p1 on the t-th frame image in the t+1 frame image, and obtain the well-matched feature point pair (p1, p2) between the two frames; ②Calculate the basic matrix Based on the known inter-frame camera pose transformation rotation matrix R and translation vector t, the basic matrix F between the two frames is obtained: F=K -T ·[t] × ·R·K -1 (11) Where K is the camera intrinsic parameter matrix, [t] × is the antisymmetric matrix of the translation vector; ③ Judgment of internal and external points based on epipolar distance Based on the feature point pair (p1, p2) obtained in ①, we know that the point p1 in the t-th frame and its corresponding epipolar line l2 = Fp1. Then, in the t+1-th frame, based on the matching point p2, we can get the point-to-episode distance: If the epipolar distance d>τ, p2 is judged as an outlier, and τ is the set threshold; ④ Target dynamic state determination All feature points in the target detection frame are subjected to the above distance calculation and inside-outside point judgment, and the feature point set in the detection frame is recorded as P = {p1, p2, ..., p n }, where the external point set is P outlier , based on the proportion of external points, define the confidence C of the target as a dynamic target dynamic for: When C dynamic When ∈ is greater than the threshold, the detected target is considered to be a dynamic target; that is: For all dynamic targets in this stage, their corresponding feature points are eliminated; All the remaining feature points on the frame image that have not been eliminated are high-quality feature points u i .

8. The method according to claim 7, wherein: The specific implementation method of step four is: Select reference frame i and target frame j, and set the high-quality feature point u in reference frame i i The depth value is d i , the corresponding semantic label is in is a set of static semantic labels defined; according to the camera pose transformation (R, t) between the reference frame and the target frame, the high-quality feature points u in the reference frame are transformed into i Project it into the target frame to get the predicted position The specific calculation is as follows: Where, π(·) is the projection function of the camera model; Based on the predicted location and its semantic label l, in the local neighborhood of the image Search for semantic features of the same category within the same category to obtain matching points for: according to and The semantic consistency residual term is obtained as:

9. The method according to claim 8, wherein: The specific implementation method of step five is: Combine the semantic consistency residual e from step 4 sem Construct a joint optimization function with marginalization residual, IMU measurement residual, and visual reprojection residual: where {r p ,H p } represents the prior information passed down by the marginalization process, and denote the IMU measurement residual and visual reprojection residual, ρ(·) is the robust kernel function, and λ sem is the weight coefficient of the semantic error term, which is used to adjust its influence in the overall optimization; the set Represents all IMU measurements in the sliding window, the set Represents the set of feature points that have been observed at least twice in the current sliding window, set N s All semantic feature points in the current sliding window; is the current sliding window, expressed as: in Indicates the IMU state at the time of image acquisition of the kth frame. The IMU state includes the position of the IMU in the world coordinate system. speed and posture And the acceleration bias b in the body coordinate system a and gyroscope bias b g ; Represents the parameters from IMU to camera, including the position of the camera in the IMU coordinate system and the rotation of the camera relative to the IMU n represents the total number of key frames, m represents the total number of feature points in the sliding window, and λ i represents the inverse depth of the i-th feature point since the first observation, l i Represents the semantic label category of the i-th feature point since the first observation.

10. The method according to claim 9, wherein: In step 3, ∈ is adjusted according to the characteristics of the dataset and is set between 0.3 and 0.6.

Citation Information

Cited By

  • Dynamic scene three-dimensional reconstruction method and device based on hydrogen energy unmanned aerial vehicle survey

    CN121120947A

  • A method and equipment for dynamic scene 3D reconstruction based on hydrogen-powered UAV reconnaissance

    CN121120947B

  • Unmanned aerial vehicle monocular vision inertial positioning method based on under-forest geometric representation

    CN121297824A

  • Robot positioning method and device

    CN121482155A

  • Dynamic environment rapid positioning method and system based on visual odometer

    CN121564680A