Multi-target tracking information enhanced laser radar visual inertia high-precision positioning method

By integrating multimodal detection of bounding boxes and geometric features from LiDAR and vision, and combining IMU pre-integration measurements, the carrier pose and dynamic target trajectory are optimized, solving the problem of performance degradation of multi-target tracking and positioning in dynamic scenes by a single sensor, and achieving high-precision multi-target tracking and autonomous positioning.

CN121069416APending Publication Date: 2025-12-05WUHAN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511614386.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing single-sensor SLAM systems suffer from problems such as target occlusion, lighting effects, and insufficient fusion of multimodal observation data when handling multi-target tracking and localization in dynamic scenes, leading to a decline in localization performance.

Method used

A high-precision LiDAR-visual-inertial positioning method with multi-target tracking information enhancement is designed. By fusing multimodal detection bounding boxes and geometric features from LiDAR and vision, combined with IMU pre-integration measurements, the carrier pose and dynamic target trajectory are optimized, and a factor graph optimization framework is used for joint optimization.

Benefits of technology

It improves the adaptability, robustness and accuracy of multi-target tracking and SLAM localization systems in complex dynamic scenarios, and achieves robust scene tracking and high-precision autonomous localization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121069416A_ABST
    Figure CN121069416A_ABST
Patent Text Reader

Abstract

The invention provides a multi-target tracking information enhanced laser radar visual inertia high-precision positioning method, and belongs to the technical field of environmental perception and navigation positioning. The positioning method comprises the following steps: acquiring measurement data of a laser radar, a camera and an IMU (Inertial Measurement Unit), and performing data preprocessing, including laser radar point cloud data preprocessing, image data preprocessing and IMU pre-integration; based on the preprocessed data, establishing a target comprehensive association scale fusing a laser radar and a visual multi-modal detection bounding box and geometric features, and associating multi-modal target data; the continuous and reliable multi-target tracking and self-localization performance is realized by combining the multi-modal detection bounding box and the geometrical characteristics of the target and the IMU pre-integration measurement optimization carrier pose and dynamic target trajectory.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of environment perception and navigation positioning, and particularly relates to a laser radar vision inertial high-precision positioning method with multi-target tracking information enhancement. BACKGROUND

[0002] With the continuous development of emerging intelligent technologies such as autonomous driving, robots, virtual reality (VR), augmented reality (AR), etc., the industrial applications present the characteristics of diversified forms and diversified scenes. In the face of increasingly complex intelligent carrier application scenarios, accurate and reliable environment perception and high-precision positioning services have become an urgent need for intelligent industrial applications.

[0003] SLAM (Simultaneous Localization and Mapping) technology has become the mainstream positioning technology in weak / GNSS signal scenarios, and can provide continuous, reliable and high-precision autonomous positioning services for intelligent unmanned systems. In the past two decades, SLAM technology based on a single sensor (such as laser radar or vision) has achieved real-time positioning and mapping in unknown environments and has made remarkable achievements. Laser radar-based SLAM technology can accurately capture the geometric structure information of the environment, but usually performs poorly in unstructured environments (such as long corridors or flat wide areas). On the contrary, vision-based SLAM technology does not rely on structural features such as edges and planes, and performs well in textured environments, but its positioning performance will be severely affected in weak texture or poor lighting conditions. Therefore, laser radar-based or vision-based SLAM systems are usually combined with inertial measurement units (IMUs) to improve the accuracy and robustness of the system positioning in complex environments. Specifically, laser inertial SLAM systems use high-frequency IMU measurements to eliminate the motion distortion of point clouds and correct the inter-frame matching errors of laser radar point clouds. While vision inertial SLAM systems recover scale information by using IMU observations to assist the system to achieve high-precision vision inertial positioning in unknown environments.

[0004] To further improve the positioning performance of intelligent unmanned systems in complex dynamic urban environments, laser radar / visual / inertial multi-source fusion SLAM technology has attracted more and more attention due to its stronger robustness in sensor degradation scenarios. The loosely coupled laser radar / visual / inertial SLAM system usually performs laser odometry and visual odometry independently and shares the positioning results, but cannot realize data fusion at the level of original observations, thereby limiting the positioning performance of the system. The tightly coupled laser radar / visual / inertial SLAM system can fully utilize the observation information of multi-source sensors, realize the optimal estimation of system state, and effectively improve the accuracy and robustness of the system positioning in complex urban environments. However, the SLAM system usually follows the assumption of static environment, and it is difficult to handle a large number of moving targets in dynamic scenes, which leads to the degradation of SLAM positioning performance and cannot meet the high-precision positioning requirements of intelligent unmanned systems.

[0005] On the other hand, 3D multi-target tracking is a key component of the environment perception module. Accurate and reliable target tracking can provide key information such as the position, direction and speed of the surrounding moving targets, providing important basic data for the safe navigation and motion planning of the carrier. Early multi-target tracking systems are mainly designed based on visual sensors and rely on rich visual texture information to achieve stable tracking in 2D space. However, 2D visual multi-target tracking systems are easily affected by target occlusion, environmental light, etc., and lack 3D perception of the surrounding environment, limiting their wide application in the field of autonomous driving, etc. In recent years, multi-target tracking methods based on lidar have attracted much attention due to their ability to accurately perceive the geometric structure of the environment through dense 3D point clouds. However, due to the lack of target texture information in the point cloud, especially in the scene of irregular target motion and fast motion, its tracking performance faces major challenges. In order to overcome the limitations of single sensor multi-target tracking systems, researchers are working on 3D multi-target tracking technology based on lidar / vision fusion to fully integrate the advantages of the two sensing technologies and improve the accuracy and robustness of multi-target tracking in complex dynamic scenes. Current mainstream lidar / vision fusion multi-target tracking systems mainly adopt the detection-based tracking paradigm, that is, first fuse the target detection information of different sensors, and then use the association matrix to associate the current target detection bounding box with the historical trajectory. Although these systems successfully utilize the complementary characteristics of lidar and visual sensors and exhibit impressive tracking performance, the current research methods still have some obvious limitations. In the data association stage, most systems use the intersection over union (IoU) metric of the target detection bounding box to construct the association matrix to achieve accurate 'detection-tracking' association. However, the IoU index only focuses on the overlap of the target detection bounding box, ignoring other key information of the target (such as the appearance of the target), which may cause mis-matching when the target moves irregularly. In the target trajectory updating stage, most methods use Kalman filters to independently update 2D and 3D target trajectories. Although such methods can improve the real-time performance of the system, the independent updating of multi-modal trajectories ignores the correlation between multi-modal observation data, failing to fully utilize the complementary characteristics of multi-source sensors. In addition, current multi-target tracking systems usually perform target data association based on high-precision carrier pose, however, in highly dynamic scenes, moving target objects dominate the sensor observation field of view, leading to false association of moving target information and static scenes, which degrades the autonomous positioning performance of the SLAM system, and further affects the multi-target tracking performance.

[0006] To overcome the technical difficulties of the above research, the present application proposes a multi-target tracking information enhanced lidar-vision-inertial high-precision positioning method. SUMMARY

[0007] The present application aims to:

[0008] 1. Design a multi-target tracking information enhanced laser radar vision inertial high-precision positioning algorithm. The algorithm effectively fuses the multi-modal detection bounding box, geometric feature information and IMU pre-integral measurement of the target, and jointly optimizes the carrier pose and dynamic target trajectory, realizes continuous and reliable multi-target tracking and autonomous positioning performance.

[0009] 2. A target data association index is proposed, which effectively fuses the multi-modal detection bounding box and geometric feature information of the target, and significantly improves the multi-modal target data association performance in complex dynamic scenes.

[0010] 3. A multi-modal target trajectory joint optimization method is proposed, which is based on a unified factor graph optimization framework, and realizes the joint optimization of 2D trajectory and 3D trajectory.

[0011] In order to achieve the above purpose, the specific technical scheme of the present application is as follows:

[0012] In one aspect of the present application, a multi-target tracking information enhanced laser radar vision inertial (LiDAR / Vision / INS) high-precision positioning method is provided, wherein the LiDAR is a laser radar, the Vision is a vision, and the INS is an inertial navigation. The method comprises the following steps:

[0013] Step 1: Obtain the measurement data of laser radar, camera and IMU, and perform data preprocessing, which mainly includes laser radar point cloud data preprocessing, image data preprocessing and IMU pre-integration.

[0014] Step 2: In view of the inherent limitation of single sensor in target data association, based on the preprocessed data, a target comprehensive association scale is designed to fuse the multi-modal detection bounding box and geometric feature of laser radar and vision, and the multi-modal target data association is realized.

[0015] Step 3: Jointly optimize the carrier pose and dynamic target trajectory based on the multi-modal detection bounding box and geometric feature of the target and the IMU pre-integral measurement.

[0016] In order to overcome the inherent limitation of single sensor in target data association, the multi-modal detection bounding box and geometric feature of laser radar and vision are fully utilized to realize stable scene tracking. Then, by effectively fusing the multi-modal target detection bounding box and geometric feature information and the IMU pre-integral measurement, the carrier pose and dynamic target trajectory are jointly optimized to improve the adaptability, robustness and accuracy of the multi-target tracking and SLAM positioning system in dynamic scenes.

[0017] As a further technical solution of the present application, in step 1, the system receives data from IMU, lidar, camera output. Once the new IMU, lidar, camera measurement data is recorded, these data are first sent to the data preprocessing module to provide basic data for subsequent scene tracking and nonlinear optimization. Among them, the original point cloud data of the lidar will be motion compensated according to the high-frequency IMU prediction value.

[0018] The data preprocessing process of the system of the present application mainly includes lidar point cloud data preprocessing, image data preprocessing and IMU pre-integration. For high-frequency IMU data, the present application integrates it to obtain pre-integrated observation values of adjacent frames. For lidar data, a target detection method combining deep learning and geometric segmentation clustering is used to identify and initialize 3D target objects. The point cloud features of background and target are extracted by calculating the local smoothness of the point cloud. For image data, a 2D target detection network is first used to identify and initialize 2D target objects. Then, FAST corner points are extracted from the masks of background and target, and target feature descriptors (Binary Robust Independent Elementary Features, BRIEF) are calculated at the same time for matching and tracking, so as to realize the association of multi-modal target data.

[0019] In order to avoid the repetition of the same target object information, the 3D target detection bounding box of the lidar is projected to the image plane, and the projected bounding box is matched with the 2D target detection bounding box in the image, so as to realize the consistent description of the same target. At the same time, in order to effectively compress the observation information of the target lidar point cloud feature and visual feature, the present application converts the target lidar feature point cloud to the camera coordinate system and projects it to the normalized coordinate system to obtain the depth information of the corresponding visual feature point, and finally forms the multi-modal feature point observation information of the target. Among them, the multi-modal feature points of the target are mainly visual feature points. If the target lidar feature point cloud projected to the final normalized coordinate system can accurately recover the depth information of the target visual feature point, the visual feature point and its lidar feature point depth information of the target are retained at the same time, forming a multi-modal target feature point with visual information and lidar point cloud depth. Conversely, if the target lidar feature point projection to the image plane cannot accurately recover the depth of the target visual feature point, the multi-modal target feature point only retains its visual feature point information. It is worth noting that due to the sparsity of the lidar point cloud, the corresponding relationship between the lidar feature point and the visual feature point cannot be directly established by point cloud projection to form the multi-modal feature point information of the target. After the above steps, each target object can be represented by a 2D detection bounding box, a 3D detection bounding box or both, and each target object has multi-modal feature point information.

[0020] As a further technical scheme of the present application, the step 2 comprises performing 4D (3D+time) scene tracking, mainly including background feature association and multi-modal target data association. The background feature association includes visual feature point tracking and feature association of laser radar background point cloud, wherein KLT sparse optical flow algorithm is used for visual feature point tracking, and spatial nearest neighbor is used to realize edge feature and plane feature association of laser radar background point cloud (LiDAR).

[0021] As a further technical scheme of the present application, in the step 2, a target comprehensive association scale is established based on target multi-modal feature tracking, 2D target detection bounding box distance and 3D target detection bounding box distance, so as to realize multi-modal target data association. In the system proposed in the present application, a target object can be represented by a 2D target detection bounding box and a 3D target detection bounding box. Specifically, the present application uses a 2D target detection bounding box to represent a 2D target in an image, and uses a 3D target detection bounding box to represent a 3D target in a field of view of a laser radar sensor. For the target objects of different modalities, a 3D target motion model and a 2D target motion model are respectively used to perform motion prediction on the targets in the field of view of the laser radar and the targets on the image plane, so as to realize accurate target data association. The 3D target motion model is used to predict the state of a target at a next time point according to the linear acceleration and angular velocity of the target at a current time point; and the 2D target motion model is used to predict the state of a target according to uniform linear motion.

[0022] Specifically, for the 3D target motion model, the state of a target at a next time point can be predicted according to the linear acceleration and angular velocity of the target at a current time point, and has: ; in the formula, , , are respectively the position, velocity and pose of the target in a world coordinate system (w) at a current time point; are respectively the position, velocity and pose of the target in a world coordinate system at a next time point; is a rotation matrix representation of the pose; is a time interval between two laser radar scanning frames; is a quaternion multiplication.

[0023] For the 2D target motion model, the present application uses uniform linear motion for description, that is: ; in the formula, , ​​​​​​​The position and velocity in the image pixel coordinate system.

[0024] It is worth noting that for the 3D-2D fusion target on the successful matching, the motion state thereof is predicted by using the above two formulas respectively.

[0025] In order to effectively fuse the 2D, 3D detection bounding box and multi-modal feature point information of the target, the application proposes an aggregated target data association index to realize accurate target data association. Specifically, the target association cost function is described as a linear combination of 2D detection bounding box distance, 3D detection bounding box distance and multi-modal feature matching rate, and the specific expression is: ; in the formula, , , and is the set of current detection target objects, is the set of historical tracking trajectories; , and respectively represent the 2D detection bounding box distance, the 3D detection bounding box distance and the multi-modal feature matching rate between the first target detection object and the first tracking trajectory; , and are the corresponding weight coefficients. Next, the specific definition of the 2D detection bounding box distance, the 3D detection bounding distance and the multi-modal feature matching rate will be introduced.

[0026] Due to the sparsity of laser radar point cloud, the perception distance of the camera is usually farther than that of the laser radar. Therefore, making full use of the 2D target detection bounding box information can realize the robust tracking of the long-distance target object. The application adopts the intersection over union (2D IoU) of the 2D target detection bounding box as its distance function, that is: ; in the formula, , respectively represent the 2D detection bounding box distance, the 3D detection bounding box distance and the multi-modal feature matching rate between the first target object and the first tracking trajectory;

[0027] In the complex dynamic environment of the city, the targets will be mutually occluded, and it is difficult to accurately associate the occluded targets according to the 2D IoU. Sometimes, in the image field of view, the front vehicles exist mutual occlusion, resulting in temporary loss of the occluded target. However, in the field of view of the laser radar sensor, the 3D point cloud data can still detect and track all targets. Therefore, making full use of the 3D target detection bounding box information can better understand the position and shape of the target to solve the problem of target trajectory loss due to target occlusion. The application adopts the intersection over union (3D IoU) of the 3D target detection bounding box as its distance function, that is: wherein, , are the 3D detection bounding boxes of the th target detection object and the th tracking trajectory, respectively.

[0028] In the long-term data association task of targets, target feature points play a crucial role. As unique identifiers and descriptors of targets, target feature points provide significant features of targets at different time instants or data sources. Continuous extraction and matching of target feature points can achieve continuous tracking of targets, thereby ensuring the stability and accuracy of tracking. Even in the face of challenges such as target occlusion, appearance change, or environmental interference, the matching results of feature points can ensure the stable operation of the target tracking task. Therefore, the present application uses a multi-modal feature point matching rate to further improve the performance of multi-target tracking, which is defined as follows: wherein, is true, otherwise false; is the number of multi-modal feature points of the target in the current frame; is the 3D position coordinates of the feature point in the corresponding target coordinate system, wherein, is the 3D position coordinates of the feature point in the laser radar coordinate system, is the pose of the target detection object in the laser radar coordinate system; is defined in the same way as ; and is the Euclidean distance threshold of multi-modal feature points. It is worth noting that multi-modal feature matching not only provides important clues for target data association, but also helps to control the observation quality of multi-modal feature point data, thereby preventing poor observation data quality from affecting the performance of the tightly coupled state estimation system.

[0029] By integrating the above multi-modal target data association indicators and according to the target association cost function, the cost of each pair of target detection object and tracking trajectory matching can be obtained, and a final target data association matrix can be constructed. Finally, the Hungarian algorithm is used to solve the multi-modal target data association problem to obtain the best target matching pair.

[0030] Target lifecycle management is crucial for reducing false negative and false positive tracking. This section mainly introduces four target states: tracking failure, temporary tracking, valid tracking, and reappearance. Specifically, tracking failure refers to the situation where a target trajectory fails to track within a certain number of frames (exceeding a pre-set threshold ​​​​) still cannot be associated with any target detection object. Temporary tracking refers to a target object that is only detected by the 2D target detector, or a target object whose tracking frame number is less than a pre-set threshold. Effective tracking refers to a target object that is detected by both the 2D and 3D target detectors, or a target object that is tracked in consecutive frames (more than a pre-set threshold ) can be associated with the target detection object. Reappearance refers to a situation that a target in effective tracking appears to lose tracking in some frames due to occlusion between targets, but can be successfully matched to the target in subsequent frames. If it cannot be associated in a certain number of subsequent frames (more than a pre-set threshold ), it is considered as tracking failure; if it is successfully associated in subsequent frames, its state becomes effective tracking. It is worth noting that for temporary tracking or reappearance targets, the present application will predict the target trajectory according to its 3D target motion model to generate temporary target objects and associate them with the current detection, thereby reducing the false negative index of multi-target tracking.

[0031] As a further technical solution of the present application, in step 3, the multi-modal feature point state of the target of the present application is maintained in the world coordinate system. In addition, the system additionally maintains the motion state of the 2D target trajectory. Based on the results of 4D scene tracking, the tracked target trajectory and navigation state are then executed by a sliding window-based nonlinear optimization to achieve robust and high-precision joint state estimation.

[0032] The specific definition of the system state is as follows: ; wherein, ; ; ; ; ; ; In the formula, is the number of moving targets; is the size of the sliding window; is the number of background visual feature points; is the number of feature points of the target ; , are the states of the carrier and the first target in the sliding window, respectively; , , are the position, velocity and attitude of the carrier in the world coordinate system at time ; , , respectively moment target position, velocity and attitude in the world coordinate system; , respectively moment target acceleration and angular velocity; respectively moment target position and velocity in the image pixel coordinate system; target detect the length, width and height of the bounding box; , is the set of background visual feature points and target multi-modal feature points, and is represented by inverse depth. It is worth noting that the present application adds the position and velocity of the target in the image pixel coordinate system, i.e. the state of the 2D target object, and optimizes it jointly with the 3D target state. In addition, the state of the target feature point is compressed from 3D position to inverse depth, effectively reducing the dimension of the to-be-estimated parameters.

[0033] The present application maintains the observation of the multi-modal feature points of the target, and at the same time maintains the multi-modal information of the target detection bounding box. Next, the multi-modal detection bounding box and the geometric feature point information of the target will be fully utilized for optimal estimation of the system state. This state estimation problem is still a maximum a posteriori estimation problem, and can be converted into a nonlinear least squares problem, i.e. ; in the formula, Huber robust kernel function is represented; is the prior information obtained by marginalization; is the IMU pre-integration residual; is the surface feature residual of the background point cloud of the lidar; is the visual background feature point residual; is the multi-modal feature residual of the target; is the 2D detection bounding box residual of the target; is the 3D detection bounding box residual of the target; is the multi-modal detection bounding box residual of the target; is the 2D motion model residual of the target; is the 3D motion model residual of the target; is the IMU pre-integration measurement set; is the surface feature measurement set of the background point cloud of the lidar; is the visual background feature point measurement set; is the multi-modal feature point measurement set of the target; , are respectively the 2D and 3D detection bounding box measurement sets of the target; A set of bounding box measurements for multimodal detection targeting; , A collection of measurements for the target's 2D and 3D motion models; The set of targets being tracked; , , , , , , , and These are information matrices representing IMU pre-integration observations, lidar background point cloud surface feature observations, visual background feature point observations, target multimodal feature point observations, target 2D detection bounding box observations, target 3D detection bounding box observations, target multimodal detection bounding box observations, target 2D motion model observations, and target 3D motion model observations, respectively. From the above equation, the entire state-to-state joint optimization problem can be decomposed into multiple factors related to the state and measurements, making the problem easier to handle and solve.

[0034] The proposed system's factor graph mainly includes IMU pre-integration residuals, LiDAR background point cloud surface feature residuals (LiDAR point cloud matching residuals), visual background feature point residuals (visual reprojection residuals), target multimodal feature residuals, target multimodal target detection bounding box residuals (target 2D detection bounding box residuals, target 3D detection bounding box residuals, and target multimodal detection bounding box residuals), and target motion model residuals (target 2D motion model residuals and target 3D motion model residuals). The specific construction forms of these residuals will be described below.

[0035] To obtain a smooth and consistent target state estimate, this invention employs a motion model to constrain the motion state at adjacent time points. Given two consecutive time points... and The residual of the 3D motion model is defined as follows: .

[0036] Information matrix corresponding to the motion model It can be defined as: .

[0037] Pre-integration observations from IMU , , ,definition The IMU pre-integration residuals of adjacent image frames at time points are: In the formula, , , For IMU pre-integration terms; a relative rotation error representing a 3D Euclidean space; a quaternion representing an extracted vector element of correspondence; , are the bias vectors of the gyroscope and accelerometer, respectively; is the gravity vector in the world coordinate system. The information matrix of the IMU pre-integrated observation can be recursively computed according to the first-order discrete-time covariance update, i.e., where is an identity matrix; is the state transition matrix; is the noise driven matrix; is the process noise covariance matrix, where the initial value is 0.

[0038] Unlike the traditional pinhole camera model, the present application defines a static visual feature point residual error on the unit sphere. For any static visual feature point , assuming that it is first observed at time , the reprojection residual (visual reprojection residual) of the feature point observed at time can be expressed as: ; and ; where is the pixel coordinate of the first observation of the static visual feature point ; is the pixel coordinate of the observation of the static visual feature point at time ; is the depth of the static visual feature point in the camera coordinate system; is the rotation matrix representation of the carrier attitude at time ; is the back-projection function of the camera. is the extrinsic parameter of the camera to the IMU sensor, which can be obtained by an offline calibration algorithm. Considering that the static visual feature point can be stably tracked in the time dimension and the left and right space dimensions, the residual error construction in the above formula is applicable to both the front and rear frame images and the binocular images. In addition, since the final residual error construction form is on the unit sphere, the information matrix of the corresponding feature point can be determined by the pixel observation noise variance transmission, i.e., ; where is the pixel observation noise variance, is the back-projection function of the camera.

[0039] The existing research results show that the edge feature points cannot improve the precision of the laser inertial odometer, therefore, the present application only considers the planar feature point observation, and maintains a static point cloud map in a world coordinate system. The observation of any static planar feature point of the current laser radar frame is converted to the world coordinate system to obtain , so as to realize the space coordinate reference unification of the current frame observation and the global static point cloud map. The feature point is aligned with the global static point cloud map through the spatial nearest neighbor, the corresponding face element is obtained, and the distance of the planar feature point to the corresponding plane is constructed to represent the matching residual (LiDAR point cloud matching residual) of the planar feature point, that is: ; wherein, ; in the formula, is the normal vector and the center point of the face element of the global static point cloud map corresponding to the planar feature point . is the external parameter of the laser radar sensor to the IMU, which can be obtained through offline calibration. The information matrix of the corresponding feature point is defined as the unit weight.

[0040] Since the pose truth value of the carrier trajectory cannot be obtained, the target point cloud detection bounding box residual constructed herein needs to consider the carrier pose. Taking a target as an example, the 3D bounding box detected at the time can be represented as . According to the target state of the current frame, it is projected into the laser radar coordinate system, and then .

[0041] The target 3D detection bounding box residual can be represented as the similarity matrix of two vectors, that is: .

[0042] The information matrix of the corresponding target 3D detection bounding box observation is defined according to the confidence of the target detection bounding box, and the elements of the main diagonal are set to: .

[0043] As can be seen from the above formula, the target feature point matching residual and the target detection bounding box residual can constrain the carrier pose and the target state, therefore, the motion target information can effectively overcome the limitations of the laser inertial odometer in dynamic scenes. It is worth noting that for the target with high-quality observation, the carrier pose in the residual equation is regarded as a parameter to be estimated; for the target with low-quality observation, the carrier pose in the residual equation is regarded as a known parameter.

[0044] The target multi-modal feature point of the application can be tracked and matched between frames by a visual BRIEF descriptor, and the depth information thereof can be accurately recovered according to a laser radar point cloud. For any target multi-modal feature point , assuming that it is first observed at and is matched at , according to the rigidity assumption of the target vehicle, there is: ; in the formula, is the position of the target multi-modal feature point in the target coordinate system; and are the positions of the target multi-modal feature point in the world coordinate system at and , and can be obtained according to the current visual and laser radar observation models. Taking as an example, there is: .

[0045] According to the above two formulas, the re-projection residuals at and are: ; wherein, ; in the formula, is the inverse depth of the feature point ; is the pixel coordinate of the target multi-modal feature point observed at ; is the pixel coordinate of the target multi-modal feature point first observed. It is worth noting that the re-projection residual of the target multi-modal feature point fully utilizes the laser radar point cloud to recover the depth of the visual feature point, and adopts the inverse depth to represent the state of the target multi-modal feature point, effectively reducing the dimension of the to-be-estimated parameters. Finally, the information matrix of the corresponding feature point can be determined by the pixel observation noise variance transmission, that is: ; in the formula, is the pixel observation noise variance.

[0046] The application also adopts a more efficient 2D visual target detection network for target detection of images, and uses the 2D target detection bounding box output by the network to constrain the system state. Similar to the 3D detection bounding box residual of the target, the 2D detection bounding box residual of the target is constructed by calculating the similarity matrix of two vectors, that is: ; the information matrix of the target 2D detection bounding box observation can be determined according to the confidence of the detection bounding box output by the 2D target detection network.

[0047] Current multi-modal multi-target tracking systems usually optimize 2D and 3D target trajectories independently. The present invention makes full use of the multi-modal detection bounding box information of the target to realize the joint optimization of multi-modal trajectories. To this end, a hypothesis is introduced, that is, the projection of the 3D target bounding box on the image should closely fit the 2D target bounding box. Based on this hypothesis, the correlation between the 2D target bounding box and the vertices of the 3D target bounding box can be established.

[0048] Taking the target as an example, the eight vertices of its 3D bounding box can be represented as: ; in the formula, is the eight vertex coordinates of the 3D bounding box of the target ; is the size of the target bounding box. Projecting its eight vertices to the image plane, we have: ; in the formula, is the camera projection function, which can project the point in the camera coordinate system to the image pixel coordinate system.

[0049] According to the above formula, the maximum and minimum pixel coordinates of the eight vertices of the 3D target bounding box can be obtained, that is: ; ; in the formula, , are the maximum and minimum functions respectively. According to the above two formulas, the target multi-modal detection bounding box residual error can be represented as: ; in the formula, , are the minimum and maximum values of the visual 2D detection bounding box at time . The information matrix corresponding to the 2D target detection bounding box observation can be determined according to the confidence of the detection frame output by the 2D target detection network and the target depth , that is: ; in the formula, 70 m is the truncation distance of the target detection frame.

[0050] According to the target 2D motion model, the target 2D motion model residual error at two consecutive times is defined as: ; in the formula, is the time interval of adjacent image frames.

[0051] As a further technical solution, through the above factor graph framework, the carrier pose and dynamic target trajectory are optimized based on a sliding window optimization strategy and combined with marginalization technology. With the continuous accumulation of multi-source sensor measurement information, complete global optimization will greatly occupy the computing resources of the system, making it difficult to meet the real-time requirements of the perception and positioning system. In order to solve this problem, the present technology adopts a sliding window optimization strategy for local optimal state estimation, that is, only the sliding window of the latest state is retained, and the past state is removed. In the window sliding process, the constraint information related to the removed state is not discarded, but is used to construct the prior constraint of the state within the sliding window. The above process can be realized by marginalization technology, which can effectively improve the calculation efficiency and maintain the accuracy of state estimation.

[0052] Finally, the present patent adopts Ceres to solve the above least square problem.

[0053] The present application also provides a multi-target tracking information enhanced laser radar vision inertial (LiDAR / Vision / INS) high-precision positioning system, comprising:

[0054] A data preprocessing module is used to acquire measurement data of a laser radar, a camera and an IMU, and perform data preprocessing, including laser radar point cloud data preprocessing, image data preprocessing and IMU pre-integration.

[0055] A multi-modal target data association module is used to establish a target comprehensive association scale that fuses laser radar and vision multi-modal detection bounding boxes and geometric features based on the preprocessed data, and associate multi-modal target data.

[0056] A nonlinear optimization module is used to optimize the carrier pose and dynamic target trajectory in combination with the multi-modal detection bounding boxes and geometric features of the target, IMU pre-integration measurement, realize autonomous positioning and target trajectory output.

[0057] Further, the above multi-target tracking information enhanced laser radar vision inertial (LiDAR / Vision / INS) high-precision positioning system further comprises a background feature association module, and the background feature association module and the multi-modal target data association module jointly constitute a 4D scene tracking module. The background feature association module is used to realize laser radar background point cloud feature association by adopting spatial nearest neighbor, and realize background vision feature point tracking association by KLT sparse optical flow tracking technology.

[0058] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores program instructions executed by the processor, and the processor invokes the program instructions to execute the above multi-target tracking information enhanced laser radar vision inertial (LiDAR / Vision / INS) high-precision positioning method.

[0059] The application also provides a non-transitory computer readable storage medium storing computer instructions, which cause the computer to execute the multi-target tracking information enhanced laser radar vision inertial high-precision positioning method.

[0060] The one or more technical solutions provided in the application have at least the following technical effects or advantages:

[0061] The application provides a multi-target tracking information enhanced laser radar vision inertial high-precision positioning algorithm. Firstly, in view of the inherent limitation of a single sensor in target data association, the application designs a comprehensive association scale that fuses laser radar and vision multi-modal detection bounding boxes and geometric features, so as to improve the multi-modal target data association performance in a complex scene. Secondly, the multi-modal detection bounding boxes and geometric features of the target, the IMU pre-integral measurement, the optimized carrier pose and the dynamic target trajectory are combined, so as to improve the adaptability, robustness and accuracy of the multi-target tracking and SLAM positioning system in a dynamic scene. BRIEF DESCRIPTION OF DRAWINGS

[0062] Figure 1 It is a multi-target tracking information enhanced LiDAR / Vision / INS high-precision positioning algorithm framework;

[0063] Figure 2 It is a 2D and 3D target detection bounding box distance diagram (the red box is the real bounding box of the target);

[0064] Figure 3 It is a multi-target tracking information enhanced LiDAR / Vision / INS high-precision positioning system observation and state diagram;

[0065] Figure 4 It is a multi-target tracking and SLAM positioning system structure diagram based on factor graph optimization. DETAILED DESCRIPTION

[0066] The technical solutions of the application will be described clearly and completely below with reference to the drawings. Obviously, the described embodiments are only some of the embodiments of the application, but not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the application.

[0067] The solution based on laser inertia is obviously superior to the method based on visual inertia, but the positioning performance of the laser radar degrades in scenes with weak structural features. In order to make full use of the advantages of the laser radar and the visual sensor, such as high ranging accuracy of the laser radar and immunity to light, and good visibility and rich texture information of the visual sensor, the present application provides a laser radar visual inertia high-precision positioning method with multi-target tracking information enhancement, so as to improve the precision and robustness of multi-target tracking and autonomous positioning of the system in a dynamic scene.

[0068] The present embodiment provides a laser radar visual inertia high-precision positioning method with multi-target tracking information enhancement, and the general idea is as follows:

[0069] In the system provided in the present application, the target object can be represented by a 2D or 3D target detection bounding box. In order to overcome the inherent limitations of a single sensor in target data association, the proposed system makes full use of the laser radar and visual multi-modal detection bounding boxes and geometric features to realize robust scene tracking. Then, by effectively fusing the multi-modal target detection bounding box and geometric feature information and the IMU pre-integral measurement, the carrier pose and dynamic target trajectory are jointly optimized to improve the adaptability, robustness and accuracy of the multi-target tracking and autonomous positioning system in a dynamic scene.

[0070] The method of the present application, as shown in Figure 1 , comprises:

[0071] Step 1: Obtain the measurement data of the laser radar, camera and IMU, and perform data preprocessing. The process mainly includes laser radar point cloud data preprocessing, image data preprocessing and IMU pre-integration.

[0072] Specifically, in step 1, the system receives data output from the IMU, laser radar and camera. Once new IMU, laser radar and camera measurement data are recorded, the data are first sent to the data preprocessing module to provide basic data for subsequent 4D scene tracking and nonlinear optimization. Among them, the original point cloud data of the laser radar are motion compensated according to the high-frequency IMU prediction value.

[0073] The data preprocessing process of the system mainly includes laser radar point cloud data preprocessing, image data preprocessing and IMU pre-integration. For high-frequency IMU data, the present application integrates it to obtain pre-integrated observation values of adjacent frames. For laser radar data, a target detection method combining deep learning and geometric segmentation clustering is used to identify and initialize 3D target objects. Then, the point cloud features of the background and the target are extracted by calculating the local smoothness of the point cloud. For image data, a 2D target detection network is first used to identify and initialize 2D target objects. Then, FAST corner points are extracted from the image background and target masks, and target feature descriptors (Binary Robust Independent Elementary Features, BRIEF) are calculated at the same time for matching and tracking, so as to realize multi-modal target data association.

[0074] In order to avoid the repetition of information of the same target object, the 3D target detection bounding box of the laser radar is projected to the image plane, and the projected bounding box is matched with the 2D target detection bounding box in the image, so as to realize consistent description of the same target. At the same time, in order to effectively compress the observation information of the target laser radar point cloud feature and the visual feature, the present application converts the target laser radar feature point cloud to the camera coordinate system, and projects it to the normalized coordinate system to obtain the depth information of the corresponding visual feature point, and finally forms the multi-modal feature point observation information of the target. Among them, the multi-modal feature points of the target are mainly visual feature points. If the target laser radar feature point cloud projected to the final normalized coordinate system can accurately recover the depth information of the target visual feature point, the visual feature point and the depth information of the target laser radar feature point are retained at the same time, forming a multi-modal target feature point with visual information and laser radar point cloud depth. Otherwise, if the target laser radar feature point projection to the image plane cannot accurately recover the depth of the target visual feature point, the multi-modal target feature point only retains its visual feature point information. It is worth noting that due to the sparsity of laser radar point cloud, the correspondence between laser radar feature points and visual feature points cannot be directly established by point cloud projection to form the multi-modal feature point information of the target. After the above steps, each target object can be represented by a 2D detection bounding box, a 3D detection bounding box or both, and each target object has multi-modal feature point information.

[0075] Step 2: 4D scene tracking, mainly including background feature association and multi-modal target data association.

[0076] The background feature association includes visual feature point tracking and feature association of laser radar background point cloud, wherein KLT sparse optical flow algorithm is used for visual feature point tracking, and spatial nearest neighbor is used to realize edge feature and plane feature association of LiDAR background point cloud.

[0077] The multi-modal target data association: for the inherent limitations of single sensor in target data association, based on the pre-processed data, a target comprehensive association scale is designed to fuse the laser radar and visual multi-modal detection bounding box and geometric features, and multi-modal target data association is realized.

[0078] Specifically, in step 2, based on target multi-modal feature tracking, 2D target detection bounding box distance, 3D target detection bounding box distance, a target comprehensive association scale is established, and multi-modal target data association is realized. In the system proposed in the application, the target object can be represented by 2D and 3D target detection bounding box. Specifically, the application uses 2D target detection bounding box to represent 2D target in image, and uses 3D target detection bounding box to represent 3D target in the field of view of laser radar sensor. For the target objects of different modalities, 3D target motion model and 2D target motion model are used respectively to predict the motion of the target in the field of view of laser radar and the target in the image plane respectively, so as to realize accurate target data association. The 3D target motion model is used to predict the state of the target at the next time according to the linear acceleration and angular velocity of the target at the current time; the 2D target motion model is used to predict the state of the target according to uniform linear motion.

[0079] For the 3D target motion model, we have: ; in the formula, , , are respectively the position, velocity and attitude of the target at the time in the world coordinate system (w); , , are respectively the position, velocity and attitude of the target at the time in the world coordinate system; is the rotation matrix representation of the attitude; is the time interval between two laser radar scanning frames; is the quaternion multiplication.

[0080] For the 2D target motion model, the application uses uniform linear motion to describe, that is: ; in the formula, , is the position and velocity in the image pixel coordinate system.

[0081] It is worth noting that for the 3D-2D fusion target matched successfully, the motion state of the target is predicted by using the above two formulas respectively.

[0082] ​​To effectively fuse the 2D, 3D detection bounding box and multi-modal feature point information of the target, the application proposes an aggregated target data association indicator to realize accurate target data association. Specifically, the target association cost function is described as a linear combination of 2D detection bounding box distance, 3D detection bounding box distance and multi-modal feature matching rate, and the specific expression is: ; in the formula, , , and is the set of current detected target objects, is the set of historical tracking trajectories; , and respectively represent the 2D detection bounding box distance, the 3D detection bounding box distance and the multi-modal feature matching rate between the first target detection object and the first tracking trajectory; , and are the corresponding weight coefficients. Next, the specific definition of 2D detection bounding box distance, 3D detection bounding distance and multi-modal feature matching rate will be introduced.

[0083] Due to the sparsity of laser radar point cloud, the perception distance of the camera is usually farther than that of the laser radar. Therefore, making full use of 2D target detection bounding box information can realize robust tracking of long-distance target objects. The application adopts the intersection over union (2D IoU) of the 2D target detection bounding box as its distance function, that is: ; in the formula, , respectively represent the 2D detection bounding box distance, the 3D detection bounding box distance and the multi-modal feature matching rate between the first target object and the first tracking trajectory;

[0084] In the complex dynamic environment of the city, the targets will be mutually occluded, and it is difficult to accurately associate the occluded targets according to the 2D IoU. As shown in (a) of Figure 2 , in the image field of view, the vehicles in front of each other are mutually occluded, resulting in temporary loss of the occluded target. However, in the field of view of the laser radar sensor, 3D point cloud data can still detect and track all targets, as shown in (b) of Figure 2 . Therefore, making full use of 3D target detection bounding box information can better understand the position and shape of the target to solve the problem of target trajectory loss due to target occlusion. The application adopts the intersection over union (3DIoU) of the 3D target detection bounding box as its distance function, that is: ; in the formula, , respectively represent the 2D detection bounding box distance, the 3D detection bounding box distance and the multi-modal feature matching rate between the first target detection object and the first A 3D detection bounding box for tracking trajectory.

[0085] In long-term target data association tasks, target feature points play a crucial role. As unique identifiers and descriptors of the target, target feature points provide salient characteristics of the target at different times or data sources. Continuously extracting and matching target feature points enables continuous target tracking, thereby ensuring the stability and accuracy of the tracking. Even in the face of challenges such as target occlusion, appearance changes, or environmental interference, the matching results of feature points can guarantee the stable operation of the target tracking task. Therefore, this invention employs a multimodal feature point matching rate to further improve the performance of multi-target tracking, specifically defined as follows: In the formula, ;if Feature points on and If a match is found, then If true, then false; For the target of the current frame The number of multimodal feature points; For feature points The 3D position coordinates in the corresponding target coordinate system, where... The coordinates of this feature point in the lidar coordinate system are 3D position coordinates. For the target detection object Pose in the lidar coordinate system; Definition method and Consistent; The threshold is the Euclidean distance between the multimodal feature points. It is worth noting that multimodal feature matching not only provides important clues for the association of target data, but also helps control the observation quality of multimodal feature point data, thereby preventing poor observation data quality from affecting the performance of the tightly coupled state estimation system.

[0086] By combining the above multimodal target data association metrics and obtaining the cost of pairwise matching pairs between the detected target and the tracking trajectory based on the target association cost function, the final target data association matrix is ​​constructed. Finally, the Hungarian algorithm is used to solve the multimodal target data association problem to obtain the optimal target matching pairs.

[0087] Target lifecycle management is crucial for reducing false negatives and false positives in tracking. This section mainly introduces four target states: tracking failure, temporary tracking, effective tracking, and reappearance. Specifically, tracking failure refers to the target trajectory failing to track within a certain number of frames (exceeding a pre-set threshold). ) still cannot be associated with any target detection object. Temporary tracking refers to a target object that is only detected by the 2D target detector, or a target object whose tracking frame number is less than a pre-set threshold. Effective tracking refers to a target object that is detected by both the 2D and 3D target detectors, or a target object whose tracking frame number is greater than a pre-set threshold ) in a certain number of subsequent frames (greater than a pre-set threshold ) cannot be associated, it is considered as tracking failure; if it is successfully associated in the subsequent frames, its state becomes effective tracking. It is worth noting that for the temporarily tracked or reappeared target, the present application will predict the target trajectory according to its 3D target motion model to generate a temporary target object and associate it with the current detection, thereby reducing the false negative index of multi-target tracking.

[0088] Step 3: Nonlinear optimization is performed to jointly optimize the target multi-modal detection bounding box and geometric features, IMU pre-integral measurement, carrier pose and dynamic target trajectory.

[0089] Specifically, in step 3, the multi-modal feature point state of the target is maintained in the world coordinate system. In addition, the system additionally maintains the motion state of the 2D target trajectory. Based on the results of 4D scene tracking, the tracked target trajectory and navigation state are then executed based on a sliding window nonlinear optimization to achieve robust and high-precision joint state estimation.

[0090] Figure 3 The observation and state of the system are shown, wherein the specific definition of the system state is as follows: ; wherein, ; ; ; ; ; ; In the formula, is the number of moving targets; is the size of the sliding window; is the number of background visual feature points; is the number of feature points of the target ; , are the states of the carrier and the first target in the sliding window, respectively; , , are respectively the position, velocity and attitude of the carrier in the world coordinate system at time t; , , are respectively the position, velocity and attitude of the target in the world coordinate system at time t; , , are respectively the acceleration and angular velocity of the target at time t; , are respectively the position and velocity of the target in the image pixel coordinate system at time t; is the target the length, width and height of the detection bounding box; , , are the sets of background visual feature points and target multi-modal feature points, and are represented by inverse depth. It is worth noting that the present application adds the position and velocity of the target in the image pixel coordinate system, i.e., the state of the 2D target object, and optimizes it jointly with the 3D target state. In addition, the state of the target feature points is compressed from 3D position to inverse depth, effectively reducing the dimension of the to-be-estimated parameters.

[0091] The present application maintains the observation of the target multi-modal feature points, and at the same time maintains the multi-modal target detection bounding box information. Next, the multi-modal detection bounding box and geometric feature point information of the target will be fully utilized for optimal estimation of the system state. The state estimation problem is still the maximum a posteriori probability estimation problem, and can be converted into a nonlinear least squares problem, i.e., ; in the formula, Huber robust kernel function is represented by Huber; is the IMU pre-integration residual; is the surface feature residual of the background point cloud of the lidar; is the visual background feature point residual; is the target multi-modal feature residual; is the target 2D detection bounding box residual; is the target 3D detection bounding box residual; is the multi-modal detection bounding box residual of the target; is the target 2D motion model residual; is the target 3D motion model residual; is the IMU pre-integration measurement set; is the surface feature measurement set of the background point cloud of the lidar; is the visual background feature point measurement set; A multimodal feature point measurement set for the target; , These are sets of bounding box measurements for 2D and 3D target detection, respectively. A set of bounding box measurements for multimodal detection targeting; , A collection of measurements for the target's 2D and 3D motion models; The set of targets being tracked; , , , , , , , and The information matrices are respectively: IMU pre-integration observation, lidar background point cloud surface feature observation, visual background feature point observation, target multimodal feature point observation, target 2D detection bounding box observation, target 3D detection bounding box observation, target multimodal detection bounding box observation, target 2D motion model observation, and target 3D motion model observation. This is prior information obtained from marginalization. From the above equation, the entire joint state optimization problem can be decomposed into multiple factors related to the state and measurement, making the problem easier to handle and solve.

[0092] Figure 4 The factor graph of the proposed system is presented, mainly including the IMU pre-integration residual, the surface feature matching residual of the LiDAR background point cloud, the visual reprojection residual, the target multimodal feature residual, the target multimodal target detection bounding box residual, and the target motion model residual. The specific construction forms of these residuals will be described below.

[0093] To obtain a smooth and consistent target state estimate, this invention employs a motion model to constrain the motion state at adjacent time points. Given two consecutive time points... and The residual of the 3D motion model is defined as follows: .

[0094] Information matrix corresponding to the motion model It can be defined as: .

[0095] Pre-integrated observations from IMU , , ,definition The IMU pre-integration residuals of adjacent image frames at time points are: In the formula, Represents the relative rotation error in 3D Euclidean space; represents extracting quaternion corresponding vector elements. Information matrix of IMU pre-integrated observation The information matrix can be calculated according to a first-order discrete-time covariance update recursion, that is: ; wherein, is a state transition matrix; is a noise driving matrix; is a process noise covariance matrix. Wherein The initial value is 0.

[0096] Unlike the traditional pinhole camera model, the application defines the static visual feature point residual on the unit sphere. For any static visual feature point , assuming that it is first observed at , the re-projection residual of the feature point observed at can be expressed as: ; and ; wherein, is the first observed pixel coordinate of the static visual feature point ; and is the observed pixel coordinate of the static visual feature point at ; and is the back-projection function of the camera. is the extrinsic parameter of the camera to the IMU sensor, which can be obtained by an offline calibration algorithm. Considering that the static visual feature point can be stably tracked in the time dimension and the left and right space dimensions, the residual construction in the above formula is suitable for both the front and rear frame images and the binocular images. In addition, since the final residual construction form is on the unit sphere, the information matrix of the corresponding feature point can be determined by the pixel observation noise variance transmission, that is: ; wherein, is the pixel observation noise variance.

[0097] Existing research results show that edge feature points cannot improve the accuracy of the laser inertial odometer. Therefore, the application only considers planar feature point observation and maintains a static point cloud map in the world coordinate system. For any static planar feature point observation of the current laser radar frame, it is converted to the world coordinate system to obtain , so as to unify the spatial coordinate reference of the current frame observation and the global static point cloud map. By spatial nearest neighbor, the feature point is aligned with the global static point cloud map, the corresponding face element is obtained, and the distance from the planar feature point to the corresponding plane is constructed to represent the matching residual of the planar feature point, that is: ; wherein, ; wherein, is the planar feature point The normal vectors and center points of the corresponding global static point cloud map elements; The extrinsic parameters from the lidar sensor to the IMU can be obtained through offline calibration. The information matrix corresponding to the feature points. Defined as unit weight.

[0098] Since the true pose of the carrier trajectory cannot be obtained, the residual of the target point cloud detection bounding box constructed here needs to consider the carrier pose. For example, it is in The 3D bounding box that is detected at any given time can be represented as Based on the target state of the current frame, projecting it onto the lidar coordinate system, we have: The target point cloud detection bounding box residual (target 3D detection bounding box residual) can be represented as a similarity matrix between two vectors, i.e.: .

[0099] Information matrix corresponding to target 3D detection bounding box observation Based on the confidence definition of the object detection bounding box, the elements of its main diagonal are set as follows: .

[0100] As shown in the above formulas, the target feature point matching residual and the target detection bounding box residual can constrain the carrier pose and the target state. Therefore, the moving target information can effectively overcome the limitations of laser inertial odometry in dynamic scenes. It is worth noting that for high-quality observed targets, the carrier pose in this residual equation... The parameter is considered to be estimated; for targets with low-quality observations, the carrier pose in the residual equation is considered to be... Treat them as known parameters.

[0101] The target multimodal feature points of this invention can be tracked and matched between frames using the visual BRIEF descriptor, and their depth information can be accurately recovered from the lidar point cloud. For any target multimodal feature point... Assuming it is in The moment was first observed, and simultaneously at In terms of timing matching, based on the rigidity assumption of the target vehicle, we have: In the formula, For target multimodal feature points Position in the target coordinate system; and They are respectively Time and Multimodal feature points of target at different times The position in the world coordinate system can be determined based on the current visual and lidar observation models. For example, we have: .

[0102] According to the above two formulas, then The re-projection residual at time is: ; wherein, ; in the formula, is the inverse depth of the feature point ; is the pixel coordinate of the target multi-modal feature point observed at time; is the pixel coordinate of the target multi-modal feature point first observed. It is worth noting that the multi-modal feature point re-projection residual of the target fully utilizes the laser radar point cloud to recover the depth of the visual feature point, and adopts the inverse depth to represent the state of the multi-modal feature point of the target, which effectively reduces the dimension of the to-be-estimated parameter. The information matrix of the corresponding feature point can be determined by pixel observation noise variance transmission, that is: ; in the formula, is the pixel observation noise variance.

[0103] The present application simultaneously adopts a more efficient 2D visual target detection network for image target detection, and uses the 2D target detection bounding box output by the network to constrain the system state. Similar to the 3D detection bounding box residual of the target, the 2D detection bounding box residual of the target is constructed by calculating the similarity matrix of two vectors, that is: ; the information matrix of the 2D detection bounding box observation of the target can be determined according to the confidence of the detection box output by the 2D target detection network.

[0104] The current multi-modal multi-target tracking system usually optimizes the 2D and 3D target trajectories independently. The present application fully utilizes the multi-modal detection bounding box information of the target to realize joint optimization of multi-modal trajectories. For this purpose, a hypothesis is introduced, that is, the projection of the 3D target bounding box on the image should closely fit the 2D target bounding box. Based on this hypothesis, the correlation between the 2D target bounding box and the vertices of the 3D target bounding box can be established. Taking the target as an example, the eight vertices of the 3D bounding box of the target ; in the formula, are the eight vertex coordinates of the 3D bounding box of the target ; is the size of the target bounding box. Projecting the eight vertices to the image plane, we have: ; in the formula, is the camera projection function, which can project the point in the camera coordinate system to the image pixel coordinate system.

[0105] According to the above formula, the maximum and minimum of the projection pixel coordinates of the eight vertices of the 3D target bounding box can be obtained, that is: ; in the formula, , are maximum and minimum functions respectively. According to the above two formulas, the target multi-modal detection bounding box residual can be expressed as: ; in the formula, , are the minimum and maximum values of the visual 2D detection bounding box at the moment . The information matrix corresponding to the 2D target detection bounding box observation can be determined according to the confidence of the detection frame output by the 2D target detection network and the target depth , that is: ; in the formula, 70 m is the truncation distance of the target detection frame.

[0106] According to the target 2D motion model, the target 2D motion model residual at two consecutive moments is defined as follows: ; in the formula, is the time interval of adjacent image frames.

[0107] Through the above factor graph framework, the carrier pose and dynamic target trajectory are optimized based on the sliding window optimization strategy and combined with the marginalization technique. With the continuous accumulation of multi-source sensor measurement information, complete global optimization will greatly occupy the computing resources of the system, and it is difficult to meet the real-time demand of the perception positioning system. In order to solve this problem, the technology adopts a sliding window optimization strategy for local optimal state estimation, that is, only the sliding window of the latest state is retained, and the past state is removed. In the window sliding process, the constraint information related to the removed state is not discarded, but is used to construct the prior constraint of the state in the sliding window. The above process can be realized by the marginalization technique, which can effectively improve the calculation efficiency and maintain the accuracy of state estimation.

[0108] Finally, the patent adopts Ceres to solve the above least square problem.

[0109] The embodiment of the application also provides a multi-target tracking information enhanced laser radar-vision-inertial (LiDAR / Vision / INS) high-precision positioning system, which comprises:

[0110] A data preprocessing module is configured to acquire measurement data of a laser radar, a camera and an IMU, and perform data preprocessing, including laser radar point cloud data preprocessing, image data preprocessing and IMU pre-integration.

[0111] ​A multi-modal target data association module is configured to associate multi-modal target data based on preprocessed data, establish a target comprehensive association scale that fuses laser radar and visual multi-modal detection bounding boxes and geometric features, and associate multi-modal target data;

[0112] A non-linear optimization module is configured to jointly optimize carrier pose and dynamic target trajectory based on multi-modal detection bounding boxes and geometric features of targets, IMU pre-integral measurements, to realize autonomous positioning and target trajectory output.

[0113] Further, the multi-target tracking information enhanced laser radar, vision and inertial (LiDAR / Vision / INS) high-precision positioning system comprises a background feature association module, and the background feature association module and the multi-modal target data association module jointly constitute a 4D scene tracking module.

[0114] The embodiment of the present application also provides an electronic device comprising a memory and a processor, wherein the memory stores program instructions executed by the processor, and the processor invokes the program instructions to execute the multi-target tracking information enhanced laser radar, vision and inertial (LiDAR / Vision / INS) high-precision positioning method.

[0115] The embodiment of the present application also provides a non-transitory computer readable storage medium storing computer instructions, wherein the computer instructions enable the computer to execute the multi-target tracking information enhanced laser radar, vision and inertial (LiDAR / Vision / INS) high-precision positioning method.

[0116] The above detailed description of the embodiments of the present application, but the present application is not limited to the specific details of the above embodiments. Within the scope of the claims and technical concepts of the present application, the technical solutions of the present application can be modified and changed in many simple ways, and these simple modifications all belong to the protection scope of the present application.

Claims

1. A multi-target tracking information enhanced laser radar vision inertial high-precision positioning method, characterized in that, The method comprises the steps of: acquiring measurement data of a laser radar, a camera and an IMU, and performing data preprocessing, including laser radar point cloud data preprocessing, image data preprocessing and IMU pre-integration; based on the preprocessed data, establishing a target comprehensive correlation scale for fusing laser radar and visual multi-modal detection bounding boxes and geometric features, and correlating multi-modal target data; optimizing carrier pose and dynamic target trajectory by jointly using multi-modal detection bounding boxes and geometric features of the target, and IMU pre-integrated measurement.

2. The multi-target tracking information-augmented lidar vision-inertial high precision positioning method according to claim 1, wherein, The data preprocessing comprises: IMU pre-integration, including integrating high-frequency IMU data to obtain pre-integrated observation values of adjacent frames; laser radar point cloud data preprocessing, including: using a target detection method combining deep learning and geometric segmentation clustering to identify and initialize 3D target objects; extracting point cloud features of the background and the target by calculating local smoothness of the point cloud; image data preprocessing, including using a 2D target detection network to identify and initialize 2D target objects, and extracting visual feature points of the target and the background based on target masks.

3. The multi-target tracking information-augmented lidar vision-inertial high precision positioning method according to claim 2, wherein, The data preprocessing further comprises projecting the 3D target detection bounding box of the laser radar to an image plane, and matching the projected bounding box with the 2D target detection bounding box in the image; converting the laser radar feature point cloud of the target to a camera coordinate system, and projecting it to a normalized coordinate system to obtain depth information of the corresponding visual feature points, thereby forming multi-modal feature point information of the target.

4. The multi-target tracking information-augmented lidar vision-inertial high precision positioning method according to claim 1, wherein, The method further comprises background feature correlation, which comprises: using spatial nearest neighbor to realize laser radar background point cloud feature correlation, and using KLT sparse optical flow tracking technology to realize background visual feature point tracking correlation.

5. The multi-target tracking information-augmented lidar vision-inertial high precision positioning method according to claim 1, wherein, The multi-modal target data correlation comprises: using 3D target motion models and 2D target motion models to respectively perform motion prediction on targets within the field of view of the laser radar and targets on the image plane.

6. The multi-target tracking information-augmented lidar vision-inertial high precision positioning method according to claim 1, wherein, The target comprehensive correlation scale is constructed by aggregating multi-modal target data correlation indicators including 2D detection bounding box distance, 3D detection bounding box distance and multi-modal feature matching rate, obtaining the cost of matching pairs of target detection objects and historical tracking trajectories according to a target correlation cost function, constructing a target data correlation matrix, and solving the multi-modal target data correlation problem using the Hungarian algorithm to obtain the best target matching pair.

7. The multi-target tracking information-augmented lidar vision-inertial high precision positioning method according to claim 6, characterized in that, The specific expression of the target correlation cost function is: ; In the formula, , , and is a set of currently detected target objects, is a set of historical tracking trajectories; , and respectively represent the 2D detection bounding box distance, the 3D detection bounding box distance and the multi-modal feature matching rate between the first target detection object and the first tracking trajectory; , and are corresponding weight coefficients.

8. The multi-target tracking information-augmented lidar vision-inertial high precision positioning method according to claim 1, wherein, The joint optimization of carrier pose and dynamic target trajectory by using multi-modal detection bounding boxes and geometric features of the target, and IMU pre-integrated measurement comprises: constructing a factor graph by jointly using IMU pre-integrated residuals, surface feature residuals of laser radar background point clouds, visual background feature point residuals, target multi-modal feature residuals, target multi-modal target detection bounding box residuals and target motion model residuals, to optimize carrier pose and dynamic target trajectory.

9. A multi-target tracking information enhanced laser radar vision inertial high-precision positioning system, characterized in that, The method comprises: a data preprocessing module for acquiring measurement data of a laser radar, a camera and an IMU, and performing data preprocessing, including laser radar point cloud data preprocessing, image data preprocessing and IMU pre-integration; A multi-modal target data association module is configured to associate multi-modal target data based on preprocessed data, establish a target comprehensive association scale that fuses laser radar and visual multi-modal detection bounding boxes and geometric features, and associate multi-modal target data; A nonlinear optimization module is configured to optimize carrier pose and dynamic target trajectory based on joint target multi-modal detection bounding boxes and geometric features and IMU pre-integral measurement.

10. An electronic device, comprising: The application also provides a computer readable storage medium storing a program instruction, wherein the program instruction is executed by a processor to perform the multi-target tracking information enhanced laser radar visual inertial high-precision positioning method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Multi-source sensor fusion positioning method and system based on dynamic target enhancement

    CN121677723A

  • Target tracking method and system adapting to scene and cascading data and related equipment

    CN121880963A

  • Multi-target tracking method based on tight coupling laser radar-inertial odometer

    CN121956024A