Rapid hierarchical association and motion enhancement type vehicle-mounted LiDAR multi-target tracking method
Through the fast hierarchical correlation and motion-enhanced vehicle-mounted LiDAR multi-objective tracking method, state prediction and update is used to use CTRA motion model and adaptive Kalman filter to solve the problem of performance degradation and excessive computational burden of three-dimensional multi-objective tracking in the prior art under extreme conditions, and achieve efficient and reliable target tracking.
Patent Information
- Application Number
- CN202510340666.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
The existing three-dimensional multi-objective tracking method has deteriorated performance under extreme lighting and harsh weather conditions. The detection-based tracking method fails when the target changes drastically or temporarily disappears. The current correlation method has too much burden on computing or ignores geometric features, making it difficult to meet the requirements of real-time and robustness.
The rapid hierarchical correlation and motion-enhanced vehicle-mounted LiDAR multi-objective tracking method are adopted, and the state prediction is performed using CTRA motion model and Kalman filter, and the adaptive mechanism is extended Kalman filter for updates. The data correlation is performed by fast hierarchical paired cost calculation and dynamic correlation threshold, and the paired cost is calculated using point cloud target distance and shape correlation attributes.
While ensuring real-time and accuracy, the robustness and adaptability of the three-dimensional multi-objective tracking task are improved, the processing capability of nonlinear scenarios is enhanced, and the accuracy and reliability of target tracking are improved.
Smart Images

Figure CN120275985A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of lidar multi-target tracking, and particularly to a fast hierarchical association and motion enhanced vehicle-mounted LiDAR multi-target tracking method. Background Art
[0002] 3D multi-target tracking is a key intermediate technology for extracting dynamic environmental information from road environments and is widely used in various intelligent transportation systems such as autonomous driving and traffic monitoring. This technology provides crucial dynamic scene perception capabilities for autonomous driving systems by real-time monitoring and accurately identifying dynamic targets in three-dimensional space. In complex traffic scenarios, 3D multi-target tracking can not only effectively track vehicles, pedestrians, cyclists, and other obstacles but also predict their future movement trajectories, thus providing timely and accurate information support for intelligent vehicles to ensure driving safety.
[0003] Most existing multi-target tracking methods rely on camera sensors to locate and track targets by analyzing image sequences captured by cameras. However, solutions based on camera sensors have certain limitations in practical applications. They are often affected by extreme lighting conditions such as darkness, high beam lights, and strong sunlight, as well as adverse weather conditions such as rain, snow, and fog. Therefore, environmental factors can cause the quality of images captured by cameras to decline, thereby affecting the performance and reliability of multi-target tracking systems.
[0004] Lidar uses active light sources, making it insensitive to the influence of environmental light and capable of perception even in a completely dark environment. Moreover, the anti-interference ability and penetration ability of lidar are stronger than those of cameras and are not easily affected by factors such as shadows, reflections, and occlusion objects such as rain, snow, and fog. Currently, 3D multi-target tracking methods are divided into detection-based tracking and joint detection-based tracking. Since the joint detection-based tracking algorithm has high complexity, poor real-time performance, and generally lower robustness than detection-based tracking. Therefore, most 3D multi-target tracking methods follow the detection-based tracking architecture, that is, by analyzing the detection information and context information of objects in the starting frame, establishing corresponding motion models and Kalman filter models for motion prediction and update, and maintaining continuous positioning of objects in subsequent frames.
[0005] In the prior art, most of the detection-based tracking methods use predictors based on constant velocity motion models or constant acceleration motion models to estimate possible future states. When the target has a drastic speed change or the target temporarily disappears in multiple consecutive frames, the predictor may fail, resulting in tracking failure. Most of the object association methods in the current 3D multi-object tracking framework use various types of intersection over union (IoU) or distance for measurement. Among them, the association based on IoU cannot associate two distant states in a non-overlapping state, and the two-stage IoU calculation burden of some methods is too large to meet real-time requirements; the distance-based association method will ignore a large number of geometric features. Summary of the Invention
[0006] In view of this, the present invention provides a fast hierarchical association and motion enhanced vehicle-mounted LiDAR multi-object tracking method to solve the above problems.
[0007] The present invention provides a fast hierarchical association and motion enhanced vehicle-mounted LiDAR multi-object tracking method, including: using a 3D object detection algorithm to perform object detection on the original lidar point cloud data and positioning data to obtain the 3D object detection results in the current frame; performing data preprocessing and coordinate transformation processing on the detection results to obtain the dimensional information in the global coordinate system; establishing a CTRA motion model based on the dimensional information in the global coordinate system; using a Kalman filter based on the CTRA motion model to perform state prediction to obtain a prediction result; performing data association on the detection results and the prediction results based on fast hierarchical pairwise cost calculation and dynamic association threshold to determine the matching detection results and prediction results; using an adaptive mechanism extended Kalman filter to update the matching detection results and prediction results and output the tracking results.
[0008] In another implementation manner of the present invention, the 3D object detection results include candidate object 3D bounding boxes and corresponding confidence scores; the performing data preprocessing and coordinate transformation processing on the detection results to obtain the dimensional information in the global coordinate system includes: quickly screening according to the confidence scores of the candidate object 3D bounding boxes to remove the candidate object 3D bounding boxes with low confidence; removing the candidate object 3D bounding boxes with high similarity based on the non-maximum suppression method to obtain the final candidate object 3D bounding box set; using the vehicle pose information extracted from the positioning data and motion data to convert the bounding boxes in the set from the sensor coordinate system to the global coordinate system to obtain the dimensional information.
[0009] In another implementation manner of the present invention, the dimensional information includes the coordinate information of the target geometric center in the global coordinate system, the target length, width and height, the target speed, the target acceleration, the target orientation angle, and the target turning rate.
[0010] In another implementation of the present invention, the prediction result of the Kalman filter based on the CTRA motion model is expressed as:
[0011]
[0012] where σ represents the interval between two adjacent frames of lidar scans; v(τ) represents the speed; θ(τ) represents the angle; w(τ) represents the turning rate; a represents the acceleration; and t represents the time step.
[0013] In another implementation of the present invention, the data association of the detection result and the prediction result is performed based on the fast hierarchical pairwise cost calculation and the dynamic association threshold to determine the matching detection result and prediction result, including: defining the targets in the prediction result through a preset distance threshold, calculating the pairwise cost by distance for the targets not exceeding the distance threshold, calculating the set shape pairwise cost while calculating the distance pairwise cost for the targets exceeding the distance threshold, and weighting the two costs; integrating multiple quantile costs of the cost data with a static threshold to obtain a dynamic association threshold; determining the corresponding relationship between the detection result and the prediction result according to the dynamic association threshold and the greedy algorithm as the association strategy to obtain the matching detection result and prediction result.
[0014] In another implementation of the present invention, the distance pairwise cost is expressed as:
[0015]
[0016] The shape pairwise cost is expressed as:
[0017]
[0018] The total cost weighting is expressed as:
[0019] Cost all =λ dis Cost dis +μ shape Cost shape
[0020] where N(·) is the normalization function, pose is the target center position (x, y, z), λ dis is the distance cost weight, μ shape is the shape cost weight, score pre is the prediction confidence.
[0021] In another implementation of the present invention, it further includes: setting a consecutive missed detection threshold to distinguish the reasons for unmatched prediction results; when the consecutive missed detection threshold fails to effectively detect the target in the prediction results, it is regarded as naturally disappearing, and the current prediction results will be deleted; when the target in the prediction results reappears in future frames, the current prediction state will be continued to be retained.
[0022] On the other hand, the present invention provides a fast hierarchical association and motion enhanced vehicle-mounted LiDAR multi-target tracking system, including: a target detection module: using a three-dimensional target detection algorithm to perform target detection on the original lidar point cloud data and positioning data to obtain the three-dimensional target detection results within the current frame; a data processing module: performing data preprocessing and coordinate transformation processing on the detection results to obtain the dimensional information in the global coordinate system; a model establishment module: establishing a CTRA motion model based on the dimensional information in the global coordinate system; a target prediction module: using a Kalman filter based on the CTRA motion model to perform state prediction to obtain prediction results; a data matching module: performing data association on the detection results and the prediction results based on fast hierarchical pairwise cost calculation and dynamic association thresholds to determine the matching detection results and prediction results; a result output module: using an adaptive mechanism extended Kalman filter to update the matching detection results and prediction results and output the tracking results.
[0023] On the other hand, the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the steps of a fast hierarchical association and motion enhanced vehicle-mounted LiDAR multi-target tracking method as described in any one of the above.
[0024] On the other hand, the present invention provides a computer storage medium, characterized in that a computer program is stored on the computer storage medium, and when the computer program is executed by a processor, it implements the steps in a fast hierarchical association and motion enhanced vehicle-mounted LiDAR multi-target tracking method as described in any one of the above.
[0025] The fast hierarchical association and motion enhanced vehicle-mounted LiDAR multi-target tracking method of the present invention uses a CTRA motion model to predict the motion state, achieving enhancement in non-linear scenarios such as turning while ensuring accuracy; applying an adaptive mechanism to the extended Kalman filter update stage, improving the adaptive ability of the extended Kalman filter and its processing ability for model changes; based on confidence guidance, quickly calculating pairwise costs using the target distance and shape correlation attributes in the point cloud, and at the same time proposing a scene-based dynamic association threshold, better realizing data association; being able to effectively improve the robustness and adaptive ability of the three-dimensional multi-target tracking task while ensuring real-time performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. By reading the detailed description of the following embodiments, the advantages and benefits in the solutions will become clear to those skilled in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. In the drawings:
[0027] Figure 1 Schematic diagram of the process of a fast hierarchical association and motion-enhanced vehicle-mounted LiDAR multi-target tracking method according to an embodiment of the present invention.
[0028] Figure 2 Schematic diagram of the global coordinate system setting and CTRA motion model definition according to an embodiment of the present invention.
[0029] Figure 3 Schematic diagram of the paired cost distribution of some typical scenarios for the necessity of dynamic association threshold according to an embodiment of the present invention.
[0030] Figure 4 Schematic diagram of the process of a fast hierarchical target association method according to an embodiment of the present invention.
[0031] Figure 5 Schematic diagram of the process of a dynamic threshold target association method according to an embodiment of the present invention.
[0032] Figure 6 Schematic diagram of the tracking effect according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the following will clearly and detailedly describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art shall fall within the protection scope of the embodiments of the present invention.
[0034] Figure 1 Schematic diagram of the process of a fast hierarchical association and motion-enhanced vehicle-mounted LiDAR multi-target tracking method provided by an embodiment of the present invention. As Figure 1 shown, this embodiment mainly includes:
[0035] S101. Use a three-dimensional object detection algorithm to perform object detection on the original lidar point cloud data and positioning data, and obtain the three-dimensional object detection results in the current frame.
[0036] Exemplarily, taking the original lidar point cloud data and GPS / IMU data as inputs, a 3D object detection algorithm based on deep learning is used to generate the 3D object detection results within the current frame.
[0037] S102. Perform data preprocessing and coordinate transformation on the detection results to obtain the dimensional information in the global coordinate system.
[0038] S103. Based on the dimensional information in the global coordinate system, establish a CTRA motion model.
[0039] S104. Use a Kalman filter based on the CTRA motion model to perform state prediction to obtain a prediction result.
[0040] S105. Based on fast hierarchical pairwise cost calculation and a dynamic association threshold, perform data association on the detection results and the prediction results to determine the matching detection results and prediction results.
[0041] S106. Use an adaptive mechanism to extend the Kalman filter to update the matching detection results and prediction results, and output the tracking results.
[0042] The fast hierarchical association and motion enhanced vehicle-mounted LiDAR multi-object tracking method of the present invention uses a CTRA motion model to predict the motion state, achieving enhancement in non-linear scenarios such as turning while ensuring accuracy; applying an adaptive mechanism to the extended Kalman filter update stage, improving the adaptive ability of the extended Kalman filter and the processing ability for model changes; based on confidence guidance, quickly calculating pairwise costs using the target distance and shape correlation attributes in the point cloud, and at the same time proposing a scene-based dynamic association threshold, better realizing data association; being able to effectively improve the robustness and adaptive ability of the 3D multi-object tracking task while ensuring real-time performance.
[0043] In another implementation manner of the present invention, the 3D object detection results include candidate target 3D bounding boxes and corresponding confidence scores; the performing data preprocessing and coordinate transformation on the detection results to obtain the dimensional information in the global coordinate system includes: quickly screening according to the confidence scores of the candidate target 3D bounding boxes, removing the candidate target 3D bounding boxes with low confidence; based on the non-maximum suppression method, removing the candidate target 3D bounding boxes with high similarity to obtain the final candidate target 3D bounding box set; using the vehicle pose information extracted from the positioning data and motion data to transform the bounding boxes in the set from the sensor coordinate system to the global coordinate system to obtain the dimensional information.
[0044] Exemplarily, the 3D object detection results include a set D of candidate target 3D bounding boxes t' and the corresponding confidence scores for subsequent screening and association. Apply a score filtering mechanism to quickly screen based on the confidence scores of candidate targets and remove candidate targets with low confidence. Use the non-maximum suppression method to remove bounding boxes with high similarity while ensuring the target recall rate, forming the final candidate target set D t As Figure 2 shown, in order to accurately estimate the motion state of the target, the vehicle pose information extracted from GPS / IMU data is used to transform the detected 3D bounding box from the sensor coordinate system to the global coordinate system.
[0045] In another implementation of the present invention, the dimension information includes the coordinate information of the target geometric center in the global coordinate system, the target length, width and height, the target speed, the target acceleration, the target orientation angle, and the target turning rate.
[0046] In another implementation of the present invention, based on the prediction result of the Kalman filter of the CTRA motion model, that is, the tracking state at time step t predicted using the CTRA model is expressed as:
[0047]
[0048] where σ represents the interval between two adjacent frames of lidar scans, which depends on the lidar frequency. For example, if the lidar is 10Hz, then σ takes the value of 0.1s; v(τ) represents the speed; θ(τ) represents the angle; w(τ) represents the turning rate; a represents the acceleration; and t represents the time step.
[0049] Exemplarily, as Figure 2 shown, the CTRA model is initialized in the global coordinate system, and the target trajectory state is expressed as [x, y, z, l, w, h, v, a, θ, ω], where (x, y, z) represents the position of the target geometric center in the global coordinate system, (l, w, h) represents the target length, width and height, v is the target speed, a is the target acceleration, θ is the target orientation angle, and ω is the target turning rate. Among them, the turning rate ω and acceleration a of the target are considered constants. For each tracking state, during the prediction process using the CTRA model, the variables z, l, w, h, a, ω remain unchanged.
[0050] The conversions of the speed, angle and turning rate of the model are shown as follows respectively:
[0051] v(τ) = v t-1 + a[τ - (t - 1)σ]
[0052] θ(τ) = θ t-1 + ω(τ)[τ - (t - 1)σ]
[0053] ω(τ) = ω(t - 1)
[0054] Among them, when the turning rate is close to 0 (ω < 0.01), taking the time interval σ as an example:
[0055] T t = f(T t-1 ) = T t-1 + [dis×cosθ, dis×sinθ, 0, 0, 0, 0, aσ, 0, ω(τ)σ, 0] T
[0056]
[0057] The proposed adaptive extended Kalman filter is used to optimize the predicted state to obtain better accuracy. The extended Kalman filter (EKF) is a non-linear filtering algorithm to solve the problem that the Kalman filtering algorithm has poor estimation effect when the state dimension is high and estimating a non-linear motion model. Based on the adaptive EKF algorithm, the motion state prediction and covariance matrix are shown as follows:
[0058]
[0059] Among them, P t-1 is the covariance matrix at the previous moment t - 1; Q is the process noise, and the process noise Q adopts an adaptive strategy:
[0060]
[0061] Among them, D calculates the variance between the predicted state and the state at the previous moment, and α AF is the adaptive factor.
[0062] In another implementation manner of the present invention, the data association of the detection result and the prediction result is performed based on the fast hierarchical pairwise cost calculation and the dynamic association threshold to determine the matching detection result and prediction result, including: As Figure 3 and Figure 4 shown, the targets in the prediction result are defined by a preset distance threshold. For the targets not exceeding the distance threshold, the pairwise cost is calculated by distance. For the targets exceeding the distance threshold, the set shape pairwise cost is calculated while calculating the distance pairwise cost, and the two costs are weighted by weights; based on the integration of multiple quantile costs of the cost data and the static threshold, the dynamic association threshold is obtained; according to the dynamic association threshold and the greedy algorithm as the association strategy, the corresponding relationship between the detection result and the prediction result is determined to obtain the matching detection result and prediction result.
[0063] In another implementation manner of the present invention, the distance pairwise cost is expressed as:
[0064]
[0065] The paired cost of the said shape is expressed as:
[0066]
[0067] The weighted total cost is expressed as:
[0068] Cost all = λ dis Cost dis + μ shape Cost shape
[0069] where N(·) is a normalization function, pose is the target center position (x, y, z), λ dis is the distance cost weight, μ shape is the shape cost weight, and score pre is the prediction confidence.
[0070] Exemplarily, as Figure 5 shown, multiple quantile costs based on cost data are integrated with a static threshold to obtain a dynamic association threshold. The recommended quantiles are the four quantiles of 5, 10, 15, and 20. The greedy algorithm is used as the association strategy to obtain the object correspondence, and the matching pairs (D t , T t ) are obtained, which respectively represent the matched detected state and the corresponding predicted state. In addition, the unmatched detected state D m and some unmatched predicted states T m are also obtained.
[0071] In another implementation manner of the present invention, after data association is completed, the extended Kalman filter with the covariance matrix adaptive mechanism added with residuals is used to update the matched detection and prediction states. The update process is expressed as the following formula:
[0072] R = (1 - α)×R + α×cov(state det - Astate pre )
[0073]
[0074] where α is a scaling factor with a value of 0.05.
[0075] In another implementation manner of the present invention, it further includes: setting a continuous missed detection threshold to distinguish the reasons for the unmatched prediction results; when the continuous missed detection threshold fails to effectively detect the target in the prediction results, it is regarded as disappearing naturally, and the current prediction results will be deleted; when the target in the prediction results reappears in the future frames, the current prediction state will continue to be retained.
[0076] Exemplarily, for two reasons of unmatched prediction states: the first is that the target normally leaves the visual range; the second is that the target is temporarily occluded and the detector misses the object. By setting a consecutive missed detection threshold N m to make a distinction, when the target state cannot be effectively detected for consecutive N m , it is regarded as natural disappearance and the prediction state will be deleted. Otherwise, once the target reappears in future frames, the prediction state will continue to be retained.
[0077] Preferably, the method of jointly detecting confidence based on consecutive hit frame counting initializes a new target as a new trajectory. At the same time, if the difference between the trajectory length and the number of consecutive losses is small, it will be judged as a false detection of the target and the trajectory will also be deleted.
[0078] In another implementation manner of the present invention, as Figure 6 shown, the multi-target tracking result is output, and the dimensions include target ID, target category, position, speed, acceleration, size, and turning rate.
[0079] On the other hand, the present invention provides a fast hierarchical association and motion enhanced vehicle-mounted LiDAR multi-target tracking system, including:
[0080] Target detection module: using a three-dimensional target detection algorithm to perform target detection on the original lidar point cloud data and positioning data, and obtaining the three-dimensional target detection result in the current frame.
[0081] Data processing module: performing data preprocessing and coordinate transformation processing on the detection result to obtain the dimensional information in the global coordinate system.
[0082] Model establishment module: establishing a CTRA motion model based on the dimensional information in the global coordinate system.
[0083] Target prediction module: using a Kalman filter based on the CTRA motion model to perform state prediction and obtaining a prediction result;
[0084] Data matching module: performing data association on the detection result and the prediction result based on fast hierarchical pairwise cost calculation and dynamic association threshold to determine the matching detection result and prediction result.
[0085] Result output module: using an adaptive mechanism extended Kalman filter to update the matching detection result and prediction result and outputting the tracking result.
[0086] The fast hierarchical association and motion enhanced vehicle-mounted LiDAR multi-target tracking system of the present invention uses the CTRA motion model to predict the motion state, achieving enhancement in non-linear scenarios such as turning while ensuring accuracy; applying an adaptive mechanism to the extended Kalman filter update stage, improving the adaptive ability of the extended Kalman filter and its processing ability for model changes; based on confidence guidance, quickly calculating pairwise costs using the target distance and shape correlation attributes in the point cloud, and at the same time proposing a scene-based dynamic association threshold to better achieve data association; capable of effectively improving the robustness and adaptive ability of the 3D multi-target tracking task while ensuring real-time performance.
[0087] On the other hand, an electronic device of the present invention includes: a processor, a memory, and a communication bus and a communications interface.
[0088] Wherein:
[0089] The processor, the memory, and the communication interface communicate with each other through the communication bus.
[0090] The communication interface is used to communicate with other electronic devices or servers.
[0091] The processor is used to execute a program, specifically, it can execute the steps of any one of the fast hierarchical association and motion enhanced vehicle-mounted LiDAR multi-target tracking methods in the above embodiments.
[0092] Specifically, the program may include program code, and the program code includes computer operation instructions.
[0093] The processor may be a central processing unit (CPU), or a specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0094] The memory is used to store the program. The memory may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.
[0095] The program can be specifically used to enable a processor to execute steps for implementing any of the fast hierarchical association and motion enhanced vehicle-mounted LiDAR multi-target tracking methods described in the embodiments. For the specific implementation of each step in the program, reference can be made to the corresponding descriptions in the steps and units executed by any of the fast hierarchical association and motion enhanced vehicle-mounted LiDAR multi-target tracking methods described above, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the devices and modules described above can refer to the corresponding process descriptions in the foregoing method embodiments.
[0096] An exemplary embodiment of the present application also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the methods of the embodiments of the present application.
[0097] The methods according to the embodiments of the present invention described above can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and to be stored in a local recording medium, so that the methods described herein can be processed by such software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as a RAM, a ROM, a flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods described herein are implemented. In addition, when a general-purpose computer accesses the code for implementing the methods shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the methods shown herein.
[0098] So far, specific embodiments of the present invention have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result.
[0099] It should be noted that all directional indications (such as up, down, left, right, back...) in the embodiments of the present invention are only used to explain the relative positional relationship between components in a specific order (as shown in the drawings). If the specific order changes, the directional indications will also change accordingly.
[0100] In the description of the present invention, the terms "first" and "second" are only used for conveniently describing different components or names, and cannot be construed as indicating or implying an order relationship, relative importance, or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features.
[0101] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used in the description of the present invention herein are for the purpose of describing specific embodiments only and are not intended to limit the present invention.
[0102] It should be noted that although the specific embodiments of the present invention have been described in detail in conjunction with the accompanying drawings, it should not be construed as a limitation on the protection scope of the present invention. Within the scope described in the claims, various modifications and variations that can be made by those skilled in the art without creative efforts still fall within the protection scope of the present invention.
[0103] The examples of the embodiments of the present invention are intended to briefly illustrate the technical features of the embodiments of the present invention, so that those skilled in the art can intuitively understand the technical features of the embodiments of the present invention, and do not serve as an improper limitation on the embodiments of the present invention.
[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A fast hierarchical association and motion enhanced vehicle-mounted LiDAR multi-object tracking method, characterized in that including: Performing object detection on the original lidar point cloud data and positioning data using a 3D object detection algorithm to obtain the 3D object detection results within the current frame; Performing data preprocessing and coordinate transformation processing on the detection results to obtain the dimensional information in the global coordinate system; Establishing a CTRA motion model based on the dimensional information in the global coordinate system; Performing state prediction using a Kalman filter based on the CTRA motion model to obtain a prediction result; Performing data association on the detection results and the prediction results based on fast hierarchical pairwise cost calculation and dynamic association threshold to determine the matching detection results and prediction results; Using an adaptive mechanism to extend the Kalman filter to update the matching detection results and prediction results and output the tracking results.
2. The method according to claim 1, wherein The 3D object detection results include candidate object 3D bounding boxes and corresponding confidence scores; The performing data preprocessing and coordinate transformation processing on the detection results to obtain the dimensional information in the global coordinate system includes: Performing fast screening according to the confidence scores of the candidate object 3D bounding boxes to remove the candidate object 3D bounding boxes with low confidence; Based on the non-maximum suppression method, removing the candidate object 3D bounding boxes with high similarity to obtain the final set of candidate object 3D bounding boxes; Using the vehicle pose information extracted from the positioning data and motion data to transform the bounding boxes in the set from the sensor coordinate system to the global coordinate system to obtain the dimensional information.
3. The method according to claim 2, wherein The dimensional information includes the coordinate information of the target geometric center in the global coordinate system, the target length, width and height, the target speed, the target acceleration, the target orientation angle, and the target turning rate.
4. The method according to claim 3, characterized in that, The prediction result of the Kalman filter based on the CTRA motion model is expressed as: where σ represents the interval between two adjacent frames of lidar scans; v(τ) represents the speed; θ(τ) represents the angle; w(τ) represents the turning rate; a represents the acceleration; and t represents the time step.
5. The method according to claim 1, wherein The performing data association on the detection results and the prediction results based on fast hierarchical pairwise cost calculation and dynamic association threshold to determine the matching detection results and prediction results includes: Defining the targets in the prediction results through a preset distance threshold. For targets not exceeding the distance threshold, calculating the pairwise cost by distance. For targets exceeding the distance threshold, calculating the pairwise cost of the set shape while calculating the pairwise cost of distance, and weighting the two costs; Integrating the multiple quantile costs of the cost data with the static threshold to obtain the dynamic association threshold; According to the dynamic association threshold and the greedy algorithm as the association strategy, determining the corresponding relationship between the detection results and the prediction results to obtain the matching detection results and prediction results.
6. The method according to claim 5, wherein The pairwise cost of distance is expressed as: The pairwise cost of shape is expressed as: The weighted total cost is expressed as: Cost all = λ dis Cost dis + μ shape Cost shape Among them, N(·) is the normalization function, pose is the target center position (x, y, z), λ dis is the distance cost weight, μ shape is the shape cost weight, score pre is the prediction confidence.
7. The method according to claim 5, characterized in that It also includes: Setting a continuous missed detection threshold to distinguish the reasons for unmatched prediction results; When the continuous missed detection threshold cannot effectively detect the targets in the prediction results, it is regarded as a natural disappearance, and the current prediction results will be deleted; When the targets in the prediction results reappear in future frames, the current prediction state will continue to be retained.
8. A fast hierarchical association and motion enhanced vehicle-mounted LiDAR multi-object tracking system, characterized in that, including: Target detection module: Using a three-dimensional target detection algorithm to perform target detection on the original lidar point cloud data and positioning data, and obtaining the three-dimensional target detection results within the current frame; Data processing module: Performing data preprocessing and coordinate transformation processing on the detection results to obtain the dimensional information in the global coordinate system; Model establishment module: Based on the dimensional information in the global coordinate system, establishing a CTRA motion model; Target prediction module: Using a Kalman filter based on the CTRA motion model to perform state prediction and obtaining the prediction results; Data matching module: Based on fast hierarchical pairwise cost calculation and dynamic association thresholds, performing data association on the detection results and the prediction results to determine the matching detection results and prediction results; Result output module: Using an adaptive mechanism extended Kalman filter to update the matching detection results and prediction results, and outputting the tracking results.
9. An electronic device, characterized in that, Including: A memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of a fast hierarchical association and motion enhanced vehicle-mounted LiDAR multi-target tracking method according to any one of claims 1 to 7 are implemented.
10. A computer storage medium, characterized in that, A computer program is stored on the computer storage medium. When the computer program is executed by a processor, the steps in a fast hierarchical association and motion enhanced vehicle-mounted LiDAR multi-target tracking method according to any one of claims 1 to 7 are implemented.