Three-dimensional multi-target tracking method based on asynchronous collaboration of multi-modal data
By using timestamp alignment and multi-stage association, the problem of asynchronous perception between LiDAR and camera was solved, achieving efficient 3D multi-target tracking and improving target tracking accuracy and adaptability in occluded scenarios.
Patent Information
- Application Number
- CN202511284059.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-12-09
AI Technical Summary
Existing multi-target tracking methods fail to effectively utilize 2D information due to the asynchronous sensing frequency between LiDAR and camera, resulting in low target tracking accuracy. Furthermore, they fail to fully leverage the complementary advantages of 2D information in target appearance and occlusion scenarios, affecting the accuracy and continuity of state estimation.
By aligning 2D camera data and LiDAR data with timestamps, the data is divided into low-frequency aligned frames and high-frequency image frames. Kalman filtering is used to extend the pure 2D observations to the 3D state space. Combined with multi-stage association strategies and detection score management, asynchronous collaborative 3D multi-target tracking is achieved.
It improves target tracking accuracy, enhances adaptability in occluded scenarios, reduces trajectory switching frequency, and improves the balance between computational efficiency and tracking accuracy.
Smart Images

Figure CN121095286A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of target tracking, in particular to a three-dimensional multi-target tracking method based on multi-modal data asynchronous cooperation. BACKGROUND
[0002] The three-dimensional multi-target tracking technology is currently widely used in automatic driving, embodied intelligence and other fields, and is an important perception technology. In the field of automatic driving and robot perception, vehicles and robots are often equipped with multiple sensors. The existing multi-target tracking technology mainly realizes target tracking through multi-sensor fusion of cameras, radars and the like. Due to the differences in the structural characteristics of cameras and laser radars, it is naturally difficult to achieve frequency synchronization. Compared with the commonly used rotating laser radar which needs to be aware of the scanning, the frequency of the camera is usually higher and the processing efficiency is faster. Therefore, how to solve the problem of asynchronous perception frequency between laser radars and cameras has become the research focus in the field.
[0003] The current three-dimensional multi-target tracking method based on multi-modal cooperation is implemented based on the multi-stage association framework of the TBD framework, including a preprocessing module, a data association module, a motion module, and a lifecycle management module. The preprocessing module is mainly for the original data of target detection to be preprocessed based on confidence or high-precision, etc. to achieve the effect of data shunting. The data association module matches the trajectories of the front and back frames of the three-dimensional target information. The motion module predicts and updates the next frame target trajectory prediction through Kalman filtering and performs trajectory matching in the data association module. The lifecycle management module aims to manage the generated trajectory throughout the process. Multi-stage association refers to three layers of data association in the data association module, including a hybrid detection box generated by fusing the 2D detection box and the 3D detection box result, projecting the 3D detection box to the 2D image plane through coordinate projection, matching the trajectories in the unified coordinate system, and matching the trajectories of the pure 3D detection box after removing the redundant detection boxes through NMS (Non-Maximum Suppression). However, this method directly discards non-key frames of 2D images, which ignores the influence of non-key frame data on the accuracy improvement of multi-target tracking, thereby failing to effectively constrain the target motion trajectory, resulting in low target tracking accuracy. At the same time, it also leads to a waste of a large amount of high-frame-rate 2D time sequence information, fails to fully utilize the complementary advantages of 2D information in target appearance and partial occlusion scenarios, restricts the accuracy and continuity of state estimation, and results in low multi-target tracking accuracy. SUMMARY
[0004] The present application aims to solve the problem of low tracking accuracy in the existing multi-target tracking method, and proposes a three-dimensional multi-target tracking method based on multi-modal data asynchronous cooperation.
[0005] A three-dimensional multi-target tracking method based on multi-modal data asynchronous cooperation, specifically:
[0006] Step one, after time stamp alignment of 2D camera data and laser radar data, the corresponding frames of the time stamp approximately aligned 2D camera data and laser radar data are taken as low-frequency alignment frames, and the corresponding frames of the redundant 2D camera data are taken as high-frequency picture frames; the laser radar data is input into a 3D detection algorithm to obtain a 3D detection result , category information and 3D detection result detection score; the 2D camera data is input into a 2D detection algorithm to obtain a 2D detection result , category information and 2D detection result detection score; according to the 2D camera data and laser radar data time stamp alignment result, the aligned 3D detection result and 2D detection result under the low-frequency alignment frame and the 2D detection result under the high-frequency picture frame are obtained ;
[0007] The 2D detection result is the position information of the 2D detection frame ; the 3D detection result is the position information of the 3D detection frame ;
[0008] Among them, is the center point coordinate of the t-th frame 2D detection frame in the image coordinate system, is the width of the t-th frame 2D detection frame, is the height of the t-th frame 2D detection frame; is the center point coordinate of the t-th frame 3D detection frame in the world coordinate system, is the width of the t-th frame 3D detection frame, is the height of the t-th frame 3D detection frame, is the length of the t-th frame 3D detection frame, is the orientation angle of the t-th frame 3D detection frame, is the velocity of the target at the t-th time;
[0009] The category information includes: pedestrian, car, truck, trailer, motorcycle, bicycle and bus;
[0010] Step two, according to the existing trajectory, the state of all trajectories in the next frame is predicted, so as to obtain a predicted trajectory, the predicted trajectory is associated with the 2D detection result and the 3D detection result obtained in step one, and an association result is obtained;
[0011] Step three, according to the association result, the predicted trajectory state is updated by Kalman, and the updated predicted trajectory state is obtained;
[0012] Step four, initializing the detection box of the failed association to a new track, then updating the detection score of the detection box of the successful association and the score of the track, and judging whether to terminate the track tracking according to the updated detection score of the detection box and the score of the track, if the track tracking is terminated, the process is ended; otherwise, return to step two for the next frame association.
[0013] Further, the step one is specifically as follows:
[0014] When the interval between the 2D camera data timestamp and the laser radar timestamp is less than the preset interval, the current 2D camera data timestamp is aligned with the current laser radar timestamp.
[0015] Further, the step two is specifically as follows:
[0016] Step two one, predicting all track states of the next frame according to the existing track, thereby obtaining a predicted track, and using Kalman prediction to represent the track state prediction process, if the current frame is a low-frequency alignment frame, step two two is executed; if the current frame is a high-frequency picture frame, step two four is executed.
[0017] Step two two, pre-associating the 2D detection result and the 3D detection result under the low-frequency alignment frame to obtain a pre-association result, and eliminating false detection from the pre-association result to obtain a false detection eliminated pre-association result.
[0018] Step two three, performing multi-stage association on the predicted track obtained in step two one and the false detection eliminated pre-association result obtained in step two two to obtain a multi-stage association result.
[0019] Step two four, performing single-stage association on the predicted track obtained in step two one and the 2D detection result under the high-frequency picture frame obtained in step one to obtain a single-stage association result.
[0020] Further, the step two one is specifically as follows:
[0021] First, based on the state of the (t-1) th frame track predicting the state of all tracks of the t th frame, specifically as follows:
[0022]
[0023]
[0024]
[0025] in, It is a motion prediction model function. These are the state parameters of the trajectory at frame t-1. These are the state parameters of the trajectory at frame t. It is the time difference between two adjacent frames. These are the position coordinates of the trajectory at frame t. It is the width of the trajectory in frame t. It is the length of the trajectory at frame t. It is the height of the trajectory at frame t. It is the orientation of the trajectory at frame t. It is the angular velocity of the trajectory at frame t. It is the resultant velocity of the trajectory at frame t. It is the resultant acceleration of the trajectory at frame t;
[0026] The motion prediction model is a constant speed model (CV), a constant acceleration model (CA), a single vehicle model (Bicycle), or a constant acceleration and rotation speed model (CTRA).
[0027] Then, obtain the trajectory of frame t based on the state of all trajectories in frame t. ;
[0028] Finally, the motion transfer equation A is obtained by using the motion prediction model to predict the trajectory state. The Kalman trajectory state prediction process is then constructed using the motion transfer equation A, specifically as follows:
[0029]
[0030]
[0031]
[0032]
[0033] in, Is the trajectory in The posterior estimated state at time 10:00. Is the trajectory in The prior estimate of the state at time t. Is the trajectory in Prior estimation of the time-varying covariance matrix. It is the equation of motion transfer. It is preset moving noise. The trajectory is Posterior estimate of the time-varying covariance matrix.
[0034] Further, the step two in the two results of 2D detection and 3D detection under the low frequency alignment frame are pre-associated to obtain a pre-association result, and the pre-association result is false detection eliminated to obtain a false detection eliminated pre-association result, specifically:
[0035] Step two one, the 2D detection result and the 3D detection result under the low frequency alignment frame are pre-associated to obtain a pre-association result, specifically:
[0036] First, the 3D detection result under the low frequency alignment frame is projected onto the camera plane, and the two-dimensional intersection union ratio of the 3D detection result projected onto the camera plane and the 2D detection result is obtained, specifically:
[0037]
[0038]
[0039] Wherein, is the two-dimensional intersection union ratio of the 3D detection result projected onto the camera plane and the 2D detection result, is the 3D detection result projected onto the camera plane, is the camera intrinsic parameter matrix, R represents the camera external rotation matrix, and T represents the camera external translation matrix;
[0040] Then, the 3D detection result and the 2D detection result are pre-associated by using the Hungarian algorithm and the two-dimensional intersection union ratio of the 3D detection result projected onto the camera plane and the 2D detection result;
[0041] The pre-association result of the 3D detection result and the 2D detection result includes: the matching pair of successful association , the unassociated 3D detection result , and the unassociated 2D detection result ;
[0042] Step two two, the unassociated 3D detection result is filtered by score and non-maximum suppression processing to obtain a false detection eliminated unassociated 3D detection result , and the matching pair of successful association , the false detection eliminated unassociated 3D detection result , and the unassociated 2D detection result are used to form a false detection eliminated pre-association result.
[0043] Further, the step two in the two results of 2D detection and 3D detection under the low frequency alignment frame are pre-associated to obtain a pre-association result, and the pre-association result is false detection eliminated to obtain a false detection eliminated pre-association result, specifically:
[0044] Step 231: Match the predicted trajectory with the successfully associated pairs. 3D detection box in The first phase of association is as follows:
[0045]
[0046]
[0047]
[0048] in, This is the trajectory of successful association in the first stage. These are the detection boxes that were successfully associated in the first phase. This is the trajectory of the first-stage association failure. This is a detection box indicating a failure in the first stage of association. It's the Hungarian algorithm. yes and Three-dimensional generalized intersection and union ratio between them yes Three-dimensional generalized intersection and union ratio between them and It is a parameter variable. yes and The minimum closed convex set between yes and The associated cost matrix, This is the first predicted trajectory. It is the nth 3D detection box. This is the m-th predicted trajectory, where m is the total number of predicted trajectories and n is the number of successfully matched pairs. The total number of 3D detection boxes in the data;
[0049] Step 232: Connect the failed correlation trajectories from the first stage with the uncorrelated 3D detection results after false detection elimination. The second phase of association is as follows:
[0050]
[0051]
[0052] in, This is the trajectory of successful association in the second stage. These are the detection boxes that were successfully associated in the second phase. This is the trajectory of the second-stage association failure. This is a detection box indicating a failed association in the second stage. yes and a cost matrix associated with is the jth track of the first stage association failure, is the kth unassociated 3D bounding box after false positive elimination, is the total number of tracks of the first stage association failure, is the total number of unassociated 3D bounding boxes after false positive elimination;
[0053] Step two three, the tracks of the second stage association failure and the unassociated 2D detection results perform the third stage association to obtain the multi-stage association result, specifically:
[0054]
[0055]
[0056] wherein, is the track of the third stage association success, is the bounding box of the third stage association success, is the track of the third stage association failure, is the bounding box of the third stage association failure, is and a cost matrix associated with is the jth track of the second stage association failure, is the kth unassociated 2D bounding box, is the total number of tracks of the second stage association failure, is the total number of unassociated 2D bounding boxes, is the total number of tracks of the second stage association failure, is the total number of unassociated 2D bounding boxes, is and the 2D intersection over union.
[0057] Further, the step two four, the predicted track obtained in step two one is associated with the 2D detection result under the high frequency picture frame obtained in step one to obtain the single stage association result, specifically:
[0058]
[0059]
[0060] wherein, is the track of the single stage association success, is the bounding box of the single stage association success, is the track of the single stage association failure, is the bounding box of the single stage association failure, is with a cost matrix associated with is the th predicted trajectory, is the th two-dimensional bounding box under the high-frequency picture frame, is the total number of two-dimensional bounding boxes under the high-frequency picture frame.
[0061] Further, the step three is to update the predicted trajectory state according to the association result, and obtain an updated predicted trajectory state, specifically:
[0062] Step three, if the current frame is a low-frequency alignment frame, step three two is executed; if the current frame is a high-frequency picture frame, step three three is executed;
[0063] Step three two, obtain an observation matrix according to the result of multi-stage association, and update the trajectory successfully associated in the multi-stage association by using the observation matrix, to obtain an updated trajectory state, specifically:
[0064] Step three two one, obtain the observation matrix H of the first stage association and the second stage association, and update the trajectory successfully associated in the first stage and the second stage by using the observation matrix H, to obtain an updated trajectory state:
[0065]
[0066]
[0067]
[0068]
[0069] wherein, is an observation matrix, is an identity matrix, and K is a Kalman gain, is a priori estimation of the covariance matrix of the trajectory at time is a preset observation noise, is a posteriori estimation of the trajectory at time is a priori estimation of the trajectory at time is an observation state of the trajectory at the is a posteriori estimation of the covariance matrix of the trajectory at time is a priori estimation of the covariance matrix of the trajectory at time
[0070] When performing Kalman update on the track successfully associated in the first stage and the second stage ;
[0071] Step three, obtaining the observation matrix of the third stage association , and using the observation matrix to perform Kalman update on the track successfully associated in the third stage association, to obtain the updated track state, specifically:
[0072]
[0073]
[0074]
[0075]
[0076] wherein, is the observation state of the track at the first time association;
[0077] When performing Kalman update on the track successfully associated in the third stage association ;
[0078] wherein, is the 2D detection box position information under the low-frequency alignment frame;
[0079] Step three, using the track successfully associated to perform two-dimensional Kalman update to obtain the updated predicted track state.
[0080] Further, the step three in the method for using the track successfully associated to perform two-dimensional Kalman update to obtain the updated predicted track state, specifically:
[0081]
[0082]
[0083]
[0084]
[0085] wherein, is the observation state of the track at the first time association;
[0086] When performing Kalman update on the track successfully associated in the single stage ;
[0087] wherein, is the 2D detection box position information under the high-frequency picture frame.
[0088] Further, the step four is to initialize the detection frame of the failed association to a new track, then update the detection score of the detection frame of the successful association and the score of the track, and determine whether to terminate the track tracking according to the updated detection score of the detection frame and the score of the track, if yes, end; otherwise, return to step two for the next frame association, specifically:
[0089] Step four, if the current frame is a low-frequency alignment frame, step four two is executed; if the current frame is a high-frequency picture frame, step four four is executed;
[0090] Step four two, the detection frame of the failed association in the first stage and the second stage is initialized to a new track, and the initial score of the new track is obtained, specifically:
[0091]
[0092] wherein, is the detection score of the three-dimensional detection frame of the failed association in the first stage and the second stage, is the detection score of the two-dimensional detection frame corresponding to the three-dimensional detection frame of the failed association, is the initial score of the new track;
[0093] wherein, the two-dimensional detection frame corresponding to the three-dimensional detection frame of the failed association in the first stage is the two-dimensional detection frame in the first stage, and the three-dimensional detection frame of the failed association in the second stage has no corresponding two-dimensional detection frame, then is 0;
[0094] Step four three, update the score of the track of the successful association in the low-frequency alignment frame and the detection score of the detection frame, and determine whether to terminate the estimation tracking, if yes, end; otherwise, let t=t+1, and return to step two, specifically:
[0095] Step four three, update the score of the track of the successful association in the low-frequency alignment frame and the detection score of the detection frame, and determine whether to terminate the estimation tracking, if yes, end; otherwise, let t=t+1, and return to step two, specifically:
[0096]
[0097] wherein, is the score of the track of the t-1 frame, is the score of the track of the t frame, 0.7, is the decay rate;
[0098] Step four three two, update the detection score of the detection frame of the successful association in the low-frequency alignment frame, specifically:
[0099] For the first stage association and the second stage association, the detection score of the detection frame is specifically:
[0100]
[0101]
[0102] wherein, is the detection score of the detection frame under the low-frequency alignment frame, is the detection score weight of the 3D detection frame, is the detection score of the 3D detection frame associated successfully, is the detection score weight of the 2D detection frame, is the detection score of the 2D detection frame associated successfully;
[0103] For the third stage association, the detection score of the detection frame is updated by the following way:
[0104]
[0105] wherein, is the detection score of the 2D detection frame under the low-frequency alignment frame associated successfully in the third stage, is the detection score of the 2D detection frame under the low-frequency alignment frame associated successfully in the third stage after updating, is the decay coefficient;
[0106] Step four three, the updated track score and the updated detection frame detection score are fused to obtain the fusion detection score of the current frame, specifically:
[0107]
[0108]
[0109] wherein, is the detection score of the detection frame of the current frame, taken or ;
[0110] Step four three four, the average value of the fusion detection scores from the first frame to the current frame is obtained, if the average value of the fusion detection scores is less than the deletion threshold or the number of continuous unmatched frames is greater than the preset threshold max-age, the track tracking is terminated; otherwise, t=t+1, and the step two is returned;
[0111] Step four four, the track score and the detection frame detection score under the high-frequency picture frame are updated, and it is judged whether to terminate the estimation tracking, if the track tracking is terminated, the process is ended; otherwise, t=t+1, and the step two is returned, specifically:
[0112] Step four four one, the track score under the high-frequency picture frame is updated, specifically:
[0113]
[0114] Step four four two, match the updated track under the high-frequency picture frame with the 2D detection frame under the high-frequency picture frame, if the matching is unsuccessful, discard the 2D detection frame under the high-frequency picture frame, if the matching is successful, obtain the fusion score of the matched track under the high-frequency picture frame, specifically:
[0115]
[0116]
[0117] Wherein, is the detection score of the decayed 2D detection frame, is the detection score of the 2D detection frame under the high-frequency picture frame, is the decay weight;
[0118] Step four four three, obtain the average value of the fusion score of the track from the first frame to the current frame under the high-frequency picture frame, if the average value of the fusion detection score is less than the deletion threshold Or the number of continuous unmatched frames is greater than the preset threshold max-age, terminate the track tracking, otherwise, let t=t+1, and return to step two.
[0119] The beneficial effects of the present application are:
[0120] The present application effectively utilizes pure 2D observation by extending Kalman filter (EKF), projects 2D observation to 3D state space, realizes state update, so as to achieve deep fusion of camera and laser radar. Meanwhile, the present application fully utilizes the detection score obtained from the detection algorithm to manage the track score, which can more flexibly manage the track. After introducing the high-frequency 2D detection data of the non-aligned frame, the present application reduces the frequency of the track switching phenomenon of the track under the occlusion scene. The high-frequency data compensates the uncertainty caused by the sensor asynchrony through the high decay factor strategy, enhances the adaptability of the tracker to the mismatching scene such as occlusion and short-time loss, verifies the effective utilization mechanism of the high-frequency data in the asynchronous fusion framework, enhances the robustness of the asynchronous scene tracking, and improves the accuracy of the target tracking. The present application combines 3D detection and 2D monitoring information to construct the cost matrix of motion information and expression information, constitutes a three-stage association strategy, reduces redundant calculation through hierarchical matching, and realizes timely elimination of invalid tracks through dynamic adjustment of the track score, so as to balance the calculation efficiency and tracking accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0121] Figure 1 It is an asynchronous tracking mode diagram;
[0122] Figure 2 It is an aligned frame tracking framework diagram;
[0123] Figure 3 Frame tracking framework for non-aligned frames. DETAILED DESCRIPTION
[0124] Specific implementation one: as shown in the figure, the specific process of the three-dimensional multi-target tracking method based on multi-modal data asynchronous collaboration in the embodiment is as follows: Figure 1
[0125] Step one, after time stamping approximation alignment of high-frequency 2D camera data and low-frequency lidar data, the corresponding frames of the time-stamped approximation-aligned high-frequency 2D camera data and low-frequency lidar data are taken as low-frequency alignment frames, and the corresponding frames of the redundant high-frequency 2D camera data are taken as high-frequency picture frames; then the low-frequency lidar data is input into a 3D detection algorithm to obtain 3D detection results , category information and detection scores of 3D detection results, the high-frequency 2D camera data is input into a 2D detection algorithm to obtain 2D detection results , category information and detection scores of 2D detection results; according to the time stamping alignment results of high-frequency 2D camera data and low-frequency lidar data, the aligned 3D detection results and 2D detection results under low-frequency alignment frames and the 2D detection results under high-frequency picture frames are obtained ;
[0126] The time stamping approximation alignment of high-frequency 2D camera data and low-frequency lidar data is specifically as follows:
[0127] When the interval between the 2D camera data time stamp and the lidar time stamp is less than the preset interval 10 ms, the 2D camera data time stamp and the lidar time stamp are aligned, and it is considered that this group of data is perceived at the same time;
[0128] The 2D detection result is the position information of the 2D detection frame ;
[0129] The 2D detection result under the low-frequency alignment frame is denoted as , and the 2D detection result under the high-frequency picture frame is denoted as ;
[0130] wherein, is the center point coordinate of the 2D detection frame in the t-th frame in the image coordinate system, is the width of the 2D detection frame in the t-th frame, is the height of the 2D detection frame in the t-th frame, is the center point coordinate of the 2D detection frame in the t-th frame under the low-frequency alignment frame, is the width of the 2D detection frame in the t-th frame under the low-frequency alignment frame, is the height of the 2D detection frame in the t-th frame under the low-frequency alignment frame, is the center point coordinate of the 2D detection frame under the high-frequency alignment frame, is the width of the 2D detection frame under the high-frequency alignment frame, is the height of the 2D detection frame under the high-frequency alignment frame;
[0131] The 3D detection result is the position information of the 3D detection frame ;
[0132] wherein, is the center point coordinate of the 3D detection frame under the world coordinate system, is the width of the 3D detection frame, is the height of the 3D detection frame, is the length of the 3D detection frame, is the orientation angle of the 3D detection frame, is the speed of the target at the tth moment.
[0133] The category information includes pedestrians, cars, trucks, trailers, motorcycles, bicycles and buses;
[0134] In this step, due to the high sensing frequency of the camera and the low frequency of the laser radar, the two are approximately time-stamped. The three-dimensional detection algorithm of the present application adopts Centerpoint, and the two-dimensional detection algorithm adopts Cascade RCNN. In the present application, one frame corresponds to one moment.
[0135] Step two, according to the existing trajectory, predict the state of all trajectories in the next frame, so as to obtain the predicted trajectory, and use the predicted trajectory to associate with the 2D detection result and the 3D detection result obtained in step one, to obtain the association result, which is specifically:
[0136] Step two, according to the existing trajectory, predict the state of all trajectories in the next frame, so as to obtain the predicted trajectory, and use the predicted trajectory to associate with the 2D detection result and the 3D detection result obtained in step one, to obtain the association result, which is specifically:
[0137] The existing trajectory is used to predict the state of all trajectories in the next frame, so as to obtain the predicted trajectory, and the Kalman prediction is used to represent the trajectory state prediction process, which is specifically:
[0138] First, based on the state of the t-1th frame trajectory itself , the time difference between two frames and the different motion prediction models for different categories of trajectories are adapted, so as to predict the state of all trajectories in the tth frame, which is specifically:
[0139]
[0140]
[0141]
[0142] wherein, is a motion prediction model function, is a state parameter of the trajectory at the t-1th frame, is a state parameter of the trajectory at the tth frame, is a time difference between adjacent two frames, is a position coordinate of the trajectory at the tth frame, is a width of the trajectory at the tth frame, is a length of the trajectory at the tth frame, is a height of the trajectory at the tth frame, is an orientation of the trajectory at the tth frame, is an angular velocity of the trajectory at the tth frame, is a resultant velocity of the trajectory at the tth frame, is a resultant acceleration of the trajectory at the tth frame;
[0143] the motion prediction model is a constant velocity model (CV), a constant acceleration model (CA), a bicycle model (Bicycle) or a constant acceleration and turning rate model (CTRA);
[0144] the motion prediction model is selected according to a category of an existing trajectory, and specifically,
[0145] if the existing trajectory is a pedestrian trajectory, the constant acceleration and turning rate model (CTRA) is selected;
[0146] if the existing trajectory is a car trajectory, the constant acceleration and turning rate model (CTRA) is selected;
[0147] if the existing trajectory is a trailer trajectory, the constant acceleration model (CA) is selected;
[0148] if the existing trajectory is a motorcycle trajectory, the bicycle model (Bicycle) is selected;
[0149] if the existing trajectory is a bicycle trajectory, the bicycle model (Bicycle) is selected;
[0150] if the existing trajectory is a bus trajectory, the constant acceleration model (CA) is selected;
[0151] then, the trajectory at the tth frame is obtained according to states of all trajectories at the tth frame ;
[0152] Finally, the process of predicting the trajectory state by using the motion prediction model obtains the motion transition equation A, takes the state of the t-th frame trajectory itself as the trajectory estimation state at the t time in the Kalman prediction, and constructs the Kalman trajectory state prediction process by using A and the trajectory estimation state at the t time, which is specifically:
[0153]
[0154]
[0155]
[0156]
[0157] wherein, is the trajectory posterior estimation state at the t time, is the trajectory prior estimation state at the t time, represents the prior estimation of the covariance matrix of the trajectory at the t time, is the motion transition equation, is the preset movement noise, represents the posterior estimation of the covariance matrix of the trajectory at the t time.
[0158] wherein, the initial covariance matrix is a preset quantity.
[0159] Step two, pre-correlate the 2D detection result and the 3D detection result under the low-frequency alignment frame to obtain a pre-correlation result, and eliminate false detection from the pre-correlation result to obtain a false detection-eliminated pre-correlation result, which is specifically:
[0160] Step two, pre-correlate the 2D detection result and the 3D detection result under the low-frequency alignment frame to obtain a pre-correlation result, which is specifically:
[0161] First, project the 3D detection result under the low-frequency alignment frame to the camera plane, and obtain the two-dimensional intersection union ratio of the 3D detection result projected to the camera plane and the 2D detection result, which is specifically:
[0162]
[0163]
[0164] wherein, is the two-dimensional intersection union ratio of the 3D detection result projected to the camera plane and the 2D detection result, is the 3D detection result projected on the camera plane, is a camera intrinsic matrix, R represents a camera extrinsic rotation matrix, and T represents a camera extrinsic translation matrix;
[0165] Then, a pre-association result of the 3D detection result and the 2D detection result is obtained by using a Hungarian algorithm and a two-dimensional intersection over union of the 3D detection result projected to a camera plane and the 2D detection result;
[0166] The pre-association result of the 3D detection result and the 2D detection result includes a matched pair of successful association a 3D detection result of unassociated a 2D detection result of unassociated .
[0167] In this step, under the low-frequency alignment frame, the detection time of the 3D detection and the 2D detection is the same, and the 3D detection result obtained by the laser radar needs to be associated with the 2D detection result on the camera plane to obtain observation results of different modalities, which are used for association with the subsequent trajectory.
[0168] Step two two, 3D detection result of unassociated Fractional filtering and non-maximum suppression (NMS) processing are performed to obtain the 3D detection result of unassociated after false detection elimination the matched pair of successful association the 3D detection result of unassociated after false detection elimination the 2D detection result of unassociated to form a pre-association result after false detection elimination.
[0169] In this step, since there is no sensing information of the laser radar, only the 2D detection result can be obtained, and therefore the 2D detection result under the high-frequency frame is not processed and is directly used for subsequent trajectory association.
[0170] Step two three, the predicted trajectory obtained in step two one is associated with the pre-association result after false detection elimination obtained in step two two in multiple stages to obtain a multiple-stage association result, and the specific process is as follows:
[0171] Step two three one, the predicted trajectory is associated with the three-dimensional detection box in the matched pair of successful association in a first stage, assuming that there are predicted trajectories and n detection boxes, and the specific matching process is as follows:
[0172]
[0173]
[0174]
[0175] wherein, This is the trajectory of successful association in the first stage. These are the detection boxes that were successfully associated in the first phase. This is the trajectory of the first-stage association failure. This is a detection box indicating a failure in the first stage of association. It's the Hungarian algorithm. yes and Three-dimensional generalized intersection and union ratio between them yes and Three-dimensional generalized intersection and union ratio between them and It is a parameter variable. yes and The minimum closed convex set between yes and The associated cost matrix, This is the first predicted trajectory. It is the nth 3D detection box. This is the m-th predicted trajectory, where m is the total number of predicted estimates and n is the number of successfully matched pairs. The total number of 3D detection boxes in the data;
[0176] Step 232: Connect the failed correlation trajectories from the first stage with the uncorrelated 3D detection results after false detection elimination. Perform the second stage of association, assuming the number of failed association trajectories is... The number of unassociated 3D detection boxes after false positive elimination is The association process is as follows:
[0177]
[0178]
[0179] in, This is the trajectory of successful association in the second stage. These are the detection boxes that were successfully associated in the second phase. This is the trajectory of the second-stage association failure. This is a detection box indicating a failed association in the second stage. yes and The associated cost matrix, It is the j-th trajectory where the first-stage association failed. It is the unassociated 3D detection box after the k-th false positive elimination. It is the total number of trajectories that failed to associate in the first stage. It represents the total number of unassociated 3D bounding boxes after false positives have been eliminated;
[0180] Step 233: Connect the failed trajectories from the second stage with the unconnected 2D detection results. Perform a third-stage association to obtain multi-stage association results. Assume the number of failed trajectories in the second-stage association is [number missing]. The number of observations containing only two-dimensional detection is The association process is as follows:
[0181]
[0182]
[0183] in, This is the trajectory of successful association in the third stage. These are the detection boxes that were successfully associated in the third stage. This is the trajectory of the third-stage association failure. This is a detection box indicating a failed association in the third stage. yes and The associated cost matrix, This is the second phase of the association failure. A trajectory, It is the first An unassociated two-dimensional detection box, It is the total number of trajectories that failed to associate in the second phase. It is the total number of associated 2D detection boxes. yes and The two-dimensional intersection and union ratio.
[0184] Step 24: Perform a single-stage association between the predicted trajectories obtained in Step 21 and the 2D detection results under the high-frequency image frames obtained in Step 1 to obtain the single-stage association result. Assume the number of predicted trajectories is... The number of two-dimensional detection boxes is The association process is as follows:
[0185]
[0186]
[0187] in, It is a single-stage associated success trajectory. It is a detection box for successful single-stage association. It is the trajectory of a single-stage association failure. It is a detection box for single-stage association failure. yes and a cost matrix associated with the association, is the predicted trajectory, is the two-dimensional detection frame under the high-frequency picture frame, is the total number of two-dimensional detection frames under the high-frequency picture frame.
[0188] Step three, Kalman update is performed on the predicted trajectory state according to the association result, and an updated predicted trajectory state is obtained, specifically as follows:
[0189] Step three, if the current frame is a low-frequency alignment frame, step three two is executed; if the current frame is a high-frequency picture frame, step three three is executed.
[0190] Step three two, an observation matrix is obtained according to the result of multi-stage association, and Kalman update is performed on the trajectory successfully associated in the multi-stage association by using the observation matrix, so as to obtain an updated trajectory state, specifically as follows:
[0191] Step three two one, an observation matrix H of the first-stage association and the second-stage association is obtained, and Kalman update is performed on the trajectory successfully associated in the first-stage association and the second-stage association by using the observation matrix H, so as to obtain an updated predicted trajectory state:
[0192]
[0193]
[0194]
[0195]
[0196] wherein, is an observation matrix, is an identity matrix, K denotes a Kalman gain, denotes a priori estimation of a covariance matrix of the trajectory at a time, is a preset observation noise, denotes a posteriori estimation state of the trajectory at a time, is a priori estimation state of the trajectory at a time, is an observation state of the trajectory at a time, is a posteriori estimation of a covariance matrix of the trajectory at a time, is a priori estimation of a covariance matrix of the trajectory at a time;
[0197] When Kalman updating the track of the first stage and the second stage association success ;
[0198] Step three, obtain the third stage association observation matrix , and use the observation matrix to perform Kalman updating on the track of the third stage association success, specifically:
[0199] The observation matrix of the third stage association is obtained in the following way:
[0200] When there is only 2D observation update (non-linear projection), the perspective projection function is constructed using the predicted track state:
[0201]
[0202] Wherein, is the camera intrinsic parameter, , is the size scaling factor, is the projection function;
[0203] Then, the first-order Taylor expansion linearizes the nonlinear relationship through the projection function, and the Jacobian matrix is constructed as the observation matrix :
[0204]
[0205] Finally, Kalman updating is performed using the observation matrix to obtain the updated track state, specifically:
[0206]
[0207]
[0208]
[0209] Wherein, is the observation state of the track at the time association;
[0210] Step three, use the associated track to perform two-dimensional Kalman updating to obtain the updated predicted track state, specifically:
[0211]
[0212]
[0213]
[0214]
[0215] wherein, is the observation state of the trajectory on the association at the time instant.
[0216] In this step, under the low-frequency alignment frame, the possible observation results are: the first-stage association successful observation: , containing three-dimensional bounding boxes and two-dimensional bounding boxes; the second-stage association successful observation: , containing only three-dimensional bounding boxes; and the third-stage association successful observation: , containing only two-dimensional bounding boxes. Based on these association results, the Kalman update strategy is adopted. For the association results with three-dimensional observation information, the update is directly based on the three-dimensional bounding box attribute, because the observation result is three-dimensional, and it has a one-to-one mapping relationship with the state space, so the observation matrix H can be directly obtained according to the observation result. When there is only a 2D bounding box, the 2D bounding box is used for updating, but at this time and the state space do not belong to one-to-one mapping, so the 2D bounding box cannot be directly used for updating, therefore, the projection matrix is used to update the Jacobian function H, which represents the conversion process of the 2D attribute to , and the calculation formula is expressed as: The perception results based on high and low frequency inputs actually cause the observation obtained in the update step to have multiple modalities, therefore, the application designs a hybrid Kalman update mechanism suitable for multiple modalities.
[0217] Step five, initializing the detection boxes that fail to be associated into new trajectories, then updating the detection scores of the associated detection boxes and the scores of the trajectories, and judging whether to terminate the trajectory tracking according to the updated detection scores of the detection boxes and the scores of the trajectories, if yes, ending; otherwise, returning to step two for the next frame association, specifically:
[0218] Step four, if the current frame is a low-frequency alignment frame, step four two is executed; if the current frame is a high-frequency picture frame, step four four is executed.
[0219] Step four two, initializing the detection boxes that fail to be associated with the predicted trajectories in the first stage and the second stage into a new trajectory, and obtaining the initial score of the new trajectory, specifically:
[0220]
[0221] wherein, is the detection score of the three-dimensional detection box that fails to be associated in the first stage and the second stage, is the detection score of the two-dimensional detection box corresponding to the three-dimensional detection box that fails to be associated. is the initial score of the new track;
[0222] wherein the two-dimensional detection frame corresponding to the three-dimensional detection frame of the first stage association failure is the two-dimensional detection frame in , the three-dimensional detection frame of the second stage association failure has no corresponding two-dimensional detection frame, then the default is 0; the two-dimensional detection frame of the third stage association failure is directly discarded and will not be initialized as a new track.
[0223] Step four three, update the score of the track and the detection score of the detection frame associated successfully under the low-frequency alignment frame, and judge whether to terminate the estimation tracking, if the track tracking is terminated, end; otherwise, let t = t + 1, and return to step two, specifically:
[0224] Step four three one, update the score of the track according to the exponential decay model, specifically:
[0225]
[0226] wherein, is the score of the track of the t-1 frame, is the score of the track of the t frame, 0.7, is the decay rate;
[0227] In this step, the track score is the detection score of the detection frame when setting a new track; the score of the historical track is weakened through the decay factor to avoid the problem of excessive trust of the track which has not been updated for a long time.
[0228] Step four three two, update the detection score of the detection frame associated successfully under the low-frequency alignment frame, specifically:
[0229] For the first stage association and the second stage association, the detection score of the detection frame is specifically:
[0230]
[0231]
[0232] wherein, is the detection score of the detection frame associated successfully under the low-frequency alignment frame, is the detection score weight of the 3D detection frame, is the detection score of the 3D detection frame associated successfully, is the detection score weight of the 2D detection frame, is the detection score of the 2D detection frame associated successfully;
[0233] In this step, the detection scores of the 3D detection frame and the 2D detection frame are dynamically adjusted according to the sensor reliability;
[0234] For the third stage association, the detection score of the detection box is updated by the following way:
[0235]
[0236] wherein, is the detection score of the two-dimensional detection box under the low-frequency alignment frame of the third stage association success, is the detection score of the two-dimensional detection box under the low-frequency alignment frame of the third stage association success after updating, is the attenuation coefficient;
[0237] In this step, in the single-mode detection, the detection score of the detection box is subjected to attenuation processing, which is used to compensate the uncertainty of the single-mode detection.
[0238] Step four three, the updated track score and the updated detection score of the detection box are fused to obtain the fusion detection score of the current frame, specifically:
[0239]
[0240]
[0241] wherein, is the detection score of the detection box of the current frame, which is taken according to the detection type , ;
[0242] In this step, the historical detection score and the current detection score are fused in the form of product, so as to avoid the overconfidence problem caused by updating.
[0243] Step four three four, the average value of the fusion detection scores from the first frame to the current frame is obtained, if the average value of the fusion detection scores is less than the deletion threshold or the number of continuous unmatched frames is greater than the preset threshold max-age, the track tracking is terminated; otherwise, t=t+1, and the step two is returned.
[0244] This step can improve the robustness of the track to occlusion and false detection.
[0245] Step four four, the score of the track under the high-frequency picture frame and the detection score of the detection box are updated, and it is judged whether to terminate the estimation tracking, if the track tracking is terminated, the process is ended; otherwise, t=t+1, and the step two is returned, specifically:
[0246] Step four four one, the score of the track under the high-frequency picture frame is updated by the following formula, specifically:
[0247]
[0248] wherein, is the track score of the high-frequency picture frame under t-1 frame, is the decay factor, is the updated track score of the high-frequency picture frame under t-1 frame;
[0249] Step four four two, match the updated track under the high-frequency picture frame with the 2D detection frame under the high-frequency picture frame (based on IoU or Euclidean distance), if the matching is unsuccessful, discard the 2D detection frame under the high-frequency picture frame; if the matching is successful, obtain the fusion score of the matched track under the high-frequency picture frame, specifically:
[0250]
[0251]
[0252] wherein, is the result of the 2D detection score after special attenuation for non-aligned scenes, is the detection score of the 2D detection frame under the high-frequency picture frame, is the special attenuation weight for non-aligned scenes, artificially set to 0.3;
[0253] This step ensures the rationality of the track score in the asynchronous scene through the reinforcement of the attenuation and matching mechanism, and avoids the accumulation of errors caused by mis-matching.
[0254] Step four four three, obtain the average value of the fusion score of the track from the first frame to the current frame under the high-frequency picture frame, if the average value of the fusion detection score is less than the deletion threshold or the number of consecutive unmatched frames is greater than the preset threshold max-age, terminate the track tracking; otherwise, let t=t+1, and return to step two.
[0255] In this step, a life cycle management method of multi-modal score fusion is designed, and the core goal is to realize robust tracking of dynamic targets and timely elimination of invalid tracks by dynamically adjusting the track score (confidence score). This step designs differentiated score updating rules for aligned frames (sensor timestamp synchronization) and non-aligned frames (sensor timestamp asynchronous), to ensure a balance between calculation efficiency and tracking accuracy.
[0256] The application designs differentiated trajectory score attenuation rules for synchronous / non-aligned frames, compensates for asynchronous uncertainty through an attenuation factor, dynamically updates trajectories using high-frequency 2D detection, protects non-key frames by interpolating between two frames to input the motion model for prediction and performing projection correlation on the intermediate result, thereby improving the tracking robustness of the occlusion scene. The application is based on extended Kalman filtering (EKF) and effectively utilizes pure 2D observation. Through a nonlinear projection model, 2D observation is associated with a 3D state space to update state estimation, thereby significantly enhancing the tightness of multi-modal fusion. The strategy of multi-stage data association in the aligned frame / non-aligned frame state of the application uses composite modal, pure 3D modal and pure 2D modal detection information to construct a cost matrix in sequence in order to minimize matching failures, and combines Kalman filter prediction with the Hungarian algorithm to realize accurate trajectory matching in all scenes. The aligned frame tracking process is as shown in Figure 2 , and the non-aligned frame tracking process is as shown in Figure 3 .
Claims
1. A three-dimensional multi-target tracking method based on asynchronous multimodal data collaboration, characterized in that... The specific process of the method is as follows: Step 1: After aligning the timestamps of the 2D camera data and the LiDAR data, take the frames corresponding to the approximately timestamp-aligned 2D camera data and LiDAR data as low-frequency aligned frames, and take the frames corresponding to the excess 2D camera data as high-frequency image frames. Input LiDAR data into a 3D detection algorithm to obtain 3D detection results. The system calculates the detection score based on category information and 3D detection results; it also inputs 2D camera data into a 2D detection algorithm to obtain 2D detection results. The system calculates the detection score based on category information and 2D detection results; it also obtains the 3D and 2D detection results for low-frequency aligned frames and the 2D detection results for high-frequency image frames based on the timestamp alignment results of 2D camera data and LiDAR data. ; The 2D detection result is the position information of the 2D detection box. The 3D detection result is the position information of the 3D detection box. ; in, These are the coordinates of the center point of the 2D detection box in the t-th frame in the image coordinate system. It is the width of the 2D detection box in frame t. It is the height of the 2D detection box in frame t; These are the coordinates of the center point of the 3D detection box in frame t of the world coordinate system. It is the width of the 3D detection box in frame t. It is the height of the 3D detection box in frame t. It is the length of the 3D detection box in frame t. It is the orientation angle of the 3D detection box in frame t. It is the velocity of the target at time t; The category information includes: pedestrians, cars, trucks, trailers, motorcycles, bicycles, and buses; Step 2: Predict the state of all trajectories in the next frame based on the existing trajectories to obtain the predicted trajectories. Then, associate the predicted trajectories with the 2D and 3D detection results obtained in Step 1 to obtain the association results. Step 3: Perform Kalman update on the predicted trajectory state based on the association results to obtain the updated predicted trajectory state; Step 4: Initialize the failed detection boxes as new trajectories, then update the detection scores of the successfully associated detection boxes and the scores of the trajectories. Based on the updated detection scores of the detection boxes and the scores of the trajectories, determine whether to terminate trajectory tracking. If trajectory tracking is terminated, the process ends; otherwise, return to Step 2 to perform the association for the next frame.
2. The three-dimensional multi-target tracking method based on asynchronous multimodal data collaboration according to claim 1, characterized in that: The time-stamp alignment of the 2D camera data and LiDAR data in step one specifically involves: When the interval between the 2D camera data timestamp and the LiDAR timestamp is less than a preset interval, the current 2D camera data timestamp and the current LiDAR timestamp will be aligned.
3. The three-dimensional multi-target tracking method based on asynchronous multimodal data collaboration according to claim 2, characterized in that: In step two, the predicted trajectory is obtained by predicting the state of all trajectories in the next frame based on the existing trajectories. The predicted trajectory is then correlated with the 2D and 3D detection results obtained in step one to obtain the correlation result. Specifically: Step 21: Predict the state of all trajectories in the next frame based on the existing trajectories to obtain the predicted trajectories. Use Kalman scattering to represent the trajectory state prediction process. If the current frame is a low-frequency aligned frame, proceed to Step 22; if the current frame is a high-frequency image frame, proceed to Step 24. Step 22: Perform pre-association on the 2D and 3D detection results under the low-frequency aligned frame to obtain the pre-association result, and perform false detection elimination on the pre-association result to obtain the false detection eliminated pre-association result; Step 23: Perform multi-stage association between the predicted trajectory obtained in Step 21 and the pre-association result after false detection elimination obtained in Step 22 to obtain multi-stage association results; Step 24: Perform a single-stage association between the predicted trajectory obtained in Step 21 and the 2D detection results under the high-frequency image frames obtained in Step 1 to obtain the single-stage association result.
4. The three-dimensional multi-target tracking method based on asynchronous multimodal data collaboration according to claim 3, characterized in that: Step 2.1, which involves predicting the states of all trajectories in the next frame based on existing trajectories to obtain predicted trajectories, and using Kalman prediction to represent the trajectory state prediction process, specifically includes: First, based on the state of the trajectory in frame t-1. Predict the state of all trajectories in frame t, specifically as follows: in, It is a motion prediction model function. These are the state parameters of the trajectory at frame t-1. These are the state parameters of the trajectory at frame t. It is the time difference between two adjacent frames. These are the position coordinates of the trajectory at frame t. It is the width of the trajectory in frame t. It is the length of the trajectory at frame t. It is the height of the trajectory at frame t. It is the orientation of the trajectory at frame t. It is the angular velocity of the trajectory at frame t. It is the resultant velocity of the trajectory at frame t. It is the resultant acceleration of the trajectory at frame t; The motion prediction model is a constant speed model (CV), a constant acceleration model (CA), a single vehicle model (Bicycle), or a constant acceleration and rotation speed model (CTRA). Then, obtain the trajectory of frame t based on the state of all trajectories in frame t. ; Finally, the motion transfer equation A is obtained by using the motion prediction model to predict the trajectory state. The Kalman trajectory state prediction process is then constructed using the motion transfer equation A, specifically as follows: in, Is the trajectory in The posterior estimated state at time 10:
00. Is the trajectory in The prior estimate of the state at time t. Is the trajectory in Prior estimation of the time-varying covariance matrix. It is the equation of motion transfer. It is preset moving noise. The trajectory is Posterior estimate of the time-varying covariance matrix.
5. A three-dimensional multi-target tracking method based on asynchronous multimodal data collaboration according to claim 4, characterized in that: In step two, the 2D and 3D detection results under the low-frequency aligned frame are pre-correlated to obtain pre-correlated results. Then, false detection elimination is performed on the pre-correlated results to obtain false detection eliminated pre-correlated results. Specifically: Step 221: Perform pre-association on the 2D and 3D detection results under the low-frequency aligned frame to obtain the pre-association result, specifically as follows: First, the 3D detection results from the low-frequency aligned frames are projected onto the camera plane, and the two-dimensional intersection-over-union ratio (IoU) of the 3D detection results projected onto the camera plane and the 2D detection results is obtained, specifically: in, It is the two-dimensional intersection-over-union ratio of the 3D detection results projected onto the camera plane and the 2D detection results. It is the 3D detection result projected onto the camera plane. R represents the camera intrinsic parameter matrix, R represents the camera extrinsic parameter rotation matrix, and T represents the camera extrinsic parameter translation matrix. Then, using the Hungarian algorithm and the two-dimensional intersection-union ratio of the 3D detection results projected onto the camera plane and the 2D detection results, the pre-association results of the 3D detection results and the 2D detection results are obtained. The pre-association results of the 3D detection results and 2D detection results include: successfully associated matching pairs. Unrelated 3D detection results Unrelated 2D detection results ; Step 2.2.
2. Unrelated 3D detection results Fraction filtering and nonmaximum suppression are performed to obtain uncorrelated 3D detection results after false detection elimination. Successful matching pairs using association Unassociated 3D detection results after false positive elimination Unrelated 2D detection results The pre-correlation results are composed after false detection elimination.
6. A three-dimensional multi-target tracking method based on asynchronous multimodal data collaboration according to claim 5, characterized in that: In steps two and three, the predicted trajectory obtained in step two-one is correlated with the pre-correlation result obtained in step two-two after false detection elimination in a multi-stage manner to obtain a multi-stage correlation result, specifically as follows: Step 231: Match the predicted trajectory with the successfully associated pairs. 3D detection box in The first phase of association is as follows: in, This is the trajectory of successful association in the first stage. These are the detection boxes that were successfully associated in the first phase. This is the trajectory of the first-stage association failure. This is a detection box indicating a failure in the first stage of association. It's the Hungarian algorithm. yes and Three-dimensional generalized intersection and union ratio between them yes Three-dimensional generalized intersection and union ratio between them and It is a parameter variable. yes and The minimum closed convex set between yes and The associated cost matrix, This is the first predicted trajectory. It is the nth 3D detection box. This is the m-th predicted trajectory, where m is the total number of predicted trajectories and n is the number of successfully matched pairs. The total number of 3D detection boxes in the data; Step 232: Connect the failed correlation trajectories from the first stage with the uncorrelated 3D detection results after false detection elimination. The second phase of association is as follows: in, This is the trajectory of successful association in the second stage. These are the detection boxes that were successfully associated in the second phase. This is the trajectory of the second-stage association failure. This is a detection box indicating a failed association in the second stage. yes and The associated cost matrix, It is the j-th trajectory where the first-stage association failed. It is the unassociated 3D detection box after the k-th false positive elimination. It is the total number of trajectories that failed to associate in the first stage. It represents the total number of unassociated 3D bounding boxes after false positives have been eliminated; Step 233: Connect the failed trajectories from the second stage with the unconnected 2D detection results. Perform the third-stage association to obtain multi-stage association results, specifically: in, This is the trajectory of successful association in the third stage. These are the detection boxes that were successfully associated in the third stage. This is the trajectory of the third-stage association failure. This is a detection box indicating a failed association in the third stage. yes and The associated cost matrix, This is the second phase of the association failure. A trajectory, It is the first An unassociated two-dimensional detection box, It is the total number of trajectories that failed to associate in the second phase. It is the total number of associated 2D detection boxes. yes and The two-dimensional intersection and union ratio.
7. A three-dimensional multi-target tracking method based on asynchronous multimodal data collaboration according to claim 6, characterized in that: In step two-four, the predicted trajectory obtained in step two-one is correlated with the 2D detection results under the high-frequency image frames obtained in step one in a single-stage manner to obtain the single-stage correlation result. Specifically: in, It is a single-stage associated success trajectory. It is a detection box for successful single-stage association. It is the trajectory of a single-stage association failure. It is a detection box for single-stage association failure. yes and The associated cost matrix, It is the first A predicted trajectory, It is the first high-frequency image frame. A two-dimensional detection box, It represents the total number of two-dimensional detection boxes in high-frequency image frames.
8. A three-dimensional multi-target tracking method based on asynchronous multimodal data collaboration according to claim 7, characterized in that: Step three, which involves performing a Kalman update on the predicted trajectory state based on the association results to obtain the updated predicted trajectory state, specifically involves: Step 31: If the current frame is a low-frequency aligned frame, proceed to Step 32; if the current frame is a high-frequency image frame, proceed to Step 33. Step 3.2: Obtain the observation matrix based on the results of the multi-stage association, and use the observation matrix to perform Kalman update on the successfully associated trajectories to obtain the updated trajectory state, specifically: Step 321: Obtain the observation matrix H of the first-stage association and the second-stage association, and use the observation matrix H to perform Kalman update on the trajectories that have been successfully associated in the first and second stages to obtain the updated trajectory state: in, It is the observation matrix. It is the identity matrix, and K is the Kalman gain. Is the trajectory in Prior estimation of the time-varying covariance matrix. It is the preset observation noise. Is the trajectory in The posterior estimated state at time 10:
00. Is the trajectory in The prior estimate of the state at time t. Is the trajectory in the first... The observation status is linked in real time. Is the trajectory in Posterior estimate of the time-varying covariance matrix. Is the trajectory in Prior estimation of the time-varying covariance matrix; When performing Kalman updates on the trajectories that were successfully associated in the first and second phases ; Step 322: Obtain the observation matrix associated with the third stage. and using the observation matrix Perform a Kalman update on the successfully associated trajectories in the third stage to obtain the updated trajectory state, specifically: in, Is the trajectory in the first... The observation status in real time; When performing Kalman updates on the successfully associated trajectories in the third stage ; in, It is the 2D detection box position information under low-frequency aligned frames; Step 3: Use the successfully associated trajectories to perform a two-dimensional Kalman update to obtain the updated predicted trajectory state.
9. A three-dimensional multi-target tracking method based on asynchronous multimodal data collaboration according to claim 8, characterized in that: Step 33, which involves performing a two-dimensional Kalman update using the successfully associated trajectories to obtain the updated predicted trajectory state, specifically involves: in, Is the trajectory in the first... The observation status in real time; When performing Kalman updates on trajectories that have successfully achieved single-stage association ; in, It is the 2D detection box position information under high-frequency image frames.
10. A three-dimensional multi-target tracking method based on asynchronous multimodal data collaboration according to claim 9, characterized in that: In step four, the failed detection boxes are initialized as new trajectories, and then the detection scores of the successfully associated detection boxes and the scores of the trajectories are updated. Based on the updated detection scores of the detection boxes and the scores of the trajectories, it is determined whether to terminate trajectory tracking. If trajectory tracking is terminated, the process ends; otherwise, it returns to step two to perform the association for the next frame. Specifically: Step 41: If the current frame is a low-frequency aligned frame, proceed to step 42; if the current frame is a high-frequency image frame, proceed to step 44. Step 4.2: Initialize all detection boxes that failed to be associated with the predicted trajectory in the first and second stages into a new trajectory, and obtain the initial score of the new trajectory. Specifically: in, It represents the detection scores of 3D bounding boxes that failed to be associated between the first and second stages. It represents the detection score of the 2D detection box corresponding to the 3D detection box that failed to associate. This is the initial score of the new trajectory; Among them, the two-dimensional detection boxes corresponding to the three-dimensional detection boxes that failed in the first stage of association are: If a 2D bounding box is found in the data, and a 3D bounding box that failed to be associated in the second stage does not have a corresponding 2D bounding box, then... =0; Step 4.3: Update the scores of successfully associated trajectories and detection scores of bounding boxes in the low-frequency aligned frames, and determine whether to terminate the estimated tracking. If trajectory tracking is terminated, the process ends; otherwise, let t = t + 1, and return to step 2, specifically: Step 431: Update the trajectory score according to the exponential decay model, specifically as follows: in, It is the score of the trajectory in frame t-1. It is the fraction of the trajectory of frame t. 0.7, It is the attenuation rate; Step 432: Update the detection scores of successfully associated detection boxes in low-frequency aligned frames, specifically as follows: For the first-stage association and the second-stage association, the detection scores of the detection boxes are as follows: in, It is the detection score of successfully associated bounding boxes under low-frequency aligned frames. These are the detection score weights of the 3D bounding box. It is the detection score of the successfully associated 3D bounding boxes. These are the detection score weights of the 2D bounding boxes. It is the detection score of the successfully associated 2D bounding boxes; For the third-stage association, the detection score of the detection box is updated in the following way: in, It is the detection score of the 2D detection box under the low-frequency aligned frame in the third stage of successful association. It is the detection score of the 2D detection box under the low-frequency aligned frame in the third stage of the updated system. It is the attenuation coefficient; Step 433: Merge the updated trajectory score and the updated bounding box detection score to obtain the fused detection score for the current frame, specifically as follows: in, It is the detection score of the detection box in the current frame, taken as... or ; Step 434: Obtain the average fusion detection score from frame 1 to the current frame. If the average fusion detection score is less than the deletion threshold... If the number of consecutive unmatched frames exceeds the preset threshold max-age, then terminate trajectory tracking; otherwise, set t=t+1 and return to step two. Step 4: Update the trajectory score and detection box score in the high-frequency image frame, and determine whether to terminate the estimated tracking. If trajectory tracking is terminated, the process ends; otherwise, let t = t + 1, and return to step 2, specifically: Step 441: Update the score of the trajectory under the high-frequency image frame, specifically as follows: Step 442: Match the updated trajectory under the high-frequency image frame with the 2D detection box under the high-frequency image frame. If the match fails, discard the 2D detection box under the high-frequency image frame; if the match succeeds, obtain the fusion score of the successfully matched trajectory under the high-frequency image frame, specifically: in, It is the detection score of the attenuated 2D detection box. It is the detection score of 2D bounding boxes in high-frequency image frames. For decay weights; Step 443: Obtain the average fusion score of the trajectory from frame 1 to the current frame in high-frequency image frames. If the average fusion detection score is less than the deletion threshold... If the number of consecutive unmatched frames exceeds the preset threshold max-age, then terminate trajectory tracking; otherwise, set t=t+1 and return to step two.