Multi-target tracking method and system for monocular vision assisted by 4D millimeter wave radar
By combining data fusion methods with YOLOX, ByteTrack, DBSCAN, Kalman filtering and Hungarian algorithms, the problem of insufficient multi-target tracking accuracy of a single vision sensor and millimeter wave radar in complex environments is solved, and the target is accurately detected and state prediction is achieved, and the tracking performance is improved.
Patent Information
- Application Number
- CN202510437692.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-25
AI Technical Summary
It is difficult for a single vision sensor to accurately obtain the three-dimensional spatial position and velocity information of the target in complex environments. The millimeter-wave radar point cloud data are sparse and lack of visual characteristics, resulting in insufficient tracking accuracy and robustness of multi-targets.
Combining the YOLOX target detection algorithm and the visual tracker ByteTrack, the millimeter-wave radar point cloud data is processed using DBSCAN clustering and Kalman filtering, the fusion of visual and radar data is achieved through the global nearest neighbor matching algorithm, and the target matching is used to assist in the completion of visual tracking results.
It improves the accuracy and robustness of multi-target tracking, especially in complex environments, which can accurately obtain target position and speed information, improving tracking performance.
Smart Images

Figure CN120374680A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and sensor data fusion, and mainly relates to a multi-object tracking method and system for 4D millimeter-wave radar assisted monocular vision. Background Art
[0002] Multi-Object Tracking (MOT), as an important research topic in the fields of computer vision and sensor data fusion, has undergone multiple evolutions and technological innovations. Most object tracking methods are mainly based on visual sensors. The initial visual object tracking methods mostly adopted template matching methods (such as the Mean-Shift algorithm), and achieved object tracking by searching for the region in the image that is most similar to the object template. Such methods are relatively intuitive and have small computational complexity, but are vulnerable to object occlusion, scale changes, and appearance feature changes in complex scenes, resulting in tracking failures. With the development of machine learning and deep learning technologies, detection-based object tracking has gradually become the mainstream method. This method first uses an object detection algorithm to identify possible objects in each frame, and then matches these objects in adjacent frames through a data association method to form a complete object trajectory. Common object detection algorithms include traditional methods such as HOG (Histogram of Oriented Gradients) and DPM (Deformable Part Model). In recent years, deep learning models such as the YOLO series and Faster R-CNN have further improved the accuracy and robustness of object detection and tracking.
[0003] In practical applications, the performance of a single visual sensor method drops sharply in environments such as strong light, backlight, low light, rain, and fog, and it is difficult to accurately obtain the three-dimensional spatial position and velocity information of the target. As an active sensor, a millimeter-wave radar can work stably under complex climate conditions and output the distance, velocity, and angle information of the target. However, due to the sparse point cloud data of the millimeter-wave radar and the lack of measurement of visual features such as the texture and color of the target, when only relying on the millimeter-wave radar for object detection and tracking, it is often impossible to determine the category of the target and it is difficult to accurately distinguish densely distributed targets. Summary of the Invention
[0004] In view of the limitations of the single-sensor multi-target tracking technology in practical applications in the prior art, the present invention provides a multi-target tracking method and system assisted by a 4D millimeter-wave radar and a monocular vision. First, in the monocular camera image, the YOLOX target detection algorithm is used to detect the position, size, and category of the object; then, the output of the YOLOX detector is combined with the visual tracker ByteTrack to obtain the preliminary visual multi-target tracking results on the image; in the 4D millimeter-wave radar point cloud, the DBSCAN clustering algorithm is first used to filter out the noise points, and the position and velocity information of the center points of each group of clustered point clouds are calculated. Then, a Kalman filter is designed to process each center point cloud to estimate the motion state of the target in the radar coordinate system; through the spatio-temporal alignment of the sensors, the image and radar data are unified into the pixel coordinate system; then, the global nearest neighbor matching algorithm is adopted to complete the matching of the two-sensor data; finally, according to the matching results, the radar data is used to assist in complementing the corresponding visual tracking results, and the final multi-target tracking results are output. The method of the present invention makes full use of the spatial positioning advantages of the millimeter-wave radar and the rich semantic features of the visual sensor, and combines the data association strategy based on the Hungarian algorithm to achieve accurate detection, tracking, and state prediction of the target, significantly improving the tracking performance in complex environments.
[0005] To achieve the above object, the technical solution adopted by the present invention is: a multi-target tracking method assisted by a 4D millimeter-wave radar and a monocular vision, including the following steps:
[0006] S1. Visual target detection: Based on the YOLOX target detection algorithm, obtain a set of visual detection targets from the input monocular camera image;
[0007] S2. Visual multi-target tracking: Input the set of visual detection targets obtained in step S1 into the visual tracking algorithm ByteTrack to obtain a set of visual multi-target tracking trajectories, which contains the coordinates of the bounding rectangle of each target at different times;
[0008] S3. Radar point cloud clustering: Perform the DBSCAN clustering algorithm on the collected 4D millimeter-wave radar point cloud data to complete the clustering and denoising of the point cloud, and obtain a set of target clustering clusters;
[0009] S4. Radar target center point calculation: According to the set of target clustering clusters obtained in step S3, calculate the center point position and velocity of each target clustering cluster, and construct the motion state vector of each target;
[0010] S5. Radar point cloud state estimation: Input the motion state vector constructed in step S4 into the Kalman filter, construct the radar system state space model, and obtain the motion state estimation of each target in the radar coordinate system;
[0011] S6, Sensor spatio-temporal alignment: The time and space of 4D millimeter-wave radar data and monocular vision data are aligned by using the nearest neighbor timestamp matching and spatial coordinate transformation respectively;
[0012] S7, Data matching: In the pixel coordinate system, calculate the Euclidean distance between the center point of the target bounding rectangle obtained by visual tracking in step S2 and the projection of the center point of the radar motion state estimation target obtained in step S5, and use the global nearest neighbor matching algorithm to complete the data matching of the two sensors;
[0013] S8, Data fusion and result output: According to the data matching result obtained in step S7, use the radar motion state information to assist in complementing the corresponding visual tracking result, and output the final multi-target tracking trajectory set.
[0014] As an improvement of the present invention, in the step S1, the visual detection target set obtained from the monocular camera image is specifically:
[0015] D vision ={D i |D i =[u i ,v i ,w i ,h i ,c i T}
[0016] where D vision represents the overall visual detection target set, D i is the i-th visual detection target vector, [u i ,v i is the upper left coordinate of the detection box corresponding to the i-th visual detection target, [w i ,h i is the width and height of the detection box corresponding to the i-th visual detection target, and c i is the object category of the i-th visual detection target.
[0017] As an improvement of the present invention, the visual multi-target tracking trajectory set obtained in the step S2 is specifically:
[0018] T vision ={T j |T j =[u j ,v j ,w j ,h j ,c j ,ID j T}
[0019] where T vision Denote the overall set of visual multi-object tracking trajectories as T j is the j-th visual tracking trajectory vector, and [u j , v j is the upper left corner coordinates of the detection box corresponding to the j-th visual tracking trajectory, and [w j , h j is the width and height of the detection box corresponding to the j-th visual tracking trajectory, and c j is the object category of the j-th visual tracking trajectory, and ID j is the unique tracking identifier of the j-th visual tracking trajectory.
[0020] As another improvement of the present invention, in the step S3, the DBSCAN clustering algorithm is a density-based clustering method that clusters points with high density into the same target cluster, and the specific clustering result is:
[0021] C radar ={C1, C2, …, C n}
[0022] where C radar represents the overall set of target clustering clusters, and C1, C2, …, C n are respectively the specific target clusters.
[0023] As another improvement of the present invention, in the step S4, the calculation formulas for the center point position and speed of each target clustering cluster are specifically:
[0024]
[0025] where x n , y n , z n represent the respective position components of the center point of the n-th target clustering cluster, |C n | represents the number of point clouds included in the n-th target clustering cluster, and x m , y m , z m represent the respective position components of the m-th point cloud in the target clustering cluster, and v xn , v yn , v zn represent the respective speed components of the center point of the n-th target clustering cluster, and v xm , v ym , v zm represent the respective speed components of the m-th point cloud in the target clustering cluster;
[0026] The motion state vector S n of each target is specifically:
[0027] S n =[x n , yn , z n , v xn , v yn , v zn T 。
[0028] As another improvement of the present invention, the radar system state space model constructed in step S5 is specifically:
[0029] S k = Φ k / k-1 S k-1 + W k-1
[0030] Z k = H k S k + V k
[0031] Wherein, S k and S k-1 respectively represent the target motion state vectors of the radar system in the k-th frame and the (k - 1)-th frame, Φ k / k-1 is the state one-step transition matrix, W k-1 is the system noise vector, Z k represents the measurement vector with the closest Euclidean distance to S k in the position component, H k is the measurement matrix, and V k is the measurement noise vector;
[0032] Using the Kalman filter for this state space model, the filtering result is as follows:
[0033]
[0034] P k = (I - K k H k ) P k / k-1
[0035] Wherein, and respectively represent the radar target motion state estimation vectors in the (k - 1)-th frame and the k-th frame, is the state one-step prediction vector in the k-th frame, Q k-1 and R k are respectively the process noise covariance matrix and the observation noise covariance matrix of the system, P k / k-1 and P k are respectively the state one-step prediction error covariance matrix and the state estimation error covariance matrix of the system, and K k represents the Kalman filter gain.
[0036] As another improvement of the present invention, in the step S6, the radar target state estimation data is projected onto the visual image through the following coordinate transformation formula:
[0037]
[0038] wherein, [u r , v r is the projection coordinate of the radar target state estimation point cloud in the pixel coordinate system, and [x r , y r , z r is the coordinate of the radar target state estimation point cloud in the radar coordinate system, K c is the camera internal parameter matrix, and R cr and t cr are respectively the rotation matrix and the translation vector between the radar coordinate system and the camera coordinate system.
[0039] As a further improvement of the present invention, the Euclidean distance calculation formula in the step S7 is specifically:
[0040]
[0041] wherein, d jn is the matching cost between the j-th visual tracking target and the n-th radar state estimation target, and [u rn , v rn represents the projection of the n-th radar target state estimation point cloud in the pixel coordinate system, and [u vj , v vj and [w vj , h vj respectively represent the upper left corner coordinates, width, and height of the detection box corresponding to the j-th visual tracking target.
[0042] As a further improvement of the present invention, the multi-target tracking trajectory set output in the step S8 is specifically:
[0043] T final = {T l | T l = [x l , y l , z l , v xl , v yl , v zl , u l , v l , w l , h l , c l , ID l T}
[0044] wherein, Tfinal is the overall set of final multi - object tracking trajectories, T l is the l - th output trajectory, [x l , y l , z l , v xl , v yl , v zl represents the position and velocity of the l - th output trajectory in space, [u l , v l , w l , h l , c l respectively represent the upper - left corner coordinates, width, height, and object category of the detection box of the l - th output trajectory in the image. ID l is the unique tracking identifier of the l - th output trajectory.
[0045] To achieve the above object, the technical solution adopted by the present invention is also: a multi - object tracking system for 4D millimeter - wave radar - assisted monocular vision, including a computer program, and when the computer program is executed by a processor, it implements the steps of any of the above - mentioned methods.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] (1) When dealing with monocular vision data, the method proposed by the present invention combines the deep - learning object - detection algorithm YOLOX with the visual - tracking algorithm ByteTrack to accurately calculate the position and size of the object in the image, and obtain the category and tracking ID of the object, realizing multi - object tracking at the image level and making full use of the rich semantic information of the image.
[0048] (2) When dealing with 4D millimeter - wave radar data, the method proposed by the present invention removes the outlier noise points in the point - cloud data through the DBSCAN clustering method, calculates the position and velocity information of the center of each clustering cluster of points, and completes the target - state estimation through the Kalman - filtering algorithm, effectively improving the effectiveness and reliability of the radar data.
[0049] (3) The present invention proposes to use the Hungarian algorithm to complete the matching between the radar - state - estimation target and the visual - tracking target. Compared with the traditional point - by - point matching method, the Hungarian algorithm can effectively reduce the matching error and ensure the globally optimal matching result, and has higher matching efficiency and accuracy in complex multi - object scenarios.
[0050] (4) The method finally proposed by the present invention to use the matched millimeter - wave radar data to assist in complementing the visual - tracking information can obtain the information of the tracking target in space and in the image simultaneously, realize the complementary advantages of the two sensors, and improve the accuracy and comprehensiveness of the multi - object tracking algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is the flowchart of the steps of the method of the present invention;
[0052] Figure 2 is the UAV platform for reading data of a monocular camera and a 4D millimeter-wave radar in the method of the present invention;
[0053] Figure 3 is the schematic diagram of the visual multi-object tracking result obtained in step S2 of the method of the present invention;
[0054] Figure 4 is the schematic diagram of the millimeter-wave radar point cloud clustering result in step S3 of Embodiment 1 of the method of the present invention;
[0055] Figure 5 is the schematic diagram of the calculation result of the center points of each target clustering cluster in step S4 of Embodiment 1 of the method of the present invention;
[0056] Figure 6 is the schematic diagram of the final multi-object tracking result output in step S8 of Embodiment 1 of the method of the present invention;
[0057] Figure 7 is the comparison chart of the evaluation curves of the visual tracking algorithm and the method of the present invention in the test example of the present invention. Specific Embodiments
[0058] The present invention will be further clarified below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0059] Embodiment 1
[0060] A multi-object tracking method assisted by a 4D millimeter-wave radar and monocular vision, as Figure 1 shown, specifically includes the following steps:
[0061] Step S1: Read monocular camera images from the Intel D435i camera equipped on the UAV as Figure 2 shown. Use the YOLOX object detection algorithm to obtain the following set of visual detection targets from the input monocular camera images:
[0062] D vision ={D i |D i =[u i ,v i ,w i ,h i ,c i T}
[0063] where D vision Denote the overall set of visual detection targets as D i is the vector of the i-th visual detection target, [u i , v i is the upper left coordinate of the detection box corresponding to the i-th visual detection target, [w i , h i is the width and height of the detection box corresponding to the i-th visual detection target, c i is the object category of the i-th visual detection target.
[0064] Step S2: Input the set of visual detection targets D vision into the visual tracking algorithm ByteTrack to obtain the following initial set of visual multi-target tracking trajectories on the image:
[0065] T vision ={T j |T j =[u j , v j , w j , h j , c j , ID j T}
[0066] where, T vision represents the overall set of visual multi-target tracking trajectories, T j is the vector of the j-th visual tracking trajectory, [u j , v j is the upper left coordinate of the detection box corresponding to the j-th visual tracking trajectory, [w j , h j is the width and height of the detection box corresponding to the j-th visual tracking trajectory, c j is the object category of the j-th visual tracking trajectory, ID j is the unique tracking identifier of the j-th visual tracking trajectory.
[0067] After the visual multi-target tracking in this step of this embodiment, an example of the detection box corresponding to the visual multi-target tracking trajectory is shown as Figure 3 shown.
[0068] Step S3: Read the radar point cloud data from the OCULii EAGLE-min 4D millimeter-wave radar of the UAV equipment as shown in Figure 2 . Perform the DBSCAN clustering algorithm on the collected 4D millimeter-wave radar point cloud data to complete the clustering and denoising of the point cloud and obtain the set of target clustering clusters.
[0069] The DBSCAN clustering algorithm used adopts a density-based clustering method to remove the noise points in the millimeter-wave radar point cloud and cluster the points with a high enough density into the same target cluster. The specific clustering results are as follows:
[0070] C radar ={C1,C2,…,C n}
[0071] Among them, C radar represents the overall set of target clustering clusters, and C1, C2, …, C n are the respective specific target clusters. An example of the clustering result of the millimeter-wave radar point cloud is shown as Figure 4 shown.
[0072] Step S4: Calculate the center point position and velocity of each target clustering cluster, and construct the motion state vectors of each target. For each target clustering cluster, the calculation formulas for its center point position and velocity are specifically as follows:
[0073]
[0074] Among them, x n , y n , z n represent the respective position components of the center point of the nth target clustering cluster, |C n | represents the number of point clouds included in the nth target clustering cluster, x m , y m , z m represent the respective position components of the mth point cloud in the target clustering cluster, v xn , v yn , v zn represent the respective velocity components of the center point of the nth target clustering cluster, v xm , v ym , v zm represent the respective velocity components of the mth point cloud in the target clustering cluster. An example of the calculation result of the center point of each target clustering cluster is shown as Figure 5 shown. Thus, the motion state vector S n is:[[]]
[0075] S n =[x n ,y n ,z n ,v xn ,v yn ,v zn T
[0076] Step S5: The motion state vector S n Input into the Kalman filter to achieve the estimation of the motion state of each target in the radar coordinate system (space). When performing state estimation, the target motion state prediction adopts a uniform motion model, and the constructed radar system state space model is:
[0077] S k = Φ k / k-1 S k-1 + W k-1
[0078] Z k = H k S k + V k
[0079] Among them, S k and S k-1 respectively represent the motion state vectors of a certain target of the radar system in the k-th frame and the (k - 1)-th frame. Φ k / k-1 is the one-step state transition matrix, W k-1 is the system noise vector, Z k represents the measurement vector with the closest Euclidean distance to S k in the position component. H k is the measurement matrix, and V k is the measurement noise vector; the specific forms of Φ k / k-1 and H k are:
[0080]
[0081] H k = I6
[0082] Among them, △t represents the time difference between the radar data in the k-th frame and the (k - 1)-th frame, and I6 is a 6×6 identity matrix; using the Kalman filter for this state space model, the obtained filtering results are as follows:
[0083]
[0084] P k = (I - K k H k )P k / k-1
[0085] Among them, and respectively represent the radar target motion state estimation vectors in the (k - 1)-th frame and the k-th frame, is the one-step state prediction vector in the k-th frame, Q k-1 and R k are respectively the process noise covariance matrix and the observation noise covariance matrix of the system, P k / k-1 and P kare the one-step prediction error covariance matrix and the state estimation error covariance matrix of the system, respectively, and K k represents the Kalman filter gain.
[0086] Step S6: Use the nearest neighbor timestamp matching and spatial coordinate transformation respectively to achieve the time and space alignment of the 4D millimeter-wave radar data and the monocular vision data.
[0087] The principle of the method for achieving the time alignment of the two-sensor data is: use the nearest neighbor matching for data time alignment to find the radar target state estimation data with the closest timestamp for each visual image. The principle of the method for achieving the spatial alignment of the two-sensor data is: use the spatial coordinate transformation model to perform spatial alignment on the radar target state estimation data and the visual image; project the radar target state estimation data (point cloud) onto the visual image through the following coordinate transformation formula:
[0088]
[0089] where, [u r , v r are the projected coordinates of the radar target state estimation point cloud in the pixel coordinate system, [x r , y r , z r are the coordinates of the radar target state estimation point cloud in the radar coordinate system, K c is the camera internal parameter matrix, obtained by the Zhang Zhengyou calibration method, R cr and t cr are the rotation matrix and translation vector between the radar coordinate system and the camera coordinate system respectively, obtained by manual measurement.
[0090] Step S7: In the pixel coordinate system, use the Euclidean distance between the projection of the radar target state estimation point cloud and the center point of the visual tracking detection box as the data matching cost, and the calculation formula is:
[0091]
[0092] where, d jn is the matching cost between the j-th visual tracking target and the n-th radar state estimation target, [u rn , v rn represents the projection of the n-th radar target state estimation point cloud in the pixel coordinate system, [u vj , v vj and [w vj , h vj represent the upper left corner coordinates and width and height of the detection box corresponding to the j-th visual tracking target respectively; after calculating the matching costs between all data pairs, construct the matching cost matrix as follows:
[0093] Cmatch = [d jm
[0094] where C match is the overall matching cost matrix. For this matching cost matrix, the global nearest neighbor matching algorithm is used to complete the matching of the two sensor data. When searching for the global nearest neighbor matching result, the Hungarian algorithm is adopted to minimize the sum of the matching costs between all pairs of the two sensor data.
[0095] Step S8: Perform data fusion according to the data matching result in Step S7. For the m-th radar state estimation target and the j-th visual tracking target that are matched, use the radar data to assist in complementing the visual tracking information. After completing the information complementation operation for all the matched data pairs, the final multi-object tracking trajectory set is output as:
[0096] T final = {T l | T l = [x l , y l , z l , v xl , v yl , v zl , u l , v l , w l , h l , c l , ID l T}
[0097] where T final is the overall final multi-object tracking trajectory set, T l is the l-th output trajectory, [x l , y l , z l , v xl , v yl , v zl represents the position and velocity of the l-th output trajectory in the space (radar coordinate system), which is sourced from the matched radar data, [u l , v l , w l , h l , c l respectively represent the upper left corner coordinates, width, height, and object category of the detection box of the l-th output trajectory in the image, which are sourced from the matched visual data, and ID l is the unique tracking identifier of the l-th output trajectory, which is also sourced from the matched visual data.
[0098] An example of the top view of the finally output multi-object tracking trajectory in the radar coordinate system (space) is as shown in Figure 6 As shown in the figure. Combining Figure 3 Examples of multi-object tracking trajectories in images, and Figure 6 Examples of multi-object tracking trajectories in space. By comparing the two, it can be seen that the final multi-object tracking trajectory has information about the object at both the image and space levels, and can achieve complementary advantages of the two sensors by means of a 4D millimeter-wave radar assisting a monocular camera.
[0099] Test case
[0100] When quantitatively evaluating the performance of the method of the present invention, control Figure 2 The shown unmanned aerial vehicle to cross and track a set of synthetic road signs in Figure 3 The shown open campus scene. During the flight of the unmanned aerial vehicle, the data acquisition frequency of the monocular camera is 30Hz, and the data acquisition frequency of the millimeter-wave radar is 12Hz. Execute the steps of the multi-object tracking method proposed by the present invention in sequence. At the time stamp (time step) after synchronization of each camera and radar, output the set T vision of the visual multi-object tracking trajectories at this time and the set T final of the final multi-object tracking trajectories. For the visual tracking result T vision , a common monocular depth estimation method is used to complete its spatial position information. The method is as follows:
[0101]
[0102] Among them, is the monocular estimated depth of the j-th visual tracking trajectory, f is the camera focal length obtained by the Zhang Zhengyou calibration method, H is the average height of the synthetic road signs, and the value is 1.5m in the experiment, h j , u j and v j are the height and the center point coordinates of the image detection box of the j-th visual tracking trajectory respectively, both taken from T vision , represents the complete spatial position information of the j-th visual tracking trajectory.
[0103] The OSPA (Optimal Sub-Pattern Assignment) distance is a metric widely used for evaluating the performance of multi-object tracking. It can be used to measure the distance between two finite point sets. This metric can be calculated by the following formula, and the smaller the value, the better the effect of the tracking algorithm.
[0104]
[0105] Among them, is the average value of the OSPA distances, and the horizontal line represents the average. d C (x i,y π (i)) p Indicates the truncation distance between the true target set and the estimated target set. c p (n - m) cardinality error term, where (n - m) represents the number of targets in the estimated target set that are not matched, and c is a penalty parameter used to penalize cardinality mismatches.
[0106] In the experiment, the ground truth of each trajectory was marked manually. At each time step, the OSPA distances between the visual trajectory set that complements the spatial position information and the trajectory set finally output by the method of the present invention and the ground truth trajectory set were calculated respectively, and the evaluation curve as shown in Figure 7 was plotted. It can be seen from Figure 7 that compared with the visual tracking algorithm, the method of the present invention has a smaller OSPA distance at all time steps, and the comprehensive tracking performance is better, which reflects the superiority of the method.
[0107] In summary, the method of the present invention provides a multi - target tracking method assisted by 4D millimeter - wave radar and monocular vision, which can combine the characteristics of 4D millimeter - wave radar point cloud data and monocular vision image data, realize the complementary characteristics of the two sensors, and significantly improve the accuracy and comprehensiveness of multi - target tracking.
[0108] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.
Claims
1. A multi-object tracking method assisted by 4D millimeter-wave radar and monocular vision, characterized in that , including the following steps: S1. Visual target detection: Based on the YOLOX target detection algorithm, obtain a set of visual detection targets from the input monocular camera image; S2. Visual multi-target tracking: Input the set of visual detection targets obtained in step S1 into the visual tracking algorithm ByteTrack to obtain a set of visual multi-target tracking trajectories, where the set contains the bounding rectangle coordinates of each target at different times; S3. Radar point cloud clustering: Execute the DBSCAN clustering algorithm on the collected 4D millimeter-wave radar point cloud data to complete the clustering and denoising of the point cloud, and obtain a set of target clustering clusters; S4. Radar target center point calculation: According to the set of target clustering clusters obtained in step S3, calculate the center point position and velocity of each target clustering cluster, and construct the motion state vector of each target; S5. Radar point cloud state estimation: Input the motion state vector constructed in step S4 into the Kalman filter to construct a radar system state space model, and obtain the motion state estimation of each target in the radar coordinate system; S6. Sensor spatio-temporal alignment: Use the nearest neighbor timestamp matching and spatial coordinate transformation respectively to align the time and space of the 4D millimeter-wave radar data and the monocular vision data; S7. Data matching: In the pixel coordinate system, calculate the Euclidean distance between the center point of the target bounding rectangle of the visual tracking obtained in step S2 and the projection of the center point of the radar motion state estimation target obtained in step S5, and use the global nearest neighbor matching algorithm to complete the data matching of the two sensors; S8. Data fusion and result output: According to the data matching result obtained in step S7, use the radar motion state information to assist in complementing the corresponding visual tracking result, and output the final set of multi-target tracking trajectories.
2. The multi-object tracking method using 4D millimeter-wave radar to assist monocular vision according to claim 1, characterized in that: In step S1, the set of visual detection targets obtained from the monocular camera image is specifically: D vision = {D i | D i = [u i , v i , w i , h i , c i T} Among them, D vision represents the overall set of visual detection targets, D i is the i-th visual detection target vector, [u i , v i is the upper left corner coordinates of the detection box corresponding to the i-th visual detection target, [w i , h i is the width and height of the detection box corresponding to the i-th visual detection target, c i is the object category of the i-th visual detection target.
3. A multi-object tracking method using 4D millimeter-wave radar to assist monocular vision as described in claim 1, characterized in that: The set of visual multi-target tracking trajectories obtained in step S2 is specifically: T vision = {T j | T j = [u j , v j , w j , h j , c j , ID j T} Among them, T vision represents the overall set of visual multi-object tracking trajectories, and T j is the j-th visual tracking trajectory vector, where [u j , v j are the upper left coordinates of the detection box corresponding to the j-th visual tracking trajectory, and [w j , h j are the width and height of the detection box corresponding to the j-th visual tracking trajectory, c j is the object category of the j-th visual tracking trajectory, and ID j is the unique tracking identifier of the j-th visual tracking trajectory.
4. A multi-object tracking method assisted by 4D millimeter-wave radar and monocular vision according to claim 1, characterized in that: In step S3, the DBSCAN clustering algorithm is a density-based clustering method that clusters points with high density into the same target cluster. The specific clustering result is: C radar = {C1, C2, …, C n} Among them, C radar represents the overall set of target clustering clusters, and C1, C2, …, C n are respectively each specific target cluster.
5. A multi-object tracking method using 4D millimeter-wave radar to assist monocular vision as claimed in claim 1, characterized in that: In step S4, the calculation formulas for the center point position and velocity of each target clustering cluster are specifically: where x n , y n , z n represent the respective position components of the center point of the nth target clustering cluster, |C n | represents the number of point clouds included in the nth target clustering cluster, x m , y m , z m represent the respective position components of the mth point cloud in the target clustering cluster, v xn , v yn , v zn represent the respective velocity components of the center point of the nth target clustering cluster, v xm , v ym , v zm represent the respective velocity components of the mth point cloud in the target clustering cluster; The motion state vector S of each target n Specifically: S n = [x n , y n , z n , v xn , v yn , v zn T . 6. The multi-target tracking method using 4D millimeter-wave radar to assist monocular vision according to claim 1, wherein: The radar system state space model constructed in step S5 is specifically: S k = Φ k / k-1 S k-1 + W k-1 Z k = H k S k + V k where, S k and S k-1 represent the target motion state vectors of the k-th frame and the (k-1)-th frame of the radar system respectively, Φ k / k-1 is the one-step state transition matrix, W k-1 is the system noise vector, Z k represents the measurement vector that is the closest to S k in the position component, H k is the measurement matrix, V k is the measurement noise vector; Using the Kalman filter for this state space model, the filtering result is as follows: P k = (I - K k H k )P k / k-1 Wherein, and represent the radar target motion state estimation vectors of the (k - 1)-th frame and the k-th frame respectively, is the one-step state prediction vector of the k-th frame, Q k-1 and R k are the process noise covariance matrix and the observation noise covariance matrix of the system respectively, P k / k-1 and P k are the one-step state prediction error covariance matrix and the state estimation error covariance matrix of the system respectively, K k represents the Kalman filter gain.
7. A multi-object tracking method using a 4D millimeter-wave radar to assist a monocular vision as claimed in claim 1, characterized in that: In step S6, the radar target state estimation data is projected onto the visual image through the following coordinate transformation formula: Among them, [u r , v r are the projected coordinates of the radar target state estimation point cloud in the pixel coordinate system, and [x r , y r , z r are the coordinates of the radar target state estimation point cloud in the radar coordinate system. K c is the camera intrinsic matrix, and R cr and t cr are respectively the rotation matrix and the translation vector between the radar coordinate system and the camera coordinate system.
8. A multi-object tracking method using a 4D millimeter-wave radar to assist monocular vision as claimed in claim 1, characterized in that: The Euclidean distance calculation formula in step S7 is specifically: where d jn is the matching cost between the j-th visual tracking target and the n-th radar state estimation target, and [u rn , v rn represents the projection of the n-th radar target state estimation point cloud in the pixel coordinate system, and [u vj , v vj and [w vj , h vj represent the upper left corner coordinates, width, and height of the detection box corresponding to the j-th visual tracking target, respectively.
9. A multi-object tracking method using a 4D millimeter-wave radar to assist monocular vision as claimed in claim 1, characterized in that: The set of multi-target tracking trajectories output in step S8 is specifically: T final = {T l | T l = [x l , y l , z l , v xl , v yl , v zl , u l , v l , w l , h l , c l , ID l T} Among them, T final is the overall set of final multi-object tracking trajectories, and T l is the l-th output trajectory. [x l , y l , z l , v xl , v yl , v zl represents the position and velocity of the l-th output trajectory in space. [u l , v l , w l , h l , c l respectively represent the upper left corner coordinates, width, height, and object category of the detection box of the l-th output trajectory in the image. ID l is the unique tracking identifier of the l-th output trajectory.
10. A multi-object tracking system assisted by 4D millimeter-wave radar and monocular vision, including a computer program, characterized in that: When the computer program is executed by the processor, it implements the steps of any of the above methods.