A roadside perception method integrating millimeter-wave radar and camera
By using the data fusion method of millimeter-wave radar and camera in the road-side perception system, combined with YOLOv3 and Deep SORT algorithms, adaptive weights and Kalman filtering are used to solve the problem of insufficient positioning accuracy in the existing technology, and achieve all-weather and high-precision target detection and tracking.
Patent Information
- Application Number
- CN202111313286.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-08
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2041-11-08
AI Technical Summary
The existing roadside perception scheme based on millimeter wave radar and camera fails to effectively utilize the advantages of each sensor during data fusion, resulting in insufficient positioning accuracy and unstable single-frame detection results of the camera, and there are missed detection and false alarms.
Data is obtained through millimeter wave radar and static targets and noise are removed, camera target detection and tracking is performed by YOLOv3 and Deep SORT algorithms, data fusion is used to fusion with adaptive weights, and Kalman filtering is used to achieve accurate matching and fusion between millimeter wave radar and camera data.
It improves the positioning accuracy and detection and tracking of the roadside perception system, reduces the calculation amount and effectively filters out noise, and realizes all-weather and high-precision target detection and tracking.
Smart Images

Figure CN114280611B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of roadside perception technology, and in particular to a roadside perception method integrating millimeter-wave radar and a camera. Background Art
[0002] Currently, most vehicles rely on onboard sensors and GPS positioning systems to perceive their surroundings and locate them in real time. However, these systems suffer from high vehicle costs and difficulty handling complex traffic scenarios. To address these issues, roadside perception has been proposed. This approach integrates roadside sensor data to provide vehicles with accurate, real-time traffic information, thereby enhancing the autonomous vehicle's perception range and improving driving safety. Sensor solutions can be categorized into single-sensor and multi-sensor fusion solutions.
[0003] Sensors widely used in roadside perception systems include cameras, millimeter-wave radars, and lidars. However, single sensors often struggle with specific scenarios, such as poor nighttime camera imaging and potential failure of lidars in rainy and snowy conditions. Therefore, single-sensor solutions cannot meet the high-quality and high-precision requirements in complex traffic environments. Multi-sensor fusion solutions, by integrating data from multiple sensors, provide redundancy and address complex scenarios. Existing solutions generally include lidar-based and camera-based perception solutions and millimeter-wave radar-based and camera-based perception solutions. The former, due to its use of lidar, is not only costly but also requires significant computing resources to process point cloud data. The latter, on the other hand, is less expensive, operates 24 / 7, and requires less data.
[0004] Existing millimeter-wave radar and camera fusion solutions generally match single-frame camera detection results with multi-target tracking results from the millimeter-wave radar. Because single-frame camera detection results are unstable and often result in missed detections and false alarms, simply matching single-frame detection results can mismatch incorrect results to the millimeter-wave radar tracking target, making it difficult to achieve stable detection and tracking results after fusion. Furthermore, existing millimeter-wave radar and camera fusion solutions directly use weighted summation for data fusion, ignoring the characteristics of different sensors and failing to leverage the strengths of each. Cameras offer higher lateral position measurement accuracy than millimeter-wave radars at close range, while millimeter-wave radars have advantages in longitudinal distance measurement. Therefore, an adaptive weighting approach can be used to fuse the data from these two approaches to achieve higher positioning accuracy. Summary of the Invention
[0005] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a roadside perception method that integrates millimeter-wave radar and camera.
[0006] The purpose of the present invention can be achieved by the following technical solutions:
[0007] A roadside perception method integrating millimeter-wave radar and camera, the method comprising the following steps:
[0008] Step 1: Obtain millimeter-wave radar data through the millimeter-wave radar, extract effective targets, and remove stationary targets and random noise;
[0009] Step 2: Obtain image data through a camera installed on a roadside pole, perform target detection and multi-target tracking based on the YOLOv3 and Deep SORT tracking algorithms, and obtain camera data;
[0010] Step 3: Perform multi-sensor fusion of millimeter-wave radar data and camera data based on a fusion algorithm to enable real-time detection, tracking, and positioning of targets by roadside equipment.
[0011] In step 1, the process of extracting effective targets specifically includes the following steps:
[0012] Step 101: Acquire millimeter-wave radar data through a millimeter-wave radar installed on a roadside pole. The millimeter-wave radar data includes the target's ID, longitudinal distance, lateral distance, longitudinal speed, lateral speed, and reflection cross-sectional area.
[0013] Step 102: Obtaining road location information of a set road area based on fixed roadside poles, and filtering out targets located within the road area based on the road location information;
[0014] Step 103: removing stationary objects within the road area based on a set speed threshold;
[0015] Step 104: Set a threshold for the number of times a target exists within a continuous period of time, index the number of times the target exists in the first four scanning cycles based on the target ID, remove index results with the number of times the target exists less than three times, and then obtain valid targets.
[0016] In step 2, the process of target detection and tracking based on YOLOv3 and Deep SORT tracking algorithms specifically includes the following steps:
[0017] Step 201: Acquire image data through a camera installed on a roadside pole;
[0018] Step 202: Annotate the image data and create a dataset. Train the YOLOv3 model based on the dataset. First, freeze the backbone extraction network and train only the prediction network. After 50 cycles of training, train the entire model and terminate the training using Early Stopping.
[0019] Step 203: Use the trained YOLOv3 model as the detector of the Deep SORT algorithm to perform multi-target tracking. The final output is [id c ,score,x min ,y min ,x max ,y max ,class] m , that is, camera data, id c is the ID of the target detected by the camera, score is the confidence, (x min ,y min ) is the upper left vertex of the detection box, (x max ,y max ) is the lower right vertex of the detection box, class is the category, and m is the number of targets detected by the camera.
[0020] In step 3, the process of fusing the millimeter-wave radar data with the camera data based on the fusion algorithm specifically includes the following steps:
[0021] Step 301: spatially fuse the millimeter-wave radar data and the camera data. The points in the radar coordinate system are converted to the pixel coordinate system corresponding to the camera through the coordinate system to achieve spatial fusion of the two. Specifically:
[0022] Let X w -Y w -Z w is the world coordinate system, O w is the origin of the world coordinate system, X r -Y r -Z r is the radar coordinate system, O r is the origin of the radar coordinate system, X c -Y c -Z c is the camera coordinate system, O c is the origin of the camera coordinate system, plane O w -X w -Y w Coincident with the ground, axis O w -Z w Along the roadside pole upwards, the origin of the camera coordinate system is O c To the origin O of the world coordinate system w The vertical height is h c , the origin of the camera coordinate system O c To the origin O of the world coordinate system w The longitudinal distance is l c , axis O c -Z c Coincident with the optical axis of the camera, axis Oc -Z c With axis O w -X w The angle between the plane O and c -Y c -Z c With plane O w -X w -Z w Coincident, the origin of the radar coordinate system O r To the origin O of the world coordinate system w The vertical height is h r , the origin of the radar coordinate system O r To the origin O of the world coordinate system w The longitudinal distance is l r , plane O r -X r -Z r With plane O w -X w -Z w Coincident, axis O r -X r With axis O w -X w Parallel, uv is the pixel coordinate system, the point (X c ,Y c ,Z c ) is converted to the pixel coordinate system using the following formula:
[0023]
[0024] Among them, f x and f y is the equivalent focal length of the camera, (u0,v0) is the pixel coordinate of the center of the image, f x 、f y , u0 and v0 are the internal parameters of the camera, (X c ,Y c ,Z c ) is the three-dimensional coordinate of the camera coordinate system, and (u, v) is the two-dimensional coordinate of the pixel coordinate system;
[0025] According to the relative position relationship between the radar coordinate system and the camera coordinate system, the point (X r ,Y r ,Z r ) is converted to the pixel coordinate system using the following formula:
[0026]
[0027] Among them, (X r ,Y r,Z r ) is the three-dimensional coordinate of the radar coordinate system, h r is the origin of the radar coordinate system O r To the origin O of the world coordinate system w The vertical height, l r is the origin of the radar coordinate system O r To the origin O of the world coordinate system w The longitudinal distance, h c is the origin O of the camera coordinate system c To the origin O of the world coordinate system w The vertical height, l c is the origin O of the camera coordinate system c To the origin O of the world coordinate system w The longitudinal distance of the axis O c -Z c With axis O w -X w Angle;
[0028] The point (X w ,Y w ,Z w ) is converted to the pixel coordinate system using the following formula:
[0029]
[0030] Among them, (X w ,Y w ,Z w ) is the three-dimensional coordinate of the world coordinate system;
[0031] Step 302: Fusing the millimeter-wave radar and camera time, aligning the timestamps of the millimeter-wave radar data and the camera data for time fusion, i.e., using an adaptive algorithm to align based on the timestamp of the camera data, to complete the time fusion of the millimeter-wave radar data and the camera data.
[0032] Step 303: Fuse the millimeter-wave radar data and the camera data based on a fusion algorithm.
[0033] In step 303, the process of fusing the millimeter-wave radar data and the camera data based on the fusion algorithm specifically includes the following steps:
[0034] Step 303A: performing data expansion on the millimeter-wave radar data and the camera data through coordinate transformation;
[0035] Step 303B: Perform preliminary matching of the millimeter-wave radar data and the camera data;
[0036] Step 303C: Data association is performed on the millimeter-wave radar data after the initial matching data selection and the expanded camera data. The cost of each millimeter-wave radar data belonging to each detection frame is calculated. Specifically, the Mahalanobis distance between the millimeter-wave radar data and the camera data in the pixel coordinate system and the Euclidean distance between the millimeter-wave radar data and the camera data in the world coordinate system are used as the cost matrix. The matching method with the lowest global cost is obtained based on the Hungarian algorithm.
[0037] Step 303D: Based on the visual tracking results, match the millimeter-wave radar data about the target, that is, based on the camera data, match and fuse the camera data with the millimeter-wave radar data, and update based on the Kalman filter to obtain the track of the fused data.
[0038] In step 303A, the data expansion process is specifically as follows:
[0039] Step 303A1: The output after extracting valid targets from the millimeter-wave radar data is:
[0040] [id r ,x r ,y r ,v xr ,v yr ] n
[0041] Among them, id r is the ID of the target detected by the millimeter radar wave, (x r ,y r ) is the radar coordinate of the target detected by the millimeter radar wave in the radar coordinate system, v xr and v yr are the longitudinal velocity and lateral velocity of the target in the radar coordinate system, respectively, and n represents the number of targets detected by the millimeter-wave radar;
[0042] Step 303A2: Radar coordinates (x r ,y r ) is projected into the pixel coordinate system to obtain the pixel coordinate (x r_c ,y r_c ), and then the millimeter wave radar data is expanded to:
[0043] [id r ,x r ,y r ,v xr ,v yr ,x r_c ,y r_c ] n ;
[0044] Among them, (x r_c,y r_c ) is the radar coordinate (x r ,y r ) is projected to the pixel coordinates in the pixel coordinate system;
[0045] Step 303A3: The output of target detection and tracking performed on the image data of the target acquired by the roadside camera is:
[0046] [id c ,score,x min ,y min ,x max ,y max ,class] m
[0047] Among them, id c is the ID of the target detected by the camera, score is the confidence, x min and y min are the horizontal and vertical coordinates of the upper left vertex of the detection box, x max and y max are the horizontal and vertical coordinates of the lower right vertex of the detection box, class is the category, and m is the number of targets detected by the camera;
[0048] Step 303A4: Use the center point of the lower boundary of the detection frame ((x min +x max ) / 2,y max ) as the reference point, the coordinates of the world coordinate system corresponding to this reference point are in Z w = 0, the calculation formula for the coordinates of the point in the world coordinate system is:
[0049]
[0050] Among them, (X w ,Y w ,Z w ) is the three-dimensional coordinate of the world coordinate system, (x c_w ,y c_w ) is the coordinate of the target detected by the camera in the world coordinate system;
[0051] Step 303A5: Expand the camera data to:
[0052] [id c ,score,x min ,y min ,x max ,y max ,class,x c_w ,y c_w ] m .
[0053] In step 303B, the process of performing preliminary matching of the millimeter-wave radar data and the camera data is specifically as follows:
[0054] The pixel coordinates of a set of millimeter wave radar data after spatial fusion and temporal fusion are (x r_c ,y r_c ), the coordinate of the upper left vertex of the detection box is (x min ,y min ), the lower right vertex is (x max ,y max ), the coordinate of the upper left vertex of the extended detection box is (x min -Δx,y min -Δy), the coordinates of the lower right vertex are (x max +Δx,y max +Δy), that is, based on the detection frame, an extended detection frame is obtained along the u-axis and v-axis. The matched millimeter-wave radar data is filtered through the expanded detection extension frame. After filtering, M targets remain in the camera data and N targets remain in the millimeter-wave data. That is, the number of camera data is M and the number of millimeter-wave radar data is N.
[0055] In step 303C, the process of data association specifically includes the following steps:
[0056] Step 303C1: At least one millimeter-wave radar data falls within the detection frame output by each camera. That is, the number of camera data, M, is less than or equal to the number of millimeter-wave radar data, N. Suppose the expressions of camera data and millimeter-wave radar data are:
[0057] [id c ,score,x min ,y min ,x max ,y max ,class,x c_w ,y c_w ] M
[0058] [id r ,x r ,y r ,v xr ,v yr ,x r_c ,y r_c ] N
[0059] Among them, id c is the ID of the target detected by the camera, score is the confidence, x min and y minare the horizontal and vertical coordinates of the upper left vertex of the detection box, x max and y max They are the horizontal and vertical coordinates of the lower right vertex of the detection box, class is the category, (x c_w ,y c_w ) is the coordinate of the center point of the lower boundary of the detection box in the world coordinate system, M is the number of targets detected by the camera, that is, the number of camera data, id r is the ID of the target detected by the millimeter radar wave, (x r ,y r ) is the radar coordinate of the target detected by the millimeter radar wave in the radar coordinate system, v xr and v yr are the longitudinal velocity and lateral velocity of the target detected by the millimeter radar wave in the radar coordinate system, (x r_c ,y r_c ) is the pixel coordinate of the target detected by the millimeter radar wave, M is the number of targets detected by the camera, that is, the number of camera data, and N is the number of targets detected by the millimeter radar wave, that is, the number of millimeter radar wave data;
[0060] Step 303C2: Calculate the Mahalanobis distance between the millimeter-wave radar data and the camera data in the pixel coordinate system:
[0061] For the i-th target detected by the camera, if the target is detected for the first time, the pixel coordinates of the camera data are initialized with the center of the detection frame, and the unit matrix I is used. 2×2 Initialize the covariance matrix. The expressions of the pixel coordinates and covariance matrix of the camera data are:
[0062]
[0063] in, is the center coordinate of the detection box, d i is the pixel coordinate corresponding to the i-th target detected by the camera, S i is the covariance matrix corresponding to the i-th target detected by the camera;
[0064] If the target belongs to the track of the existing fusion data, the covariance matrix of the observation variable and the observation space is calculated. Since the speed information is not involved, the pixel coordinates (x w_c ,y w_c ) for pixel coordinate d i Assign values and take the position covariance to covariance matrix S in the observation space covariance matrix i Assignment:
[0065]
[0066]
[0067] Among them, S k+1 is the observation space covariance matrix at time k+1, h is a 2×4 matrix, and T represents the transpose of the matrix;
[0068] According to the pixel coordinate d i and the covariance matrix S i Calculate the Mahalanobis distance between the millimeter-wave radar data and the camera data in the pixel coordinate system. The calculation formula of the Mahalanobis distance is:
[0069]
[0070] d j =[x r_c ,y r_c ]
[0071] Among them, dis image is the Mahalanobis distance between the millimeter-wave radar data and the camera data in the pixel coordinate system, d j is the pixel coordinate corresponding to the jth target detected by the millimeter-wave radar;
[0072] Step 303C3: Calculate the Euclidean distance between the millimeter-wave radar data and the camera data in the world coordinate system. The calculation formula of the Euclidean distance is:
[0073]
[0074] Among them, dis world is the Euclidean distance between the millimeter-wave radar data and the camera data in the world coordinate system, (x r +l r ,y r ) is the world coordinate corresponding to the jth target detected by the millimeter wave radar, (x c_w ,y c_w ) is the world coordinate corresponding to the i-th target detected by the camera;
[0075] Step 303C4: Normalize the obtained Mahalanobis distance and Euclidean distance respectively, and then perform weighted summation to obtain the cost of matching the j-th millimeter-wave radar data with the i-th camera data. The cost calculation formula is:
[0076]
[0077] Among them, task_matrix ij is the cost of matching the j-th radar data with the i-th camera data, and λ is the weight between the Mahalanobis distance and the Euclidean distance;
[0078] Step 303C5: Obtain a global cost matrix, and obtain a matching result with the best global cost based on the Hungarian algorithm. For a matching result, if the corresponding cost is greater than the set cost threshold, the matching result is considered to be wrong and is removed.
[0079] In step 303D, if the target ID does not belong to the track of the existing fusion data, the process of updating based on the Kalman filter specifically includes the following steps:
[0080] Step 303D1: The target ID does not belong to the existing fusion data track, that is, the camera data does not have matching millimeter radar wave data. The state variables are initialized, and the expressions of the state variables and state covariance matrix are:
[0081]
[0082] Among them, I 4×4 is a 4×4 unit matrix;
[0083] Process noise covariance matrix q k The expression is:
[0084] q k =0.1I 4×4 ;
[0085] Observation noise covariance matrix r k The expression is:
[0086] r k =0.1I 4×4
[0087] State transition matrix f k The expression is:
[0088]
[0089] Where t is the time difference between the current moment and the previous moment;
[0090] Observation matrix h k The expression is:
[0091] h k =I 4×4 ;
[0092] Step 303D2: Perform prediction and update to obtain the predicted state vector and predicted state covariance matrix:
[0093]
[0094] Among them, X k+1|k is the predicted state vector at time k+1, fk+1 is the state transition matrix at time k+1, P k+1|k is the predicted state covariance matrix at time k+1, q k+1 is the process noise covariance matrix at time k+1, T represents the transpose;
[0095] According to the predicted state vector and the predicted state covariance matrix, the observation vector and the observation space covariance matrix are obtained:
[0096]
[0097] Among them, Y k+1 is the observation vector at time k+1, h k+1 is the state transition matrix at time k+1, S k+1 is the observation space covariance matrix at time k+1, r k+1 is the observation noise covariance matrix at time k+1;
[0098] The actual measured value Z at time k+1 k+1 and the observation noise covariance matrix r k+1 The expression is:
[0099] Z k+1 =[x c_w ,y c_w ,0,0] T
[0100] r k+1 =I 4×4
[0101] Update the state vector and state covariance matrix to obtain the posterior estimates of the state vector and state covariance matrix:
[0102]
[0103] K k+1 is the gain at time k+1, Z k+1 is the actual measurement value at time k+1, X k+1 is the posterior estimate of the state vector at time k+1, P k+1 is the posterior estimate of the state covariance matrix at time k+1;
[0104] If the target ID belongs to the existing fusion data track, the update process based on Kalman filtering includes the following steps:
[0105] Step 303D3: The target ID belongs to the track of the existing fusion data, that is, the camera data has matching millimeter radar wave data, then the expressions of the state variable, state covariance matrix and observation variable are:
[0106]
[0107] Among them, x r +l r is the longitudinal coordinate of the millimeter radar wave data in the world coordinate system, λ1 and λ2 are weights, X k is the state variable, dis and vel are the position and velocity of the target in the world coordinate system respectively, dis x 、dis y 、vel x and vel y are the longitudinal distance, lateral distance, longitudinal speed and lateral speed of the target in the world coordinate system, Y k is the observed variable, I is the unit matrix;
[0108] The expression of weight λ1 is:
[0109]
[0110] Among them, x r It is the longitudinal coordinate of the target detected by millimeter radar wave in the radar coordinate system;
[0111] Step 303D4: Perform prediction and update to obtain the predicted state vector and predicted state covariance matrix:
[0112]
[0113] Among them, X k+1|k is the predicted state vector at time k+1, f k+1 is the state transition matrix at time k+1, P k+1|k is the predicted state covariance matrix at time k+1, q k+1 is the process noise covariance matrix at time k+1, T represents the transpose;
[0114] According to the predicted state vector and the predicted state covariance matrix, the observation vector and the observation space covariance matrix are obtained:
[0115]
[0116] Among them, Y k+1 is the observation vector at time k+1, h k+1 is the state transition matrix at time k+1, S k+1 is the observation space covariance matrix at time k+1, r k+1 is the observation noise covariance matrix at time k+1;
[0117] The actual measured value Z at time k+1 k+1 The expression is:
[0118] Zk+1 =[λ1(x r +l r )+(1-λ1)x c_w ,λ2y r +(1-λ2)x c_w ,v xr ,v yr ] T
[0119] Update the state vector and state covariance matrix to obtain the posterior estimates of the state vector and state covariance matrix:
[0120]
[0121] K k+1 is the gain at time k+1, Z k+1 is the actual measurement value at time k+1, X k+1 is the posterior estimate of the state vector at time k+1, P k+1 is the posterior estimate of the state covariance matrix at time k+1;
[0122] Step 303D5: Get the latest fusion data of the target:
[0123] [id,class,score,x min ,y min ,x max ,y max ,dis x ,dis y ,vel x ,vel y ] M
[0124] Among them, id is the ID of the fusion target, dis x 、dis y , vel and vel y is the updated posterior estimate based on Kalman filtering, dis x 、dis y , vel and vel y They represent the longitudinal distance, lateral distance, longitudinal speed and lateral speed of the fusion target in the world coordinate system respectively;
[0125] Step 303D6: Obtain the track of the fused data, and use age and moment to manage the generation and disappearance of the track of the fused data.
[0126] In step 303D6, the generation and disappearance of the track of the fused data is managed by using age and moment as follows:
[0127] After the camera detects a target, a track is generated. The fused track is initialized based on the camera's track. At this time, the age of all tracks is set to 1, and the moment of all tracks is set to the timestamp of the current data. When a new camera track appears in the camera data received in the next cycle, the new camera track is added to the track of the fused data, and the age of the track is set to 1, and the moment of the track is set to the timestamp of the corresponding camera data. If the camera track has the same id as the track corresponding to the previous cycle, the new camera track is used to update it, the age is still 1, and the moment is updated to the timestamp of the camera data corresponding to the new camera track. If the track that existed in the previous cycle does not have a corresponding track in this cycle, the age of the track is increased by 1, and the moment remains unchanged. When the age of a track reaches the set threshold, it is considered that the track has left the scene and is deleted from the track set.
[0128] Compared with the prior art, the present invention has the following advantages:
[0129] 1. The millimeter-wave radar effective target extraction method effectively filters out a large amount of noise and false alarm signals, reducing the amount of subsequent calculations.
[0130] 2. By using a self-built roadside dataset, we effectively solved the difficulties of high roadside camera installation positions, large visual ranges, and a large number of small targets.
[0131] 3. Use weighted distance features for data association, and perform weighted fusion of the associated data as measurement values for Kalman filter update, thereby improving the distance tracking accuracy of the fusion model.
[0132] 4. When fusing data, the characteristics of millimeter-wave radar and camera are taken into consideration. When measuring at close range, the lateral position measurement accuracy is higher than that of millimeter-wave radar, while millimeter-wave radar has more advantages in longitudinal distance measurement. Therefore, an adaptive weighting method is used to fuse the data of the two to achieve higher positioning accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0133] Figure 1 It is a structural block diagram of the present invention.
[0134] Figure 2 A schematic diagram of the relative position of the coordinate system.
[0135] Figure 3 Schematic diagram of the time fusion algorithm.
[0136] Figure 4 This is the structural block diagram of the fusion algorithm.
[0137] Figure 5 Schematic diagram of the preliminary selection of matching data. DETAILED DESCRIPTION
[0138] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0139] Example
[0140] like Figure 1 As shown, the present invention provides a roadside perception method integrating millimeter-wave radar and camera, which includes the following steps:
[0141] Step 1: Obtain millimeter-wave radar data through the millimeter-wave radar and extract effective targets to remove stationary targets and random noise;
[0142] Step 2: Obtain image data from a roadside camera, perform target detection and multi-target tracking based on the YOLOv3 and Deep SORT tracking algorithms, and obtain camera data.
[0143] Step 3: Perform multi-sensor fusion of millimeter-wave radar data and camera data based on a fusion algorithm to enable real-time detection, tracking, and positioning of targets by roadside equipment.
[0144] In step 1, stationary targets are removed by setting the region of interest and speed thresholds, and random noise is removed by setting a threshold for the number of times a target appears in a continuous period of time, so as to extract possible valid targets. That is, millimeter-wave radar data is obtained by a millimeter-wave radar installed on a roadside pole. The millimeter-wave radar data includes the target's ID, longitudinal distance, lateral distance, longitudinal speed, lateral speed, and reflection cross-sectional area. Since the roadside pole is fixed, the road position information can be obtained in advance, thereby screening out targets located in the road area, and removing stationary targets by setting a speed threshold. There is a small amount of random noise and false alarm results in the millimeter-wave radar data. According to the number of times the ID index target appears in the first four scanning cycles, the index results with a number of appearances less than 3 times are filtered out to obtain valid targets.
[0145] In step 2, target detection and tracking are performed based on machine vision. Cameras installed on roadside poles are used to obtain image data. The image data is annotated and a dataset is created. The YOLOv3 model is trained based on the dataset. The training process is as follows:
[0146] First, freeze the backbone extraction network and only train the prediction network. After 50 cycles of training, train the entire model and end the training using Early Stopping.
[0147] The trained YOLOv3 model is used as the detector of the Deep SORT algorithm to perform multi-target tracking. The final output is [id c ,score,xmin ,y min ,x max ,y max ,class] m , that is, camera data, id c is the ID of the target detected by the camera, score is the confidence, (x min ,y min ) is the upper left vertex of the detection box, (x max ,y max ) is the lower right vertex of the detection box, class is the category, and m is the number of targets detected by the camera.
[0148] In step 3, the process of fusing the millimeter-wave radar data with the camera data based on the fusion algorithm specifically includes the following steps:
[0149] Step 301: Perform spatial fusion on the millimeter-wave radar data and the camera data. The points in the radar coordinate system are converted to the pixel coordinate system corresponding to the camera through the coordinate system to achieve spatial fusion of the two.
[0150] X w -Y w -Z w is the world coordinate system, O w is the origin of the world coordinate system, X r -Y r -Z r is the radar coordinate system, O r is the origin of the radar coordinate system, X c -Y c -Z c is the camera coordinate system, O c is the origin of the camera coordinate system, plane O w -X w -Y w Coincident with the ground, axis O w -Z w Along the roadside pole upwards, the origin of the camera coordinate system is O c To the origin O of the world coordinate system w The vertical height is h c , the origin of the camera coordinate system O c To the origin O of the world coordinate system w The longitudinal distance is l c , axis O c -Z c Coincident with the optical axis of the camera, axis O c -Z c With axis O w -X w The angle between the plane O and c -Yc -Z c With plane O w -X w -Z w Coincident, the origin of the radar coordinate system O r To the origin O of the world coordinate system w The vertical height is h r , the origin of the radar coordinate system O r To the origin O of the world coordinate system w The longitudinal distance is l r , plane O r -X r -Z r With plane O w -X w -Z w Coincident, axis O r -X r With axis O w -X w Parallel, uv is the pixel coordinate system, the point (X c ,Y c ,Z c ) is converted to the pixel coordinate system using the following formula:
[0151]
[0152] Among them, f x and f y is the equivalent focal length of the camera, (u0, v0) is the pixel coordinate of the center of the image, f x 、f y , u0 and v0 are the internal parameters of the camera, (X c ,Y c ,Z c ) is the three-dimensional coordinate of the camera coordinate system, and (u, v) is the two-dimensional coordinate of the pixel coordinate system;
[0153] According to the relative position relationship between the radar coordinate system and the camera coordinate system, the point (X r ,Y r ,Z r ) to pixel coordinates:
[0154]
[0155] Among them, (X r ,Y r ,Z r ) is the three-dimensional coordinate of the radar coordinate system, h r is the origin of the radar coordinate system O r To the origin O of the world coordinate system wThe vertical height, l r is the origin of the radar coordinate system O r To the origin O of the world coordinate system w The longitudinal distance, h c is the origin O of the camera coordinate system c To the origin O of the world coordinate system w The vertical height, l c is the origin O of the camera coordinate system c To the origin O of the world coordinate system w The longitudinal distance of the axis O c -Z c With axis O w -X w Angle;
[0156] The point (X w ,Y w ,Z w ) is converted to pixel coordinates:
[0157]
[0158] Among them, (X w ,Y w ,Z w ) is the three-dimensional coordinate of the world coordinate system;
[0159] Step 302: Fusing the millimeter-wave radar and the camera time, aligning the timestamps of the millimeter-wave radar data and the camera data for time fusion, that is, using an adaptive algorithm to align according to the timestamp of the camera data, completing the time fusion of the millimeter-wave radar data and the camera data, such as Figure 3 As shown in the figure, the black straight line represents the data stream released by the camera and millimeter-wave radar, and the black dots represent the data released by the camera and millimeter-wave radar in a complete cycle. It is judged whether the timestamp difference of the two data is within the set time difference threshold. If so, the two are considered to be aligned and connected by a dotted line.
[0160] Step 303: Fuse the millimeter-wave radar data and the camera data based on a fusion algorithm:
[0161] The output after parsing and filtering the millimeter-wave radar data in step 1, that is, the output after extracting valid targets from the millimeter-wave radar data, is:
[0162] [id r ,x r ,y r ,v xr ,v yr ] n
[0163] Among them, id r is the ID of the target detected by the millimeter radar wave, (x r ,y r ) is the radar coordinate of the target in the radar coordinate system, v xr and v yr are the longitudinal velocity and lateral velocity of the target in the radar coordinate system, respectively, and n represents the number of targets detected by the millimeter-wave radar;
[0164] The radar coordinates (x r ,y r ) is projected into the pixel coordinate system to obtain the pixel coordinate (x r_c ,y r_c ), and then the millimeter wave radar data is expanded to:
[0165] [id r ,x r ,y r ,v xr ,v yr ,x r_c ,y r_c ] n ;
[0166] Among them, (x r_c ,y r_c ) is the radar coordinate (x r ,y r ) is projected to the pixel coordinates in the pixel coordinate system;
[0167] Step 2: Target detection and tracking are performed on the image data of the target obtained by the drive test camera. The output of the camera data is:
[0168] [id c ,score,x min ,y min ,x max ,y max ,class] m
[0169] Among them, id c is the ID of the target detected by the camera, score is the confidence, x min and y min are the horizontal and vertical coordinates of the upper left vertex of the detection box, x max and y max are the horizontal and vertical coordinates of the lower right vertex of the detection box, class is the category, and m is the number of targets detected by the camera;
[0170] To detect the center point of the lower boundary of the box ((x min +x max) / 2,y max ) as the reference point, the coordinates of the world coordinate system corresponding to this reference point are in Z w = 0, the calculation formula for the coordinates of the point in the world coordinate system is:
[0171]
[0172] Among them, (X w ,Y w ,Z w ) is the three-dimensional coordinate of the world coordinate system, (x c_w ,y c_w ) is the coordinate of the target detected by the camera in the world coordinate system;
[0173] Then expand the camera data to:
[0174] [id c ,score,x min ,y min ,x max ,y max ,class,x c_w ,y c_w ] m
[0175] Among them, (x c_w ,y c_w ) is the coordinate of the center point of the lower boundary of the detection box in the world coordinate system;
[0176] Roadside cameras have a wide field of view, a small occlusion area, and good detection and tracking performance. They mainly use camera data, match it with millimeter-wave radar data about the target, and update it based on Kalman filtering to complete data fusion.
[0177] like Figure 4 The framework flow chart of the fusion algorithm shown in the figure expands the millimeter-wave radar data and camera data through coordinate transformation, performs preliminary matching data selection on the millimeter-wave radar data and camera data, then associates the expanded millimeter-wave radar data with the expanded camera data, matches and fuses the camera data with the millimeter-wave radar data, and updates them based on Kalman filtering to obtain the track of the fused data.
[0178] The process of updating based on Kalman filtering specifically includes the following steps:
[0179] Step a: Initialize. If the camera data has matching millimeter radar wave data, then
[0180]
[0181] Among them, xr +l r is the longitudinal coordinate of the millimeter radar wave data in the world coordinate system, l r =0.16m, λ1 and λ2 are weights, take λ2=0.8, X k is the state variable, dis x 、dis y 、vel x and vel y They are the longitudinal distance, lateral distance, longitudinal speed and lateral speed of the fusion target in the world coordinate system, Y k =X k is the observed variable;
[0182] The weight λ1 and the target's longitudinal coordinate x in the radar coordinate system r The relational expression is:
[0183]
[0184] If the camera data does not have matching millimeter radar wave data, the state variables are initialized, and the expressions of the state variables and state covariance matrix are:
[0185]
[0186] Among them, I 4×4 is a 4×4 unit matrix;
[0187] Process noise covariance matrix q k The expression is:
[0188] q k =0.1I 4×4 ;
[0189] Observation noise covariance matrix r k The expression is:
[0190] r k =0.1I 4×4
[0191] State transition matrix f k The expression is:
[0192]
[0193] Where t is the time difference between the current moment and the previous moment;
[0194] Observation matrix h k The expression is:
[0195] h k =I 4×4 ;
[0196] Step b: Perform prediction and update to obtain the predicted state vector and predicted state covariance matrix:
[0197]
[0198] Among them, X k+1|k is the predicted state vector at time k+1, f k+1 is the state transition matrix at time k+1, P k+1|k is the predicted state covariance matrix at time k+1, q k+1 is the process noise covariance matrix at time k+1, T represents the transpose;
[0199] According to the predicted state vector and the predicted state covariance matrix, the observation vector and the observation space covariance matrix are obtained:
[0200]
[0201] Among them, Y k+1 is the observation vector at time k+1, h k+1 is the state transition matrix at time k+1, S k+1 is the observation space covariance matrix at time k+1, r k+1 is the observation noise covariance matrix at time k+1;
[0202] Update the state vector and state covariance matrix to obtain the posterior estimates of the state vector and state covariance matrix:
[0203]
[0204] K k+1 is the gain at time k+1, Z k+1 is the actual measurement value at time k+1, X k+1 is the posterior estimate of the state vector at time k+1, P k+1 is the posterior estimate of the state covariance matrix at time k+1;
[0205] By performing effective target extraction on the millimeter-wave radar data, the interference signal can be effectively filtered out. However, irrelevant signals still exist, that is, there are signals that are not pedestrians or vehicles, and pedestrians or vehicles may not have radar signals. Therefore, it is necessary to screen out the camera data and radar data involved in data association, that is, to perform preliminary matching data selection on the millimeter-wave radar data and camera data, such as Figure 5 As shown in the figure, the pixel coordinates of a set of millimeter wave radar data after spatial fusion and temporal fusion are (x r_c ,y r_c ), the coordinate of the upper left vertex of the detection box is (x min ,y min), the lower right vertex is (x max ,y max ), the coordinate of the upper left vertex of the extended detection box is (x min -Δx,y min -Δy), the coordinates of the lower right vertex are (x max +Δx,y max +Δy), the echo signal received by the millimeter radar wave is reflected by a part of the target, so the radar coordinates are projected to the pixel coordinates. The coordinates are within the target detection frame. Due to errors in the calibration process and the deviation between the target detection frame and the minimum rectangular frame surrounding the target, it is necessary to expand the detection frame. You can choose Δx = 2Δy = 20 pixels, that is, based on the detection frame, an extended detection frame is obtained along the u-axis and v-axis. The matching millimeter wave radar data is filtered through the expanded detection extension frame.
[0206] After the initial selection of matching data, M camera data and N millimeter-wave radar data are obtained. Suppose the expressions of camera data and millimeter-wave radar data are:
[0207] [id c ,score,x min ,y min ,x max ,y max ,class,x c_w ,y c_w ] M
[0208] [id r ,x r ,y r ,v xr ,v yr ,x r_c ,y r_c ] N
[0209] Among them, id c is the ID of the target detected by the camera, score is the confidence, x min and y min are the horizontal and vertical coordinates of the upper left vertex of the detection box, x max and y max They are the horizontal and vertical coordinates of the lower right vertex of the detection box, class is the category, (x c_w ,y c_w ) is the coordinate of the center point of the lower boundary of the detection box in the world coordinate system, M is the number of targets detected by the camera, that is, the number of camera data, id r is the ID of the target detected by the millimeter radar wave, (x r ,y r) is the radar coordinate of the target detected by the millimeter radar wave in the radar coordinate system, v xr and v yr are the longitudinal velocity and lateral velocity of the target detected by the millimeter radar wave in the radar coordinate system, (x r_c ,y r_c ) is the pixel coordinate of the target detected by the millimeter radar wave, M is the number of targets detected by the camera, that is, the number of camera data, and N is the number of targets detected by the millimeter radar wave, that is, the number of millimeter radar wave data;
[0210] Since at least one millimeter-wave radar data falls within the detection expansion box output by each camera, the camera data is always less than or equal to the millimeter-wave radar data, that is, M≤N, as shown in Figure 5 As shown, a detection expansion box may contain multiple radar data points, and a millimeter-wave radar data may fall within multiple detection expansion boxes. Therefore, it is necessary to further calculate the cost of each millimeter-wave radar data belonging to each detection expansion box. The present invention uses the Mahalanobis distance between the millimeter-wave radar data and the camera data in the pixel coordinate system and the Euclidean distance between the millimeter-wave radar data and the camera data in the world coordinate system as the cost matrix, and obtains the matching method with the lowest global cost based on the Hungarian algorithm:
[0211] 1) Calculate the Mahalanobis distance between the millimeter-wave radar data and the camera data in the pixel coordinate system:
[0212] For the i-th target detected by the camera, if the target is detected for the first time, the pixel coordinates of the camera data are initialized with the center of the detection frame, and the unit matrix I is used. 2×2 Initialize the covariance matrix. The expressions of the pixel coordinates and covariance matrix of the camera data are:
[0213]
[0214] in, is the center coordinate of the detection box, d i is the pixel coordinate corresponding to the i-th target detected by the camera, S i is the covariance matrix corresponding to the i-th target detected by the camera;
[0215] If the target belongs to the track of the existing fusion data, the covariance matrix of the observation variable and the observation space is calculated. Since the speed information is not involved, the pixel coordinates corresponding to the position information of the observation variable are assigned to the pixel coordinate di, and the position covariance in the observation space covariance matrix is assigned to the covariance matrix Si:
[0216]
[0217]
[0218] Among them, S k+1 is the observation space covariance matrix at time k+1, h is a 2×4 matrix, and T represents the transpose of the matrix;
[0219] According to d i and S i Calculate the Mahalanobis distance between the millimeter-wave radar data and the camera data in the pixel coordinate system. The calculation formula of the Mahalanobis distance is:
[0220]
[0221] d j =[x r_c ,y r_c ]
[0222] Among them, dis image is the Mahalanobis distance between the millimeter-wave radar data and the camera data in the pixel coordinate system, d j is the pixel coordinate corresponding to the jth target detected by the millimeter-wave radar;
[0223] 2) Calculate the Euclidean distance between the millimeter-wave radar data and the camera data in the pixel coordinate system:
[0224] The coordinates of the jth target detected by the millimeter wave radar in the world coordinate system are (x r +l r ,y r ), the coordinates of the i-th target detected by the camera in the world coordinate system are (x c_w ,y c_w ), the formula for calculating the Euclidean distance between the two is:
[0225]
[0226] Among them, dis world is the Euclidean distance between the millimeter-wave radar data and the camera data in the world coordinate system, (x r +l r ,y r ) is the world coordinate corresponding to the jth target detected by the millimeter wave radar, (x c_w ,y c_w ) is the world coordinate corresponding to the i-th target detected by the camera;
[0227] The obtained Mahalanobis distance and Euclidean distance are normalized respectively, and then the weighted sum is taken as the cost of the j-th radar data belonging to the i-th camera data. The cost expression is:
[0228]
[0229] Among them, task_matrix ij is the cost of matching the j-th radar data with the i-th camera data, λ is the weight between the Mahalanobis distance and the Euclidean distance, which is 0.9;
[0230] Obtain the global cost matrix, and obtain the matching result with the best global cost based on the Hungarian algorithm according to the global cost matrix. For a certain matching result, if the corresponding cost is greater than the set cost threshold, the matching result is considered to be wrong and removed. For example, if the corresponding cost is greater than 0.99995, the matching result is considered to be wrong and removed. This is because there is a possibility that the camera detects the target, the millimeter radar wave does not detect the target, and there is a possibility that an unfiltered interference signal falls into the detection expansion box.
[0231] After the matching data is initially selected and the data is associated, it is integrated and updated:
[0232] First, determine whether the target ID belongs to an existing track. If not, initialize the state variable. If so, that is, for the camera data that matches the millimeter radar wave data, the actual measurement value Z k+1 is [λ1(x r +l r )+(1-λ1)x c_w ,λ2y r +(1-λ2)x c_w ,v xr ,v yr ] T Update and get the posterior estimate X k+1 ;
[0233] If not, that is, for the camera data that does not match the millimeter radar wave data, [x c_w ,y c_w ,0,0] T As the actual measured value Z k+1 However, due to the large error in camera distance estimation, the observation noise covariance matrix r needs to be increased. k , take r k =I 4×4 , and obtain the posterior estimate X k+1 .
[0234] 3) After fusion and updating, the latest fusion data of the target is:
[0235] [id,class,score,x min ,y min ,x max ,y max ,disx ,dis y ,vel x ,vel y ] M
[0236] Among them, dis x 、dis y 、vel x and vel y is the updated posterior estimate based on Kalman filtering, dis x 、dis y 、vel x and vel y They represent the longitudinal distance, lateral distance, longitudinal speed and lateral speed of the fusion target in the world coordinate system respectively;
[0237] 4) Use age and moment to manage the generation and disappearance of the track of fused data. When the camera detects a target, a track is generated. The fused track is initialized based on the camera's track. At this time, the age of all tracks is set to 1, and the moment of all tracks is set to the timestamp of the current data. When a new camera track appears in the camera data received in the next cycle, the new camera track is added to the track of the fused data, and the age of the track is set to 1, and the moment of the track is set to the timestamp of the corresponding camera data. If the camera track has the same id as the track corresponding to the previous cycle, the new camera track is used to update it, the age is still 1, and the moment is updated to the timestamp of the camera data corresponding to the new camera track. If the track that existed in the previous cycle does not have a corresponding track in this cycle, the age of the track is increased by 1, and the moment remains unchanged. When the age of a track reaches 30, it is considered that the track has left the scene and is deleted from the track set.
[0238] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A roadside perception method integrating millimeter-wave radar and camera, characterized in that: The method comprises the following steps: Step 1: Obtain millimeter-wave radar data through the millimeter-wave radar, extract effective targets, and remove stationary targets and random noise; Step 2: Obtain image data through a camera installed on a roadside pole, perform target detection and multi-target tracking based on the YOLOv3 and Deep SORT tracking algorithms, and obtain camera data; Step 3: Perform multi-sensor fusion of millimeter-wave radar data and camera data based on a fusion algorithm to enable real-time detection, tracking, and positioning of targets by roadside equipment. In step 3, the process of fusing the millimeter-wave radar data with the camera data based on the fusion algorithm specifically includes the following steps: Step 301: spatially fuse the millimeter-wave radar data and the camera data by converting the points in the radar coordinate system into the pixel coordinate system corresponding to the camera through the coordinate system, thereby achieving spatial fusion of the two; Step 302: Fusing the millimeter-wave radar and camera time, aligning the timestamps of the millimeter-wave radar data and the camera data for time fusion, i.e., using an adaptive algorithm to align based on the timestamp of the camera data, to complete the time fusion of the millimeter-wave radar data and the camera data. Step 303: Fusing the millimeter-wave radar data and the camera data based on a fusion algorithm; In step 303, the process of fusing the millimeter-wave radar data and the camera data based on the fusion algorithm specifically includes the following steps: Step 303A: performing data expansion on the millimeter-wave radar data and the camera data through coordinate transformation; Step 303B: Perform preliminary matching of the millimeter-wave radar data and the camera data; Step 303C: Data association is performed on the millimeter-wave radar data after the initial matching data selection and the expanded camera data. The cost of each millimeter-wave radar data belonging to each detection frame is calculated. Specifically, the Mahalanobis distance between the millimeter-wave radar data and the camera data in the pixel coordinate system and the Euclidean distance between the millimeter-wave radar data and the camera data in the world coordinate system are used as the cost matrix. The matching method with the lowest global cost is obtained based on the Hungarian algorithm. Step 303D: Based on the visual tracking results, the millimeter-wave radar data about the target is matched. That is, based on the camera data, the camera data and the millimeter-wave radar data are matched and fused, and updated based on the Kalman filter to obtain the track of the fused data; In step 303C, the process of data association specifically includes the following steps: Step 303C1: At least one millimeter-wave radar data falls within the detection frame output by each camera, that is, the number of camera data M Less than or equal to the number of millimeter wave radar data N , let the expressions of camera data and millimeter wave radar data be: in, The ID of the target detected by the camera. is the confidence level, and are the horizontal and vertical coordinates of the upper left vertex of the detection box, and are the horizontal and vertical coordinates of the lower right vertex of the detection frame, For categories, is the coordinate of the center point of the lower boundary of the detection box in the world coordinate system, The number of targets detected by the camera, that is, the number of camera data, The ID of the target detected by the millimeter radar wave. is the radar coordinate of the target detected by the millimeter radar wave in the radar coordinate system, and are the longitudinal velocity and lateral velocity of the target detected by the millimeter radar wave in the radar coordinate system, is the pixel coordinate of the target detected by the millimeter radar wave, The number of targets detected by millimeter radar waves, that is, the number of millimeter radar wave data; Step 303C2: Calculate the Mahalanobis distance between the millimeter-wave radar data and the camera data in the pixel coordinate system: For the first i If the target is detected for the first time, the pixel coordinates of the camera data are initialized with the center of the detection frame, and the unit matrix is used. I 2×2 Initialize the covariance matrix. The expressions of the pixel coordinates and covariance matrix of the camera data are: in, is the center coordinate of the detection frame, The first The pixel coordinates corresponding to the target, The first The covariance matrix corresponding to the targets; If the target belongs to the track of the existing fusion data, the covariance matrix of the observation variable and the observation space is calculated. Since the speed information is not involved, the pixel coordinates corresponding to the position information of the observation variable are taken. Pixel coordinates d i Assign values and take the position covariance to covariance matrix in the observation space covariance matrix S i Assignment: in, for The observation space covariance matrix at time , for The matrix, Represents the transpose of a matrix; According to pixel coordinates d i and covariance matrix S i Calculate the Mahalanobis distance between the millimeter-wave radar data and the camera data in the pixel coordinate system. The calculation formula of the Mahalanobis distance is: in, is the Mahalanobis distance between the millimeter-wave radar data and the camera data in the pixel coordinate system, The first The pixel coordinates corresponding to the target; Step 303C3: Calculate the Euclidean distance between the millimeter-wave radar data and the camera data in the world coordinate system. The calculation formula of the Euclidean distance is: in, is the Euclidean distance between the millimeter-wave radar data and the camera data in the world coordinate system, The first j The world coordinates corresponding to the target, The first i The world coordinates corresponding to the target; is the origin of the radar coordinate system To the origin of the world coordinate system The longitudinal distance; Step 303C4: Normalize the obtained Mahalanobis distance and Euclidean distance respectively, and then perform weighted summation to obtain the first j The millimeter-wave radar data matches the i The cost of camera data is calculated as follows: in, For the j The radar data matches the i The cost of camera data, λ is the weight between Mahalanobis distance and Euclidean distance; Step 303C5: Obtain a global cost matrix, and obtain a matching result with the best global cost based on the Hungarian algorithm. For a matching result, if the corresponding cost is greater than the set cost threshold, the matching result is considered to be wrong and is removed.
2. The roadside perception method integrating millimeter-wave radar and camera according to claim 1, characterized in that: In step 1, the process of extracting effective targets specifically includes the following steps: Step 101: Acquire millimeter-wave radar data through a millimeter-wave radar installed on a roadside pole. The millimeter-wave radar data includes the target's ID, longitudinal distance, lateral distance, longitudinal speed, lateral speed, and reflection cross-sectional area. Step 102: Obtaining road location information of a set road area based on fixed roadside poles, and filtering out targets located within the road area based on the road location information; Step 103: removing stationary objects within the road area based on a set speed threshold; Step 104: Set a threshold for the number of times a target exists within a continuous period of time, index the number of times the target exists in the first four scanning cycles based on the target ID, remove index results with the number of times the target exists less than three times, and then obtain valid targets.
3. The roadside perception method integrating millimeter-wave radar and camera according to claim 1, characterized in that: In step 2, the process of target detection and tracking based on YOLOv3 and Deep SORT tracking algorithms specifically includes the following steps: Step 201: Acquire image data through a camera installed on a roadside pole; Step 202: Annotate the image data and create a dataset. Train the YOLOv3 model based on the dataset. First, freeze the backbone extraction network and train only the prediction network. After 50 cycles of training, train the entire model and terminate the training using Early Stopping. Step 203: Use the trained YOLOv3 model as the detector of the Deep SORT algorithm to perform multi-target tracking. The final output is , that is, camera data, The ID of the target detected by the camera. score is the confidence level, is the upper left vertex of the detection box, is the lower right vertex of the detection box, class For categories, m Indicates the number of targets detected by the camera.
4. The roadside perception method integrating millimeter-wave radar and camera according to claim 1, characterized in that: The step 301 is specifically as follows: set up is the world coordinate system, is the origin of the world coordinate system, is the radar coordinate system, is the origin of the radar coordinate system, is the camera coordinate system, is the origin of the camera coordinate system, the plane Coincident with the ground, axis Along the roadside pole facing upwards, the origin of the camera coordinate system To the origin of the world coordinate system The vertical height is , the origin of the camera coordinate system To the origin of the world coordinate system The longitudinal distance is ,axis Coincident with the optical axis of the camera, axis With axis The angle is ,flat With plane Coincident, the origin of the radar coordinate system To the origin of the world coordinate system The vertical height is , the origin of the radar coordinate system To the origin of the world coordinate system The longitudinal distance is ,flat With plane coincidence, axis With axis parallel, For pixel coordinate system, the point in the camera coordinate system is The calculation formula for the coordinates converted to the pixel coordinate system is: in, and is the equivalent focal length of the camera, is the pixel coordinate of the center of the image, 、 、 u 0 and v 0 are all internal parameters of the camera, is the three-dimensional coordinate of the camera coordinate system, is the two-dimensional coordinate of the pixel coordinate system; According to the relative position relationship between the radar coordinate system and the camera coordinate system, the point in the radar coordinate system is The calculation formula for the coordinates converted to the pixel coordinate system is: in, is the three-dimensional coordinate of the radar coordinate system, is the origin of the radar coordinate system To the origin of the world coordinate system The vertical height, The origin of the camera coordinate system To the origin of the world coordinate system The vertical height, The origin of the camera coordinate system To the origin of the world coordinate system The longitudinal distance, Axis With axis Angle; The point in the world coordinate system The calculation formula for the coordinates converted to the pixel coordinate system is: in, is the three-dimensional coordinate of the world coordinate system.
5. The roadside perception method integrating millimeter-wave radar and camera according to claim 1, characterized in that: In step 303A, the data expansion process is specifically as follows: Step 303A1: The output after extracting valid targets from the millimeter-wave radar data is: in, The ID of the target detected by the millimeter radar wave. is the radar coordinate of the target detected by the millimeter radar wave in the radar coordinate system, and are the longitudinal velocity and lateral velocity of the target in the radar coordinate system, n Indicates the number of targets detected by the millimeter-wave radar; Step 303A2: Combine radar coordinates through spatial fusion Projected to the pixel coordinate system to get the pixel coordinates , and then expand the millimeter wave radar data to: ; in, is the radar coordinate Pixel coordinates projected into the pixel coordinate system; Step 303A3: The output of target detection and tracking performed on the image data of the target acquired by the roadside camera is: in, The ID of the target detected by the camera. is the confidence level, and are the horizontal and vertical coordinates of the upper left vertex of the detection box, and are the horizontal and vertical coordinates of the lower right vertex of the detection box, For categories, Indicates the number of targets detected by the camera; Step 303A4: Detect the center point of the lower boundary of the frame As a reference point, the coordinates of the reference point in the world coordinate system are Z w = 0, the calculation formula for the coordinates of the point in the world coordinate system is: in, is the three-dimensional coordinate of the world coordinate system, The coordinates of the target detected by the camera in the world coordinate system; Step 303A5: Expand the camera data to: 。 6. The roadside perception method integrating millimeter-wave radar and camera according to claim 1, characterized in that: In step 303B, the process of performing preliminary matching of the millimeter-wave radar data and the camera data is specifically as follows: The pixel coordinates of a set of millimeter-wave radar data after spatial fusion and temporal fusion are: , the coordinates of the upper left vertex of the detection box are , the lower right vertex is , the coordinates of the upper left vertex of the extended detection box are , the coordinates of the lower right vertex are , that is, based on the detection frame, an extended detection frame is obtained along the u-axis and v-axis, and the matching millimeter-wave radar data is filtered through the expanded detection extension frame. The remaining camera data after filtering M targets, the remaining N targets, that is, the number of camera data is M , the amount of millimeter wave radar data is N .
7. The roadside perception method integrating millimeter-wave radar and camera according to claim 1, characterized in that: In step 303D, if the target id For tracks that are not part of the existing fusion data, the process of updating based on the Kalman filter includes the following steps: Step 303D1: Target id For tracks that do not belong to the existing fusion data, that is, the camera data does not have matching millimeter radar wave data, the state variables are initialized, and the expressions of the state variables and state covariance matrix are: in, for The unit array; Process noise covariance matrix The expression is: ; Observation noise covariance matrix The expression is: State transition matrix The expression is: in, t is the time difference between the current moment and the previous moment; Observation Matrix The expression is: ; Step 303D2: Perform prediction and update to obtain the predicted state vector and predicted state covariance matrix: in, for The predicted state vector at time , for The state transition matrix at time , for The predicted state covariance matrix at time , for The process noise covariance matrix at time , represents transpose; According to the predicted state vector and the predicted state covariance matrix, the observation vector and the observation space covariance matrix are obtained: in, for The observation vector at time t, for The observation matrix at time t, for The observation space covariance matrix at time , for The observation noise covariance matrix at time t; Actual measured value at the moment and the observation noise covariance matrix The expression is: Update the state vector and state covariance matrix to obtain the posterior estimates of the state vector and state covariance matrix: for The gain of time, for The actual measured value at the moment, for The posterior estimate of the state vector at time , for The posterior estimate of the state covariance matrix at time t; If the target id The process of updating the existing fused data track based on Kalman filtering includes the following steps: Step 303D3: Target id The track belongs to the existing fusion data, that is, the camera data has matching millimeter radar wave data, then the expressions of state variables, state covariance matrix and observation variables are: in, is the vertical coordinate of the millimeter radar wave data in the world coordinate system, λ 1 and λ 2 is the weight, is the state variable, dis and vel are the position and velocity of the target in the world coordinate system, dis x 、 dis y 、 vel x and vel y are the longitudinal distance, lateral distance, longitudinal speed and lateral speed of the target in the world coordinate system respectively. is the observed variable, is a unit array; Weight λ The expression for 1 is: in, It is the longitudinal coordinate of the target detected by millimeter radar wave in the radar coordinate system; Step 303D4: Perform prediction and update to obtain the predicted state vector and predicted state covariance matrix: in, for The predicted state vector at time , for The state transition matrix at time , for The predicted state covariance matrix at time , for The process noise covariance matrix at time , represents transpose; According to the predicted state vector and the predicted state covariance matrix, the observation vector and the observation space covariance matrix are obtained: in, for The observation vector at time t, for The observation matrix at time t, for The observation space covariance matrix at time , for The observation noise covariance matrix at time t; Actual measured value at the moment The expression is: Update the state vector and state covariance matrix to obtain the posterior estimates of the state vector and state covariance matrix: for The gain of time, for The actual measured value at the moment, for The posterior estimate of the state vector at time , for The posterior estimate of the state covariance matrix at time t; Step 303D5: Get the latest fusion data of the target: in, is the target ID, dis x 、 dis y 、 vel x and vel y is the updated posterior estimate based on Kalman filtering; Step 303D6: Obtain the track of the fused data, and use age and moment to manage the generation and disappearance of the track of the fused data.
8. The roadside perception method integrating millimeter-wave radar and camera according to claim 7, characterized in that: In step 303D6, the generation and disappearance of the track of the fused data is managed by using age and moment as follows: After the camera detects a target, a track is generated. The fused track is initialized based on the camera's track. At this time, the age of all tracks is set to 1, and the moment of all tracks is set to the timestamp of the current data. When a new camera track appears in the camera data received in the next cycle, the new camera track is added to the track of the fused data, and the age of the track is set to 1, and the moment of the track is set to the timestamp of the corresponding camera data. If the camera track has the same id as the track corresponding to the previous cycle, the new camera track is used to update it, the age is still 1, and the moment is updated to the timestamp of the camera data corresponding to the new camera track. If the track that existed in the previous cycle does not have a corresponding track in this cycle, the age of the track is increased by 1, and the moment remains unchanged. When the age of a track reaches the set threshold, it is considered that the track has left the scene and is deleted from the track set.
Citation Information
Patent Citations
Image information and radar information fusion method and system for traffic scene
CN108983219A