Millimeter-wave radar and vision fused vehicle three-dimensional detection method
Through the post-fusion technology of millimeter wave radar and camera, three-dimensional detection of vehicles is realized, solving the problem of accurate identification and perception of traffic monitoring systems in complex environments, and improving the stability and accuracy of detection.
Patent Information
- Application Number
- CN202510445554.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-11
AI Technical Summary
The existing traffic monitoring systems rely on two-dimensional image processing, making it difficult to accurately identify and perceive three-dimensional vehicle information in complex environments, and there are difficulties in data synchronization and spatial calibration of millimeter-wave radar and vision sensors.
Through the post-fusion technology of millimeter-wave radar and camera, radar data and image data are used to fusion time and space, calculate the vehicle's three-dimensional frame information, and draw three-dimensional annotations through coordinate conversion algorithms to realize the vehicle's three-dimensional detection.
It improves the accuracy and reliability of vehicle detection, and can stably identify the existence, size and direction of the vehicle in complex environments, make up for the shortcomings of a single sensor, and achieves more comprehensive traffic target perception.
Smart Images

Figure BDA0005352469600000031 
Figure BDA0005352469600000041 
Figure BDA0005352469600000042
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing, and particularly relates to a three-dimensional vehicle detection method by fusing millimeter-wave radar and vision. Background Art
[0002] With the continuous development and intelligentization of urban traffic, visual traffic monitoring systems, as an important part of traffic management and safety, have broad application prospects. These systems use cameras and other sensor devices to capture real-time traffic scenes, and use computer vision algorithms to analyze and process images or videos to extract traffic information and key data. Visual traffic monitoring systems include functions such as vehicle detection, vehicle tracking, license plate recognition, and behavior analysis, which are used to extract traffic information and features. Visual traffic monitoring systems have a wide range of applications in traffic management, urban planning, traffic safety, etc. They can assist traffic management departments to better monitor and manage traffic flow, improve road usage efficiency, reduce the probability of traffic congestion and accidents, and provide a safer and more convenient traffic environment for urban residents. Traffic target detection from the roadside perspective particularly needs to consider the ability of sensors to acquire and understand environmental information. From the roadside perspective, complex traffic scenes and uncertain environmental conditions (such as vehicle occlusion, dynamic traffic flow, etc.) pose challenges to target detection.
[0003] To make up for the deficiencies of single sensors and improve the accuracy and reliability of environmental perception, multi-sensor fusion technology is usually adopted for traffic target detection. Sensor fusion technology is divided into two types: pre-fusion and post-fusion. Pre-fusion refers to data fusion at the sensor level, while post-fusion is data fusion at the feature extraction or decision-making level. Currently, most autonomous driving systems adopt pre-fusion technology to preliminarily fuse data from different sensors at the sensor level. This method can utilize the advantages of multiple sensors, but its disadvantage is that data processing is complex and it is difficult to cope with various environmental changes. In contrast, post-fusion technology can more flexibly adapt to environmental changes, improve the robustness and accuracy of the system by separately processing the data of each sensor and then fusing at the feature or decision-making level.
[0004] Applying the post - fusion of millimeter - wave radar and camera to target detection can make full use of the advantages of the two sensors and make up for their respective deficiencies. By using sensor fusion technology, combining millimeter - wave radar and vision sensors can make up for their respective deficiencies, thus improving the overall detection effect. Millimeter - wave radar: It has strong penetration and can still provide accurate distance, speed and angle information in complex weather (such as rain, snow, fog, etc.) and low - light environments. It is suitable for target detection in scenarios such as long - distance, low - light, and bad weather. The disadvantage of millimeter - wave radar is the lack of detailed information about the target, especially the inability to provide the appearance or semantic features of the target. Vision sensor (camera): It can provide rich image information, especially clearly identify the appearance features of the target. However, in poor lighting, rain, fog and other harsh environments, the performance of the vision sensor will be significantly affected. Through the post - fusion algorithm, the advantages of the sensors can be combined to more accurately perceive the surrounding environment and improve the accuracy and reliability of detection.
[0005] Traditional traffic monitoring systems mainly rely on two - dimensional image processing, which can only provide limited information, thus restricting the accurate recognition and stereo perception ability of traffic targets. To solve this problem, researchers have begun to integrate 3D object detection technology in the field of computer vision into roadside monitoring viewpoints to provide more comprehensive and accurate traffic object information. As one of the methods, 3D traffic target detection from images aims to estimate the 3D position and orientation of traffic targets based on image information. By using 3D detection, the system can accurately perceive the position, size and orientation information of vehicles, thus achieving accurate perception of vehicles in 3D space. This helps to better understand vehicle motion behavior and spatial layout.
[0006] Although the post - fusion technology of vision and millimeter - wave radar has significant advantages, there are still some challenges in the implementation process. The first is data time synchronization and spatial alignment. Due to the different working principles of vision sensors and millimeter - wave radars, the data acquisition frequencies and delays are also different. How to achieve accurate synchronization and spatial calibration of the two data is a difficult point. In 3D target detection, the actual position, yaw angle and ground - plane information of the target need to be considered. The design and implementation of the algorithm need to consider various factors, such as the real - time nature of data processing, the complexity of the algorithm and the requirements of computing resources. In addition, a large number of experiments and tests are required to verify the effectiveness and reliability of the algorithm. Summary of the Invention
[0007] Aiming at the problems in the prior art, the present invention provides a method for three - dimensional vehicle detection by fusing millimeter - wave radar and vision.
[0008] A three-dimensional vehicle detection method integrating millimeter-wave radar and vision of the present invention obtains the pixel position information corresponding to the actual position of the vehicle in the image by performing spatio-temporal fusion on millimeter-wave radar data and camera images, calculates the three-dimensional frame information of the vehicle using vision algorithms, locates it to the target position according to the coordinate conversion algorithm, judges the target through the radar-vision fusion algorithm, and draws three-dimensional annotations at the location of the target to achieve three-dimensional vehicle detection; specifically includes the following steps:
[0009] Step 1: Simultaneously collect the image data of the camera and the point cloud data of the millimeter-wave radar, record the acquisition time of each piece of data, and use the vision detection algorithm to perform target detection on the image captured by the camera to obtain the position reference of the vehicle target in the image.
[0010] Step 2: Analyze the original radar point cloud data, calculate the horizontal and vertical distances of the target relative to the radar, use the Cartesian coordinate system to construct the two-dimensional coordinate point position of the target in the radar coordinate system, and at the same time perform filtering processing on the radar data to remove noise interference and retain effective target information.
[0011] Step 3: Complete the radar data frame and the camera image frame by the frame filling method, obtain the transformation matrix by the least square method, and align the radar data and the camera data frame by frame in space and time.
[0012] Step 4: Map the radar data obtained in Step 3 and the prior target detection framework obtained in Step 1 to each other to obtain the final target detection result.
[0013] Step 5: Draw the target 3D frame at the target position obtained in Step 4.
[0014] Further, in Step 1, the target detection algorithm is used to obtain the prior 2D detection framework, obtain the prior frameworks (x1, y1) and (x2, y2) of the target, and further obtain the center point coordinates (x center , y center ) of the target in the image.
[0015] Further, in Step 2, according to the position, scattering area, and life cycle of the radar-detected target, invalid targets and ghost targets are filtered; invalid targets include targets in unconcerned areas, false targets, and empty targets.
[0016] Filter in different ways according to experience:
[0017] For targets in unconcerned areas: Set the horizontal distance threshold as Y and the vertical distance threshold as X, compare the horizontal coordinate y of the target with the threshold Y, and compare the vertical distance x of the target with the threshold X. If |x| > X or |y| > Y for the target, then filter out the target.
[0018] For false targets: According to the consecutive radar frame rates in chronological order, a threshold λ is set. When |d t+1 - d t | > λ, it is considered a false target and filtered out. d is the position vector d(x, y) of the target in the radar coordinate system, and d t+1 and d t represent the position vectors of the target at frames t + 1 and t respectively.
[0019] For empty targets: Targets with a radar cross - section (RCS) of zero are filtered out.
[0020] For ghost targets: The life cycle of the target is divided into three stages: appearance, persistence, and disappearance. When a new target appears, detect = 1 and lost = 0; for each frame that appears, detect += 1 and lost = 0; for each frame that disappears, lost is decremented by 1; the parameter for the appearance state is detect >= D, indicating that the target has appeared for multiple consecutive frames; the parameter for the persistence state is detect >= D and lost <= S, indicating that the number of frames the target has appeared is greater than or equal to the threshold D and the number of disappearances is less than or equal to S, and at this time the target is a target to be tracked; lost >= S indicates that the target has disappeared for multiple consecutive frames, and at this time the target is in the disappearance state; a ghost target is a target that appears for only a few frames and then disappears. At this time, by adjusting the thresholds D and S, ghost targets are filtered out, and only valid targets are output.
[0021] Furthermore, in step 3, the collected radar data and camera data are aligned respectively. Using the frame - filling method, the x and y between two radar frames for each target are calculated respectively, and the position of the target at the moment of the image timestamp is calculated through the timestamps between the two frames, thereby achieving frame - filling time alignment.
[0022] The plane scanned by the meter - wave radar is parallel to the ground and the height is the installation height of the radar. A coordinate system is constructed with the sensor center as the origin. The world coordinate system is O W (X W , Y W , Z W ), the radar coordinate system is O R (X R , Y R , Z R ), and the camera coordinate system is O C (X C , Y C , Z C ); The distances between the origins of the world coordinate system and the radar coordinate system on each axis are D(D X , D Y , D Z ), then the representation of the target in the world coordinate system is:
[0023]
[0024] The conversion from the world coordinate system to the camera coordinate system uses the pinhole imaging model. Assume that the camera coordinate system is O C (X C ,Y C ,Z C ). The coordinates of the target point in the camera coordinate system are (X0, Y0, Z0), and the projected coordinates in the image are (X img ,Y img ). The world coordinate system and the camera coordinate system only differ in the origin position and the axis directions. The world coordinate system can be transformed into the camera coordinate system through rotation and translation. A 3×3 is the rotation matrix, B 3×1 is the translation column vector, and O 1×3 is a row vector of three 0s. The conversion formula from world coordinates to camera coordinates is:
[0025]
[0026] The image coordinate system O img (X img ,Y img ) is a two-dimensional coordinate system. Let the camera focal length be f. Through the principle of similar triangles and pinhole imaging, the conversion formula between camera coordinates and image coordinates is:
[0027]
[0028] Written in matrix form as:
[0029]
[0030] The pixel coordinate system O(U, V) is a two-dimensional coordinate system. During the alignment of the image coordinate system and the pixel coordinate system, the origin of the image coordinate system is the center of the image, while the origin of the pixel coordinate system is the upper left vertex of the image. The horizontal axis is to the right horizontally, and the vertical axis is downward vertically. Assume that the projection of the origin of the image coordinate system in the pixel coordinate system is (U0, V0), and the projection of the target (X img ,Y img ) in the pixel coordinate system is (U, V). The values represented by each pixel in the X img axis and the Y img axis directions are d X and d Y , respectively. Then the conversion formula between the pixel coordinate system and the image coordinate system is:
[0031]
[0032] Through all the above transformations, the conversion formula between the radar coordinate system and the pixel coordinate system is:
[0033]
[0034] The above formula indicates that the radar coordinate system and the pixel coordinate system can correspond through a matrix. The radar itself used does not return pitch information. By introducing the radar installation height, ZR becomes a fixed value. Regarding the inverse matrix of the product of the middle three transformation formulas as T, the radar vector as R, and the pixel vector as P, it can be known that P = TR. The formula for solving matrix T by the least squares method is: T = (RR T ) -1 RP T .
[0035] Furthermore, in step 4, the target of the millimeter-wave radar signal obtained at the same timestamp in the radar coordinate system is mapped to the prior data center position of the target obtained in step 1, and the Euclidean distance d is calculated mutually. Let be the Euclidean distance between the i-th target in the image at time t and the j-th target in the millimeter-wave radar. Retain the minimum distance between the target in each image and the millimeter-wave radar target, divide the distances, and determine whether the detection results of the two sensors are the same target to complete the fusion detection results.
[0036] Furthermore, if the millimeter-wave radar detects a target but the vision algorithm does not recognize the target, it is regarded as an unsatisfactory situation of the camera. At this time, only the target information detected by the millimeter-wave radar is output.
[0037] Outside the filtering range of the millimeter-wave radar or after the number of targets reaches the upper limit, if the vision algorithm recognizes a target but the millimeter-wave radar does not detect it, the target position recognized by the vision algorithm and the 3D framework are output.
[0038] If both the millimeter-wave radar and the camera recognize a target and the Euclidean distance is less than the threshold, the two sensors recognize the same target, and a 3D framework is drawn at the position of the millimeter-wave radar through the vision algorithm.
[0039] Furthermore, in step 5, the algorithm calculates the traffic targets in the image to provide 3D bounding boxes, where each bounding box B consists of seven degrees of freedom:
[0040] B = (x, y, z, l, w, h, θ)
[0041] where (x, y, z) represents the position of the center point of each 3D bounding box, (l, w, h) represent the length, width, and height of the cuboid respectively, and θ represents the global direction of each traffic target in space, that is, the angle between the heading direction of the object and the X-axis of the camera coordinate system.
[0042] For the target B = (x, y, z, l, w, h, θ), the 3D framework drawn during actual detection needs to use the ground plane equation axG + by G + cz G + d = 0, where the normal vector of the ground plane is Z G (a, b, c) determines the direction of the normal vector of the ground plane in the camera coordinate system, where a 2 + b 2 + c 2 = 1; The relationship between the ground coordinate system and the camera coordinate system is:
[0043] X G = X C -(X C · Z G )· Z G
[0044] Y G = Z G × X G
[0045] Normalize the above two equations to obtain the transformation matrix representing the ground coordinate system in the camera coordinate system:
[0046] R C2G = [X G , Y G , Z G T
[0047] The 8 corner points of the 3D bounding box in the ground coordinate system are as follows:
[0048]
[0049] At this time, the frame has not been translated to the target position, nor rotated to the target forward direction, so it is represented by W 3D ; The center point in the ground coordinate system is: G center = R C2G × [x, y, z] T , and the yaw angle in the ground coordinate system is: R C2G × [cosθ, 0, -sinθ] T Take the second row number and the first row number as vectors to calculate the radian G θ ; The calculation formula for the 8 corner points of the 3D bounding box in the ground coordinate system is:
[0050]
[0051] Then transform G 3D through and the camera internal parameter K into pixels:
[0052] P 3D = K × R G2C × G3D
[0053] The obtained P at this time 3D The matrix is a 3×8 matrix. The first two columns are divided by the third column row by row to obtain the actual pixel values, and only the first two columns are retained and transposed to obtain P 2D which are the coordinates of the 8 corner points of the target in the pixel coordinate system.
[0054] P 2D is an 8×2 matrix. Each row is numbered from 0 to 7, representing a total of 8 corner points. Twelve lines are connected according to (0,1)(0,3)(0,4)(1,2)(1,5)(2,3)(2,6)(3,7)(4,5)(4,7)(5,6)(6,7) respectively to obtain a 3D framework.
[0055] The beneficial technical effects of the present invention are as follows:
[0056] 1. Using the radar-vision post-fusion algorithm to determine the existence and position information of the target: As there are more and more vehicles on the road, it means that traffic monitoring has become increasingly important. In the case of a single camera, due to weather, lighting, target occlusion and other situations, the performance of target detection based only on images may be unstable. Therefore, the present invention introduces millimeter-wave radar and camera for post-fusion to make up for the fluctuations in target detection by the camera and improve the accuracy of target detection. The fusion detection of millimeter-wave radar and camera can effectively improve the detection of target existence.
[0057] 2. Using a three-dimensional target detection algorithm to better obtain the boundary, size and direction of traffic targets: In the existing target detection methods, most of them adopt planar detection methods during the process of detecting target information. The planar detection method only returns the information of the target in the image and cannot effectively obtain the actual size boundary and motion direction of the target. Therefore, the present invention calculates the three-dimensional information of the target through image information and obtains the target position information based on radar-vision fusion, which can better and more accurately obtain various information of the target.
[0058] 3. Using the ground plane equation to transform the three-dimensional framework, which helps to obtain a more accurate framework in the image: The main difficulties in the process of drawing the three-dimensional framework are positioning and rotation. The positioning information can be solved through radar-vision fusion, and the rotation angle on the plane can be solved through the predicted target rotation angle. The bottom surface of the three-dimensional framework drawn by the target should be parallel to the ground. At this time, the ground plane equation is needed to make the three-dimensional framework of the target present a better effect in the image. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is a schematic flow chart of the vehicle three-dimensional detection method for the fusion of millimeter-wave radar and vision of the present invention.
[0060] Figure 2 It is a schematic diagram of the spatial conversion for the fusion of millimeter-wave radar and vision.
[0061] Figure 3 It is a schematic diagram for drawing a 3D framework.
[0062] Figure 4 It is the actual prediction effect. Specific implementation manner
[0063] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0064] A three-dimensional vehicle detection method for the fusion of millimeter-wave radar and vision according to the present invention has a process as Figure 1 shown. By using the millimeter-wave radar data and the camera image for spatio-temporal fusion, the pixel position information corresponding to the actual position of the vehicle in the image is obtained. The three-dimensional framework information of the vehicle is calculated using a vision algorithm, and it is located at the target position according to the coordinate conversion algorithm. The target is judged through the radar-vision fusion algorithm, and a three-dimensional annotation is drawn at the location of the target to achieve three-dimensional vehicle detection; specifically, it includes the following steps:
[0065] Step 1: Simultaneously collect the image data of the camera and the point cloud data of the millimeter-wave radar, and record the acquisition time of each piece of data. Use the vision detection algorithm to perform target detection on the image captured by the camera to obtain the position reference of the vehicle target in the image.
[0066] The camera and the millimeter-wave radar collect data and stamp timestamps. The prior two-dimensional detection framework is obtained by using the target detection algorithm, and the prior frameworks (x1, y1) and (x2, y2) of the target are obtained, and then the center point coordinates (x center , y center ) of the target in the image are obtained.
[0067] Step 2: Analyze the original radar point cloud data, calculate the horizontal and vertical distances of the target relative to the radar, use the Cartesian coordinate system to construct the two-dimensional coordinate point position of the target in the radar coordinate system, and at the same time perform filtering processing on the radar data to remove noise interference and retain effective target information.
[0068] Filter out invalid targets and ghost targets according to the position, scattering area, and life cycle of the radar-detected target, etc.; invalid targets include targets in unconcerned areas, false targets, and empty targets.
[0069] Filter in different ways according to experience:
[0070] For targets in unconcerned areas: Set the horizontal distance threshold as Y and the vertical distance threshold as X. Compare the horizontal coordinate y of the target with the threshold Y and compare the vertical distance x of the target with the threshold X. If |x| > X or |y| > Y for the target, then filter out the target.
[0071] For false targets: According to consecutive radar frame rates in chronological order, a threshold λ is set. When |d t+1 - d t | > λ, it is considered a false target and filtered out. d is the position vector d(x, y) of the target in the radar coordinate system, and d t+1 and d t represent the position vectors of the target at frames t + 1 and t respectively.
[0072] For empty targets: Targets with a radar cross - section (RCS) of zero are filtered out.
[0073] For ghost targets: The life cycle of the target is divided into three stages: appearance, duration, and disappearance. When a new target appears, detect = 1 and lost = 0; for each frame that appears, detect += 1 and lost = 0; for each frame that disappears, lost is decremented by 1; the parameter for the appearance state is detect >= D, indicating that the target has appeared for multiple consecutive frames; the parameter for the duration state is detect >= D and lost <= S, indicating that the number of frames the target has appeared is greater than or equal to the threshold D and the number of disappearances is less than or equal to S, and at this time the target is the target to be tracked; lost >= S indicates that the target has disappeared for multiple consecutive frames, and at this time the target is in the disappearance state. Ghost targets are those that appear for only a few frames and then disappear. At this time, by adjusting the thresholds D and S, ghost targets are filtered out and only valid targets are output.
[0074] Step 3: Complete the radar data frames and camera image frames through frame filling methods, obtain the transformation matrix through the least - squares method, and align the radar data and camera data frame - by - frame in space and time.
[0075] Align the collected radar data and camera data respectively. Using the frame filling method, since the time interval between two radar frames is less than 0.1 seconds, it can be considered that all targets between the two radar frames are moving at a constant speed or stationary. Calculate the x and y of each target between the two radar frames respectively, and calculate the position of the target at the moment of the image timestamp through the timestamps between the two frames, thereby achieving frame filling time alignment.
[0076] The radar - vision fusion spatial transformation is as Figure 2 shown. The plane scanned by the millimeter - wave radar is parallel to the ground and the height is the height at which the radar is installed. A coordinate system is constructed with the sensor center as the origin. The world coordinate system is O W (X W , Y W , Z W ), the radar coordinate system is O R (X R , Y R , Z R ), and the camera coordinate system is OC (X C , Y C , Z C ); The distances between the origin of the world coordinate system and the radar coordinate system on each axis are D (D X , D Y , D Z ). Then the representation of the target in the world coordinate system is:
[0077]
[0078] The conversion from the world coordinate system to the camera coordinate system uses the pinhole imaging model. Assume the camera coordinate system O C (X C , Y C , Z C ). The coordinates of the target point in the camera coordinate system are (X0, Y0, Z0), and the projected coordinates in the image are (X img , Y img ); The world coordinate system and the camera coordinate system only differ in the origin position and the axis directions. The world coordinate system can be transformed into the camera coordinate system through rotation and translation; A 3×3 is the rotation matrix, B 3×1 is the translation column vector, O 1×3 is a row vector of three zeros. The conversion formula from the world coordinates to the camera coordinates is:
[0079]
[0080] The image coordinate system O img (X img , Y img ) is a two-dimensional coordinate system. Let the camera focal length be f. Through the principle of similar triangles and pinhole imaging, the conversion formula between the camera coordinates and the image coordinates is:
[0081]
[0082] Written in matrix form as:
[0083]
[0084] The pixel coordinate system O(U, V) is a two-dimensional coordinate system. During the alignment process of the image coordinate system and the pixel coordinate system, the origin of the image coordinate system is the center of the image, while the origin of the pixel coordinate system is the upper left vertex of the image. The horizontal axis is horizontal to the right, and the vertical axis is vertical downward; Assume the projection of the origin of the image coordinate system in the pixel coordinate system is (U0, V0), and the projection of the target (X img , Y img ) in the pixel coordinate system is (U, V). Each pixel on the X img axis and Yimg The values represented in the axial direction are d X and d Y , then the conversion formula between the pixel coordinate system and the image coordinate system is:
[0085]
[0086] Through all the above transformations, the formula for the mutual conversion between the radar coordinate system and the pixel coordinate system is:
[0087]
[0088] The above formula indicates that the radar coordinate system and the pixel coordinate system can be corresponding through a matrix. Since the radar itself does not return pitch information, the installation height of the radar is introduced to make ZR a fixed value. Regarding the inverse matrix of the product of the middle three transformation formulas as T, the radar vector as R, and the pixel vector as P, it can be known that P = TR. The formula for solving the matrix T by the least squares method is: T = (RR T ) -1 RP T .
[0089] Step 4: Map the radar data obtained in Step 3 and the prior target detection framework obtained in Step 1 to each other to obtain the final target detection result.
[0090] Map the target of the millimeter-wave radar signal obtained at the same timestamp in the radar coordinate system and the center position of the target prior data obtained in Step 1 to each other, and calculate the Euclidean distance d. Let be the Euclidean distance between the i-th target in the image and the j-th target in the millimeter-wave radar at time t. Retain the minimum distance between the target in each image and the millimeter-wave radar target, divide the distance, and judge whether the detection results of the two sensors are the same target to complete the fused detection result.
[0091] If the millimeter-wave radar detects a target but the visual algorithm does not recognize the target, it is regarded as an unsatisfactory condition of the camera (such as weather reasons, reflection, etc.). At this time, only the target information detected by the millimeter-wave radar is output.
[0092] Outside the filtering range of the millimeter-wave radar or after the number of targets reaches the upper limit, if the visual algorithm recognizes a target but the millimeter-wave radar does not detect it, the target position recognized by the visual algorithm and the 3D framework are output.
[0093] If both the millimeter-wave radar and the camera recognize a target and the Euclidean distance is less than the threshold, the two sensors recognize the same target, and a 3D framework is drawn at the position of the millimeter-wave radar through the visual algorithm.
[0094] Step 5: Draw a target 3D framework at the target position obtained in Step 4 (such as Figure 3as shown).
[0095] The algorithm calculates 3D bounding boxes for traffic objects in the image, where each bounding box B consists of seven degrees of freedom:
[0096] B = (x, y, z, l, w, h, θ)
[0097] where (x, y, z) represents the position of the center point of each 3D bounding box, (l, w, h) represent the length, width, and height of the cuboid respectively, and θ represents the global orientation of each traffic object in space, that is, the angle between the heading direction of the object and the X-axis of the camera coordinate system.
[0098] For the target B = (x, y, z, l, w, h, θ), the 3D frame drawn during actual detection needs to use the ground plane equation ax G + by G + cz G + d = 0, and the ground plane normal vector is Z G (a, b, c) determines the direction of the ground plane normal vector in the camera coordinate system, where a 2 + b 2 + c 2 = 1; the relationship between the ground coordinate system and the camera coordinate system is:
[0099] X G = X C - (X C · Z G ) · Z G
[0100] Y G = Z G × X G
[0101] Normalize the above two equations to obtain the transformation matrix representing the ground coordinate system in the camera coordinate system:
[0102] R C2G = [X G , Y G , Z G T
[0103] The 8 corner points of the 3D bounding box in the ground coordinate system are as follows:
[0104]
[0105] At this time, the frame has neither been translated to the position of the target nor rotated to the forward direction of the target, so it is represented by W 3D ; the center point in the ground coordinate system is: G center = R C2G ×[x,y,z] T , the yaw angle in the ground coordinate system is: R C2G ×[cosθ, 0, -sinθ] T Take the second row number and the first row number among them as vectors to calculate the radian G θ ; The calculation formula for the 8 corner points of the 3D border in the ground coordinate system is:
[0106]
[0107] Then G 3D is converted into pixels through and the camera internal parameter K:
[0108] P 3D = K × R G2C × G 3D
[0109] The P obtained at this time 3D matrix is a 3×8 matrix. Divide the first two columns by the third column row by row to obtain the actual pixel values, and only keep the transpose of the first two columns to get P 2D which are the coordinates of the 8 corner points of the target in the pixel coordinate system.
[0110] P 2D is an 8×2 matrix. Each row is numbered from 0 to 7, representing a total of 8 corner points. Connect twelve lines according to (0,1)(0,3)(0,4)(1,2)(1,5)(2,3)(2,6)(3,7)(4,5)(4,7)(5,6)(6,7) respectively, and then the 3D framework can be obtained. The actual effect is as Figure 4 shown.
[0111] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention.
[0112] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment contains an independent technical solution. The narrative way of this specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A three-dimensional vehicle detection method integrating millimeter-wave radar and vision, characterized in that, By using spatio-temporal fusion of millimeter-wave radar data and camera images, the pixel position information corresponding to the actual position of the vehicle in the image is obtained. The three-dimensional frame information of the vehicle is calculated using a vision algorithm, and it is located at the target position according to the coordinate transformation algorithm. The target is judged through the radar-vision fusion algorithm, and a three-dimensional annotation is drawn at the location of the target to achieve three-dimensional vehicle detection. Specifically, it includes the following steps: Step 1: Simultaneously collect the image data of the camera and the point cloud data of the millimeter-wave radar, and record the acquisition time of each piece of data. Use the vision detection algorithm to perform target detection on the image captured by the camera to obtain the position reference of the vehicle target in the image. Step 2: Analyze the original radar point cloud data, calculate the horizontal and vertical distances of the target relative to the radar, and use the Cartesian coordinate system to construct the two-dimensional coordinate point position of the target in the radar coordinate system. At the same time, filter the radar data to remove noise interference and retain effective target information. Step 3: Complete the radar data frame and the camera image frame through the frame filling method, obtain the transformation matrix through the least squares method, and align the radar data and the camera data frame by frame in space and time. Step 4: Mutually map the radar data obtained in Step 3 and the prior target detection framework obtained in Step 1 to obtain the final target detection result. Step 5: Draw the target 3D frame at the target position obtained in Step 4.
2. A three-dimensional vehicle detection method by fusing millimeter-wave radar and vision according to claim 1, characterized in that In the said step 1, a prior 2D detection framework is obtained by using an object detection algorithm, and the prior framework (x1, y1) and (x2, y2) of the object are obtained, and then the center point coordinates (x center , y center ) of the object in the image are obtained.
3. A three-dimensional vehicle detection method by fusing millimeter-wave radar and vision according to claim 1, characterized in that, In Step 2, invalid targets and ghost targets are filtered according to the position, scattering area, and life cycle of the radar-detected target. Invalid targets include targets in unconcerned areas, false targets, and empty targets. Filter in different ways according to experience: For targets in unconcerned areas: Set the horizontal distance threshold as Y and the vertical distance threshold as X. Compare the horizontal coordinate y of the target with the threshold Y and the vertical distance x of the target with the threshold X. If |x|>X or |y|>Y for the target, then filter out the target. For false targets: According to the sequential radar frame rates in chronological order, set a threshold λ. When |d t+1 - d t | > λ, it is considered a false target and filtered out. d is the position vector d(x, y) of the target in the radar coordinate system. d t+1 and d t represent the position vectors of the target at the (t + 1)-th frame and the t-th frame respectively; For empty targets: Filter out targets with a radar cross-sectional area (RCS) of zero. For ghost targets: Divide the life cycle of the target into three stages: appearance, duration, and disappearance. When a new target appears, detect = 1 and lost = 0. For each frame that appears, detect += 1 and lost = 0; for each frame that disappears, lost is decremented by 1. The parameter for the appearance state is detect >= D, indicating that the target appears continuously for multiple frames; the parameter for the duration state is detect >= D and lost <= S, indicating that the number of frames the target appears is greater than or equal to the threshold D and the number of disappearances is less than or equal to S. At this time, the target is the target to be tracked; lost >= S indicates that the target disappears continuously for multiple frames, and at this time the target is in the disappearance state. Ghost targets are targets that only appear for a very few frames and then disappear. At this time, adjust the thresholds D and S to filter out ghost targets and only output valid targets.
4. A three-dimensional vehicle detection method by fusing millimeter-wave radar and vision according to claim 1, characterized in that In Step 3, align the collected radar data and camera data respectively. Using the frame filling method, calculate the x and y of each target between two frames of radar respectively, and calculate the position of the target at the moment of the image timestamp through the timestamps between the two frames, so as to achieve frame filling time alignment. The plane scanned by the meter-wave radar is parallel to the ground and the height is the installation height of the radar. A coordinate system is constructed with the center of the sensor as the origin, and the world coordinate system is O W (X W ,Y W ,Z W ), the radar coordinate system is O R (X R ,Y R ,Z R ), the camera coordinate system is O C (X C ,Y C ,Z C ); The distances between the origins of the world coordinate system and the radar coordinate system on each axis are D (D X ,D Y ,D Z ) Then the representation of the target in the world coordinate system is: The conversion from the world coordinate system to the camera coordinate system uses the pinhole imaging model. Assume the origin of the camera coordinate system is O C (X C , Y C , Z C ). The coordinates of the target point in the camera coordinate system are (X0, Y0, Z0), and the projected coordinates in the image are (X img , Y img ); The world coordinate system and the camera coordinate system only differ in the origin position and the axis directions. The world coordinate system can be transformed into the camera coordinate system through rotation and translation; A 3×3 is the rotation matrix, B 3×1 is the translation column vector, and O 1×3 is a row vector of three 0s. The conversion formula from world coordinates to camera coordinates is: Image coordinate system O img (X img , Y img ) is a two-dimensional coordinate system. Assuming the camera focal length is f, the conversion formula between the camera coordinates and the image coordinates can be obtained through the principle of similar triangles and pinhole imaging as follows: Written in matrix form as: The pixel coordinate system O(U, V) is a two-dimensional coordinate system. During the alignment process of the image coordinate system and the pixel coordinate system, the origin of the image coordinate system is the center of the image, while the origin of the pixel coordinate system is the upper left vertex of the image. The horizontal axis is horizontal to the right, and the vertical axis is vertical downward. Assume that the projection of the origin of the image coordinate system in the pixel coordinate system is (U0, V0), and the projection of the target (X img , Y img ) in the pixel coordinate system is (U, V). The values represented by each pixel in the X img -axis and the Y img -axis are d X and d Y , respectively. Then the conversion formula between the pixel coordinate system and the image coordinate system is: Through all the above transformations, the formula for the mutual conversion between the radar coordinate system and the pixel coordinate system is as follows: The above formula indicates that the radar coordinate system and the pixel coordinate system can be made to correspond through a matrix. Since the radar itself does not return pitch information, the installation height of the radar is introduced to make ZR a fixed value. Regarding the inverse matrix of the product of the middle three transformation formulas as T, the radar vector as R, and the pixel vector as P, it can be known that P = TR. The formula for solving matrix T by the least squares method is: T = (RR T ) - 1 RP T .
5. A three-dimensional vehicle detection method integrating millimeter-wave radar and vision according to claim 1, characterized in that, In the said step 4, the target of the millimeter-wave radar signal obtained at the same timestamp in the radar coordinate system is mutually mapped with the prior data center position of the target obtained in step 1, and the Euclidean distance d is calculated mutually. Let be the Euclidean distance between the i-th target in the image at time t and the j-th target in the millimeter-wave radar. Retain the minimum distance between the target in each image and the millimeter-wave radar target, divide the distances, and determine whether the detection results of the two sensors are the same target to complete the fused detection result.
6. A three-dimensional vehicle detection method based on the fusion of millimeter-wave radar and vision according to claim 5, characterized in that, In step 4, if the millimeter-wave radar detects a target but the vision algorithm does not recognize it, it is regarded as an unsatisfactory condition of the camera. At this time, only the target information detected by the millimeter-wave radar is output. Outside the filtering range of the millimeter-wave radar or after the target number reaches the upper limit, if the vision algorithm recognizes a target but the millimeter-wave radar does not detect it, the target position recognized by the vision algorithm and the 3D frame are output. If both the millimeter-wave radar and the camera recognize a target and the Euclidean distance is less than the threshold, the two sensors recognize the same target, and a 3D frame is drawn at the position of the millimeter-wave radar through the vision algorithm.
7. A three-dimensional vehicle detection method by fusing millimeter-wave radar and vision according to claim 4, characterized in that In step 5, the algorithm calculates the traffic target in the image to provide a 3D bounding box, where each bounding box B consists of seven degrees of freedom: B = (x, y, z, l, w, h, θ) where (x, y, z) represents the position of the center point of each 3D bounding box, (l, w, h) represent the length, width, and height of the cuboid respectively, and θ represents the global direction of each traffic target in space, that is, the angle between the heading direction of the object and the X-axis of the camera coordinate system. For the target B=(x,y,z,l,w,h,θ), the 3D frame drawn during actual detection needs to use the ground plane equation ax G +by G +cz G +d = 0, and the ground plane normal vector is Z G (a,b,c) determines the direction of the ground plane normal vector in the camera coordinate system, where a 2 +b 2 +c 2 = 1; the relationship between the ground coordinate system and the camera coordinate system is: X G = X C - (X C · Z G ) · Z G Y G = Z G × X G Normalize the above two equations to obtain the transformation matrix representing the ground coordinate system in the camera coordinate system: R C2G = [X G , Y G , Z G T The 8 corner points of the 3D bounding box in the ground coordinate system are as follows: At this time, the frame has neither been translated to the position where the target is located nor rotated to the forward direction of the target. Therefore, it is represented by W 3D The center point in the ground coordinate system is: G center = R C2G × [x, y, z] T , and the yaw angle in the ground coordinate system is: R C2G × [cosθ, 0, -sinθ] T Take the second row number and the first row number among them as vectors to calculate the radian G θ The calculation formula for the 8 corner points of the 3D border in the ground coordinate system is: Then convert G 3D through and the camera intrinsic parameter K into pixels: P 3D = K × R G2C × G 3D The obtained P at this time 3D The matrix is a 3×8 matrix. The first two columns are divided by the third column row by row to obtain the actual pixel values, and only the first two columns are transposed to obtain P 2D which are the coordinates of the 8 corner points of the target in the pixel coordinate system; P 2D It is an 8×2 matrix, with each row numbered from 0 to 7, representing a total of 8 corner points. By connecting twelve lines according to (0,1)(0,3)(0,4)(1,2)(1,5)(2,3)(2,6)(3,7)(4,5)(4,7)(5,6)(6,7) respectively, a 3D framework can be obtained.
Citation Information
Cited By
Visual target position detection method and device
CN121564329A
Visual target position detection method and apparatus
CN121564329B