A radar-vision automatic calibration method based on target tracking

Through the deep learning-based object detection and target tracking technology, the perspective transformation matrix is ​​optimized, and the accuracy and real-time problems of the existing Leising Vision automatic calibration method are solved, and efficient automatic calibration and precise fusion of Leising Vision sensors are realized.

CN115685102BActive Publication Date: 2025-05-27LIANYUNGANG JARI ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211150390.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-21
Publication Date
2025-05-27
Estimated Expiration
2042-09-21

AI Technical Summary

Technical Problem

The existing Leishi automatic calibration method has problems such as inaccurate manual point selection, inability to realize real-time automatic calibration, and inadequate scene switching, making it difficult to achieve high-precision fusion of radar and video sensors.

Method used

The target detection method based on deep learning is adopted, combining the tracking information of video and radar targets, and iteratively optimizes the perspective transformation matrix to achieve accurate association and fusion of the lightning vision target.

Benefits of technology

It realizes efficient automatic calibration of the lightning vision sensor, avoids errors in manual point selection, supports real-time detection and multi-scene adaptation, and improves the accuracy and reliability of lightning vision fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115685102B_ABST
    Figure CN115685102B_ABST
Patent Text Reader

Abstract

The present invention discloses a radar-vision automatic calibration method based on target tracking, including: collecting the state information of radar target detection points; collecting video images, performing target detection, and obtaining the state information of detected targets; performing video target tracking, and screening out dynamic tracking targets through the tracking state; performing radar target tracking, and screening out radar dynamic tracking targets through a speed threshold; and judging whether to end the calibration process according to the number of current-frame radar-vision dynamic tracking targets, the mean mapping rate, and the number of target point pairs. The method of the present invention removes the interference of static targets and false targets based on target tracking, realizes automatic coordinate point association by selecting points according to the relative position, and determines the accuracy of the selected points according to the continuous substitution of the perspective transformation matrix, improving the disadvantages of long time consumption, low efficiency, and small number of calibrated point pairs in manual calibration point selection, and also avoiding the errors caused by overly complex image processing, realizing fast and efficient automatic calibration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent transportation detection, especially the field of multi-view area correction, and particularly relates to a radar-vision automatic calibration method based on target tracking. Background Art

[0002] In recent years, with the development of "smart cities", traditional traffic monitoring means, such as inductive loops, geomagnetism, sectional microwave, etc., can no longer meet the needs of traffic managers. Existing traffic requires real-time large-area detection, timely alarm, all-weather operation, and high-precision detection. Video traffic detection sensors are area detection, but are greatly affected by weather and cannot achieve all-weather detection. Millimeter-wave radar detectors have high detection accuracy, but are inconvenient for detecting static targets. Radar-vision integrated detectors can achieve complementary advantages and achieve all-weather and all-round accurate detection. The premise of the complementary advantages of radar and video sensors is precise fusion, so the two coordinate systems need to be closely related.

[0003] The existing association methods are basically: manually select target points with the naked eye of an operator, and then calculate the perspective transformation matrix. This method has the following disadvantages: When selecting points for a small number of targets, it is relatively accurate. When multiple targets are involved, due to different detection perspectives and inaccurate detection of static targets by radar, the number of target points of the two cannot be matched, and point-to-point correspondence cannot be accurately performed; when manually clicking to select points, it is easy to have an offset of the point pair, resulting in an error in the solution of the coordinate plane; the solution of the perspective transformation matrix is obtained by fitting multiple groups of point pairs, and the point pairs are required not to be collinear. It is impossible for an operator to achieve full coverage, resulting in an unsatisfactory calibration effect, causing repeated calibration, wasting time and inaccurate results. In addition, there are also methods that use neural networks for model training and then achieve automatic calibration. This method requires re-training the model after the scene is switched, and the model training requires a large amount of data support, which does not meet the real-time requirement and cannot achieve precise association.

[0004] Therefore, it is very necessary to propose an efficient automatic calibration method. The core of this process lies in how to deal with situations such as false point detection when selecting target points, target detection specific to different sensors, target occlusion, target splitting, etc. Summary of the Invention

[0005] The object of the present invention is to provide a radar-vision automatic calibration method based on target tracking for the above-mentioned radar-vision calibration defects. A detection method based on deep learning is used for video target detection. The radar targets and video targets are respectively tracked, and the targets are screened according to the status, and then the transformation matrix is solved, and the matrix is continuously substituted back to realize matrix iterative optimization. Starting from the target status, the invention is not limited to the target data of the current frame, uses the tracking information to avoid the interference of false points, uses the speed threshold to eliminate the influence of static targets, and uses the iterative mapping rate to prevent wrong point selection. In short, this method provides a reliable idea for the precise fusion of radar-vision targets.

[0006] The technical solution for realizing the object of the present invention is: a radar-vision automatic calibration method based on target tracking, the method comprising the following steps:

[0007] Step 1, collecting the state information of radar target detection points, including position, speed and status;

[0008] Step 2, collecting video images, loading a target detection model, inputting the video images into the detection model for target detection, and obtaining the state information of the detected targets, including position, category and size;

[0009] Step 3, performing video target tracking, and screening out dynamically tracked targets through the tracking status;

[0010] Step 4, performing radar target tracking, and screening radar dynamically tracked targets through a speed threshold;

[0011] Step 5, judging whether the number of radar-vision dynamically tracked targets in the current frame is the same. If not, return to Step 1. Otherwise, associate the targets at the corresponding positions, calculate the perspective transformation matrix, and perform data back substitution on the matrix. If the mean mapping rate is less than or equal to the preset threshold TH1, then eliminate the target point pairs added in the current frame. Otherwise, judge whether the number of target point pairs is greater than the preset threshold N. If so, end the calibration process. Otherwise, return to Step 1.

[0012] Further, the state information of the radar target radar detection points in Step 1 is the structured data obtained after point cloud clustering analysis, including position, speed and status, indicating the actual relative position of the target point to the radar sensor.

[0013] Further, the detection framework of the target in Step 2 is darknet, the target detection model is the yolov3 model trained using the coco training set, and the TensorRT acceleration inference framework is used to accelerate the inference during the inference process.

[0014] Further, the video target tracking in Step 3 specifically includes:

[0015] Step 3-1, establish a video object tracker;

[0016] Step 3-2, use the detected object state information obtained in Step 2 to perform forward and backward association of objects at different times. The association method is as follows:

[0017] Step 3-2-1, calculate the intersection over union (IoU) between the object in the video object tracker and the object detection bounding box in the current frame:

[0018]

[0019] Among them, A i is the detection bounding box of the i-th object in the video object tracker, B j is the detection bounding box of the j-th object in the current frame. The value range of i is from 1 to m, where m is the total number of objects in the video object tracker, and the value range of j is from 1 to n, where n is the total number of detected objects in the current frame;

[0020] Step 3-2-2, calculate the best object matcher through the Hungarian matching method and the IoU calculated in Step 3-2-1. Specifically: use 1-I OU(i,j) as the input parameter in the Hungarian matching process to obtain the best object matcher;

[0021] Step 3-2-3, associate the video objects in the front and back frames, that is, associate the objects in the video object tracker with the best object matcher.

[0022] Furthermore, in Step 3, dynamic tracking objects are filtered out through the tracking state, specifically including: setting the IoU threshold th0. When the IoU is greater than th0, it means that the detected object is in a nearly stationary state and cannot be detected by the radar object, and such objects are excluded; when the IoU is between 0 and th0, it means that the detected object is dynamic, and it is added to the video object tracker to update the video tracking object information, including the motion state, the number of matching frames, the position, the size, and the type.

[0023] Furthermore, the specific steps for radar object tracking in Step 4 include:

[0024] Step 4-1, establish a radar object tracker;

[0025] Step 4-2, use the target detection point state information obtained in Step 1 to perform forward and backward association of objects at different times. The method is as follows:

[0026] Step 4-2-1, calculate the Euclidean distance between the object in the radar object tracker and the object in the current frame:

[0027]

[0028] Among them, (A i 'x, A i(x, y) is the coordinate of the i-th target in the radar target tracker, (B' j x, B' j (x, y) is the coordinate of the j-th target in the current frame;

[0029] Step 4-2-2: Calculate the best target matcher through the Hungarian matching method and the Euclidean distance calculated in Step 4-2-1. Specifically: Take Dis i,j as the input parameter in the Hungarian matching process to obtain the best target matcher;

[0030] Step 4-2-3: Associate the radar targets in the front and back frames, that is, associate the targets in the radar target tracker with the best target matcher.

[0031] Furthermore, in Step 4, the radar dynamic tracking targets are screened through the speed threshold. Specifically: Set the speed threshold and the matching frame number threshold. When the target speed is greater than the speed threshold, it indicates that the detected target is a dynamic target, otherwise it is a static target, and the static targets are removed; then according to the matching frame number, the false point targets are removed. Specifically: When the matching frame number is less than the matching frame number threshold, it indicates a false point target.

[0032] Furthermore, in Step 5, the association of targets at corresponding positions is specifically as follows: When the number of targets screened in the current frame of the radar-vision two coordinate systems is equal, according to the relative position, the targets are first sorted longitudinally and then horizontally, so as to solve the relative position sorting. Then the point pairs at the corresponding positions are associated and added to the coordinate point pair queue, and the perspective transformation matrix is calculated based on all the point pair data.

[0033] Furthermore, the calculation of the mean mapping rate in Step 5 specifically includes:

[0034] Map the radar points that meet the requirements in the current frame to the video coordinate system based on the perspective transformation matrix;

[0035] Calculate the mean mapping rate:

[0036]

[0037] In the formula, (C i x - D i x) 2 +(C i y - D i y) 2 represents the ratio of the distance from the radar mapping point to the center of the corresponding detection box to the distance from the upper left corner of the detection box to the center of the lower edge of the detection box.

[0038] A radar-vision automatic calibration system based on target tracking, the system includes the following steps executed sequentially:

[0039] The first module is used to collect the status information of radar target detection points, including position, speed, and status;

[0040] The second module is used to collect video images, load the target detection model, input the video images into the detection model for target detection, and obtain the status information of the detected targets, including position, category, and size;

[0041] The third module is used to perform video target tracking and filter out dynamic tracking targets through the tracking status;

[0042] The fourth module is used to perform radar target tracking and filter out radar dynamic tracking targets through the speed threshold;

[0043] The fifth module is used to determine whether the number of current-frame radar-vision dynamic tracking targets is consistent. If not, it returns to execute the first module. Otherwise, it associates the corresponding position targets, calculates the perspective transformation matrix, and performs data feedback on this matrix. If the mean mapping rate is less than or equal to the preset threshold TH1, it eliminates the target point pairs added in the current frame. Otherwise, it determines whether the number of target point pairs is greater than the preset threshold N. If so, it ends the calibration process. Otherwise, it returns to execute the first module.

[0044] Compared with the prior art, the significant advantages of the present invention are as follows:

[0045] (1) By setting the constraints of iterative solution, accurately calculating the target point coordinates, and strictly sorting the position of point pairs, the present invention avoids problems such as incomplete selection coverage by manual visual calibration, coordinate drift of target points, and point pair mismatches caused by human subjective factors.

[0046] (2) Without complex preprocessing such as screening of detection areas and training of convolutional models, the present invention updates the point pair sequence based on the tracking results of the targets, thereby calculating the transformation matrix, saving the calibration time.

[0047] (3) The present invention sets multiple threshold constraints, including speed constraints to avoid the influence of static targets; matching frame constraints to eliminate the influence of false detection point pairs; and constraints on the mean mapping rate to prevent incorrect point selection and cause the collapse of the overall matrix iterative calculation. Compared with the prior art, this method has strong plasticity. Based on the tracking of targets detected by two sensors, it realizes the automatic calibration of multiple planes, providing a new method for target association in radar-vision fusion.

[0048] The present invention will be further described in detail below with reference to the accompanying drawings. Description of the Drawings

[0049] Figure 1 It is a schematic flowchart of a radar-vision automatic calibration method based on target tracking in an embodiment.

[0050] Figure 2 Schematic diagram of the radar target tracking effect of the radar-vision automatic calibration method based on target tracking in one embodiment.

[0051] Figure 3 Schematic diagram of the video target detection of the radar-vision automatic calibration method based on target tracking in one embodiment.

[0052] Figure 4 Schematic diagram of the video target tracking process of the radar-vision automatic calibration method based on target tracking in one embodiment.

[0053] Figure 5 Schematic diagram of the radar-vision association effect of the radar-vision automatic calibration method based on target tracking in one embodiment. Detailed implementation manners

[0054] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0055] It should be noted that if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of these features. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those skilled in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.

[0056] In one embodiment, in combination with Figure 1 , a radar-vision automatic calibration method based on target tracking is provided, including the following steps:

[0057] Step 1, collect the state information of radar target detection points, including position, speed, and status;

[0058] Step 2, collect video images, load a target detection model, input the video images into the detection model for target detection, and obtain the state information of the detected targets, including position, category, and size;

[0059] Step 3, perform video target tracking, and screen out dynamic tracking targets through the tracking status;

[0060] Step 4, perform radar target tracking, and screen out radar dynamic tracking targets through a speed threshold;

[0061] Step 5, when the number of radar-vision targets in the current frame is inconsistent, return to Step 1; otherwise, associate the corresponding position targets, calculate the perspective transformation matrix, and perform data feedback on the matrix. When the mean mapping rate does not meet the threshold, remove the target point pairs added in the current frame. When the number of target point pairs reaches the threshold, end the calibration process;

[0062] In the embodiments of the present invention, the radar target detection point status information is the structured data obtained after the radar point cloud clustering analysis; the convolutional model used for target detection and recognition in the video frame is Darknet, combined with the Yolov3 method. The recognized targets include: car, bus, person, bike, truck, and TensorRT is used to accelerate the inference in target detection; the intersection over union of targets in every two frames is used to realize video target tracking and remove static targets, and false detection boxes are filtered out by the number of matching frames; for radar targets, the Euclidean distance between targets in adjacent frames is used to realize related tracking, static targets are removed by the speed threshold, and false points are filtered out by the number of matching frames; under the condition that the number of targets in the current frame of radar and video is the same, perform relevant position sorting and point pair combination, add the point pair combination of the current frame to the point pair sequence, and calculate the transformation matrix between the two planes; use the transformation matrix to back-substitute the radar target into the video coordinate system, calculate the relative position difference between the mapped target and the detected target to obtain the mean mapping rate, and thus judge whether the point pair sequence selected in the current frame is correct; set the threshold of the number of point pair sequences as the flag to end the iteration.

[0063] In view of the requirements of the all-round, real-time, accurate, and efficient integrated transportation management system in intelligent transportation, to improve the safety level of transportation and relieve traffic congestion, especially for the accurate calibration in the process of radar-vision integration equipment, the present invention proposes a radar-vision automatic calibration method based on target tracking. This method aims at two different coordinate systems of video and radar, and realizes the selection of relative target point pairs by respectively tracking the dynamic targets in different coordinate systems, so as to calculate the transformation matrix and accurately associate the two coordinate systems.

[0064] Radar Target Tracking

[0065] To perform radar target tracking, first establish a target tracker. The target information stored in the tracker includes: position, speed, motion state, number of matching frames, and selection status. Calculate the Euclidean distance between the targets with the selection status being true in the tracker and the targets in the current frame whose speed meets the threshold. The calculation formula is as follows:

[0066]

[0067] A is the target meeting the requirements in the tracker, B is the target meeting the speed requirement in the current frame, and Dis i,j is the Euclidean distance between the two.

[0068] The Hungarian matching algorithm and the calculated Euclidean distance are used to perform the association matching between the targets in the tracker and the targets in the current frame, and update various state information successfully matched in the tracker. Specifically: According to the position, state, and speed information of the matched targets in the current frame, the corresponding parameter information in the tracker is changed, and the selection state is changed by the number of matched frames. The speed threshold in this embodiment is 1 m / s, the distance threshold in the Hungarian matching process is 3, and the number-of-matched-frames threshold is 5. Figure 2 The tracking effect of the radar is shown.

[0069] Video Object Tracking

[0070] Video object tracking first requires object detection. The process of object detection includes: dataset collection, data annotation, model training, and model verification. The dataset in this embodiment is the COCO dataset, the model is Darknet, and the YoloV3 method is used for training. The sample ratio between the training set and the validation set is 7:3, the batchsize is set to 64, the learning_rate is set to 0.001, and more training samples are generated by setting the rotation angle, adjusting the saturation, exposure, and hue. The learning rate adjustment strategy selects steps, and the learning rate is decayed when the training reaches a certain number of times. The trained model file is passed into the TensorRT deep learning framework for parsing, generating an ONNX general model and then converting it into a TRT model for accelerated inference and deployment. Figure 3 is the target recognition result.

[0071] A video object tracker is established, and the stored target information includes: position, size, category, probability, number of matched frames, selection state, and motion state. The association between the target in the tracker and the target in the current frame is calculated through the intersection over union (IoU) calculation between them. The formula for calculating the IoU is as follows:

[0072]

[0073] Assume that the total number of targets in the tracker is m, and the number of detected targets in the current frame is n. Calculate the IoU I of the detection box of each target in the tracker and the target in the current frame. OU , where i ranges from 1 to m and j ranges from 1 to n;

[0074] The Hungarian matching method and the IoU calculation between the targets in the previous and current frames are used to find the best target match. Specifically, 1 - I ou is used as the input parameter in the Hungarian matching process. After associating the targets, the corresponding parameters are modified. The motion state is judged as stationary or not according to the IoU threshold. The IoU threshold in this embodiment is set to 0.8, and a value greater than the threshold represents a stationary state. In this embodiment, the targets exceeding the threshold are removed, and the selection state is modified by the number-of-matched-frames threshold and the motion state. Figure 4Display the video target tracking effect.

[0075] Lidar-camera target association

[0076] First, for target association, point pairs need to be selected. When the number of video targets in the selection state is the same as the number of lidar targets, point pair matching is performed. The matching method is based on the relative position. The targets are sorted by position. The sorting method is to sort vertically first, and then sort horizontally on the basis of vertical sorting, that is, from the upper left to the lower right, and then select point pairs with corresponding serial numbers.

[0077] Among them, the selected point of the video target is the center point of the lower edge of the detection box, as shown in the following formula:

[0078]

[0079] In the formula, x img and y img are the coordinates of the video detection point, x o and y o are the upper left corner points of the video detection box, w and h are the width and height of the detection box respectively, and the ground is used as the reference for point selection.

[0080] Add the selected number of point pairs to the point pair sequence, and use all the point pair data as input to calculate the perspective transformation matrix. The calculation formula is as follows:

[0081]

[0082] In the formula, M is the perspective transformation matrix, (x r , y r ) are the coordinates of the lidar point in the lidar coordinate system, (x img , y img ) are the coordinates of the video coordinate system, and multiple target point pairs are used as input for inverse solution.

[0083] After obtaining the perspective transformation matrix, map the lidar points of the current frame to the video coordinate system and compare their positions with the corresponding targets. The comparison method is the mapping rate, and the overall mapping effect is evaluated by the mean mapping rate. The formula for solving the mean mapping rate is as follows:

[0084]

[0085] Calculate the ratio of the distance from each lidar mapped point to the center of the corresponding detection box to the distance from the upper left corner of the detection box to the center of the lower edge of the detection box, and perform mean processing. The mean mapping rate threshold is set to 1. When the condition is met, continue to select the next point pair data. Otherwise, eliminate the current point pair data and select the next frame of point pair data.

[0086] In this embodiment, the threshold for the stop point to select the target is 5000. When the condition is met, the matrix calculation is stopped to obtain the final correlation matrix. Figure 5 It shows the radar-vision correlation effect achieved by the method of this embodiment.

[0087] In one embodiment, a radar-vision automatic calibration system based on target tracking is provided. The system includes the following steps executed sequentially:

[0088] The first module is used to collect the state information of radar target detection points, including position, speed, and status.

[0089] The second module is used to collect video images, load a target detection model, input the video images into the detection model for target detection, and obtain the state information of the detected targets, including position, category, and size.

[0090] The third module is used to perform video target tracking and screen out dynamic tracking targets through the tracking status.

[0091] The fourth module is used to perform radar target tracking and screen out radar dynamic tracking targets through a speed threshold.

[0092] The fifth module is used to determine whether the number of radar-vision dynamic tracking targets in the current frame is the same. If not, it returns to execute the first module. Otherwise, it associates the targets at the corresponding positions, calculates the perspective transformation matrix, and performs data feedback on the matrix. If the mean mapping rate is less than or equal to the preset threshold TH1, the target point pairs added in the current frame are excluded. Otherwise, it determines whether the number of target point pairs is greater than the preset threshold N. If so, the calibration process ends. Otherwise, it returns to execute the first module.

[0093] For the specific limitations of the radar-vision automatic calibration system based on target tracking, reference can be made to the limitations of the radar-vision automatic calibration method based on target tracking in the above text, which will not be elaborated here. Each module in the above radar-vision automatic calibration system based on target tracking can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0094] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.

Claims

1. A radar-vision automatic calibration method based on target tracking, characterized in that, the method comprises the following steps: Step 1, collect the state information of radar target detection points, including position, speed and status; Step 2, collect video images, load a target detection model, input the video images into the detection model for target detection, and obtain the state information of the detected targets, including position, category and size; Step 3, perform video target tracking, and filter out dynamic tracking targets through the tracking status; Step 4, perform radar target tracking, and filter radar dynamic tracking targets through a speed threshold; Step 5, determine whether the number of radar-vision dynamic tracking targets in the current frame is the same. If not, return to Step 1. Otherwise, associate the targets at the corresponding positions, calculate the perspective transformation matrix, and perform data feedback on this matrix. If the mean mapping rate is less than or equal to the preset threshold TH1, then eliminate the target point pairs added in the current frame. Otherwise, determine whether the number of target point pairs is greater than the preset threshold N. If so, end the calibration process. Otherwise, return to Step 1; The specific process of associating the targets at the corresponding positions in Step 5 is as follows: When the number of filtered targets in the current frame of the two coordinate systems of radar and vision is equal, according to the relative position, first sort the targets longitudinally, and then sort them horizontally, so as to solve the relative position sorting. Then, associate the point pairs at the corresponding positions and add them to the coordinate point pair queue, and calculate the perspective transformation matrix based on all the point pair data; The specific calculation of the mean mapping rate in Step 5 includes: Based on the perspective transformation matrix, map the radar points meeting the requirements in the current frame to the video coordinate system; Calculate the mean mapping rate: Where, (C i x - D i x) 2 +(C i y - D i y) 2 represents the ratio of the distance from the radar mapping point to the center of the corresponding detection box to the distance from the upper left corner of the detection box to the center of the lower edge of the detection box; w and h are the width and height of the detection box respectively.

2. The radar-vision automatic calibration method based on target tracking according to claim 1, characterized in that, the state information of the radar target detection points in Step 1 is the structured data obtained after point cloud clustering analysis, including position, speed and status, indicating the actual relative position of the target points to the radar sensor.

3. The radar-vision automatic calibration method based on target tracking according to claim 1, characterized in that, the detection framework of the target in Step 2 is darknet, the target detection model is the yolov3 model trained using the coco training set, and the TensorRT acceleration inference framework is used to accelerate the inference during the inference process.

4. The radar-vision automatic calibration method based on target tracking according to claim 1, characterized in that, the specific video target tracking in Step 3 includes: Step 3-1, establish a video target tracker; Step 3-2, use the state information of the detected targets obtained in Step 2 to perform forward and backward association of targets at different times. The association method is: Step 3-2-1, calculate the intersection over union of the target in the video target tracker and the target detection box in the current frame: Among them, A i is the detection box of the i-th target in the video object tracker, and B j is the detection box of the j-th target in the current frame. The value range of i is from 1 to m, where m is the total number of targets in the video object tracker, and the value range of j is from 1 to n, where n is the total number of detected targets in the current frame; Step 3-2-2: Calculate the best target matcher through the intersection over union calculated by the Hungarian matching method and Step 3-2-1. Specifically, use 1-I OU(i,j) as the input parameter in the Hungarian matching process to obtain the best target matcher; Step 3-2-3, associate the video targets in the front and back frames, that is, associate the target in the video target tracker with the best target matcher.

5. The radar-vision automatic calibration method based on target tracking according to claim 4, characterized in that, In step 3, dynamic tracking targets are filtered by tracking status, specifically including: setting the intersection over union threshold th0. When the intersection over union is greater than th0, it indicates that the detected target is in a nearly stationary state and cannot be detected by the radar target, and such targets are excluded; when the intersection over union is between 0 and th0, it indicates that the detected target is dynamic, and it is added to the video target tracker to update the video tracking target information, including motion state, matching frame number, position, size, and type.

6. The radar-vision automatic calibration method based on target tracking according to claim 1, wherein, the specific steps of radar target tracking in step 4 include: Step 4-1, establishing a radar target tracker; Step 4-2, using the target detection point status information obtained in step 1 to perform forward and backward association of targets at different times, and the method is as follows: Step 4-2-1, calculating the Euclidean distance between the target in the radar target tracker and the target in the current frame: Among them, (A i 'x, A i 'y) is the coordinate of the i-th target in the radar target tracker, and (B' j x, B' j y) is the coordinate of the j-th target in the current frame; Step 4-2-2: Calculate the best target matcher by using the Hungarian matching method and the Euclidean distance calculated in Step 4-2-1. Specifically, use Dis i,j as the input parameter in the Hungarian matching process to obtain the best target matcher; Step 4-2-3, associating the radar targets in the front and back frames, that is, associating the targets in the radar target tracker with the best target matchers.

7. The radar-vision automatic calibration method based on target tracking according to claim 1, wherein, in step 4, the radar dynamic tracking targets are filtered by the speed threshold, specifically: setting the speed threshold and the matching frame number threshold. When the target speed is greater than the speed threshold, it indicates that the detected target is a dynamic target, otherwise it is a static target, and the static targets are excluded; then, according to the matching frame number, the false point targets are excluded. Specifically: when the matching frame number is less than the matching frame number threshold, it indicates a false point target.

8. A radar-vision automatic calibration system based on target tracking according to any one of claims 1 to 7, wherein, the system includes the following executed sequentially: The first module is used to collect the status information of radar target detection points, including position, speed, and status; The second module is used to collect video images, load the target detection model, input the video images into the detection model for target detection, and obtain the detection target status information, including position, category, and size; The third module is used to perform video target tracking and filter out dynamic tracking targets by tracking status; The fourth module is used to perform radar target tracking and filter out radar dynamic tracking targets by the speed threshold; The fifth module is used to judge whether the number of radar-vision dynamic tracking targets in the current frame is consistent. If not, return to execute the first module. Otherwise, associate the corresponding position targets, calculate the perspective transformation matrix, and perform data feedback on this matrix. If the mean mapping rate is less than or equal to the preset threshold TH1, then exclude the target point pairs added in the current frame. Otherwise, judge whether the number of target point pairs is greater than the preset threshold N. If so, end the calibration process. Otherwise, return to execute the first module.

Citation Information

Patent Citations

  • Method and device for locating spherical camera

    CN104125390A

  • Vehicle detection and tracking method based on radar signal and visual fusion

    CN112991391A