Multi-target tracking method, device, equipment and storage medium

By matching and fusing the confidence of image information and radar information in multi-target tracking technology, the problem of limited multi-target detection performance caused by a single sensor is solved, and higher target tracking performance and detection accuracy are achieved.

CN114154528BActive Publication Date: 2025-05-16CHINA MOBILE SHANGHAI ICT CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202010922153.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-04
Publication Date
2025-05-16
Estimated Expiration
2040-09-04

AI Technical Summary

Technical Problem

In the existing multi-objective tracking technology, because the target detection stage relies on a single sensor, the multi-objective detection performance is limited, and the target information cannot be effectively matched and tracked due to sensor failure or missed detection.

Method used

By acquiring image information and radar information, using preset multi-objective detection algorithms and association algorithms, the confidence of the target detection information obtained by different sensors is matched and fused, deepening the sensor fusion to the target detection stage, so that the target detection and recognition performance can still be guaranteed when any sensor has problems.

Benefits of technology

It effectively reduces the missed detection problems caused by the failure of a single sensor, improves the performance of multi-objective detection, and thus ensures the target tracking performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114154528B_ABST
    Figure CN114154528B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, equipment and storage medium for multi-target tracking. Specifically, it includes: obtaining image information and radar information in the road section to be tested; determining first multi-target detection information using a preset multi-target detection algorithm based on the image information, radar information and a preset confidence threshold; correlating the first multi-target detection information and radar information using a preset first association algorithm to obtain first multi-target association information; correlating the first multi-target association information, first target tracking information and second target tracking information using a preset second association algorithm to obtain target traffic information in the road section to be tested. According to the embodiments of the present application, problems such as missed detection caused by failure of a single sensor can be effectively reduced, the performance of multi-target detection can be improved, and the target tracking performance can be guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of computer technology, and in particular, relates to a method, device, equipment and computer storage medium for multi-target tracking. Background Art

[0002] Multi-target tracking technology is an indispensable perception link in vehicle-road cooperative technology. Sensors can obtain target measurement data to accurately estimate the target state. As multi-sensor fusion gradually becomes the mainstream of traffic perception, the sensors mainly used for multi-target tracking are cameras and millimeter-wave radars. Cameras can obtain traffic image information of relevant road sections, and radars can obtain accurate relative position information of obstacles.

[0003] Multi-target tracking technology is mostly based on target detection. However, in related technologies, since the target detection stage is still completed by a single sensor, that is, the information obtained by the sensor is separately detected for the target information, the target information of different sensor types is determined, and then the fusion operation is performed, the multi-target detection performance is still limited. If the target detection operation of any sensor fails or misses, it may not be possible to effectively match and track the target information. Summary of the invention

[0004] The embodiments of the present application provide a method, apparatus, device and computer storage medium for multi-target tracking, which can effectively reduce problems such as missed detection caused by failure of a single sensor, improve the performance of multi-target detection, and thus ensure target tracking performance.

[0005] In a first aspect, an embodiment of the present application provides a method for multi-target tracking, comprising:

[0006] Obtain image information and radar information within the road section to be tested;

[0007] Determining first multi-target detection information according to the image information, the radar information and a preset confidence threshold using a preset multi-target detection algorithm;

[0008] Using a preset first association algorithm, associating the first multi-target detection information with the radar information to obtain first multi-target association information;

[0009] Using a preset second association algorithm, the first multi-target association information, the first target tracking information, and the second target tracking information are associated to obtain target traffic information within the road section to be tested;

[0010] The first target tracking information is visual multi-target tracking information determined according to the image information and the first multi-target association information, and the second target tracking information is radar multi-target tracking information determined according to the radar information.

[0011] Optionally, determining the first multi-target detection information according to the image information, the radar information and a preset confidence threshold using a preset multi-target detection algorithm includes:

[0012] Determine, according to the image information, the second multi-target detection information and the confidence level corresponding to the second multi-target detection information by using a preset multi-target detection algorithm;

[0013] Determining a first confidence set and a second confidence set according to the confidence corresponding to the second multi-target detection information and a preset first confidence threshold;

[0014] Modifying the second confidence set according to the radar information to obtain a third confidence set;

[0015] The first confidence set and the third confidence set are data-fused to obtain corresponding first multi-target detection information.

[0016] Optionally, the modifying the second confidence set in combination with the radar information to obtain a third confidence set includes:

[0017] Calculate and obtain a confidence correction result corresponding to the second confidence set according to a preset correction rule and the radar information;

[0018] The confidence correction result that reaches a preset second confidence threshold is determined to be the third confidence set.

[0019] Optionally, the using a preset first association algorithm to associate the first multi-target detection information with the radar information to obtain first multi-target association information includes:

[0020] Calculate a distance vector between an observed target in the first multi-target detection information and an observed target in the radar information according to a preset scaling factor and the first multi-target detection information and the radar information to obtain a first distance vector;

[0021] When the first distance vector satisfies a preset condition, determining an association result between the first multi-target detection information and the radar information;

[0022] The association result is used as the first multi-target association information.

[0023] Optionally, the using of a preset second association algorithm to associate the first multi-target association information, the first target tracking information, and the second target tracking information to obtain the target traffic information in the road section to be tested includes:

[0024] Determining identity information of the detected target according to the first multi-target association information;

[0025] Determining first tracking trajectory information and first target speed information of the target according to the first target tracking information;

[0026] Determining second tracking trajectory information and second target speed information of the target according to the second target tracking information;

[0027] The target traffic information in the road section to be tested is obtained by associating the target identity information, the first tracking trajectory information and the first target speed information, the second tracking trajectory information and the second target speed information by using a preset weighted data fusion algorithm.

[0028] Optionally, before determining the first multi-target detection information by using a preset multi-target detection algorithm according to the image information, the radar information and a preset confidence threshold, the method further includes:

[0029] A time alignment operation is performed on the image information and the radar information according to the timestamp of the image information and the timestamp of the radar information to obtain the time-aligned image information and radar information.

[0030] Optionally, before determining the first multi-target detection information by using a preset multi-target detection algorithm according to the image information, the radar information and a preset confidence threshold, the method further includes:

[0031] By using a four-point calibration algorithm, a spatial alignment operation is performed on the image information and the radar information to obtain spatially aligned image information and radar information.

[0032] Optionally, determining the first target tracking information includes:

[0033] According to the image information, using a preset third association algorithm, the same target in different video frames is associated to determine the first target tracking information;

[0034] Determining the second target tracking information includes:

[0035] According to the radar information, a preset fourth association algorithm is used to associate the same target in different radar frames to determine the second target tracking information.

[0036] Optionally, the acquiring of image information and radar information in the road section to be tested includes:

[0037] Obtain image data of the road section to be tested collected in real time by the road test camera;

[0038] According to the called internal measurement parameters and distortion coefficients, dedistortion operation is performed on the image data to obtain image information within the road section to be measured; and

[0039] Obtain radar point cloud data within the road section to be tested collected in real time using the road test millimeter wave radar;

[0040] The radar point cloud data is clustered and analyzed using a target clustering algorithm to obtain radar information within the road section to be tested.

[0041] Optionally, the preset scaling factor is a scaling factor when the target corresponding to the first multi-target detection information matches the target corresponding to the radar information.

[0042] In a second aspect, an embodiment of the present application provides a device for multi-target tracking, the device comprising:

[0043] An acquisition module is used to acquire image information and radar information within the road section to be tested;

[0044] A determination module, configured to determine first multi-target detection information according to the image information, the radar information and a preset confidence threshold using a preset multi-target detection algorithm;

[0045] A first association module, configured to associate the first multi-target detection information with the radar information using a preset first association algorithm to obtain first multi-target association information;

[0046] The second association module is used to use a preset second association algorithm to associate the first multi-target association information, the first target tracking information and the second target tracking information to obtain the target traffic information in the road section to be tested; wherein the first target tracking information is the visual multi-target tracking information determined based on the image information and the first multi-target association information, and the second target tracking information is the radar multi-target tracking information determined based on the radar information.

[0047] In a third aspect, an embodiment of the present application provides a device for multi-target tracking, the device comprising:

[0048] a processor and a memory storing computer program instructions;

[0049] When the processor executes the computer program instructions, the method for multi-target tracking as described in the first aspect and any optional item of the first aspect is implemented.

[0050] In a fourth aspect, an embodiment of the present application provides a computer storage medium having computer program instructions stored thereon, and when the computer program instructions are executed by a processor, the method for multi-target tracking as described in the first aspect and any optional item of the first aspect is implemented.

[0051] The multi-target tracking method, device, equipment and computer storage medium of the embodiment of the present application can deepen the sensor fusion to the target detection stage by matching and fusing the confidence of target detection information obtained by different sensors during the target detection process, and can still ensure the performance of target detection and recognition when any sensor has a problem. Based on this solution, it can effectively reduce problems such as missed detection caused by the failure of a single sensor, improve the performance of multi-target detection, and thus ensure the target tracking performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solution of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0053] Figure 1 It is a flowchart of a method for multi-target tracking provided by an embodiment of the present application;

[0054] Figure 2 is a flowchart of a method for multi-target tracking provided by another embodiment of the present application;

[0055] Figure 3 It is a schematic diagram of a framework for multi-target tracking based on roadside camera millimeter wave radar fusion provided by an embodiment of the present application;

[0056] Figure 4a is a schematic diagram of an application example of a scaling factor provided by an embodiment of the present application;

[0057] Figure 4b is a schematic diagram of an application example of another scaling factor provided by an embodiment of the present application;

[0058] Figure 5 is a schematic diagram of a multi-target recognition result of fusion perception provided by an embodiment of the present application;

[0059] Figure 6 is a schematic diagram of the structure of a multi-target tracking device provided by another embodiment of the present application;

[0060] Figure 7 It is a schematic diagram of the hardware structure of the multi-target tracking device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0061] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only configured to explain the present application and are not configured to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by illustrating the examples of the present application.

[0062] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "include..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0063] Vehicle-road collaborative technology is an important part of building the future smart transportation system. By establishing an information perception and interaction system between vehicles and roadside equipment, it can realize the dynamic collection and integration of traffic information, thereby ensuring traffic safety and improving traffic efficiency. It has far-reaching scientific research significance and application value.

[0064] Tracking technology is an indispensable perception link in vehicle-road cooperative technology. It aims to associate targets in different frames of the same observation, assign consistent identity labels, and then generate their motion trajectories.

[0065] The existing multi-target tracking technology is mainly based on target detection: first detect the target and then track the result of target detection, so the tracking performance depends largely on the multi-target detection performance. Multi-target tracking includes two key technologies: distance definition and association algorithm: define a generalized distance to quantify the proximity between two targets between the previous and next frames, or between different sensors; then follow the association algorithm to select the best association that minimizes the sum of generalized distances.

[0066] However, since the target detection stage is still completed by a single sensor, that is, the information obtained by the sensor is separately detected, the target information of different sensor types is determined, and then the fusion association operation is performed, its multi-target detection performance is still limited. If the target detection operation of any sensor fails or misses, it may not be able to effectively match and track the target information.

[0067] In order to solve the problems of the prior art, the embodiments of the present application provide a method, apparatus, device and computer storage medium for multi-target tracking, which can deepen the sensor fusion to the target detection stage by matching and fusing the confidence of target detection information obtained by different sensors during target detection, thereby effectively reducing problems such as missed detection caused by failure of a single sensor, improving the performance of multi-target detection, and thus ensuring target tracking performance.

[0068] The following describes the method, apparatus, device and computer storage medium for multi-target tracking provided according to the embodiments of the present application in conjunction with the accompanying drawings. It should be noted that these embodiments are not intended to limit the scope of the present application.

[0069] First, the multi-target tracking method provided in the embodiment of the present application is introduced.

[0070] Figure 1 FIG. 1 is a flow chart of a method for multi-target tracking provided by an embodiment of the present application. Figure 1 As shown, in the embodiment of the present application, the multi-target tracking method can be specifically implemented as follows:

[0071] S101: Obtain image information and radar information within the road section to be tested.

[0072] Here, the image information may be image information of monitoring within the road section to be tested acquired by a road test camera, and the image information may include a monitoring screen and a millisecond-level timestamp.

[0073] The radar information may be radar information within the road section to be tested obtained by using a road test millimeter-wave radar. The radar information may include radar messages such as a timestamp, an object identity document (ID), location information, and speed information.

[0074] S102: Determine first multi-target detection information according to the image information, the radar information and a preset confidence threshold using a preset multi-target detection algorithm.

[0075] S103: using a preset first association algorithm to associate the first multi-target detection information with the radar information to obtain first multi-target association information.

[0076] S104: using a preset second association algorithm, associating the first multi-target association information, the first target tracking information and the second target tracking information to obtain target traffic information in the road section to be tested.

[0077] Specifically, the first target tracking information may be visual multi-target tracking information determined based on the image information and the first multi-target association information, and the second target tracking information may be radar multi-target tracking information determined based on the radar information.

[0078] Optionally, the preset first association algorithm and the preset second association algorithm may include association algorithms such as the Hungarian algorithm, the KM algorithm, etc. It is understandable that other association algorithms may also be used here, and the present application does not impose specific restrictions on the association algorithm.

[0079] In summary, the multi-target tracking method implemented in this application can deepen sensor fusion to the target detection stage by matching and fusing the confidence of target detection information obtained by different sensors during target detection. When any sensor has a problem, the target detection and recognition performance can still be guaranteed. Based on this solution, problems such as missed detection caused by failure of a single sensor can be effectively reduced, the performance of multi-target detection can be improved, and the target tracking performance can be guaranteed.

[0080] In order to illustrate the method of multi-target tracking in the embodiment of the present application in more detail, it will be described in detail below.

[0081] like Figure 2 As shown, Figure 2 FIG. 1 is a flow chart of a method for multi-target tracking provided by another embodiment of the present application. In the embodiment of the present application, the method for multi-target tracking can be extended and implemented as follows:

[0082] S201: Obtain image information and radar information within the road section to be tested.

[0083] In some embodiments, obtaining image information in the road section to be tested may include: first, obtaining image data in the road section to be tested collected in real time by a road test camera, and then performing a dedistortion operation on the image data according to the called internal measurement parameters and distortion coefficients to obtain image information in the road section to be tested.

[0084] Specifically, the camera can be used to collect the monitoring images of the road section to be tested in real time and add millisecond timestamps. The Python open source library OpenCV is called to use the internal parameters and distortion coefficients of the camera to perform dedistortion processing on the monitoring images. OpenCV is a cross-platform computer vision and machine learning software library released under the BSD license.

[0085] Optionally, the surveillance video collected by the camera is saved and processed offline later. There is no restriction on the type, model, data collection method, and de-distortion algorithm of the camera, and you can choose according to actual needs.

[0086] In some embodiments, obtaining radar information within the road section to be tested may specifically include: first, obtaining radar point cloud data within the road section to be tested that is collected in real time using a road test millimeter-wave radar; then, using a target clustering algorithm to perform cluster analysis on the radar point cloud data to obtain radar information within the road section to be tested.

[0087] Specifically, after obtaining the radar point cloud data in the road section to be tested, the radar's own target clustering algorithm can be used to process the point cloud data into a radar message. The radar message may include at least one or more of a timestamp, an object ID, location information, and speed information.

[0088] Optionally, in actual application, the radar point cloud data can be processed into radar messages by developing a target clustering algorithm. Furthermore, the radar messages obtained after radar acquisition and analysis can be saved for subsequent offline processing.

[0089] Here, there are no restrictions on the type, model, data collection method, and target clustering algorithm of millimeter-wave radars, and users can make choices based on actual needs.

[0090] S202: Determine second multi-target detection information and a confidence level corresponding to the second multi-target detection information based on the image information and using a preset multi-target detection algorithm.

[0091] The image information may include each frame of video image information. The second multi-target detection information may be target information corresponding to the image information, for example, all vehicles, pedestrians and other traffic targets in the image.

[0092] Exemplarily, using a preset multi-target detection algorithm, each frame of video image in the acquired image information can be used as input to output the bounding box (Bounding Box) or pixel-level contour of all vehicles, pedestrians and other traffic targets in the image, as well as the confidence level of the target existence, that is, the second multi-target detection information and the confidence level corresponding to the second multi-target detection information.

[0093] It can be understood that the preset multi-target detection algorithm may include any one of a visual multi-target detection algorithm based on YOLO, a visual multi-target detection algorithm based on Faster R-CNN convolutional neural network, and a visual multi-target detection algorithm based on CornerNet.

[0094] S203: Determine a first confidence set and a second confidence set according to the confidence corresponding to the second multi-target detection information and a preset first confidence threshold.

[0095] In some embodiments, according to a preset first confidence threshold, the confidences corresponding to all the second multi-target detection information determined in S202 may be screened to obtain two confidence sets.

[0096] When the confidence level corresponding to the second multi-target detection information is higher than or reaches the preset first confidence level threshold, the corresponding confidence levels can form a first confidence level set. That is, the confidence levels corresponding to the second multi-target detection information in the first confidence level set are all higher than the preset first confidence level threshold. In other words, these detected targets are credible targets.

[0097] When the confidence corresponding to the second multi-target detection information is greater than a preset first confidence threshold, the corresponding confidences may form a second confidence set.

[0098] Optionally, the preset first confidence threshold can be selected after testing in an actual target detection system. Different equipment and algorithms used in the system test may cause the preset first confidence threshold to be different, but the preset first confidence threshold should be able to ensure that the output result does not contain any false alarms.

[0099] S204: Correct the second confidence set according to the radar information to obtain a third confidence set.

[0100] In some embodiments, the second confidence set may be corrected using radar information to obtain a corrected third confidence set.

[0101] In some embodiments, a corrected confidence result corresponding to the second confidence set may be calculated based on a preset correction rule and radar information.

[0102] Optionally, the modified confidence result can be obtained by the following modified confidence calculation formula (1):

[0103] Corrected confidence = original confidence ÷ (0.5 + scaling factor of the nearest radar point) (1)

[0104] Then, the modified confidence results may be screened according to a preset second confidence threshold, and the confidence modification results that are determined to have reached the preset second confidence threshold may be set as a third confidence set.

[0105] Optionally, the preset second confidence threshold may be selected after testing in an actual target detection system. Different devices and algorithms used in system testing may result in different preset second confidence thresholds, but different values ​​should be tested and the preset second confidence threshold that maximizes the target recognition performance of the target detection system should be selected.

[0106] In addition, in target detection, the target information data fusion of image information and radar information can make use of the advantages of radar in perceiving distant, small and overlapping objects to compensate for the disadvantages of visual perception in corresponding scenes and improve the accuracy of target detection.

[0107] S205: Fusing the first confidence set and the third confidence set to obtain corresponding first multi-target detection information.

[0108] The first confidence set and the third confidence set are subjected to data fusion to obtain corresponding first multi-target detection information, which can be used as the output of the fusion perception multi-target detection result.

[0109] Here, the first multi-target detection information includes the target detection information after the first confidence set and the third confidence set are fused. The first confidence set includes the target detection information of visual perception detection, that is, the target detection information corresponding to the image information monitored by the camera. The fused perception multi-target detection result determined based on this can ensure the target detection performance when the radar fails or misses detection.

[0110] It can be understood that S202 to S205 are the specific implementation process of S102. Optionally, before determining the first multi-target detection information based on the image information, the radar information and the preset confidence threshold using the preset multi-target detection algorithm, it also includes calibrating the acquired image information and the radar information, and the calibration operation may include time alignment and space alignment.

[0111] Firstly, according to the timestamp of the image information and the timestamp of the radar information, a time alignment operation is performed on the image information and the radar information to obtain the time-aligned image information and radar information.

[0112] By using a four-point calibration algorithm, a spatial alignment operation is performed on the image information and the radar information to obtain the spatially aligned image information and the radar information.

[0113] S206: Utilizing a preset first association algorithm, the first multi-target detection information and the radar information are associated to obtain first multi-target association information.

[0114] In some embodiments, after the first multi-target detection information is determined, the first multi-target detection information and the radar information may be associated according to a preset first association algorithm to obtain the first multi-target association information.

[0115] Optionally, the preset first association algorithm may include a KM algorithm. It is understandable that other association algorithms may also be used here, and the present application does not impose any specific restrictions on the association algorithm.

[0116] In some embodiments, a distance vector between an observed target in the first multi-target detection information and an observed target in the radar information is calculated based on a preset scaling factor, the first multi-target detection information, and the radar information to obtain a first distance vector. The preset scaling factor is a scaling factor when the target corresponding to the first multi-target detection information matches the target corresponding to the radar information.

[0117] Here, the preset scaling factor can be defined as: scaling the BoundingBox or pixel-level contour of the i-th object output by the target detection algorithm with the geometric center as the anchor point, adjusting the scaling multiple, and there must be only one moment in the adjustment structure when the j-th radar target point falls exactly on the edge of the Bounding Box or the pixel-level contour, and recording the scaling multiple at this time as the scaling factor between i and j.

[0118] For example, if the coordinates of the i-th radar target point are (X, Y), the j-th visual detection result is a rectangular BoundingBox, and its upper left corner coordinates are (x1, y1) and lower right corner (x2, y2); then their preset scaling factors can be calculated by the following formula (2):

[0119]

[0120] Optionally, the first distance vector is used to quantify the proximity between the radar observation target and the camera observation target. The first distance vector may be composed of a scaling factor, an observation target speed, and an observation target type information. The first distance vector may be used as an association criterion of a preset first association algorithm.

[0121] In the embodiment of the present application, the target scale information can be introduced into the distance definition in the image coordinate system by using the preset scaling factor. The preset scaling factor can more accurately describe the real physical distance, thereby improving the association and tracking accuracy in the image coordinate system.

[0122] When the first distance vector meets a preset condition, an association result of the first multi-target detection information and the radar information is determined.

[0123] The association result is used as the first multi-target association information.

[0124] Specifically, the preset first association algorithm is used to obtain the optimal association result of the first multi-target detection information and the radar information when the first distance vector meets the preset condition, as the association result of the first multi-target detection information and the radar information. Thus, the same target observed by the radar and the camera respectively is matched one by one.

[0125] For example, when the first association algorithm is preset as the KM algorithm, an M×N association hypothesis matrix is ​​first established, where M is the number of radar detection targets and N is the number of multi-target recognition results of fusion perception. Each element a of the association hypothesis matrix m,n The first distance vector between the target point detected by the radar in the mth row and the Bounding Box of the fusion sensing detection target in the nth column. The association hypothesis matrix is ​​input into the KM algorithm, and the optimal match that minimizes the total generalized distance sum is output, and the same target observed by the radar and the camera are matched one by one. Optionally, the preset condition can be that the total first distance vector sum is the minimum.

[0126] It is understandable that in S206, a consistent identity tag can be assigned to the same target observed by the radar and the camera respectively through the preset first association algorithm, and the data of the same target observed by different sensors can be matched together.

[0127] S207: Determine the identity information of the detected target according to the first multi-target association information.

[0128] S208: Determine first tracking trajectory information and first target speed information of the target according to the first target tracking information.

[0129] In some embodiments, the first target tracking information may be visual multi-target tracking information determined based on the image information and the first multi-target association information.

[0130] Optionally, the same target in different video frames may be associated with each other based on the acquired image information using a preset third association algorithm to determine the first target tracking information.

[0131] Using the preset third association algorithm, a consistent identity tag is assigned to the same target observed in the previous and next frames of the camera surveillance video, and the data observed by the same target in different video frames of the camera are matched together. The first target tracking information includes the first tracking trajectory information and the first target speed information of the detected target, that is, the corresponding tracking trajectory and speed of the target in the camera surveillance video.

[0132] Optionally, a second distance vector consisting of the geometric center position of the target predicted by the extended Kalman filter in the front frame and the pairwise scaling coefficients between the targets in the rear frame, the target image features extracted by the convolutional neural network, the target speed and the target type information can be selected in the camera monitoring video image coordinate system. Using the second association algorithm, the optimal match that minimizes the total second distance vector sum is output, and the same target observed by the front and rear frames of the camera are matched one by one.

[0133] Optionally, the preset second association algorithm may include a KM algorithm. It is understandable that other association algorithms may also be used here, and the present application does not impose any specific restrictions on the association algorithm.

[0134] S209: Determine second tracking trajectory information and second target speed information of the target according to the second target tracking information.

[0135] According to the radar information, the same target in different radar frames is associated using a preset fourth association algorithm to determine the second target tracking information. The second target tracking information includes the second tracking trajectory information and the second target speed information of the detected target, that is, the corresponding tracking trajectory and speed of the target in the radar.

[0136] By using the preset fourth association algorithm, a consistent identity tag can be assigned to the same target observed by the radar in the previous and next frames in time, and the data observed by the same target in different radar message frames can be matched together.

[0137] It is understandable that here, the multi-target tracking function of the radar can be used for association. Alternatively, the association criterion can be selected by oneself, such as the Hungarian algorithm, the KM algorithm or other association algorithms. This application does not impose specific restrictions on the association algorithm.

[0138] S210: using a preset weighted data fusion algorithm, correlating the target's identity information, the first tracking trajectory information and the first target speed information, the second tracking trajectory information and the second target speed information, to obtain target traffic information within the road section to be tested.

[0139] In some embodiments, the preset weighted data fusion algorithm may be a weighted fusion algorithm based on Bayesian estimation.

[0140] The weighted fusion algorithm can be expressed as the following formula (3):

[0141]

[0142] in is the numerical fusion result, z1, z2 are the measurement values ​​of the two sensors, σ1, σ2 are the noise variances measured by the two sensors, namely the camera and radar.

[0143] In some embodiments, the sensor variance determination method is: draw a circular track in the test field and determine the position of the center of the circle, select a target to be measured, such as a radar corner reflector, move around the circular track and observe it using different sensors, and measure the distance between the target to be measured and the center of the circle observed by different sensors. A group of frames is randomly selected, and the distance between the target to be measured and the center of the circle in the group of frames is used to calculate the variance of each sensor.

[0144] It is understandable that other numerical fusion methods may be used, and this application does not limit the scope of data fusion, specific algorithms, etc.

[0145] In summary, the multi-target tracking method in the embodiment of the present application can deepen the sensor fusion to the target detection stage by matching and fusing the confidence of the target detection information obtained by different sensors during the target detection process, and can still ensure the performance of target detection and recognition when any sensor has a problem. Based on this solution, it is possible to effectively reduce problems such as missed detection caused by the failure of a single sensor, improve the performance of multi-target detection, and thus ensure the target tracking performance.

[0146] In addition, a scaling factor is defined to determine the generalized distance, replacing the existing method of calculating the Euclidean distance. This is because the scaling factor introduces the target scale information into the distance definition in the image coordinate system during the scaling of the Bounding Box or pixel-level contour with the geometric center as the anchor point. Therefore, the real physical distance can be described more accurately, thereby improving the association and tracking accuracy in the image coordinate system.

[0147] In order to better understand the implementation scheme of the present invention, the method of multi-target tracking in the embodiment of the present application is described in detail in combination with the actual application scenario of multi-target tracking based on roadside camera millimeter wave radar fusion.

[0148] like Figure 3 As shown, Figure 3 It is a schematic diagram of a framework for multi-target tracking based on roadside camera millimeter wave radar fusion provided by an embodiment of the present application.

[0149] In actual application scenarios, the multi-target tracking method based on the fusion of roadside camera millimeter wave radar, first of all, on the one hand, camera data acquisition and distortion removal. The camera is used to collect the monitoring screen in the road section to be tested in real time, and a millisecond timestamp is attached. The Python open source library OpenCV is called to use the measured internal parameters and distortion coefficients to perform distortion removal on the monitoring screen. The monitoring video collected by the camera can be saved and processed offline later. The camera type, model, data acquisition method and distortion removal algorithm can be selected according to actual needs. On the other hand, millimeter wave radar data acquisition and target clustering. The millimeter wave radar is used to collect radar point cloud data in the road section to be tested in real time, and the radar's built-in target clustering algorithm is used to process the point cloud data into radar messages. The radar message should at least include timestamp, object ID, location and speed information. Here, the target clustering algorithm can also be developed to process the point cloud data into radar messages. The messages collected by the radar are saved and processed offline later. The millimeter wave radar type, model, data acquisition method and target clustering algorithm can be selected according to actual needs.

[0150] Then, align the camera and radar perception data after the above processing in time and space. That is, unify the data collected by the two sensors into the same timestamp and spatial coordinate system.

[0151] Optionally, the time alignment is based on the radar frame, and the camera image frame with the most recent timestamp is selected for each radar message frame in the received camera video frame, and paired with the radar message frame, and the camera video frame and the radar message frame are given the same timestamp.

[0152] The spatial alignment adopts the four-point calibration method. Radar corner reflectors are placed at least four different positions that can be observed by both the camera and the radar, and their coordinates in the camera image coordinate system and the radar coordinate system are recorded respectively. The eight coordinates of the above four points are brought into the transformation matrix of the coordinate system to calculate the projection matrix. The radar points in the radar coordinate system are projected into the image coordinate system using the projection matrix. All radar points are assigned coordinates in the image coordinate system.

[0153] Optionally, an offline processing solution is used to select the camera image with the most recent timestamp for each radar message frame in all camera video frames, including future video frames, and pair it with the radar message frame. Optionally, the camera extrinsic parameters are directly measured using the radar as the world coordinate origin to calculate the projection matrix. Optionally, the visual recognition results in the image coordinate system are projected into the radar coordinate system using the projection matrix. It is understandable that the embodiments of the present application do not limit the spatiotemporal alignment scheme between sensors.

[0154] Again, multi-target detection that integrates camera and radar perception. In order to detect effective traffic targets from visual images and radar messages, the camera data can be first subjected to a YOLO-based visual multi-target detection algorithm. The algorithm can take 1 frame of video image as input, and output the bounding box or pixel-level outline of all vehicles, pedestrians and other traffic targets in the image as well as the confidence level of the target's existence. Unlike existing multi-target detection algorithms that directly use a single threshold to filter the confidence level of target existence and only output results above the threshold value, this application outputs all multi-target detection results and their corresponding confidence levels.

[0155] The result set {Objects_1} with a confidence higher than the first confidence threshold G1 is selected, which is the first confidence set, and the remaining results are recorded as the set {Objects_2}, which is the second confidence set. It should be noted that even if the radar does not detect {Objects_1} at all, {Objects_1} will be used as a subset of the fusion perception multi-target detection output. This can ensure the target detection performance when the radar fails or misses detection.

[0156] Then use the radar message to correct the confidence of the results in {Objects_2}: the closer the radar message is to the visual multi-target detection result, the higher the confidence; otherwise, it decreases. In this way, the advantage of radar in perceiving small and overlapping objects can be used to make up for the disadvantage of visual perception in the above scenario.

[0157] Furthermore, a result set {Objects_3} whose corrected confidence is higher than the second confidence threshold G2 is screened in {Objects_2}, and {Objects_1}∪{Objects_3} is used as the final output of the fusion perception multi-target detection.

[0158] Again, the multi-target detection results are associated with the radar target points. Here, the association technical rules may include association criteria and association algorithms. The association criterion refers to defining a generalized distance to quantify the proximity between radar observation targets and camera observation targets. The association algorithm refers to a specific process followed to screen out an optimal association that minimizes the sum of generalized distances based on the generalized distances between targets observed by different sensors.

[0159] Specifically, the association criterion selects a generalized distance vector composed of a scaling factor, target speed, and target type information. Furthermore, the scaling factor is defined as: scaling the Bounding Box or pixel-level contour of the i-th object output by the target detection algorithm with the geometric center as the anchor point, adjusting the multiple of the scaling size, and there must be only one moment in the adjustment composition when the j-th radar target point falls exactly on the edge of the Bounding Box or pixel-level contour, and the scaling multiple at this time is recorded as the scaling factor between i and j. This generalized distance vector is compatible with the Bounding Box or pixel-level target recognition algorithm. Unlike the widely used Euclidean distance, this scaling factor introduces the target scale information into the distance definition in the image coordinate system during the scaling of the Bounding Box or pixel-level contour with the geometric center as the anchor point. As Figure 4a and Figure 4b As shown, Figure 4a and Figure 4b They are schematic diagrams of application examples of different scaling factors provided by an embodiment of the present application. Figure 4a , the target vehicle is farther away, and the displayed zoom factor is also larger (zoom factor>1), that is, the actual distance is large, and the zoom factor is large; see Figure 4b , the target vehicle is closer, and the displayed zoom factor is also smaller (zoom factor < 1), that is, the actual distance is small and the zoom factor is small. The zoom factor can more accurately describe the actual physical distance, thereby improving the association and tracking accuracy in the image coordinate system.

[0160] Next, the fusion perception multi-target detection results are subjected to visual multi-target tracking, and the radar messages are subjected to radar multi-target tracking.

[0161] Finally, according to the multi-target detection results and radar target point association results, a new ID is assigned to the successfully associated target, and its tracking trajectory and target speed are fused to finally output unified and accurate dynamic traffic information. Dynamic traffic information includes the ID, tracking trajectory and target speed information of each detected target in the monitored section.

[0162] The above multi-target tracking method can improve the performance of multi-target detection and thus ensure the target tracking performance. In challenging scenes with overlapping targets, the multi-target detection module with camera and millimeter-wave radar fusion perception has excellent performance, which provides a guarantee for subsequent accurate tracking. For details, please refer to Figure 5 , Figure 5 FIG. 1 is a schematic diagram of a multi-target recognition result of fusion perception provided by an embodiment of the present application, wherein the square represents visual sensor detection and the circle represents millimeter wave radar detection.

[0163] Based on the multi-target tracking method provided in the above embodiment, the present application also provides a specific implementation of the multi-target tracking device. Please refer to the following embodiment.

[0164] In the embodiments of the present application, Figure 6 As shown, Figure 6 : is a schematic diagram of a structure of a multi-target tracking device provided by another embodiment of the present application. The multi-target tracking device specifically includes:

[0165] An acquisition module 601 is used to acquire image information and radar information in the road section to be tested;

[0166] A determination module 602 is used to determine first multi-target detection information according to the image information, the radar information and a preset confidence threshold using a preset multi-target detection algorithm;

[0167] A first association module 603 is used to associate the first multi-target detection information with the radar information using a preset first association algorithm to obtain first multi-target association information;

[0168] The second association module 604 is used to use a preset second association algorithm to associate the first multi-target association information, the first target tracking information and the second target tracking information to obtain the target traffic information in the road section to be tested; wherein the first target tracking information is visual multi-target tracking information determined based on the image information and the first multi-target association information, and the second target tracking information is radar multi-target tracking information determined based on the radar information.

[0169] In summary, the multi-target tracking device implemented in this application can be used to implement the multi-target tracking method in the above-mentioned embodiment. This method can deepen the sensor fusion to the target detection stage by matching and fusing the confidence of the target detection information obtained by different sensors during the target detection process. When any sensor has a problem, the performance of target detection and recognition can still be guaranteed. Based on this solution, problems such as missed detection caused by the failure of a single sensor can be effectively reduced, the performance of multi-target detection can be improved, and the target tracking performance can be guaranteed.

[0170] Figure 6 Each module / unit in the multi-target tracking device has the following features: Figure 1 to Figure 2 The functions of the various steps in the corresponding figure numbers of the multi-target tracking method shown in the figure can achieve their corresponding technical effects. For the sake of concise description, they will not be repeated here.

[0171] Based on the multi-target tracking method provided in the above embodiment, the present application also provides a specific hardware structure description of a multi-target tracking device. Please refer to the following embodiment.

[0172] Figure 7 It is a schematic diagram of the hardware structure of the multi-target tracking device provided in an embodiment of the present application.

[0173] The multi-target tracking device may include a processor 701 and a memory 702 storing computer program instructions.

[0174] Specifically, the processor 701 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0175] The memory 702 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 702 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory 702 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 702 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 702 is a non-volatile solid-state memory. In a specific embodiment, the memory 702 includes a read-only memory (ROM). In appropriate cases, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM) or a flash memory or a combination of two or more of these.

[0176] The processor 701 implements any one of the multi-target tracking methods in the above embodiments by reading and executing computer program instructions stored in the memory 702 .

[0177] In one example, the multi-target tracking device may further include a communication interface 703 and a bus 710. Figure 7 As shown, the processor 701, the memory 702, and the communication interface 703 are connected via a bus 710 and communicate with each other.

[0178] The communication interface 703 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.

[0179] Bus 710 includes hardware, software or both, and the parts of the equipment of multi-target tracking are coupled to each other. For example, but not limitation, bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industrial standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industrial standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 710 may include one or more buses. Although the present application embodiment describes and shows a specific bus, the application considers any suitable bus or interconnection.

[0180] The multi-target tracking device can execute the multi-target tracking method in the embodiment of the present application, thereby realizing the combination of Figure 1 to Figure 2 Describe the method of multiple target tracking.

[0181] In addition, in combination with the multi-target tracking method in the above embodiment, the embodiment of the present application can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any multi-target tracking method in the above embodiment is implemented.

[0182] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.

[0183] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, a dedicated integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0184] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be performed simultaneously.

[0185] Aspects of the present disclosure are described above with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It can also be understood that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs a specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0186] The above is only a specific implementation of the present application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the protection scope of this application.

Claims

1. A method for multi-target tracking, characterized in that: include: Obtain image information and radar information within the road section to be tested; Determining first multi-target detection information according to the image information, the radar information and a preset confidence threshold using a preset multi-target detection algorithm; Using a preset first association algorithm, associating the first multi-target detection information with the radar information to obtain first multi-target association information; Using a preset second association algorithm, the first multi-target association information, the first target tracking information, and the second target tracking information are associated to obtain target traffic information within the road section to be tested; The first target tracking information is visual multi-target tracking information determined according to the image information and the first multi-target association information, and the second target tracking information is radar multi-target tracking information determined according to the radar information; The determining of first multi-target detection information based on the image information, the radar information and a preset confidence threshold using a preset multi-target detection algorithm includes: Determine, according to the image information, the second multi-target detection information and the confidence level corresponding to the second multi-target detection information by using a preset multi-target detection algorithm; Determining a first confidence set and a second confidence set according to the confidence corresponding to the second multi-target detection information and a preset first confidence threshold; Modifying the second confidence set according to the radar information to obtain a third confidence set; The first confidence set and the third confidence set are data-fused to obtain corresponding first multi-target detection information.

2. The method according to claim 1, characterized in that The step of correcting the second confidence set in combination with the radar information to obtain a third confidence set includes: Calculate and obtain a confidence correction result corresponding to the second confidence set according to a preset correction rule and the radar information; The confidence correction result that reaches a preset second confidence threshold is determined to be the third confidence set.

3. The method according to claim 1, characterized in that The using a preset first association algorithm to associate the first multi-target detection information with the radar information to obtain first multi-target association information includes: Calculate a distance vector between an observed target in the first multi-target detection information and an observed target in the radar information according to a preset scaling factor and the first multi-target detection information and the radar information to obtain a first distance vector; When the first distance vector satisfies a preset condition, determining an association result between the first multi-target detection information and the radar information; The association result is used as the first multi-target association information.

4. The method according to claim 1, characterized in that: The method of using a preset second association algorithm to associate the first multi-target association information, the first target tracking information, and the second target tracking information to obtain target traffic information in the road section to be tested includes: Determining identity information of the detected target according to the first multi-target association information; Determining first tracking trajectory information and first target speed information of the target according to the first target tracking information; Determining second tracking trajectory information and second target speed information of the target according to the second target tracking information; The target traffic information in the road section to be tested is obtained by associating the target identity information, the first tracking trajectory information and the first target speed information, the second tracking trajectory information and the second target speed information by using a preset weighted data fusion algorithm.

5. The method according to claim 1, characterized in that Before determining the first multi-target detection information according to the image information, the radar information and the preset confidence threshold using a preset multi-target detection algorithm, the method further includes: A time alignment operation is performed on the image information and the radar information according to the timestamp of the image information and the timestamp of the radar information to obtain the time-aligned image information and radar information.

6. The method according to claim 1 or 5, characterized in that: Before determining the first multi-target detection information according to the image information, the radar information and the preset confidence threshold using a preset multi-target detection algorithm, the method further includes: A four-point calibration algorithm is used to perform a spatial alignment operation on the image information and the radar information to obtain spatially aligned image information and radar information.

7. The method according to claim 1, characterized in that Determining the first target tracking information includes: According to the image information, using a preset third association algorithm, the same target in different video frames is associated to determine the first target tracking information; Determining the second target tracking information includes: According to the radar information, a preset fourth association algorithm is used to associate the same target in different radar frames to determine the second target tracking information.

8. The method according to claim 1, characterized in that The step of obtaining image information and radar information in the road section to be tested includes: Obtain image data of the road section to be tested collected in real time by the road test camera; According to the called internal measurement parameters and distortion coefficients, dedistortion operation is performed on the image data to obtain image information within the road section to be measured; and Obtain radar point cloud data within the road section to be tested collected in real time using the road test millimeter wave radar; The radar point cloud data is clustered and analyzed using a target clustering algorithm to obtain radar information within the road section to be tested.

9. The method according to claim 3, characterized in that: The preset scaling factor is a scaling factor when the target corresponding to the first multi-target detection information matches the target corresponding to the radar information.

10. A device for tracking multiple targets, characterized in that: The device comprises: An acquisition module is used to acquire image information and radar information within the road section to be tested; A determination module, configured to determine first multi-target detection information according to the image information, the radar information and a preset confidence threshold using a preset multi-target detection algorithm; A first association module, configured to associate the first multi-target detection information with the radar information using a preset first association algorithm to obtain first multi-target association information; a second association module, configured to use a preset second association algorithm to associate the first multi-target association information, the first target tracking information, and the second target tracking information to obtain target traffic information in the road section to be tested; wherein the first target tracking information is visual multi-target tracking information determined according to the image information and the first multi-target association information, and the second target tracking information is radar multi-target tracking information determined according to the radar information; The determining of first multi-target detection information based on the image information, the radar information and a preset confidence threshold using a preset multi-target detection algorithm includes: Determine, according to the image information, the second multi-target detection information and the confidence level corresponding to the second multi-target detection information by using a preset multi-target detection algorithm; Determining a first confidence set and a second confidence set according to the confidence corresponding to the second multi-target detection information and a preset first confidence threshold; Modifying the second confidence set according to the radar information to obtain a third confidence set; The first confidence set and the third confidence set are data-fused to obtain corresponding first multi-target detection information.

11. A device for multi-target tracking, characterized in that: The device comprises: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the method for multi-target tracking according to any one of claims 1 to 9 is implemented.

12. A computer storage medium, characterized in that: The computer storage medium stores computer program instructions, which, when executed by a processor, implement the multi-target tracking method according to any one of claims 1 to 9.