Target Tracking Method, Device, Computer Equipment and Storage Medium

Through multi-sensor data fusion, the actual measured feature information and historical detection boxes of candidate detection objects are obtained, combined with interleaving and matching and similarity, the problem of insufficient feature dimensions in traditional target tracking methods is solved, and higher accuracy and stable target tracking is achieved.

CN114359334BActive Publication Date: 2025-07-04VANJEE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011062425.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-30
Publication Date
2025-07-04
Estimated Expiration
2040-09-30

AI Technical Summary

Technical Problem

In traditional target tracking methods, the use of a single sensor leads to insufficient feature dimensions, resulting in low accuracy of target tracking results.

Method used

Various types of sensors (such as camera equipment and lidar) are used to collect data, and accurately track the candidate detection objects by obtaining the actual measured feature information of each candidate detection object and the historical detection box of the target to be tracked, combined with the interleaving and comparison and similarity matching, the accurate tracking of the candidate detection objects is achieved.

Benefits of technology

Improve the accuracy and stability of target tracking results, ensuring consistency of different sensor data and accurate positioning of detection objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114359334B_ABST
    Figure CN114359334B_ABST
Patent Text Reader

Abstract

The present application relates to a target tracking method, apparatus, computer device, and storage medium. By obtaining first current frame data of a target scene collected by a first sensor and second current frame data of the target scene collected by at least one second sensor of a different type from the first sensor, obtaining measured feature information of each detected object in the first current frame data and each candidate detected object that is successfully associated and matched with each detected object in the second current frame data in the current frame, and then tracking each candidate detected object according to the measured feature information of each candidate detected object in the current frame and the predicted detection boxes of each target to be tracked in the current frame. This method can accurately determine the target to be tracked to which each detected object belongs and effectively complete the target tracking for each frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of tracking technology, and particularly to a target tracking method, device, computer device, and storage medium. Background Art

[0002] With the development of sensor technology and computer technology, simultaneous localization and mapping solutions based on various sensors have been widely applied in fields such as robot autonomous navigation, unmanned driving, mobile measurement, and battlefield environment construction.

[0003] For example, when performing target tracking, data information of a target can be detected by sensors, and after analyzing the detected data information, tracking of the target can be achieved. Generally, data information obtained by different sensors has different dimensions, different architectures, and different emphases on features (such as contour, size, trajectory, category, color, texture, etc.). For the same target, the extracted features are different. However, in traditional technologies, a single sensor is often used to track the target, which easily causes insufficient feature dimensions when detecting the target, resulting in low accuracy of the target tracking result. Summary of the Invention

[0004] Based on this, it is necessary to provide a target tracking method, device, computer device, and storage medium that can improve the accuracy of the target tracking result for the above technical problems.

[0005] In a first aspect, an embodiment of the present application provides a target tracking method, which includes:

[0006] Obtain first current frame data of a target scene collected by a first sensor, and second current frame data of the target scene collected by at least one second sensor; the first sensor and the second sensor are sensors of different types;

[0007] According to the first current frame data and the second current frame data, obtain the measured feature information of each candidate detection object in the current frame; the candidate detection object is a detection object that is successfully associated and matched between each detection object in the first current frame data and each detection object in the second current frame data; the measured feature information represents the inherent characteristic information of each candidate detection object;

[0008] According to the historical detection frames of each target to be tracked in the previous frame, predict the predicted detection frames of each target to be tracked in the current frame;

[0009] Track each candidate detection object according to the measured feature information of each candidate detection object in the current frame, and the predicted detection frames of each target to be tracked in the current frame.

[0010] In one embodiment, the measured feature information includes a measured three-dimensional detection box and measured trajectory features; tracking each candidate detection object according to the measured feature information of each candidate detection object in the current frame and the predicted detection boxes of each target to be tracked in the current frame includes:

[0011] Obtaining the intersection over union (IoU) between the measured three-dimensional detection box of each candidate detection object in the current frame and the predicted three-dimensional detection box of each target to be tracked in the current frame;

[0012] Obtaining the similarity between the measured trajectory features of the remaining detection objects and the historical trajectory features of each target to be tracked in the previous frame; the remaining detection objects are candidate detection objects with an IoU less than a preset IoU threshold;

[0013] Tracking each candidate detection object with an IoU greater than the IoU threshold and each of the remaining detection objects with a similarity greater than a preset similarity threshold.

[0014] In one embodiment, obtaining the first current frame data of the target scene collected by the first sensor and the second current frame data of the target scene collected by at least one second sensor includes:

[0015] For the data collected by any frame of the sensor, obtaining the timestamp T1 of the first current frame data and the timestamp T2 of the second current frame data;

[0016] If the interval between T1 and T2 is less than a preset interval threshold, determining the first current frame data and the second current frame data as the data collected in the current frame;

[0017] If the interval between T1 and T2 is greater than the preset threshold, discarding the first current frame data and the second current frame data, and re-obtaining the timestamp T1 of the first current frame data and the timestamp T2 of the second current frame data for the next frame.

[0018] In one embodiment, obtaining the timestamp T1 of the first current frame data and the timestamp T2 of the second current frame data includes:

[0019] If the first current frame data and the second current frame data do not carry timestamps, converting the acquisition time of the first current frame data and the acquisition time of the second current frame data to the same time axis to obtain T1 and T2.

[0020] In one embodiment, before obtaining the first current frame data of the target scene collected by the first sensor and the second current frame data of the target scene collected by at least one second sensor, the method further includes:

[0021] Adjust the sampling frequencies of the first sensor and the second sensor to be the same, and calibrate the extrinsic parameter information between the first sensor and the second sensor.

[0022] In one embodiment, the first sensor is a camera device; the second sensor is a lidar;

[0023] Then, calibrating the extrinsic parameter information between the first sensor and the second sensor includes:

[0024] Adjust the relative pose information between the camera device and the lidar to the target pose information, and obtain the calibrated extrinsic parameter information of the camera device according to a preset calibration algorithm;

[0025] Calibrate the extrinsic parameter information of the camera device according to the calibrated extrinsic parameter information.

[0026] In one embodiment, the first current frame data is pixel data collected by the camera device, and the second current frame data is point cloud data collected by the lidar; the measured feature information includes measured 3D detection boxes and measured trajectory features;

[0027] Then, determining the measured feature information of each candidate detection object in the current frame according to the first current frame data and the second current frame data includes:

[0028] Obtain the 2D detection boxes of each detection object in the pixel data and the measured 3D detection boxes of each detection object in the point cloud data;

[0029] Match the 2D detection boxes of each detection object in the pixel data with the measured 3D detection boxes of each detection object in the point cloud data to determine candidate detection objects;

[0030] Determine the measured trajectory features of each candidate detection object according to the 2D detection boxes and feature information of each candidate detection object in the pixel data, and the measured 3D detection boxes and feature information of each detection object in the point cloud data.

[0031] In one embodiment, matching the 2D detection boxes of each detection object in the pixel data with the measured 3D detection boxes of each detection object in the point cloud data to determine candidate detection objects includes:

[0032] Map each measured 3D detection box to a corresponding 2D mapped detection box;

[0033] Obtain the intersection over union (IoU) between each 2D detection box and each 2D mapped detection box, and determine the detection objects with an IoU greater than the IoU threshold as candidate detection objects.

[0034] In one embodiment, determining the measured trajectory features of each candidate detection object based on the two-dimensional detection boxes and feature information of each candidate detection object in the pixel data, and the measured three-dimensional detection boxes and feature information of each detection object in the point cloud data includes:

[0035] Combining the two-dimensional detection boxes and feature information of each candidate detection object in the pixel data to determine the two-dimensional trajectory features of each candidate detection object; combining the measured three-dimensional detection boxes and feature information of each candidate detection object in the point cloud data to determine the three-dimensional trajectory features of each candidate detection object;

[0036] Converting the three-dimensional trajectory features of each candidate detection object into corresponding two-dimensional mapped trajectory features;

[0037] Extracting the measured trajectory features of each candidate detection object based on the fusion data of each two-dimensional mapped trajectory feature and each two-dimensional trajectory feature.

[0038] In one embodiment, the above-mentioned converting the three-dimensional trajectory features of each candidate detection object into corresponding two-dimensional mapped trajectory features includes:

[0039] Converting the three-dimensional coordinates of the point cloud points in the measured three-dimensional detection box corresponding to each three-dimensional trajectory feature into two-dimensional coordinates;

[0040] Obtaining the bird's-eye view corresponding to each measured three-dimensional detection box according to the two-dimensional coordinates of each point cloud point and the z-axis coordinate of each point cloud point in the three-dimensional coordinates;

[0041] Obtaining the intensity map corresponding to each measured three-dimensional detection box according to the two-dimensional coordinates of each point cloud point and the intensity of each point cloud point;

[0042] Obtaining the density map corresponding to each measured three-dimensional detection box according to the density of each point cloud point in the z-axis direction in the bird's-eye view;

[0043] Combining and processing the bird's-eye view, intensity map and density map to obtain the two-dimensional mapped trajectory features corresponding to each three-dimensional trajectory feature.

[0044] In one embodiment, the above-mentioned obtaining the bird's-eye view corresponding to each measured three-dimensional detection box according to the two-dimensional coordinates of each point cloud point and the z-axis coordinate of each point cloud point in the three-dimensional coordinates includes:

[0045] Performing normalization processing on the z-axis coordinates of each point cloud point in the three-dimensional coordinates, and determining the normalized z-axis coordinates as the pixel values of each point cloud point;

[0046] Using the pixel value corresponding to each point cloud point as the pixel value at the corresponding two-dimensional coordinate position to obtain the bird's-eye view corresponding to each measured three-dimensional detection box.

[0047] In one embodiment, obtaining the intensity map corresponding to each measured three-dimensional detection box according to the two-dimensional coordinates and intensities of each point cloud point includes:

[0048] Normalize the intensities of each point cloud point, and determine the normalized intensities as the pixel values of each point cloud point;

[0049] Use the pixel values corresponding to each point cloud point as the pixel values at the corresponding two-dimensional coordinate positions to obtain the intensity map corresponding to each measured three-dimensional detection box.

[0050] In one embodiment, obtaining the density map corresponding to each measured three-dimensional detection box according to the density of each point cloud point in the z-axis direction in the bird's-eye view includes:

[0051] Determine the density of the point cloud point at each position in the z-axis direction according to the number of point cloud points in the z-axis direction at the two-dimensional coordinate position of each point cloud point, the maximum value of the number of point cloud points at all coordinate positions, and the minimum value of the number of point cloud points at all coordinate positions;

[0052] Use the density of the point cloud point at each coordinate position in the z-axis direction as the pixel value at the corresponding two-dimensional coordinate position to obtain the density map corresponding to each measured three-dimensional detection box.

[0053] In one embodiment, extracting the measured trajectory features of each candidate detection object according to each two-dimensional mapping trajectory feature and the fusion data of each two-dimensional trajectory feature includes:

[0054] After compressing each two-dimensional trajectory feature and each two-dimensional mapping trajectory feature to the same proportional size, splice them to obtain a fusion data matrix;

[0055] Determine the feature information extracted from the fusion data matrix as the measured trajectory features of each candidate detection object.

[0056] In one embodiment, predicting the predicted detection box of each target to be tracked in the current frame according to the historical detection box of each target to be tracked in the previous frame includes:

[0057] Through a preset tracking algorithm model, predict the predicted three-dimensional detection box of each target to be tracked in the current frame according to the historical three-dimensional detection box of each target to be tracked in the previous frame; wherein, the tracking algorithm model is constructed based on the spatial equation of the uniformly variable motion state.

[0058] In a second aspect, an embodiment of the present application provides a target tracking device, and the device includes:

[0059] An acquisition module, configured to acquire first current frame data of a target scene collected by a first sensor and second current frame data of the target scene collected by at least one second sensor; the first sensor and the second sensor are sensors of different types;

[0060] A feature acquisition module, configured to acquire measured feature information of each candidate detection object in the current frame according to the first current frame data and the second current frame data; the candidate detection objects are detection objects that are successfully associated and matched between the detection objects in the first current frame data and the detection objects in the second current frame data; the measured feature information represents the inherent characteristic information of the candidate detection objects;

[0061] A prediction module, configured to predict predicted detection frames of each target to be tracked in the current frame according to the historical detection frames of each target to be tracked in the previous frame;

[0062] A tracking module, configured to track each candidate detection object according to the measured feature information of each candidate detection object in the current frame and the predicted detection frames of each target to be tracked in the current frame.

[0063] In a third aspect, an embodiment of the present application provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of any method provided in the first aspect embodiment are implemented.

[0064] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any method provided in the first aspect embodiment are implemented.

[0065] A target tracking method, device, computer device, and storage medium provided by an embodiment of the present application. The target tracking method, device, computer device, and storage medium obtain first current frame data of a target scene collected by a first sensor and second current frame data of the target scene collected by at least one second sensor of a different type from the first sensor, obtain measured feature information of each candidate detection object in the current frame determined by successfully associating and matching each detection object in the first current frame data with each detection object in the second current frame data, and then track each candidate detection object according to the measured feature information of each candidate detection object in the current frame and the predicted detection frames of each target to be tracked in the current frame. In this method, when tracking each detection object in the current frame, the measured feature information of each detection object is determined based on data collected by two or more different types of sensors, that is, the data collected by multiple types of sensors is integrated to determine the measured feature information of the detection objects in each frame, which can accurately and completely reflect the features of the detection objects in each frame. In this way, when matching the measured feature information of each candidate detection object with accurate data of each target to be tracked, the target to which each detection object belongs can be accurately determined, and the target tracking of each frame can be effectively completed. In addition, candidate detection objects are screened out based on the successful association and matching of each detection object in the first current frame data with each detection object in the second current frame data, ensuring the consistency of the detection objects in the data collected by each type of sensor. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 FIG. is an application environment diagram of a target tracking provided for an embodiment;

[0067] Figure 1a FIG. is the positional relationship between a lidar and a camera device in an embodiment;

[0068] Figure 1b FIG. is the internal structure diagram of a computer device in an embodiment;

[0069] Figure 2 FIG. is a flowchart of a target tracking method provided for an embodiment;

[0070] Figure 3 FIG. is a flowchart of a target tracking method provided for another embodiment;

[0071] Figure 4 FIG. is a flowchart of a target tracking method provided for another embodiment;

[0072] Figure 5 FIG. is a flowchart of a target tracking method provided for another embodiment;

[0073] Figure 6Flow schematic diagram of a target tracking method provided for another embodiment;

[0074] Figure 7 Flow schematic diagram of a target tracking method provided for another embodiment;

[0075] Figure 8 Flow schematic diagram of a target tracking method provided for another embodiment;

[0076] Figure 9 Flow schematic diagram of a target tracking method provided for another embodiment;

[0077] Figure 10 Flow schematic diagram of a target tracking method provided for another embodiment;

[0078] Figure 11 Flow schematic diagram of a target tracking method provided for another embodiment;

[0079] Figure 12 Flow chart of a target tracking method provided for another embodiment;

[0080] Figure 13 Block diagram of a target tracking device provided for one embodiment. Detailed implementation manners

[0081] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0082] A target tracking method provided by the present application can be applied to Figure 1 the application environment shown. This application environment includes a lidar 01, an imaging device 02, and a computer device 03. Among them, the lidar 01, the imaging device 02, and the computer device can communicate with each other; the lidar 01 includes but is not limited to pulsed radar, continuous wave radar, meter wave radar, decimeter wave radar, centimeter wave radar, etc., including 8-line, 16-line, 24-line, 32-line, 64-line, 128-line lidar; the imaging device 02 includes but is not limited to professional cameras, CCD cameras, network cameras, portable cameras, black and white cameras, color cameras, infrared cameras, X-ray cameras, undercover cameras, etc.; the computer device 03 includes but is not limited to servers, various terminals: personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices, etc., and the cameras include bullet cameras, dome cameras, spherical cameras.

[0083] Among them, the positions between the lidar 01 and the imaging device 02 are relatively fixed. The installation method can be that the radar and the camera are installed on a roadside pole, and the position is not limited, and both the vertical pole and the horizontal pole are acceptable. For example, Figure 1a is the installation schematic diagram of the lidar 01 and the imaging device 02 shown. Figure 1b An internal structure diagram of a computer device 03 is provided. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store relevant data for target tracking. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a target tracking method.

[0084] Embodiments of the present application provide a target tracking method, device, computer device, and storage medium, which can improve the accuracy of target tracking results. Next, the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems will be described in detail through embodiments and in conjunction with the accompanying drawings. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. It should be noted that a target tracking method provided by the present application Figures 2 - 12 has a computer device as its execution subject. Among them, its execution subject can also be a target tracking device, and the device can be implemented as a part or all of the computer device through software, hardware, or a combination of software and hardware.

[0085] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments.

[0086] In one embodiment, Figure 2 a target tracking method is provided. This embodiment involves the specific process in which the computer device performs correlation matching of detection objects based on data collected by two or more different types of sensors in the same scene, performs matching analysis on the predicted data and the measured data of the detection objects with successful correlation matching, and then corresponds the detection objects with successful matching in the matching analysis result to the target to be tracked, and then tracks the target to be tracked. As Figure 2 shown, the method includes:

[0087] S101. Obtain the first current frame data of the target scene collected by the first sensor and the second current frame data of the target scene collected by at least one second sensor. The first sensor and the second sensor are sensors of different types.

[0088] Among them, the target scene refers to the scene where the target to be tracked is located. For example, if the target is a vehicle and the vehicle is driving on a certain road, then a certain range of this road is the target scene.

[0089] The first sensor includes but is not limited to a camera device or a lidar. Similarly, the second sensor also includes but is not limited to a camera device or a lidar. However, the first sensor and the second sensor are sensors of different types. For example, if the first sensor is a lidar, then the second sensor is a camera device. Refer to Figure 1a the installation methods of the lidar and the camera device shown.

[0090] The current frame data of the target scene collected by the first sensor is the first current frame data. Taking the first sensor as a lidar as an example, the first current frame data is the three-dimensional point cloud data of the current data of the target scene collected. For example, if the target is a vehicle and the vehicle is driving on a certain road, the lidar set on the roadside scans a certain range of this road to obtain the three-dimensional point cloud data in this spatial scene. Similarly, if the second sensor is a camera device, then the second current frame data is the two-dimensional pixel data of the current frame of the target scene collected by the camera device, such as the video data of the target scene.

[0091] S102. According to the first current frame data and the second current frame data, obtain the measured feature information of each candidate detection object in the current frame. The candidate detection object is the detection object that is successfully associated and matched between each detection object in the first current frame data and each detection object in the second current frame data. The measured feature information represents the inherent characteristic information of each candidate detection object.

[0092] Among them, the candidate detection object is the detection object that is successfully associated and matched between the first current frame data and the second current frame data. For example, there are 5 detection objects in the first current frame data and 5 detection objects in the second current frame data, but only 4 are associated and matched. Then these 4 detection objects are the candidate detection objects.

[0093] In practical applications, after obtaining the first current frame data and the second current frame data, first perform association and matching on the detection objects in the first current frame data and the second current frame data to obtain the candidate detection objects, and then continue to obtain the measured feature information of each candidate detection object in the current frame from the first current frame data and the second current frame data.

[0094] In addition, when the computer device determines the measured 3D detection box and the measured trajectory features of the current frame of the candidate detection objects, the first current frame data and the second current frame data must be valid frames. Optionally, a valid frame means that the acquisition times of the first current frame data and the second current frame data are synchronized. For example, the time interval between their acquisitions is less than a preset threshold.

[0095] The measured feature information represents the inherent characteristic information of each candidate detection object. Optionally, the measured feature information includes the measured 3D detection box and the measured trajectory features. Among them, the measured 3D detection box and the measured trajectory features refer to the 3D detection boxes of each target and the trajectory features of each target calculated based on the actual information in the first current frame data and the second current frame data. Among them, the 3D detection box is the detection box of each candidate detection object in the point cloud, and the trajectory feature can be a feature such as a histogram of oriented gradients that can reflect various information of the detection object. For example, the computer device comprehensively determines the trajectory features of each candidate detection object based on the position information of the detection boxes of each candidate detection object, the color, texture information, etc. in the detection box in the first current frame data and the second current frame data, as the measured trajectory features of each candidate detection object in the current frame. Of course, in practical applications, the measured feature information may also include the 2D detection box and the 2D trajectory features of the candidate detection objects, and the embodiments of the present application do not limit this.

[0096] S103. Predict the predicted detection boxes of each target to be tracked in the current frame according to the historical detection boxes of each target to be tracked in the previous frame.

[0097] For each target to be tracked, the position change of its motion trajectory in consecutive frames is related, and some prediction methods can be used to predict the future trajectory of the target based on the trajectory that has occurred. Now, what needs to be predicted is the predicted detection boxes of each target to be tracked in the current frame, and the feature information (including the detection box and the trajectory features) of each target to be tracked in the previous frame is already known. Therefore, the predicted detection boxes of each target to be tracked in the current frame can be predicted according to the historical detection boxes of each target to be tracked in the previous frame. In practical applications, the detection box may include a 2D detection box or a 3D detection box, and the embodiments of the present application do not limit this. For example, according to the historical 3D detection boxes of each target to be tracked in the previous frame, the predicted 3D detection boxes of each target to be tracked in the current frame are predicted. Here, the historical 3D detection box refers to the known 3D detection box that has occurred for the target to be tracked in the previous frame. For example, the computer device can use a preset Kalman filter to predict the predicted 3D detection boxes of each target to be tracked in the current frame according to the historical 3D detection boxes of each target to be tracked in the previous frame.

[0098] S104. Track each candidate detection object according to the measured feature information of each candidate detection object in the current frame and the predicted detection boxes of each target to be tracked in the current frame.

[0099] Since the historical detection boxes of each target to be tracked in the previous frame are already known, the predicted detection boxes of each target to be tracked in the current frame can be predicted based on the historical detection boxes of each target to be tracked in the previous frame, which can be used as the identification basis for each target to be tracked in the current frame. For example, taking the three-dimensional detection box as an example, by determining whether the measured three-dimensional detection box of each candidate detection object in the current frame matches the predicted three-dimensional detection box of each target to be tracked in the current frame, it can be determined which target to be tracked each candidate detection object belongs to, and thus each candidate detection object can be tracked. It should be noted here that there are multiple targets to be tracked in each frame of data during target tracking, and it is necessary to determine which target to be tracked each detection object in the current frame data belongs to. After determining the target to be tracked to which each detection object in the current frame data belongs, the trajectories of each target to be tracked in the current frame data can be determined conversely. Therefore, in all embodiments of the present application, the detection object is called during the tracking process (not yet tracked successfully), and the successfully tracked one is called the target to be tracked (or target), which will not be elaborated hereinafter.

[0100] The target tracking method provided in this embodiment obtains the first current frame data of the target scene collected by the first sensor and the second current frame data of the target scene collected by at least one second sensor of a different type from the first sensor, obtains the measured feature information of each candidate detection object in the current frame determined by successfully associating and matching each detection object in the first current frame data and each detection object in the second current frame data, and then tracks each candidate detection object according to the measured feature information of each candidate detection object in the current frame and the predicted detection boxes of each target to be tracked in the current frame. In this method, when tracking each detection object in the current frame, the measured feature information of each detection object is determined based on the data collected by two or more different types of sensors, that is, the data collected by multiple types of sensors are integrated to determine the measured feature information of the detection object in each frame, which can accurately and completely reflect the features of the detection object in each frame. In this way, when matching the measured feature information of each candidate detection object with the accurate data of each target to be tracked, the target to be tracked to which each detection object belongs can be accurately determined, and the target tracking of each frame can be effectively completed. In addition, candidate detection objects are screened out based on the successful association and matching of each detection object in the first current frame data and each detection object in the second current frame data, ensuring the consistency of the detection objects in the data collected by each type of sensor.

[0101] An embodiment is provided to illustrate the process of predicting the predicted detection boxes of each target to be tracked in the current frame based on the feature information of each target to be tracked in the previous frame in step S103. This embodiment is illustrated by taking a three-dimensional detection box as an example. This embodiment includes: using a preset tracking algorithm model to predict the predicted three-dimensional detection boxes of each target to be tracked in the current frame based on the historical three-dimensional detection boxes of each target to be tracked in the previous frame; wherein, the tracking algorithm model is constructed based on the spatial equation of a uniformly variable motion state.

[0102] Among them, the state space equation is established according to the different motion states of the target in space, and can reflect the changes of the target's trajectory at different times, motion information, etc. expressions; a tracking algorithm model constructed based on this state space equation, such as a Kalman filter, can be closer to the real information when the target moves in space. After constructing the tracking algorithm model, using this tracking algorithm model to predict the predicted three-dimensional detection boxes of each target to be tracked in the current frame based on the historical three-dimensional detection boxes of each target can make the predicted three-dimensional detection boxes of each target to be tracked more accurate.

[0103] Since the spatial equation of the uniformly variable motion state couples the motion of the target in space with time and takes into account the influence of acceleration, it can make the trajectory prediction error smaller and improve the tracking effect of large speed changes. Therefore, the tracking algorithm model constructed based on the spatial equation of the uniformly variable motion state is very accurate in predicting the predicted three-dimensional detection boxes of each target to be tracked in the current frame based on the historical three-dimensional detection boxes of each target to be tracked in the previous frame.

[0104] Exemplarily, taking the detection box of each target at each moment as a rectangular box, the detection result of each frame of data is a target box (for example, the point cloud is a three-dimensional target box, and the image is a two-dimensional target box); regarding the motion of the target in the image as a uniformly variable motion, the following state space equation (1) is constructed to reflect the information change of the target when moving uniformly variably in the image.

[0105]

[0106] Among them, in the above formula, x′ and y′ represent the coordinates of the center point of the detection box of the target in the current frame image on the x-axis and y-axis of the image, x and y represent the coordinates of the center point of the detection box of the target before time t on the x-axis and y-axis of the image, represents the velocity of the same target in the x-axis and y-axis directions of the image before time t, represents the acceleration of the same target in the x-axis and y-axis directions of the image before time t; α′ and h′ represent the aspect ratio and height of the trajectory of the target in the current frame image, α and h represent the aspect ratio and height of the trajectory of the target before time t, represents the aspect ratio change rate and height change rate of the same target before time t.

[0107] Among them, the above-mentioned t represents the moment of tracking failure. Then, the center point coordinates of the detection box before t are based on the last image that has been successfully tracked in the collected data. The speed and acceleration refer to the average speed and average acceleration during a certain period before the moment t. Then, based on the above state space equation (1), the detection box of the target in each image (each frame of data collected by each sensor) can be predicted to perform real-time position tracking of each target. Since the state space equation (1) is coupled with time and the acceleration is considered, the influence of acceleration can be taken into account during the trajectory prediction stage, which can make the trajectory prediction error smaller and improve the tracking effect for large speed changes.

[0108] Based on the above embodiments, the embodiments of the present application further provide a target tracking method, which involves the specific process of a computer device tracking each candidate detection object according to the measured feature information of each candidate detection object in the current frame and the predicted detection boxes of each target to be tracked in the current frame. This embodiment takes the measured feature information including the measured three-dimensional detection box and the measured trajectory feature as an example, and predicts the predicted three-dimensional detection box of each target to be tracked in the current frame according to the historical three-dimensional detection boxes of each target to be tracked in the previous frame, as Figure 3 shown, the above step S104 includes:

[0109] S201, obtain the intersection over union (IoU) between the measured three-dimensional detection boxes of each candidate detection object in the current frame and the predicted three-dimensional detection boxes of each target to be tracked in the current frame.

[0110] This embodiment illustrates the matching process between the measured three-dimensional detection boxes of each candidate detection object in the current frame and the predicted three-dimensional detection boxes of each target to be tracked in the current frame. Calculating the IoU between the two can essentially be regarded as matching the similarity between the measured three-dimensional detection boxes of each candidate detection object in the current frame and the predicted three-dimensional detection boxes of each target to be tracked in the current frame through the IoU.

[0111] For example, if the measured three-dimensional detection box of each candidate detection object in the current frame is A1 and the predicted three-dimensional detection box of each target to be tracked in the current frame is B1, then the IoU of A1 and B1 is calculated to determine the similarity between A1 and B1. Among them, the IoU is the ratio of the intersection area and the union area of A1 and B1.

[0112] After determining the IoU, an IoU higher than the preset IoU threshold indicates that the two are very similar, and it can be determined that the two match successfully. An IoU lower than the preset IoU threshold indicates that the two are quite different and the two match fails.

[0113] S202. Obtain the similarity between the measured trajectory features of the remaining detection objects and the historical trajectory features of each target to be tracked in the previous frame. The remaining detection objects are candidate detection objects with an intersection over union (IoU) less than a preset IoU threshold.

[0114] After calculating the IoU between the measured 3D detection boxes of each candidate detection object and the predicted 3D detection boxes of each target to be tracked, the candidate detection objects corresponding to the IoU less than the preset IoU threshold are called the remaining detection objects.

[0115] Each target to be tracked is tracked frame by frame in consecutive frames of the video. When tracking each target in the current frame, the trajectory features of each target that has been tracked in the previous frame are known. Then, the trajectory features of each target in the previous frame are called historical trajectory features. For the remaining detection objects, it is determined whether they match by calculating the similarity between the measured trajectory features of the remaining detection objects in the current frame and the historical trajectory features of each target to be tracked in the previous frame. For example, obtain the cosine similarity between the two, or it can also be through distance metrics, such as calculating the Euclidean distance, etc. This embodiment does not limit this. For example, the trajectory features can be Histogram of Oriented Gradient (HOG) features, or other features, and this embodiment also does not limit this.

[0116] After determining the similarity, a similarity higher than the preset similarity threshold indicates that the two are very similar, and it can be determined that the two match successfully. While a similarity lower than the preset similarity threshold indicates that the two are quite different and the two match fails.

[0117] S203. Track each candidate detection object with an IoU greater than the IoU threshold and each remaining detection object with a similarity greater than the preset similarity threshold.

[0118] Both the above IoU greater than the preset IoU threshold and the similarity greater than the preset similarity threshold indicate successful matching. For all candidate detection objects with successful IoU matching, it can be determined that the identifier of the target to be tracked corresponding to the candidate detection object is the identifier of the known target to be tracked corresponding to the predicted 3D detection box that matches it successfully; while for the similarity greater than the preset similarity threshold, it can be determined that the identifier of the target to be tracked corresponding to the candidate detection object is the identifier of the known target to be tracked corresponding to the historical trajectory feature that matches it successfully. After determining the identifiers of each candidate detection object with successful matching, associate each candidate detection object with the corresponding identifier. After association, it is possible to determine which target to be tracked each candidate detection object is based on the associated identifier, thus completing the marking and tracking of the candidate detection objects with successful matching.

[0119] Of course, in practical applications, it is also possible to determine the candidate detection objects that match successfully only by using the intersection over union method, or only by using the similarity method to determine the candidate detection objects that match successfully. The embodiments of the present application do not limit this.

[0120] Optionally, after marking and tracking the candidate detection objects that match successfully, the detection frame and trajectory information of the target to be tracked to which each candidate detection object belongs in the current frame data can be updated according to the measured three-dimensional detection frame and the measured trajectory features of each candidate detection object. It can be understood that the measured three-dimensional detection frame and the measured trajectory features of the target to be tracked in the updated current frame data can be used as the known historical three-dimensional detection frame and historical trajectory features in the previous frame data of the next frame data.

[0121] The object tracking method provided in this embodiment first matches the measured three-dimensional detection frame of each candidate detection object in the current frame with the predicted three-dimensional detection frame of the target to be tracked in the current frame based on the intersection over union, and then performs a similarity match between the measured trajectory features of the candidate detection objects that do not match successfully and the historical trajectory features of the target to be tracked in the previous frame. The entire process uses a mixed match of different matching methods, that is, matching from different dimensions, which can effectively match each candidate detection object in the current frame to the corresponding target to be tracked, avoiding the loss of tracking caused by the target not being successfully matched when the target is occluded as a detection object, so that each target to be tracked can be stably tracked during the tracking process. And because the three-dimensional detection frame can more accurately reflect the physical shape of the target, in this embodiment, the intersection over union match is performed with the three-dimensional detection frame of the detection object, which improves the accuracy of target matching and further increases the stability of target tracking.

[0122] It was previously mentioned that the data of any frame (any moment) acquired by the first sensor and the second sensor needs to be a valid frame, and a valid frame refers to the time synchronization between the first current frame data and the second current frame data. Therefore, it is necessary to judge whether the time between the first current frame data and the second current frame data in each frame is synchronized. Based on this, an embodiment is provided for illustration, as Figure 4 shown, this embodiment includes:

[0123] S301, for the data collected by the sensor of any frame, obtain the timestamp T1 of the first current frame data and the timestamp T2 of the second current frame data.

[0124] Specifically, it can be determined by obtaining the timestamp T1 of the first current frame data and the timestamp T2 of the second current frame data. In practical applications, since there may or may not be timestamps on the data collected by the first sensor and the second sensor, for the case where the first current frame data and the second current frame data carry timestamps, for example, both the first current frame data and the second current frame data carry the time of the data collected by the GPS module of the sensor. In this case, the computer device can directly obtain the timestamp T1 of the first current frame data and the timestamp T2 of the second current frame data.

[0125] Optionally, for the case where the first current frame data and the second current frame data do not carry timestamps, after converting the acquisition times of the first current frame data and the second current frame data to the same time axis, the timestamp T1 of the first current frame data and the timestamp T2 of the second current frame data are obtained respectively.

[0126] Specifically, if there is no GPS module in the sensor and the accurate time given by the GPS module cannot be obtained, then the system time of each sensor itself is converted into the system time of the computer device (i.e., the processing platform). Under the system time of the processing platform, the timestamps of the first current frame data and the second current frame data can be accurately judged. For example, at the same moment, the system time of the processing platform is 4:00, the system time of the first sensor is 4:03, and the system time of the second sensor is 4:02. After conversion, the time of the first current frame data collected by the first sensor at the current moment is converted to 4:00 (timestamp T1), and the time of the second current frame data collected by the second sensor is converted to 4:00 (timestamp T2), so as to obtain the timestamp T1 of the first current frame data and the timestamp T2 of the second current frame data.

[0127] S302, if the interval between T1 and T2 is less than the preset interval threshold, it is determined that the first current frame data and the second current frame data are the data collected in the current frame; if the interval between T1 and T2 is greater than the preset threshold, the first current frame data and the second current frame data are discarded, and the timestamp T1 of the next frame of the first current frame data and the timestamp T2 of the second current frame data are obtained again.

[0128] After obtaining T1 and T2, determine the relationship between the interval between T1 and T2 (the time interval is the absolute value |T1 - T2|) and a preset interval threshold (for example, 10 ms). If the interval between T1 and T2 is less than the preset interval threshold, then it can be determined that the first current frame data and the second current frame data are time-synchronized and are valid frames, and the first current frame data and the second current frame data can be directly determined as the data at the current moment (current frame). However, if the interval between T1 and T2 is greater than the preset threshold, it means that the first current frame data and the second current frame data are not time-synchronized and are invalid frames, then discard the first current frame data and the second current frame data, and re-obtain the time stamps T1 of the next frame of the first current frame data and the time stamp T2 of the second current frame data, and perform the judgment process. For example, find the next frame of data at a certain frame rate (such as 10 Hz).

[0129] Optionally, if the interval between T1 and T2 is less than the preset interval threshold, then discard the smaller value of T1 and T2 and retain the larger value. Then obtain the time stamp T in the next frame of data collected by the sensor with the smaller value, and compare the new T with the retained larger value to determine whether the interval between the two is less than the preset interval threshold. If so, retain the new T value and the larger value. Otherwise, continue to compare the next new T with the retained larger value until the interval between the latest T and the retained larger value is less than the preset interval threshold. For example, the time stamp of the data collected by the lidar is T1, the time stamp of the data collected by the camera device is T2, and the preset interval threshold is 10 Hz; initially T1 > T2, then discard T2, and then compare the new T2 collected by the camera device with T1. If the interval between the new T2 and T1 < 10 Hz, then retain the new T2 and T1; but if the interval between the new T2 and T1 > 10 Hz, then continue to check the interval between the latest T2 obtained by the camera device and T1.

[0130] In this embodiment, by using the time stamps of the data collected by each sensor, it is determined whether the data collected by each sensor is synchronized. The unsynchronized data is discarded, and only the synchronized data is retained, making the data more accurate when performing target tracking.

[0131] In addition, before obtaining the first current frame data of the target scene collected by the first sensor and the second current frame data of the target scene collected by at least one second sensor as described above, some preprocessing work can also be performed to further ensure the synchronization of the data collected by each sensor. Optionally, in one embodiment, the method further includes: adjusting the sampling frequencies of the first sensor and the second sensor to be the same, and calibrating the external parameter information between the first sensor and the second sensor.

[0132] The preprocessing preparation work includes adjusting the sampling frequencies of each sensor and calibrating the external parameter information of each sensor.

[0133] Among them, adjusting the sampling frequencies of the sensors can be to make the sampling frequencies of the sensors consistent when installing the sensors; it can also be to pre-embed an adjustment program in the computer device and send the adjustment program (which can carry a specified sampling frequency) to each sensor regularly to instruct each sensor to adjust its own sampling frequency.

[0134] Optionally, the first sensor is a camera device; the second sensor is a lidar; calibrating the extrinsic information between the first sensor and the second sensor includes: adjusting the relative pose information between the camera device and the lidar to the target pose information, and obtaining the calibrated extrinsic information of the camera device according to a preset calibration algorithm; calibrating the extrinsic information of the camera device according to the calibrated extrinsic information.

[0135] The extrinsic information includes: pose information (relative position and relative angle) and the extrinsic information of the camera device

[0136] In practical applications, in order for the lidar and the camera device to comprehensively and effectively obtain the surrounding environment information, when installing the lidar and the camera device, there should be a suitable relative position and relative angle between them. For example, refer to Figure 1a the installation angles of the lidar and the camera device shown, so it is necessary to ensure the relative position and relative angle between the lidar and the camera device to ensure that the relative position and relative angle between the lidar and the camera device can comprehensively and effectively obtain the surrounding environment information.

[0137] For example, obtain the target pose information, which includes a preset target relative position and relative angle, then the computer device instructs the camera device and the lidar to adjust their own poses according to the target relative position and target relative angle.

[0138] Based on the lidar and the camera device with adjusted pose information, the computer device calibrates the extrinsic information of the camera device. For example, the extrinsic information of the camera device can be calibrated according to the mapping relationship between the point cloud and the image. Among them, the mapping relationship between the point cloud and the image represents the relationship between the world coordinate system (the coordinate system used by the lidar) and the pixel coordinate system (the coordinate system used by the pixels in the image of the camera device): Among them, in this mapping relationship, is the internal parameter matrix of the camera device, is the extrinsic parameter matrix of the camera device, is the coordinate matrix of each point in the world coordinate system, is the coordinate matrix of the points in the pixel coordinate system.

[0139] Then, it is required to obtain the coordinate matrix of each point in the world coordinate system and the coordinate matrix of the point in the pixel coordinate system according to the point cloud data collected by the actual lidar and the pixel data collected by the camera device. The internal parameter information of the camera device can be directly obtained, and then the external parameter matrix of the camera device can be obtained from this mapping relationship. Then, using the obtained external parameter matrix as the new external parameter information of the camera device, the calibration of the camera device is completed.

[0140] In this embodiment, the external parameter information of the camera device and the lidar is calibrated through the preprocessing preparation work. In this way, the camera device and the lidar can comprehensively and effectively obtain the surrounding environment information, so that the data of the collected target scene is more accurate. In addition to calibrating the external parameter information of the camera device as described above, the external parameter information (relative position and relative angle) between the camera device and the lidar can also be calibrated. In this way, by calibrating the external parameter information between the camera device and the lidar, it can be ensured that when the data collected by the camera device and the lidar are spatially transformed, the spatial unity is maintained and the accuracy of the data conversion is improved.

[0141] The process of "determining the measured feature information of each candidate detection object in the current frame according to the first current frame data and the second current frame data" in the above S102 step is described through the following embodiment. This embodiment still takes the measured feature information including the measured three-dimensional detection frame and the measured trajectory feature as an example for illustration, as Figure 5 shown, in one embodiment, the S102 step includes:

[0142] S401, obtain the two-dimensional detection frame of each detection object in the pixel data and the measured three-dimensional detection frame of each detection object in the point cloud data.

[0143] This embodiment takes the first current frame data as the pixel data collected by the camera device and the second current frame data as the point cloud data collected by the lidar as an example for illustration.

[0144] Obtain the two-dimensional detection frame of each detection object from the pixel data collected by the camera device. For example, through a preset deep learning algorithm model, such as the YOLOv3 model, etc., output information such as the pixel rectangle frame positioning, target confidence, and target classification of each detection object; among them, the rectangle frame positioning can reflect the two-dimensional detection frame of each detection object, and the target confidence can reflect the accuracy of each detection object relative to the target to be tracked; the target classification information reflects the category of each detection object. For example, the target is a vehicle, a person, or an animal, etc.

[0145] Obtain the three-dimensional detection boxes (measured three-dimensional detection boxes) of each detection object from the point cloud data. For example, through a preset deep learning algorithm model, such as the SECOND model, etc., output information such as the positioning, size, and heading angle of each detection object in the point cloud coordinate system. Among them, the positioning, size, heading angle, etc. of each detection object in the point cloud coordinate system can reflect the measured three-dimensional detection box of each detection object. It should be noted that the measured three-dimensional detection boxes and two-dimensional detection boxes in all embodiments of this application refer to the actually measured three-dimensional or two-dimensional detection boxes of the detection objects. However, the three-dimensional detection boxes are called measured three-dimensional detection boxes to maintain unity with the previous description, and the previous ones are called measured three-dimensional detection boxes to distinguish them from the predicted three-dimensional detection boxes.

[0146] S402. Match the two-dimensional detection boxes of each detection object in the pixel data with the measured three-dimensional detection boxes of each detection object in the point cloud data to determine the candidate detection objects.

[0147] After obtaining the two-dimensional detection boxes of each detection object in the pixel data and the measured three-dimensional detection boxes of each detection object in the point cloud data, first perform associative matching on the detection objects in the pixel data and the detection objects in the point cloud data, and determine the detection objects with successful associative matching as the candidate detection objects.

[0148] Optionally, as Figure 6 shown, an embodiment of determining the candidate detection objects includes:

[0149] S501. Map each measured three-dimensional detection box to a corresponding two-dimensional mapped detection box.

[0150] Map the measured three-dimensional detection box of each detection object to a corresponding two-dimensional mapped detection box. For example, each measured three-dimensional detection box in the point cloud data can be used as input data and input into a pre-trained conversion network model, and the output result is the two-dimensional image corresponding to each measured three-dimensional detection box in the point cloud data. Another example is to first convert the three-dimensional coordinates of each point cloud point in each measured three-dimensional detection box into two-dimensional coordinates based on the mapping relationship between the preset three-dimensional coordinate system and the two-dimensional coordinate system, then obtain the two-dimensional coordinates of each point cloud point in each measured three-dimensional detection box, and then determine the pixel value corresponding to the point cloud point at each two-dimensional coordinate. After determining the pixel values at each two-dimensional coordinate position, the two-dimensional image corresponding to each measured three-dimensional detection box is obtained. In this way, the pixel values of each point at each two-dimensional coordinate position in the obtained two-dimensional detection box are also different, and the characteristics of the three-dimensional point cloud points are also retained.

[0151] S502. Obtain the intersection over union (IoU) between each two-dimensional detection box and each two-dimensional mapped detection box, and determine the detection objects with an IoU greater than the IoU threshold as the candidate detection objects.

[0152] After obtaining the two-dimensional mapped detection frames corresponding to the measured three-dimensional detection frames of each detection object, calculate the intersection over union (IoU) between each two-dimensional mapped detection frame and the two-dimensional detection frame of each detection object in the pixel data, and then determine the candidate detection objects as those with an IoU greater than a preset IoU threshold. It can be understood that the candidate detection object is essentially a pair of detection frames. It should be noted that the preset IoU threshold here can be the same as or different from the preset IoU threshold for matching in the previous embodiments, and this embodiment does not limit this.

[0153] By associating and matching the two-dimensional detection frames of each detection object in the pixel data with the measured three-dimensional detection frames of each detection object in the point cloud data, candidate detection objects are screened out, ensuring the consistency of the detection objects in the data collected by various types of sensors.

[0154] S403. Determine the measured trajectory features of each candidate detection object based on the two-dimensional detection frame and feature information of each candidate detection object in the pixel data, and the measured three-dimensional detection frame and feature information of each detection object in the point cloud data.

[0155] After determining each candidate detection object, further determine the measured trajectory features of each candidate detection object. For example, based on the two-dimensional detection frame of each candidate detection object, extract the feature information in this two-dimensional detection frame; for example, extract the Histogram of Oriented Gradient (HOG) feature of the two-dimensional detection frame as the feature of the corresponding two-dimensional detection frame; based on the measured three-dimensional detection frame of each candidate detection object, extract the feature information in this measured three-dimensional detection frame; then comprehensively determine the measured trajectory features of each candidate detection object according to the feature information in the two-dimensional detection frame and the feature information in the measured three-dimensional detection frame. Among them, the measured three-dimensional detection frame of each candidate detection object is the three-dimensional detection frame of each detection object in the point cloud data.

[0156] In this embodiment, first determine the candidate detection objects based on the two-dimensional detection frames of each detection object in the pixel data and the measured three-dimensional detection frames of each detection object in the point cloud data, and then determine the measured trajectory features of each candidate detection object according to the feature information in the two-dimensional detection frame and the feature information in the measured three-dimensional detection frame of each candidate detection object. When determining the measured trajectory features of each detection object, it is determined based on the data collected by more than two different types of sensors, and the data collected by multiple types of sensors is comprehensively used to determine the measured trajectory features of the detection objects in each frame, which can accurately and completely reflect the features of the detection objects in each frame. In this way, when matching the measured trajectory features of each candidate detection object with the historical trajectory features of each target to be tracked, the target to which each detection object belongs can be accurately determined, effectively completing the target tracking of each frame.

[0157] Provide an implementable way to determine the measured trajectory features of each candidate detection object in the above step S403, such as Figure 7 As shown, this embodiment includes:

[0158] S601, comprehensively determine the two-dimensional detection boxes and feature information of each candidate detection object in the pixel data as the two-dimensional trajectory features of each candidate detection object; comprehensively determine the measured three-dimensional detection boxes and feature information of each candidate detection object in the point cloud data as the three-dimensional trajectory features of each candidate detection object.

[0159] Comprehensively determine the feature information extracted from the two-dimensional detection boxes of each candidate detection object in the pixel data and the two-dimensional detection boxes as the two-dimensional trajectory features of each candidate detection object. That is, the two-dimensional trajectory features can be regarded as the two-dimensional detection boxes of each candidate detection object and carry the feature information of the pixel points in the two-dimensional detection boxes.

[0160] Comprehensively determine the measured three-dimensional detection boxes of each candidate detection object in the point cloud data and the feature information extracted from the measured three-dimensional detection boxes as the three-dimensional trajectory features of each candidate detection object. Similarly, the three-dimensional trajectory features can be regarded as the measured three-dimensional detection boxes of each candidate detection object and carry the feature information of the point cloud points in the measured three-dimensional detection boxes.

[0161] S602, convert the three-dimensional trajectory features of each candidate detection object into corresponding two-dimensional mapped trajectory features.

[0162] Since the three-dimensional trajectory features are the measured three-dimensional detection boxes in terms of the display form, when converting the three-dimensional trajectory features of each candidate detection object into the corresponding two-dimensional mapped trajectory features, the conversion can be performed according to the mapping relationship between the points in the point cloud and the image. Among them, after converting the three-dimensional coordinates of each point cloud point of the measured three-dimensional detection box into two-dimensional coordinates based on the mapping relationship, the two-dimensional coordinates of each point cloud point in the measured three-dimensional detection box are obtained. Then, determine the pixel value corresponding to the point cloud point at each two-dimensional coordinate, and after determining the pixel value at each two-dimensional coordinate position, the two-dimensional detection box corresponding to the measured three-dimensional detection box is obtained. In this way, the pixel values of each point at the two-dimensional coordinate positions in the obtained two-dimensional detection box are also different, and the features of the point cloud points in the measured three-dimensional detection box are also retained. In this way, the converted two-dimensional mapped trajectory features include the detection box and the feature information of the points in the detection box in the three-dimensional trajectory features.

[0163] S603, extract the measured trajectory features of each candidate detection object according to the fusion data of each two-dimensional mapped trajectory feature and each two-dimensional trajectory feature. Optionally, compress each two-dimensional trajectory feature and each two-dimensional mapped trajectory feature to the same proportional size, and then splice them to obtain a fusion data matrix; determine the feature information extracted from the fusion data matrix as the measured trajectory features of each candidate detection object.

[0164] Each two-dimensional mapping trajectory feature is the feature information of each candidate detection object determined from the point cloud data, and each two-dimensional trajectory feature is the feature information of each candidate detection object determined from the pixel data. The feature information refined from these two fused data is the measured trajectory feature of each candidate detection object.

[0165] Still taking the above example for illustration, the two-dimensional trajectory feature information and the two-dimensional mapping trajectory feature of each candidate detection object are adjusted to the same proportion. For example, the two-dimensional trajectory feature information of each candidate detection object is compressed to the size of N*N*3, and the two-dimensional mapping trajectory feature is scaled to the same size proportion. After the proportion adjustment, the two are spliced to obtain a fused data matrix. Then, the fused data matrix is input into a convolutional neural network for feature extraction, and the extracted feature is the measured trajectory feature of each candidate detection object.

[0166] In this embodiment, by converting the three-dimensional trajectory feature of each candidate detection object in the point cloud data into a two-dimensional mapping trajectory feature, and then splicing the two-dimensional mapping trajectory feature and the two-dimensional trajectory feature of each candidate detection object in the pixel data to form a fused data matrix, the feature information extracted from the fused data matrix is determined as the measured trajectory feature of each candidate detection object. Determining the measured trajectory feature of the detection object in each frame from multi-dimensional types of data can accurately and completely reflect the features of the detection object in each frame. In this way, when matching the measured trajectory feature of each candidate detection object with the accurate data of each target to be tracked, the target to which each detection object belongs can be accurately determined, and the target tracking of each frame can be effectively completed.

[0167] Next, a specific embodiment is used to illustrate the process of converting the three-dimensional trajectory feature of each candidate detection object into the corresponding two-dimensional mapping trajectory feature in S602 above, as Figure 8 shown, S602 includes:

[0168] S701, converting the three-dimensional coordinates of the point cloud points in the measured three-dimensional detection box corresponding to each three-dimensional trajectory feature into two-dimensional coordinates.

[0169] Among them, the conversion of three-dimensional coordinates to two-dimensional coordinates can be based on the mapping relationship between the three-dimensional coordinate system and the two-dimensional coordinate system. For example, the mapping relationship between the three-dimensional coordinate system and the two-dimensional coordinate system can be shown by the following formula (1):

[0170]

[0171] In the above formula (1), a and b represent the coordinates of the point cloud points in the three-dimensional coordinate system, a t 、b tIndicates the coordinates of the point cloud points after mapping in the two-dimensional coordinate system. h represents the distance from the boundary of the point cloud to the y-axis, and w represents the distance from the boundary of the point cloud to the x-axis. Here, the x-axis and y-axis are the coordinate axes with the lidar as the origin.

[0172] Based on the mapping relationship of formula (1), convert the three-dimensional coordinates of each point cloud point in the three-dimensional detection box into two-dimensional coordinates.

[0173] S702. Obtain the bird's-eye view corresponding to each measured three-dimensional detection box according to the two-dimensional coordinates of each point cloud point and the z-axis coordinate of each point cloud point in the three-dimensional coordinates; obtain the intensity map corresponding to each measured three-dimensional detection box according to the two-dimensional coordinates of each point cloud point and the intensity of each point cloud point; obtain the density map corresponding to each measured three-dimensional detection box according to the density of each point cloud point in the z-axis direction in the bird's-eye view.

[0174] In this step, the bird's-eye view, intensity map, and density map corresponding to each measured three-dimensional detection box are obtained respectively. These three maps are all two-dimensional.

[0175] Among them, obtaining the bird's-eye view corresponding to the measured three-dimensional detection box is determined according to the two-dimensional coordinates of each point cloud point and the z-axis coordinate of each point cloud point in the three-dimensional coordinates. Among them, the two-dimensional coordinates of the point cloud points are obtained by conversion in the above steps, and the two-dimensional coordinates of each point cloud point determine the position of the point cloud point on the two-dimensional plane. Among them, the z-axis coordinate of the point cloud point in the three-dimensional coordinates is the Z coordinate of the point cloud point in the three-dimensional coordinates where the point cloud data is located. The Z coordinate can also be regarded as the height of the point cloud point in the two-dimensional coordinates. Based on the measured three-dimensional detection box, the position of each point cloud point on the two-dimensional plane and the height of the point cloud point in the two-dimensional coordinates, the bird's-eye view corresponding to each measured three-dimensional detection box can be obtained.

[0176] Optionally, as Figure 9 shown, the process of obtaining the bird's-eye view corresponding to each measured three-dimensional detection box includes:

[0177] S801. Perform normalization processing on the z-axis coordinates of each point cloud point in the three-dimensional coordinates, and determine the normalized z-axis coordinates as the pixel values of each point cloud point.

[0178] First, determine the pixel value of each point cloud point in the bird's-eye view according to the z-axis coordinate of the point cloud point in the three-dimensional coordinates. Since the range of pixel values is 0 - 255, it is necessary to first perform normalization processing on the z-axis coordinate of the point cloud point in the three-dimensional coordinates and normalize it to the range of 0 - 255. The normalized value is the pixel value of each point cloud point in the bird's-eye view.

[0179] S802. Use the pixel values corresponding to each point cloud point as the pixel values at the corresponding two-dimensional coordinate positions to obtain the bird's-eye view corresponding to each measured three-dimensional detection box.

[0180] After obtaining the pixel values of each point cloud point in the bird's-eye view, the coordinate position of each point cloud point can be determined by combining the two-dimensional coordinates of each point cloud point, and then the pixel value of the corresponding point cloud point is filled into the two-dimensional coordinate position to obtain the final bird's-eye view.

[0181] Among them, obtaining the intensity map is determined according to the two-dimensional coordinates of each point cloud point and the intensity of each point cloud point. Among them, the intensity of the point cloud point is obtained from the intensity of each point cloud when the lidar collects point cloud data. Based on the position of each point cloud point in the two-dimensional plane and the intensity of the point cloud point in the measured three-dimensional detection frame, the intensity map corresponding to each measured three-dimensional detection frame can be obtained.

[0182] Optionally, as Figure 10 shown, the process of obtaining the intensity map corresponding to each measured three-dimensional detection frame includes:

[0183] S901, perform normalization processing on the intensity of each point cloud point, and determine the normalized intensity as the pixel value of each point cloud point.

[0184] First, determine the pixel value of each point cloud point in the bird's-eye view according to the intensity of the point cloud point. Similarly, the range of the pixel value is 0-255. Therefore, it is necessary to perform normalization processing on the intensity of the point cloud point first, normalize it to between 0 and 255, and the normalized value is the pixel value of each point cloud point in the intensity map.

[0185] S902, use the pixel value corresponding to each point cloud point as the pixel value of the corresponding two-dimensional coordinate position to obtain the intensity map corresponding to each measured three-dimensional detection frame.

[0186] After obtaining the pixel values of each point cloud point in the intensity map, the coordinate position of each point cloud point can be determined by combining the two-dimensional coordinates of each point cloud point, and then the pixel value of the corresponding point cloud point is filled into the two-dimensional coordinate position to obtain the final intensity map.

[0187] Among them, obtaining the density map is determined according to the density of each point cloud point in the z-axis direction in the bird's-eye view. Among them, the density of the point cloud point in the z-axis direction refers to the density of the number of point cloud points in the z-axis direction at the position determined by each x-axis and y-axis when the point cloud is in three-dimensional coordinates.

[0188] Optionally, as Figure 11 shown, the process of obtaining the density map corresponding to each measured three-dimensional detection frame includes:

[0189] S1001, determine the density of the point cloud point in the z-axis direction at each position according to the number of point cloud points in the z-axis direction at the two-dimensional coordinate position of each point cloud point, the maximum value of the number of point cloud points at all coordinate positions, and the minimum value of the number of point cloud points at all coordinate positions.

[0190] If the two-dimensional coordinate position of a pixel point is determined by the x coordinate and the y coordinate, then each point cloud point in the three-dimensional detection frame corresponds to a two-dimensional coordinate position. Obtain the number of point cloud points in the z-axis direction for each two-dimensional coordinate position. Then, for one point cloud point, based on the number of point cloud points in the z-axis direction at the two-dimensional coordinate position of this point cloud point, the maximum value of the number of point cloud points at all coordinate positions, and the minimum value of the number of point cloud points at all coordinate positions, the density of this point cloud point in the z-axis direction can be determined.

[0191] For example, the density of a point cloud point in the z-axis direction can be determined by the following formula (2).

[0192]

[0193] Where ρ i is the density of the i-th pixel point (i.e., the two-dimensional coordinate position), c i is the number of point clouds of the i-th pixel point, c min is the minimum value of the number of point clouds among all pixel points, c max represents the maximum value of the number of point clouds among all pixel points.

[0194] S1002. Take the density of the point cloud points at each coordinate position in the z-axis direction as the pixel value of the corresponding two-dimensional coordinate position, and obtain the density map corresponding to each measured three-dimensional detection frame.

[0195] After obtaining the density of each point cloud point, take this density as the pixel point of this point cloud point. Combining the two-dimensional coordinates of each point cloud point can determine the coordinate position of each point cloud point. Then, fill the pixel value of the corresponding point cloud point into the two-dimensional coordinate position to obtain the final density map.

[0196] The process of obtaining the bird's-eye view, intensity map, and density map is as above. The following is the process of merging and processing the bird's-eye view, intensity map, and density map.

[0197] S705. Merge and process the bird's-eye view, intensity map, and density map to obtain the two-dimensional mapped trajectory features corresponding to each three-dimensional trajectory feature.

[0198] After obtaining the bird's-eye view, intensity map, and density map corresponding to each measured three-dimensional detection frame in the point cloud data, merge these three two-dimensional images. Optionally, merge the bird's-eye view, intensity map, and density map as images of the R, G, and B three channels respectively for processing, and the two-dimensional mapped detection frame corresponding to each measured three-dimensional detection frame can be obtained. The characteristic information of each point cloud point in the measured three-dimensional detection frame is reflected on each pixel point in the two-dimensional mapped detection frame. For each three-dimensional trajectory feature corresponding to each measured three-dimensional detection frame, and the two-dimensional mapped trajectory feature corresponding to each two-dimensional mapped detection frame, that is, the two-dimensional mapped trajectory features corresponding to each three-dimensional trajectory feature are obtained.

[0199] In this embodiment, since the bird's-eye view, intensity map, and density map corresponding to each measured three-dimensional detection box are obtained first, the pixel value of each point in these three maps is converted based on the actual data of the point cloud data, that is, all the information that the point cloud points should have in the point cloud is retained. When these three maps are merged, the pixel values of each point in the two-dimensional mapping trajectory feature can also accurately reflect the information of the point cloud points in the three-dimensional point cloud, so that the two-dimensional mapping trajectory feature can more accurately reflect each three-dimensional trajectory feature.

[0200] In one embodiment, as Figure 12 shown, an embodiment of a target tracking method is provided, and this embodiment includes:

[0201] S1101, adjust the sampling frequencies of the first sensor and the second sensor to be the same, and calibrate the extrinsic parameter information between the first sensor and the second sensor;

[0202] S1102, for the data collected by the sensor in any frame, obtain the timestamp T1 of the first current frame data and the timestamp T2 of the second current frame data;

[0203] S1103, if the interval between T1 and T2 is less than the preset interval threshold, determine the first current frame data and the second current frame data as the data collected in the current frame; if the interval between T1 and T2 is greater than the preset threshold, discard the first current frame data and the second current frame data, and re-obtain the timestamp T1 of the first current frame data and the timestamp T2 of the second current frame data for the next frame;

[0204] S1104, obtain the two-dimensional detection boxes of each detection object in the pixel data and the measured three-dimensional detection boxes of each detection object in the point cloud data;

[0205] S1105, determine candidate detection objects according to the two-dimensional detection boxes of each detection object in the pixel data and the measured three-dimensional detection boxes of each detection object in the point cloud data;

[0206] S1106, comprehensively determine the two-dimensional trajectory features of each candidate detection object based on the two-dimensional detection boxes and feature information of each candidate detection object in the pixel data; comprehensively determine the three-dimensional trajectory features of each candidate detection object based on the measured three-dimensional detection boxes and feature information of each candidate detection object in the point cloud data;

[0207] S1107, convert the three-dimensional coordinates of the point cloud points in the measured three-dimensional detection boxes corresponding to each three-dimensional trajectory feature into two-dimensional coordinates;

[0208] S1108, obtain the bird's-eye view corresponding to each measured three-dimensional detection box; obtain the density map corresponding to each measured three-dimensional detection box; obtain the intensity map corresponding to each measured three-dimensional detection box; perform a merging process on the bird's-eye view, intensity map, and density map to obtain the two-dimensional mapped trajectory features corresponding to each three-dimensional trajectory feature;

[0209] S1109, extract the measured trajectory features of each candidate detection object according to the fusion data of each two-dimensional mapped trajectory feature and each two-dimensional trajectory feature;

[0210] S1110, through a preset tracking algorithm model, predict the predicted three-dimensional detection box of each target to be tracked in the current frame according to the historical three-dimensional detection box of each target to be tracked in the previous frame;

[0211] S1111, obtain the intersection over union (IoU) between the measured three-dimensional detection box of each candidate detection object in the current frame and the predicted three-dimensional detection box of each target to be tracked in the current frame;

[0212] S1112, obtain the similarity between the measured trajectory features of the remaining detection objects and the historical trajectory features of each target to be tracked in the previous frame; the remaining detection objects are candidate detection objects with an IoU less than a preset IoU threshold;

[0213] S1113, perform tracking on each candidate detection object with an IoU greater than the IoU threshold and each remaining detection object with a similarity greater than a preset similarity threshold.

[0214] For each step in the object tracking method provided in this embodiment, its implementation principle and technical effects are similar to those in the previous object tracking method embodiments, and will not be elaborated here. Figure 12 The implementation manners of the steps in the embodiment are only examples, and there is no limitation on each implementation manner. The order of each step can be adjusted in actual applications as long as the purpose of each step can be achieved.

[0215] It should be understood that although Figures 2 - 12 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order restriction, and these steps can be executed in other orders. Moreover, Figures 2 - 12 at least a part of the steps in

[0216] In one embodiment, as Figure 8As shown in the figure, a target tracking device is provided, including: an acquisition module 10, a determination module 11, a prediction module 12, and a tracking module 13, where,

[0217] The acquisition module 10 is configured to acquire first current frame data of a target scene collected by a first sensor and second current frame data of the target scene collected by at least one second sensor; the first sensor and the second sensor are sensors of different types;

[0218] The feature acquisition module 11 is configured to acquire measured feature information of each candidate detection object in the current frame according to the first current frame data and the second current frame data; the candidate detection object is a detection object obtained by successfully associating and matching each detection object in the first current frame data with each detection object in the second current frame data; the measured feature information represents the inherent characteristic information of each candidate detection object;

[0219] The prediction module 12 is configured to predict a predicted detection box of each target to be tracked in the current frame according to the historical detection boxes of each target to be tracked in the previous frame;

[0220] The tracking module 13 is configured to track each candidate detection object according to the measured feature information of each candidate detection object in the current frame and the predicted detection boxes of each target to be tracked in the current frame.

[0221] In one embodiment, the above-mentioned tracking module 13 includes:

[0222] The first acquisition unit is configured to acquire the intersection over union between the measured three-dimensional detection box of each candidate detection object in the current frame and the predicted three-dimensional detection box of each target to be tracked in the current frame;

[0223] The second acquisition unit is configured to acquire the similarity between the measured trajectory feature of the remaining detection objects and the historical trajectory feature of each target to be tracked in the previous frame; the remaining detection objects are candidate detection objects with an intersection over union less than a preset intersection over union threshold;

[0224] The first determination unit is configured to track each candidate detection object with an intersection over union greater than the intersection over union threshold and each remaining detection object with a similarity greater than a preset similarity threshold.

[0225] In one embodiment, the above-mentioned acquisition module 10 includes:

[0226] The third acquisition unit is configured to acquire the timestamp T1 of the first current frame data and the timestamp T2 of the second current frame data for the data collected by any frame of sensor;

[0227] A second determination unit, configured to determine the first current frame data and the second current frame data as the data collected in the current frame if the interval between T1 and T2 is less than a preset interval threshold; and discard the first current frame data and the second current frame data and re-obtain the time stamps T1 of the first current frame data and the time stamp T2 of the second current frame data of the next frame if the interval between T1 and T2 is greater than the preset threshold.

[0228] In one embodiment, the above-mentioned third acquisition unit is specifically configured to, if the time stamps are not carried in the first current frame data and the second current frame data, convert the acquisition times of the first current frame data and the second current frame data into the same time axis to obtain T1 and T2.

[0229] In one embodiment, the device further includes: an adjustment module, configured to adjust the sampling frequencies of the first sensor and the second sensor to be the same, and calibrate the external parameter information between the first sensor and the second sensor.

[0230] In one embodiment, the first sensor is a camera device; the second sensor is a lidar;

[0231] Then the above-mentioned adjustment module is specifically configured to adjust the relative pose information between the camera device and the lidar to the target pose information, and obtain the calibrated external parameter information of the camera device according to a preset calibration algorithm; calibrate the external parameter information of the camera device according to the calibrated external parameter information.

[0232] In one embodiment, the above-mentioned first current frame data is pixel data collected by the camera device, and the second current frame data is point cloud data collected by the lidar; then the above-mentioned feature acquisition module 11 includes:

[0233] A detection box acquisition unit, configured to acquire the two-dimensional detection boxes of the detection objects in the pixel data and the measured three-dimensional detection boxes of the detection objects in the point cloud data;

[0234] A candidate detection object determination unit, configured to match the two-dimensional detection boxes of the detection objects in the pixel data and the measured three-dimensional detection boxes of the detection objects in the point cloud data to determine candidate detection objects;

[0235] A measured feature determination unit, configured to determine the measured trajectory features of the candidate detection objects according to the two-dimensional detection boxes and feature information of the candidate detection objects in the pixel data, and the measured three-dimensional detection boxes and feature information of the detection objects in the point cloud data.

[0236] In one embodiment, the above-mentioned candidate detection object determination unit is specifically configured to map each three-dimensional detection box to a corresponding two-dimensional mapped detection box; obtain the intersection over union between each two-dimensional detection box and each two-dimensional mapped detection box, and determine the detection objects with the intersection over union greater than the intersection over union threshold as candidate detection objects.

[0237] In one embodiment, the above-mentioned measured feature determination unit includes:

[0238] A trajectory feature determination subunit, configured to comprehensively determine the two-dimensional trajectory features of each candidate detection object based on the two-dimensional detection boxes and feature information of each candidate detection object in the pixel data; and comprehensively determine the three-dimensional trajectory features of each candidate detection object based on the measured three-dimensional detection boxes and feature information of each candidate detection object in the point cloud data;

[0239] A conversion subunit, configured to convert the three-dimensional trajectory features of each candidate detection object into corresponding two-dimensional mapped trajectory features;

[0240] A feature extraction subunit, configured to extract the measured trajectory features of each candidate detection object according to the fusion data of each two-dimensional mapped trajectory feature and each two-dimensional trajectory feature.

[0241] In one embodiment, the above-mentioned conversion subunit includes:

[0242] A coordinate conversion subunit, configured to convert the three-dimensional coordinates of the point cloud points in the measured three-dimensional detection boxes corresponding to the three-dimensional trajectory features into two-dimensional coordinates;

[0243] An aerial view subunit, configured to obtain the aerial view corresponding to each measured three-dimensional detection box according to the two-dimensional coordinates of each point cloud point and the z-axis coordinate of each point cloud point in the three-dimensional coordinates;

[0244] An intensity map subunit, configured to obtain the intensity map corresponding to each measured three-dimensional detection box according to the two-dimensional coordinates of each point cloud point and the intensity of each point cloud point;

[0245] A density map subunit, configured to obtain the density map corresponding to each measured three-dimensional detection box according to the density of each point cloud point in the z-axis direction in the aerial view;

[0246] A shooting trajectory feature determination subunit, configured to perform a merging process on the aerial view, the intensity map, and the density map to obtain the two-dimensional mapped trajectory features corresponding to the three-dimensional trajectory features.

[0247] In one embodiment, the above-mentioned aerial view subunit is specifically configured to perform normalization processing on the z-axis coordinates of each point cloud point in the three-dimensional coordinates, and determine the normalized z-axis coordinates as the pixel values of each point cloud point; and use the pixel values corresponding to each point cloud point as the pixel values at the corresponding two-dimensional coordinate positions to obtain the aerial view corresponding to each measured three-dimensional detection box.

[0248] In one embodiment, the above-mentioned intensity map subunit is specifically configured to normalize the intensity of each point cloud point, and determine the normalized intensity as the pixel value of each point cloud point; use the pixel value corresponding to each point cloud point as the pixel value at the corresponding two-dimensional coordinate position to obtain the intensity map corresponding to each measured three-dimensional detection box.

[0249] In one embodiment, the above-mentioned density map subunit is specifically configured to determine the density of the point cloud points at each position in the z-axis direction according to the number of point cloud points in the z-axis direction in the two-dimensional coordinate position of each point cloud point, the maximum value of the number of point cloud points at all coordinate positions, and the minimum value of the number of point cloud points at all coordinate positions; use the density of the point cloud points at each coordinate position in the z-axis direction as the pixel value at the corresponding two-dimensional coordinate position to obtain the density map corresponding to each measured three-dimensional detection box.

[0250] In one embodiment, the above-mentioned feature extraction subunit is specifically configured to compress each two-dimensional trajectory feature and each two-dimensional mapped trajectory feature to the same proportional size, and then splice them to obtain a fused data matrix; determine the feature information extracted from the fused data matrix as the measured trajectory feature of each candidate detection object.

[0251] In one embodiment, the above-mentioned prediction module 12 is specifically configured to, through a preset tracking algorithm model, predict the predicted three-dimensional detection box of each target to be tracked in the current frame according to the historical three-dimensional detection box of each target to be tracked in the previous frame; wherein, the tracking algorithm model is constructed based on the spatial equation of the uniformly variable motion state.

[0252] For the specific limitations of the target tracking device, reference can be made to the limitations of the target tracking method in the above text, which will not be elaborated here. Each module in the above target tracking device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0253] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 1bAs shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a target tracking method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0254] Those skilled in the art can understand that Figure 1b the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0255] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:

[0256] Obtain the first current frame data of the target scene collected by the first sensor and the second current frame data of the target scene collected by at least one second sensor; the first sensor and the second sensor are sensors of different types;

[0257] According to the first current frame data and the second current frame data, obtain the measured feature information of each candidate detection object in the current frame; the candidate detection object is a detection object that is successfully associated and matched between each detection object in the first current frame data and each detection object in the second current frame data; the measured feature information represents the inherent characteristic information of each candidate detection object;

[0258] According to the historical detection frames of each target to be tracked in the previous frame, predict the predicted detection frames of each target to be tracked in the current frame;

[0259] According to the measured feature information of each candidate detection object in the current frame and the predicted detection frames of each target to be tracked in the current frame, track each candidate detection object.

[0260] For the computer device provided in the above embodiment, its implementation principle and technical effects are similar to those of the above method embodiment, and will not be elaborated here.

[0261] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0262] Obtain first current frame data of a target scene collected by a first sensor and second current frame data of the target scene collected by at least one second sensor; the first sensor and the second sensor are sensors of different types;

[0263] According to the first current frame data and the second current frame data, obtain measured feature information of each candidate detection object in the current frame; the candidate detection objects are the detection objects that are successfully associated and matched between the detection objects in the first current frame data and the detection objects in the second current frame data; the measured feature information represents the inherent characteristic information of each candidate detection object;

[0264] According to the historical detection frames of each target to be tracked in the previous frame, predict the predicted detection frames of each target to be tracked in the current frame;

[0265] Track each candidate detection object according to the measured feature information of each candidate detection object in the current frame and the predicted detection frames of each target to be tracked in the current frame.

[0266] The principle of implementation and technical effects of the computer-readable storage medium provided in the above embodiment are similar to those of the above method embodiment, and will not be elaborated here.

[0267] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0268] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0269] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A target tracking method, characterized in that, The method includes: Obtaining first current frame data of a target scene collected by a first sensor and second current frame data of the target scene collected by at least one second sensor; the first sensor and the second sensor are sensors of different types; Obtaining measured feature information of each candidate detection object in the current frame according to the first current frame data and the second current frame data; the candidate detection objects are detection objects that are successfully associated and matched between the detection objects in the first current frame data and the detection objects in the second current frame data; the measured feature information represents the inherent characteristic information of each candidate detection object; Predicting a predicted three-dimensional detection frame of each target to be tracked in the current frame through a preset tracking algorithm model according to the historical three-dimensional detection frames of each target to be tracked in the previous frame; wherein, the tracking algorithm model is constructed based on a spatial equation of a uniformly variable motion state; Tracking each candidate detection object according to the measured feature information of each candidate detection object in the current frame and the predicted detection frames of each target to be tracked in the current frame.

2. The method according to claim 1, characterized in that The measured feature information includes a measured three-dimensional detection frame and a measured trajectory feature; The tracking each candidate detection object according to the measured feature information of each candidate detection object in the current frame and the predicted detection frames of each target to be tracked in the current frame includes: Obtaining the intersection over union between the measured three-dimensional detection frames of each candidate detection object in the current frame and the predicted three-dimensional detection frames of each target to be tracked in the current frame; Obtaining the similarity between the measured trajectory features of the remaining detection objects and the historical trajectory features of each target to be tracked in the previous frame; the remaining detection objects are candidate detection objects with an intersection over union less than a preset intersection over union threshold; Tracking each candidate detection object with an intersection over union greater than the intersection over union threshold and each remaining detection object with a similarity greater than a preset similarity threshold.

3. The method according to claim 1 or 2, characterized in that, The first current frame data is pixel data collected by a camera device, and the second current frame data is point cloud data collected by a lidar; the measured feature information includes a measured three-dimensional detection frame and a measured trajectory feature; Then the determining the measured feature information of each candidate detection object in the current frame according to the first current frame data and the second current frame data includes: Obtaining the two-dimensional detection frames of the detection objects in the pixel data and the measured three-dimensional detection frames of the detection objects in the point cloud data; Matching the two-dimensional detection frames of the detection objects in the pixel data and the measured three-dimensional detection frames of the detection objects in the point cloud data to determine the candidate detection objects; Determining the measured trajectory features of each candidate detection object according to the two-dimensional detection frames and feature information of each candidate detection object in the pixel data and the measured three-dimensional detection frames and feature information of each detection object in the point cloud data.

4. The method according to claim 3, wherein The matching the two-dimensional detection frames of the detection objects in the pixel data and the measured three-dimensional detection frames of the detection objects in the point cloud data to determine the candidate detection objects includes: Mapping each measured three-dimensional detection frame to a corresponding two-dimensional mapped detection frame; Obtain the intersection over union (IoU) between each of the two-dimensional detection boxes and each of the two-dimensional mapped detection boxes, and determine the detection objects with an IoU greater than the IoU threshold as the candidate detection objects.

5. The method according to claim 3, wherein The determining the actual trajectory features of each of the candidate detection objects according to the two-dimensional detection boxes and feature information of each of the candidate detection objects in the pixel data, and the actual three-dimensional detection boxes and feature information of each of the detection objects in the point cloud data includes: Comprehensively determine the two-dimensional trajectory features of each of the candidate detection objects based on the two-dimensional detection boxes and feature information of each of the candidate detection objects in the pixel data; comprehensively determine the three-dimensional trajectory features of each of the candidate detection objects based on the actual three-dimensional detection boxes and feature information of each of the candidate detection objects in the point cloud data; Convert the three-dimensional trajectory features of each of the candidate detection objects into corresponding two-dimensional mapped trajectory features; Extract the actual trajectory features of each of the candidate detection objects according to the fusion data of each of the two-dimensional mapped trajectory features and each of the two-dimensional trajectory features.

6. The method according to claim 5, characterized in that, The converting the three-dimensional trajectory features of each of the candidate detection objects into corresponding two-dimensional mapped trajectory features includes: Convert the three-dimensional coordinates of the point cloud points in the actual three-dimensional detection box corresponding to each of the three-dimensional trajectory features into two-dimensional coordinates; Obtain the bird's-eye view corresponding to each of the actual three-dimensional detection boxes according to the two-dimensional coordinates of each of the point cloud points and the z-axis coordinates of each of the point cloud points in the three-dimensional coordinates; Obtain the intensity map corresponding to each of the actual three-dimensional detection boxes according to the two-dimensional coordinates of each of the point cloud points and the intensity of each of the point cloud points; Obtain the density map corresponding to each of the actual three-dimensional detection boxes according to the density of each of the point cloud points in the z-axis direction in the bird's-eye view; Perform a merging process on the bird's-eye view, the intensity map, and the density map to obtain the two-dimensional mapped trajectory features corresponding to each of the three-dimensional trajectory features.

7. The method according to claim 6, characterized in that The obtaining the bird's-eye view corresponding to each of the actual three-dimensional detection boxes according to the two-dimensional coordinates of each of the point cloud points and the z-axis coordinates of each of the point cloud points in the three-dimensional coordinates includes: Perform a normalization process on the z-axis coordinates of each of the point cloud points in the three-dimensional coordinates, and determine the normalized z-axis coordinates as the pixel values of each of the point cloud points; Use the pixel values corresponding to each of the point cloud points as the pixel values at the corresponding two-dimensional coordinate positions to obtain the bird's-eye view corresponding to each of the actual three-dimensional detection boxes.

8. The method according to claim 6, characterized in that The obtaining the intensity map corresponding to each of the actual three-dimensional detection boxes according to the two-dimensional coordinates of each of the point cloud points and the intensity of each of the point cloud points includes: Perform a normalization process on the intensity of each of the point cloud points, and determine the normalized intensity as the pixel values of each of the point cloud points; Use the pixel values corresponding to each of the point cloud points as the pixel values at the corresponding two-dimensional coordinate positions to obtain the intensity map corresponding to each of the actual three-dimensional detection boxes.

9. The method according to claim 6, wherein The obtaining the density map corresponding to each of the actual three-dimensional detection boxes according to the density of each of the point cloud points in the z-axis direction in the bird's-eye view includes: Determine the density of the point cloud points in the z-axis direction at each position according to the number of point cloud points in the z-axis direction among the two-dimensional coordinate positions of each point cloud point, the maximum value of the number of point cloud points at all coordinate positions, and the minimum value of the number of point cloud points at all coordinate positions. Use the density of the point cloud points in the z-axis direction at each coordinate position as the pixel value of the corresponding two-dimensional coordinate position to obtain the density map corresponding to each of the measured three-dimensional detection frames.

10. The method according to claim 9, wherein The extracting the measured trajectory features of each candidate detection object according to the fusion data of each of the two-dimensional mapping trajectory features and each of the two-dimensional trajectory features includes: After compressing each of the two-dimensional trajectory features and each of the two-dimensional mapping trajectory features to the same proportional size, splice them to obtain a fusion data matrix. Determine the feature information extracted from the fusion data matrix as the measured trajectory features of each candidate detection object.

11. A target tracking device, characterized in that, The device includes: An acquisition module, configured to acquire the first current frame data of the target scene collected by the first sensor and the second current frame data of the target scene collected by at least one second sensor; the first sensor and the second sensor are sensors of different types. A feature acquisition module, configured to acquire the measured feature information of each candidate detection object in the current frame according to the first current frame data and the second current frame data; the candidate detection object is a detection object whose detection objects in the first current frame data and the detection objects in the second current frame data are successfully associated and matched; the measured feature information represents the inherent characteristic information of each candidate detection object. A prediction module, configured to predict the predicted three-dimensional detection frame of each target to be tracked in the current frame through a preset tracking algorithm model according to the historical three-dimensional detection frames of each target to be tracked in the previous frame; wherein, the tracking algorithm model is constructed based on the spatial equation of the uniform variable motion state. A tracking module, configured to track each candidate detection object according to the measured feature information of each candidate detection object in the current frame and the predicted detection frame of each target to be tracked in the current frame.

12. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 10 are implemented.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Image and laser point cloud fused target tracking method, computing device and medium

    CN110472553A

  • Vehicle detection method based on laser and vision fusion

    CN110942449A