A target tracking method and device

The state amount of the target tracking object is determined through the performance map and matching results of the detector, which solves the problem of low output efficiency in the prior art and achieves more efficient and accurate target tracking.

CN113780064BActive Publication Date: 2025-06-27HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110852187.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-27
Publication Date
2025-06-27
Estimated Expiration
2041-07-27

AI Technical Summary

Technical Problem

The existing target tracking methods have low output efficiency when detectors are mis-detected or missed, making it difficult to accurately determine whether to output tracking objects.

Method used

The detector obtains the detection objects of the current frame and the previous frame, matches them to determine the position and state amount of tracked objects, and uses the detector's performance map to determine the decision to output or die.

Benefits of technology

Improves the efficiency and accuracy of target tracking, reduces output delay, and temporarily tracks when tracking objects are blocked to avoid loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113780064B_ABST
    Figure CN113780064B_ABST
Patent Text Reader

Abstract

The present application discloses a target tracking method and device in the field of artificial intelligence, which are used to efficiently and accurately determine the tracking of a target in combination with the performance of a detector, improving the tracking efficiency and tracking accuracy. The method includes: obtaining at least one detected object in the current frame through the detector, where the current frame is any frame in the input data; obtaining a tracking object, where the tracking object includes an object detected by the detector in the previous frame of the current frame; matching the tracking object and the at least one detected object to obtain a matching result; determining a first state quantity of the tracking object according to the matching result and the performance map of the detector, where the first state quantity is used to indicate whether to output the tracking object or whether to eliminate the tracking object, and the performance map includes the detection accuracies within multiple grids in the detection range of the detector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular, to an object tracking method and apparatus. Background Art

[0002] Multiple object tracking (MOT), as a very important task in many important scenarios, such as autonomous driving applications, can establish the association relationship of obstacle objects between consecutive frames. It can continue to maintain the output of the tracked object when object detection fails, and can also solve the problem of false detection of object detection to a certain extent, making up for the deficiencies of frame-by-frame object detection. In addition, multiple object tracking can obtain the motion state information of the object, so as to provide important information for more advanced tasks such as intention recognition and behavior prediction of autonomous driving.

[0003] In the object tracking method of some scenarios, the detector is used to perform object detection frame by frame to obtain an object sequence, and then data association, motion state estimation, and tracking management are performed to complete the data association and state estimation of the detected objects in consecutive frames, and output the optimal tracked object sequence. However, the detector may have false detection or missed detection. Therefore, usually after an object is detected, it is not directly output, but output after a certain time delay. Therefore, the efficiency of outputting the tracked object is low. How to efficiently and accurately determine whether to output the tracked object has become an urgent problem to be solved. Summary of the Invention

[0004] This application provides an object tracking method and apparatus, which are used to efficiently and accurately determine the tracking of an object by combining the performance of the detector, and improve the tracking efficiency and accuracy.

[0005] In view of this, in a first aspect, this application provides an object tracking method, including: obtaining at least one detected object in a current frame through a detector, where the current frame is any frame in the input data; obtaining a tracked object, where the tracked object includes an object detected by the detector in the previous frame of the current frame; matching the tracked object and the at least one detected object to obtain a detection result; determining a first state quantity of the tracked object according to the matching result and the performance map of the detector, where the first state quantity is used to indicate whether to output the tracked object or whether to eliminate the tracked object, and the performance map includes the detection accuracy in multiple grids within the detection range of the detector.

[0006] Therefore, in the embodiments of the present application, by combining the detection accuracies of the detector in various regions, the state quantity of the tracked object can be calculated more accurately. Thus, according to the detection accuracy of the detector, the confidence level of the detected object can be determined, and the effectiveness of the tracked object can be determined efficiently. It can be understood as determining whether the detected tracked object is accurate, and then accurately determining whether to output the tracked object or eliminate the tracked object, which can reduce the output delay and improve the efficiency of determining whether to output the tracked object.

[0007] In a possible embodiment, the foregoing determining the first state quantity of the tracked object according to the matching result and the performance map of the detector may include: determining the position information of the tracked object in the current frame according to the matching result; querying the detection accuracy corresponding to the tracked object in the performance map based on the position information of the tracked object in the current frame; and calculating the first state quantity according to the detection accuracy.

[0008] In this embodiment, the position of the tracked object can be determined according to the matching result between the tracked object and the detected object. Thus, the corresponding detection accuracy can be queried in the performance map based on this position, and the confidence level that the tracked object is tracked can be indirectly determined. Therefore, the state quantity of the tracked object can be determined more accurately according to this detection accuracy, reducing the delay in determining whether to output or eliminate the tracked object, and more efficiently determining whether to output the tracked object or whether to eliminate the tracked object.

[0009] In a possible embodiment, the foregoing determining the position information of the tracked object in the current frame according to the matching result may include: if there is no detected object in at least one detected object that matches the tracked object, determining the position information of the tracked object in the current frame according to the motion state information of the tracked object; if the tracked object matches the first detected object in at least one detected object, using the position information of the first detected object as the position information of the tracked object in the current frame.

[0010] Therefore, in the embodiments of the present application, if the tracked object matches the detected object, that is, the detected object and the tracked object may be the same object, the position of the detected object can be used as the position of the tracked object to realize the tracking of the tracked object, and the corresponding detection accuracy can be queried in the performance map of the detector. Thus, based on this detection accuracy, the state quantity indicating whether to output or eliminate the tracked object can be calculated more accurately, obtaining a more accurate state quantity, and realizing more efficient determination of whether to output or eliminate the tracked object. Moreover, in the case where the tracked object is occluded, the position of the tracked object can be predicted. If the detection accuracy of the tracked object is determined based on the predicted position, the tracked object can be temporarily tracked even when the tracked object is occluded to avoid the situation of losing the tracked object.

[0011] In a possible implementation, the foregoing determination of the position information of the tracking object in the current frame based on the motion state information of the tracking object may include: obtaining the predicted position of the tracking object in the current frame according to the motion state information of the tracking object; calculating the predicted distance value between the tracking object in the current frame and the acquisition device that acquires the current frame according to the predicted position; obtaining the actual distance value between the predicted position and the acquisition device in the current frame through input data; if the difference between the predicted distance value and the actual distance value is greater than a first threshold, using the predicted position as the position information of the tracking object in the current frame.

[0012] Therefore, in the implementation manner of the present application, when the tracking object is occluded, the predicted distance between the predicted tracking object and the acquisition device at the predicted position of the tracking object can be calculated, as well as the actual distance between the actually acquired obstacle and the acquisition device at the predicted position, and it is determined whether the tracking object is occluded based on the difference between the predicted distance and the actual distance. If the difference is greater than a certain value, it indicates that the tracking object is occluded by the obstacle, and the information at the predicted position in the current frame acquired by the acquisition device is the information of the obstacle. Thus, the tracking object can be tracked based on the predicted position, avoiding the problem of tracking loss caused by the temporary occlusion of the tracking object.

[0013] In a possible implementation, the first state quantity includes a first output indication state quantity and a first extinction indication state quantity. The first output indication state quantity indicates whether to output the tracking object in the current frame, and the first extinction indication state quantity indicates whether to extinguish the tracking object in the current frame. The foregoing calculation of the first state quantity according to the detection accuracy may include: obtaining a direct observation quantity, where the direct observation quantity indicates the motion state of the tracking object. For example, the direct observation quantity may include information such as the acquired motion speed and motion direction of the tracking object; calculating the first posterior probability and the second posterior probability of the tracking object in the current frame according to the detection accuracy and the direct observation quantity. The first posterior probability is used to represent the probability of outputting the tracking object in the current frame, and the second posterior probability is used to represent the probability of extinguishing the tracking object in the current frame; obtaining the first output indication state quantity based on the first posterior probability, and obtaining the first extinction indication state quantity based on the second posterior probability.

[0014] Therefore, in the implementation manner of the present application, whether to output or extinguish the tracking object can be represented by the first output indication state quantity and the first extinction indication state quantity respectively. Specifically, the first output indication state quantity and the first extinction indication state quantity can be calculated through the motion state of the tracking object and the detection accuracy of the detector, so that the first output indication state quantity and the first extinction indication state quantity can be calculated more accurately, and thus it can be known more accurately whether to output or extinguish the tracking object.

[0015] In a possible implementation, obtaining the first output indication status quantity based on the first posterior probability and obtaining the first extinction indication status quantity based on the second posterior probability may include: obtaining the second output indication status quantity and the second extinction indication status quantity of the tracked object in the previous frame. The first status quantity includes the first output indication status quantity and the first extinction indication status quantity. The second output indication status quantity is used to indicate whether the tracked object was output in the previous frame, and the second extinction indication status quantity is used to indicate whether the tracked object was extinguished in the previous frame; fusing the first posterior probability and the second output indication status quantity to obtain the first output indication status quantity, and fusing the second posterior probability and the second extinction indication status quantity to obtain the first extinction indication status quantity.

[0016] In the embodiments of the present application, by combining the status quantity of the tracked object calculated when tracking the previous frame, the tracked object can be tracked in a timely manner through an iterative method, and it can be determined whether to output or extinguish the tracked object. For example, through a recursive Bayesian inference tracking method, the efficiency of determining whether to output or extinguish the tracked object can be improved.

[0017] In a possible implementation, the first posterior probability calculated when there is no detected object matching the tracked object among at least one detected object is greater than the first posterior probability calculated when the tracked object matches the first detected object among at least one detected object; and the second posterior probability calculated when there is no detected object matching the tracked object among at least one detected object is greater than the second posterior probability calculated when the tracked object matches the first detected object among at least one detected object.

[0018] Therefore, when no detected object matching the tracked object is detected or the tracked object is occluded, a negative gain can be applied to the output indication status quantity of the tracked object, that is, the confidence of outputting the tracked object is reduced, and a positive gain can be applied to the extinction indication status quantity, that is, the confidence of extinguishing the tracked object is increased, thereby avoiding false detection.

[0019] In a possible implementation, the detection accuracy of each grid in the performance map includes the accuracies corresponding to multiple categories; the foregoing querying the detection accuracy corresponding to the tracked object in the performance map based on the position information of the tracked object in the current frame may include: obtaining the category of the tracked object; querying the detection accuracy corresponding to the tracked object in the performance map based on the position information of the tracked object in the current frame and the category of the tracked object.

[0020] Therefore, in the embodiments of the present application, the performance map of the detector can be set to include the accuracies of the detector in different regions and different categories, so that more accurate detection accuracies can be queried by combining the position and category of the tracked object. Furthermore, based on the more accurate detection accuracies, more accurate state quantities of the tracked object can be calculated, improving the efficiency of determining whether to output or eliminate the tracked object.

[0021] In a possible implementation, the above method may further include: if at least one detected object includes a second detected object that does not match the tracked object, then using the second detected object as a new tracked object and tracking the new tracked object in the next frame of the current frame.

[0022] In the embodiments of the present application, if a new detected object is detected, then using the detected object as a new tracked object to update the new tracked object in the next frame, so that the detected object can be tracked in a timely manner.

[0023] In a possible implementation, before obtaining at least one detected object in the current frame by the detector, the above method may further include: dividing the detection range of the detector into multiple grids; obtaining ground truth data, where the ground truth data includes the acquisition data collected by the acquisition device and the information of the corresponding ground truth object; using the detector to detect the acquisition data to obtain the information of the predicted object; and calculating the detection accuracy corresponding to each grid in the multiple grids according to the predicted object and the ground truth object to obtain the performance map.

[0024] Therefore, in the embodiments of the present application, the performance of the detector can be encoded in advance to obtain the performance map, so that when performing target tracking, the accuracy corresponding to the tracked object can be quickly determined according to the performance map, improving the efficiency of determining whether to output or eliminate the tracked object.

[0025] In a second aspect, an embodiment of the present application provides a target tracking device, which has the function of implementing the target tracking method in the first aspect above. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.

[0026] In a third aspect, an embodiment of the present application provides a target tracking device, including: a processor and a memory. The processor and the memory are interconnected by a line, and the processor calls the program code in the memory to execute the functions related to processing in the target tracking method shown in any one of the first aspects above. Optionally, the target tracking device may be a chip.

[0027] Fourthly, an embodiment of the present application provides a target tracking device, which may also be referred to as a digital processing chip or a chip. The chip includes a processing unit and a communication interface. The processing unit obtains program instructions through the communication interface, and the program instructions are executed by the processing unit. The processing unit is used to execute the functions related to processing in the first aspect or any optional implementation manner of the first aspect as described above.

[0028] Fifthly, an embodiment of the present application provides a computer-readable storage medium, including instructions, which when running on a computer, cause the computer to execute the method in the first aspect or any optional implementation manner of the first aspect as described above.

[0029] Sixthly, an embodiment of the present application provides a computer program product containing instructions, which when running on a computer, cause the computer to execute the method in the first aspect or any optional implementation manner of the first aspect as described above. Description of the Drawings

[0030] Figure 1 It is a schematic diagram of an artificial intelligence main framework applied in the present application;

[0031] Figure 2 It is a schematic diagram of a system architecture provided by the present application;

[0032] Figure 3 It is a schematic diagram of an application scenario of a target tracking method provided by the present application;

[0033] Figure 4 It is a schematic diagram of the process flow of a target tracking method provided by the present application;

[0034] Figure 5 It is a schematic diagram of an application framework provided by the present application;

[0035] Figure 6 It is a schematic diagram of the process flow of detector performance coding provided by the present application;

[0036] Figure 7 It is a schematic diagram of the process flow of performing target tracking provided by the present application;

[0037] Figure 8 It is a schematic diagram of an occlusion analysis method provided by the present application;

[0038] Figure 9 It is a schematic diagram of the state quantity of Bayesian tracking provided by the present application;

[0039] Figure 10 It is a schematic diagram of the structure of a target tracking device provided by the present application;

[0040] Figure 11 It is a schematic diagram of the structure of another target tracking device provided by the present application;

[0041] Figure 12 A schematic structural diagram of a chip provided for this application. Specific embodiments

[0042] The following will describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.

[0043] First, the overall working process of the artificial intelligence system will be described. Please refer to Figure 1 , Figure 1 shown is a schematic structural diagram of an artificial intelligence main framework. The above artificial intelligence theme framework will be described from two dimensions: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). Among them, the "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be the general processes of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes the refinement process of "data - information - knowledge - wisdom". The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of artificial intelligence, information (providing and processing technology implementation) to the industrial ecological process of the system.

[0044] (1) Infrastructure

[0045] The infrastructure provides computing power support for the artificial intelligence system, realizes communication with the external world, and is supported through the basic platform. Communicate with the outside through sensors; the computing power is provided by intelligent chips, such as hardware acceleration chips such as central processing unit (CPU), neural-network processing unit (NPU), graphics processing unit (GPU), application specific integrated circuit (ASIC), or field programmable gate array (FPGA); the basic platform includes relevant platform guarantees and supports such as distributed computing frameworks and networks, and may include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside to obtain data, and these data are provided to the intelligent chips in the distributed computing system provided by the basic platform for computing.

[0046] (2) Data

[0047] The data at the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voices, texts, and also involves the Internet of Things data of traditional devices, including the business data of existing systems and the sensed data such as force, displacement, liquid level, temperature, humidity, etc.

[0048] (3) Data Processing

[0049] Data processing usually includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.

[0050] Among them, machine learning and deep learning can perform symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on data.

[0051] Reasoning refers to the process of simulating the intelligent reasoning mode of humans in a computer or intelligent system, based on the reasoning control strategy, using formalized information for machine thinking and problem-solving. The typical function is search and matching.

[0052] Decision-making refers to the process of making decisions after intelligent information is reasoned, and usually provides functions such as classification, sorting, prediction, etc.

[0053] (4) General Capabilities

[0054] After the data is processed through the above-mentioned data processing, some general capabilities can be formed further based on the results of the data processing. For example, it can be an algorithm or a general system. For example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0055] (5) Intelligent Products and Industry Applications

[0056] Intelligent products and industry applications refer to the products and applications of artificial intelligence systems in various fields, which is the encapsulation of the overall artificial intelligence solution, productizing intelligent information decision-making and realizing practical applications. Its application fields mainly include: intelligent terminals, intelligent transportation, intelligent healthcare, autonomous driving, smart cities, etc.

[0057] First, an exemplary introduction is made to the system architecture applied by the target tracking method provided in this application. Refer to Figure 2 , a schematic diagram of a system architecture provided in this application. Among them, the system can include a collection device 201 and an execution device 202.

[0058] The collection device 201 can include devices for collecting depth information, such as lidar, millimeter-wave radar, image sensors, infrared sensors, etc. The collection device can transmit the collected data to the execution device.

[0059] The execution device 202 can detect the information of the objects included therein based on the data collected by the collection device 201, track each object in the data collected by the collection device 201, and output each tracked object.

[0060] For example, the Multiple Object Tracking (MOT) algorithm can establish the association relationship of the object between the front and rear frames, and basically follows the Tracking by Detection framework, that is, using the detector to perform object detection frame by frame, outputting the sequence of detected objects as the input of the multiple object tracking module. Then, the multiple object tracking module completes the data association and state estimation of the detected objects in the front and rear frames through data association, motion state estimation, and tracking management, etc. Finally, the optimal tracking object sequence is output through the tracking management module.

[0061] However, as the output interface between the multiple object tracking module and the downstream module, the tracking management not only needs to maintain the length of the internal tracking object sequence, but also needs to determine which reliable tracking objects to output externally. Usually, the detector may have false negatives (FN) and false positives (FP). Therefore, the tracking module cannot directly output all tracking objects externally and should have a certain "delay" output mechanism to reduce the false detection and missed detection problems caused by the uncertainty of the detector.

[0062] On the other hand, usually the user hopes that based on the "delay" output of the tracking objects, the "true" tracking objects can be generated as soon as possible, and the "false detection" targets can be eliminated as soon as possible, so as to realize the output of the optimal tracking object sequence and ensure the tracking performance. Moreover, the tracking management module also needs to be able to handle the problem of missed detection of the object due to short-term occlusion, and still maintain the output of the tracking object for a period of time when the detector misses detection due to occlusion.

[0063] Therefore, the tracking management plays a very important role in the multiple object tracking task. Some commonly used multiple object tracking management mechanisms can usually be divided into two categories: the fixed waiting threshold generation and extinction mechanism and the finite state machine state transition management mechanism. The former delays the generation or extinction of the tracking object by setting a fixed waiting threshold in the generation and extinction stages of the tracking object, and can solve the false detection and missed detection problems of the detector to a certain extent. However, usually the size of the threshold is not easy to determine, and it is usually set empirically. Moreover, the determination of the threshold is heuristic and lacks theoretical support. It is very difficult to balance between reducing the false tracking objects caused by false detection and reducing the loss of detected objects caused by missed detection. The latter performs tracking state transition according to the number of frames or time of continuous tracking of the tracking object. Compared with the method of directly setting the generation and extinction threshold, it can recover the temporarily lost target quickly to a certain extent, but does not adjust the threshold size according to the detector performance and tracking state, and requires more manually set thresholds.

[0064] For another example, in some commonly used target tracking methods, such as the multi-layer track file management method used in the Multiple Hypothesis Tracking (MHT) system, a multi-layer track file including a general track file and an internal track file is established, and the internal track file is updated in real time corresponding to the MHT hypothesis generation and pruning process. The corresponding logic of the general track file and the internal track file is set, and the general track file information is updated based on the internal track file information. According to the target information in the general track file, the target track is output and displayed in real time to complete the target tracking management and target information reporting. However, this method indirectly realizes the number management of the tracking object by means of the hypothesis generation and pruning operations in the multiple hypothesis tracking algorithm. However, as the tracking time increases, each tracking object needs to save a large number of hypothesis sets, and the calculation efficiency is low. In addition, this method is suitable for tracking distant targets, but for closer targets, the tracking effect may be poor due to reasons such as target occlusion or detection accuracy.

[0065] For example, the optical flow method can also be used to track the feature points of the target area to generate tracking results. According to the tracking results and the inverse mapping relationship between the feature points and the tracking object frames, it is determined whether the tracking object frame of each of the multiple tracking objects is successful in tracking. If successful, the tracking object frame is recalculated and the feature points are updated, otherwise the tracking object frame and the corresponding feature points of the tracking object are deleted. However, this method only uses the matching results of the optical flow features to decide whether to add and delete the tracking object, which may result in poor tracking effect and low robustness for scenes with complex optical flow. And if the target is temporarily lost, such as being blocked, it is impossible to maintain and restore the lost target, resulting in poor tracking effect.

[0066] Therefore, the present application provides a target tracking method, which divides the detection accuracy of different areas of the detector to quickly and accurately determine which tracking objects can be output, thereby improving the target tracking efficiency and tracking effect. The target tracking method provided by the present application is introduced below.

[0067] First, the target tracking method provided in this application can be applied to a variety of scenarios that require target detection or target tracking, such as autonomous driving, monitoring, or video recording. For ease of understanding, some possible scenarios are only introduced below by way of example, but not as a limitation.

[0068] Scenario 1: Autonomous Driving

[0069] The method provided by this application can be applied to the perception module of a vehicle. For example, in implementation, it can be combined with the software and hardware systems of an autonomous vehicle. The hardware system may include a target detection sensor, or a processor, etc. The target detection sensor may include a lidar sensor for detecting targets in the environment around the vehicle. The processor can be used to receive data from the target detection sensor, process the data, and output obstacle targets, such as a general-purpose processor, a graphics processor, etc. The software system includes an operating system, a sensor driver, a sensor data processing program, etc. This application can be deployed in the perception module of an autonomous driving software system. This application can be used as a tracking management sub-module of the multi-target tracking module in the perception module, capable of maintaining an internal sequence of tracked objects and outputting stable and reliable tracked object results, usually sent to other sub-modules in the perception module, such as a prediction module.

[0070] Specifically, for example, during autonomous driving, tracking other obstacles near the vehicle, such as other vehicles, pedestrians, roadblocks, or signs, etc., is crucial for autonomous driving and affects the driving safety and efficiency of the vehicle. For example, as Figure 3 shown, one or more lidars can be set in the vehicle, such as Figure 3 the lidar 301 installed on the roof as shown in. During the process of autonomous driving, the one or more lidars can collect environmental information near the vehicle in real time to obtain a point cloud, and then detect objects in the point cloud data through a detector and track the objects in the point cloud data in real time, so that the vehicle can know in real time the obstacles near the vehicle that affect the driving safety of the vehicle and timely plan or adjust the driving path of the vehicle. Of course, the lidar in the vehicle can also be replaced with other sensors that can obtain depth, such as an infrared sensor, an image sensor, etc., which will not be elaborated here.

[0071] Scenario Two: Monitoring

[0072] Among them, in the monitoring scenario, an image sensor can be set in the monitoring device. The monitoring video can be collected in real time through the image sensor. The monitoring device can detect each frame of the image, identify the objects in each frame of the image, and establish an association relationship between the objects in each frame of the image and the objects included in the previous frame or multiple previous frames, and continuously track the objects. When the output conditions are met, the tracked objects can be output. For example, the people in the monitoring screen can be tracked, and the orientation of the camera can be adjusted in real time according to the tracked people, so that the camera can track the people and monitor the status of the people in the scene in time.

[0073] Scenario Three: Robot

[0074] The target tracking method provided by this application can be applied to intelligent robots. A lidar or an image sensor can be set in the intelligent robot to collect data within the monitoring range in real time, identify and track the objects in the collected data, and output the tracked object after meeting the output conditions. For example, an image sensor can be set in the intelligent robot to identify the images collected by the image sensor in real time, detect the objects therein and track them, and after meeting the output conditions, the tracked object can be output, and the intelligent robot can perform tracking operations based on the tracked object, such as adjusting the orientation or traveling direction of the intelligent robot, etc.

[0075] Therefore, target tracking can be widely applied to various scenarios that require target tracking. The target tracking method provided by this application analyzes the performance of the detector to determine the detection accuracy of different detection regions within the detection range of the detector, thereby determining the accuracy of the detected object based on this accuracy, and more accurately and efficiently determining whether to output the tracked object.

[0076] Participate Figure 4 , the specific process of the target tracking method provided by this application will be introduced as follows.

[0077] 401. Obtain at least one detected object in the current frame through the detector.

[0078] Among them, the detector can be used to identify the objects in the current frame and output at least one detected object. That is, the current frame is used as the input of the detector, and the information of at least one identified object is output, such as the category, position, or motion state of the at least one object.

[0079] The current frame can be any frame in the input data, and the input data can be the data collected by the aforementioned acquisition device, such as the point cloud data collected by the lidar, the point cloud data collected by the millimeter-wave radar, or the images collected by the image sensor. For example, the point cloud data collected by the lidar can be divided according to a preset time unit, and each unit time is divided into one frame, so as to divide the point cloud data into multiple frames.

[0080] The detector may include a preselected model or a pre-trained model. For example, the detector may include a network for object detection, a classification network, a segmentation network, etc. Specifically, for example, the detector may include a convolutional neural network (CNN), a deep convolutional neural network (DCNN), a recurrent neural network (RNN), a region based convolutional neural network (RCNN), or a faster RCNN, etc.

[0081] 402. Obtain a tracking object.

[0082] Among them, the tracking object is the object that has been determined to be tracked in the input data. The tracking object may include the object detected by the detector in the previous frame, and information such as the position, classification, or motion state of the tracking object can be obtained.

[0083] For example, when processing the input data, a tracking object set can be established, which includes one or more tracking objects. When processing the first frame, the tracking object set is initialized, and the detected object recognized by the detector is added to the tracking object set. When processing each subsequent frame, the objects in the tracking object set can be tracked.

[0084] The number of tracking objects may include one or more. In this application, one of the tracking objects is taken as an example for illustrative purposes and is not intended to be limiting.

[0085] It should be understood that the previous frame referred to here is the frame arranged in a certain order before the current frame. This order may be the time order of collecting the input data or the order of detecting the input data. For example, if the input data is a video, target tracking of the video can be performed in the forward playback order of the video, or in the reverse order, etc.

[0086] 403. Match the tracking object with at least one detected object to obtain a matching result.

[0087] Among them, the tracking object can be matched with at least one detected object to determine the matching result between the tracking object and the at least one object.

[0088] It can be understood that the tracking object and at least one detected object can be matched to determine whether the tracking object and one of the at least one detected objects are the same object. For example, if the tracking object is a red car, it can be determined whether there is a car among the identified multiple cars that is the same as the tracking object. If so, it means that there is an object among the identified multiple cars that matches the tracking object.

[0089] Specifically, the method of matching the tracking object and at least one detected object can be to match according to the specific information of the tracking object and the detected object, such as matching according to information such as the category, shape, position, movement speed, or movement direction of the object, so as to identify whether the tracking object and the detected object are the same object or calculate the confidence level that the tracking object and the detected object are the same object. For example, if the size, shape, color of the tracking object and the detected object, the position of the tracking object, and the predicted position of the detected object are the same, and the predicted position can be calculated according to the movement speed or movement direction of the detected object, etc., it means that the tracking object and the detected object are the same object.

[0090] Generally, the matching result between the tracking object and the detected object can be divided into matching or not matching. Matching means that the tracking object and the detected object are the same object, and not matching means that the tracking object and the detected object are different objects. Of course, the matching result between the tracking object and the detected object can also be represented by the degree of matching. For example, when the degree of matching is higher than a certain value, it means that the tracking object and the detected object are the same object, and the continuous tracking of the tracking object is realized.

[0091] For example, the matching result can specifically include: there is a detected object that matches the tracking object, there is no detected object that matches the tracking object (i.e., there are redundant tracking objects), or there is no tracking object that matches the detected object (i.e., there are redundant detected objects), etc.

[0092] 404. Determine the first state quantity of the tracking object according to the matching result and the performance map of the detector.

[0093] After matching the tracking object and the detected object, the output tracking object or the disappearing tracking object can be determined based on the performance map of the detector, that is, the output of the tracking object or the disappearance of the tracking object is represented by the calculated first state quantity. That is, after obtaining the first state quantity, it can be determined whether to output the tracking object in the current frame or disappear the tracking object in the current frame.

[0094] Output the tracked object, which can be understood as determining the continuous existence of the tracked object in multiple frames, that is, confirming the existence of the tracked object, and outputting the tracked object for further processing based on the tracked object. For example, in an autonomous driving scenario, outputting the tracked object means determining the existence of other vehicles or pedestrians near the vehicle, etc. The information of the tracked object can be transmitted to the autonomous driving control module of the vehicle, and further processing can be performed on the tracked object, such as adjusting the driving path of the autonomous driving vehicle, accelerating or decelerating, etc.

[0095] Eliminate the tracked object, which can be understood as stopping tracking the tracked object. When targeting the next frame, the determined eliminated object is deleted from the tracked objects, that is, no further tracking is performed on the tracked object. For example, the tracked objects include a green car and a red car. If it is determined to stop tracking the red car, such as the red car driving out of the monitoring range of the vehicle or the red car being blocked, etc., then the red car can be deleted from the tracked objects. When processing the next frame, that is, no tracking is performed on the red car to reduce the tracking workload.

[0096] The performance map can include the detection accuracies of multiple grids within the detection range of the detector, and the performance map can reflect the differences in the detection accuracies of the detector in different regions or different categories, etc. Among them, the detection accuracy of the detector can be represented by recall rate, precision, or average precision, etc. The aforementioned detection range can include the effective detection range of the detector, such as the range where the detection accuracy is greater than the preset accuracy, or a pre-defined range, etc. For example, the detection range of the detector can be divided into multiple grids in advance, and the detection accuracy of the detector within each grid is calculated. The detection accuracies of the multiple grids constitute the subsequent performance map. The performance map can be saved in the form of a graph, or in the form of a table or other ways, and can be adjusted specifically according to the actual application scenario.

[0097] Generally, the performance of the detector has high reference value for the life cycle management of the tracked object, and there are differences in the detection performance of the detector in different regions. In the embodiments of the present application, the detection accuracy of the detector in different regions can be used to determine whether to output or eliminate the tracked object, so as to more accurately judge whether to output or eliminate the tracked object, and improve the efficiency of outputting the tracked object.

[0098] Specifically, the position information of the tracked object in the current frame can be determined according to the matching result, and the detection accuracy corresponding to the tracked object can be queried in the performance map based on the position information, and the first state quantity can be calculated based on the detection accuracy.

[0099] In a possible implementation, if there is no detected object in at least one detected object in the current frame output by the detector that matches the tracked object, the position information of the tracked object in the current frame can be determined according to the motion state of the tracked object. If the tracked object matches one of the at least one detected objects (referred to as the first detected object), it can be understood that the first detected object and the tracked object are the same object, and the position information of the first detected object is used as the position information of the tracked object.

[0100] Specifically, if there is no detected object in at least one detected object in the current frame output by the detector that matches the tracked object, the predicted position of the tracked object in the current frame can be obtained according to the motion state information of the tracked object. The motion state information may specifically include information such as the motion speed, motion direction, or starting position of the tracked object. The motion state information can be extracted from the input data or calculated based on the input data. Then, the predicted distance value between the tracked object in the current frame and the sensor that collects the current frame is calculated according to the predicted position, and the actual distance value between the predicted position and the collection device in the current frame is obtained through the input data. If the difference between the predicted distance value and the actual distance value is greater than the first threshold, the predicted position is used as the position information of the tracked object in the current frame.

[0101] For example, it can be understood that if the difference between the distance between the predicted position where the tracked object is located detected through the input data and the collection device and the predicted distance is greater than the first threshold, it indicates that the tracked object may be occluded, resulting in the collection device being unable to collect data of the tracked object. At this time, the predicted position can be used as the position information of the tracked object. Therefore, in the implementation manner of the present application, the problem of the tracked object being occluded can be adaptively solved. The position of the tracked object can be predicted to avoid the temporary loss of the tracked object, and continuous tracking of the tracked object can be achieved. Even if there is a temporary loss, the tracked object can be stably output.

[0102] More specifically, the first state quantity may include a first output indication state quantity and a first extinction indication state quantity. The first output indication state quantity is used to indicate whether to output the tracked object in the current frame, and the first extinction indication state quantity is used to indicate whether to extinguish the tracked object in the current frame. Generally, the first output indication state quantity and the first extinction indication state quantity are negatively correlated. The method for calculating the first state quantity may specifically include: obtaining a direct observation quantity, where the direct observation quantity is used to represent the motion state of the tracked object, or using the aforementioned motion state information as the direct observation quantity; subsequently, based on the detection accuracy and the direct observation quantity, calculating the first posterior probability and the second posterior probability of the tracked object in the current frame. The first posterior probability is used to represent the probability of outputting the tracked object in the current frame, and the second posterior probability is used to represent the probability of extinguishing the tracked object in the current frame; obtaining the first output indication state quantity based on the first posterior probability, and obtaining the first extinction indication state quantity based on the second posterior probability. For example, the first posterior probability can be transformed to generate an observation likelihood ratio related to precision and recall, and this observation likelihood ratio can be used as the first output indication state quantity. The first posterior probability can be transformed to generate an observation likelihood ratio related to precision and recall, and this observation likelihood ratio can be used as the first extinction indication state quantity.

[0103] Optionally, the second output indication state quantity and the second extinction indication state quantity of the tracked object in the previous frame can also be obtained. The first state quantity includes the first output indication state quantity and the first extinction indication state quantity. The second output indication state quantity is used to indicate whether the tracked object was output in the previous frame, and the second extinction indication state quantity is used to indicate whether the tracked object was extinguished in the previous frame; subsequently, fusing the first posterior probability and the second output indication state quantity to obtain the first output indication state quantity, and fusing the second posterior probability and the second extinction indication state quantity to obtain the first extinction indication state quantity. Therefore, in the embodiments of the present application, the state quantity of the tracked object in the input data can be determined iteratively to quickly and accurately determine whether to output the tracked object or extinguish the tracked object.

[0104] Generally, the first posterior probability calculated when there is no detected object matching the tracked object in at least one detected object is greater than the first posterior probability calculated when the tracked object matches the first detected object in at least one detected object; the second posterior probability calculated when there is no detected object matching the tracked object in at least one detected object is greater than the second posterior probability calculated when the tracked object matches the first detected object in at least one detected object.

[0105] Of course, the first posterior probability can also be directly used as the first output indication state quantity, and the second posterior probability can be directly used as the first extinction indication state quantity to reduce subsequent calculation amounts.

[0106] In a possible implementation, before step 404, a performance map of the detector can also be obtained. The performance map can be extracted from a memory or obtained by analyzing the performance of the detector.

[0107] The specific ways to obtain the performance map can include dividing the detection range of the detector into multiple grids and obtaining ground truth data. The ground truth data includes the acquisition data collected by the acquisition device and the annotation information of the corresponding ground truth objects, such as the sequence of the annotated objects. The annotation information can include the ground truth objects in each frame of the acquisition data and the matching results between the ground truth objects. Taking the acquisition data as the input of the detector to obtain the information of the predicted objects output by the detector, such as the sequence of the predicted objects, and then comparing the difference between the information of the predicted objects and the information of the ground truth objects, the detection accuracy corresponding to each grid can be calculated to obtain the performance map of the detector. The detection accuracy can be represented by recall rate, precision metric or average precision, etc.

[0108] Among them, the ways to divide the detection range can include dividing by depth or angular range, etc. For example, it can be divided according to the distance from the acquisition device. For example, the range within 10 meters can be divided into one grid, the range from 10 - 20 meters can be divided into one grid, and so on. Or, it can be divided according to the planar area. For example, in the autonomous driving scenario, a certain area facing the front of the vehicle can be divided into one grid, and the areas facing the sides of the vehicle body can be divided into grids respectively, and so on. Of course, it can also be divided into grids in three-dimensional space, that is, combining depth and planar area to divide the grids, which can be specifically adjusted according to the actual application scenario.

[0109] In addition, the detection range of the detector can be an area within a preset range of the distance from the acquisition device. For example, it can be an area within 100 meters in diameter of the acquisition device, or the effective detection range of the acquisition device can be used as the detection range. For example, the range where the accuracy of the detector exceeds 90% can be used as the detection range, etc. The detection range can be specifically selected according to the actual application scenario.

[0110] In a possible implementation, if there are detection objects in at least one detection object in the current frame output by the detector that do not match the tracking object, the detection objects that do not match the tracking object can be used as new tracking objects and tracked in the next frame. For example, when tracking the objects in the current frame, the tracking objects include objects A, B, C, and D, and the detection objects include A, B, C, D, and E. When tracking the objects in the next frame, object E can be used as the tracking object and tracked.

[0111] In addition, if the performance map represents the detection accuracy of the detector in different regions within the detection range and for different categories of objects, when querying the detection accuracy corresponding to the position information of the tracked object, the category of the tracked object is also combined to query the detection accuracy corresponding to the tracked object. That is, based on the position information of the tracked object and the category of the tracked object in the current frame, the detection accuracy corresponding to the tracked object is queried in the performance map, so as to achieve a finer-grained division of the detection accuracy of the detector. The calculated first state quantity is also more accurate, and thus the analysis result of tracking or eliminating the tracked object is also more accurate, and it is possible to efficiently determine whether to output or eliminate the tracked object.

[0112] The above has introduced the process of the target tracking method provided by the present application. For ease of understanding, the process of the target tracking method provided by the present application will be introduced in more detail below in combination with a specific application scenario.

[0113] Exemplarily, refer to Figure 5 , a schematic diagram of the architecture to which the target tracking method provided by the present application is applied.

[0114] Among them, the architecture is divided into two parts, namely Figure 5 the offline part 501 and the online part 502 shown in

[0115] The offline part 501 is to divide the detection range of the detector in the previous step, and through a ground truth sequence with pre-added annotations, that is, a sequence added to the objects in the collected ground truth data, the detection accuracy in each grid is statistically calculated, so as to obtain the performance map of the detector.

[0116] The online part 502 is to use the performance map of the detector output by the offline part to analyze the data collected by the acquisition device and output a sequence of tracked objects.

[0117] Specifically, the online part 502 may include collecting input data through an acquisition device, then using the detector to detect the objects in each frame of the input data, and outputting the detected objects in each frame. Then, data association analysis is performed based on the input data, that is, analyzing the matching result between the tracked object and the detected object, and subsequently, the tracked object is tracked and managed based on the matching result.

[0118] Further, after the acquisition device acquires the input data, a detector is used to detect the data in each frame of the input data to identify the objects in each frame. During the process of processing each frame, the objects in the tracking object set can be tracked. For example, when processing the current frame, the tracking object set can be obtained, and the tracking object set includes one or more tracking objects, including the detection objects detected by the detector in the previous frame. Then, the tracking objects in the tracking object set are matched with one or more detection objects detected in the current frame to achieve the data association between the tracking objects and the detection objects, such as adding the information of the tracking object in the current frame to the sequence of the tracking object. Among them, the matching results can include various types, such as matching the tracking object and the detection object, there is no detection object matching the tracking object, or there is no tracking object matching the detection object, etc. Each tracking object or detection object has a state quantity within one, such as a generation indication state quantity or a disappearance indication state quantity, etc. The generation indication state quantity is used to indicate whether to output the tracking object, and the disappearance indication state quantity is used to indicate whether to disappear the tracking object, that is, to stop tracking the tracking object.

[0119] Then, the detected objects can be tracked according to the matching results. If the tracking object and the detection object are successfully matched, that is, the detection object and the tracking object are the same object, then the tracking object can be continuously tracked according to the detection accuracy of the detector for the detection object, and the generation indication state quantity of the tracking object can be positively incremented, that is, the confidence of outputting the tracking object is increased. At the same time, the disappearance indication state quantity of the tracking object can be negatively incremented. If there is no tracking object matching the detection object, the detection object can be added to the tracking object set, and the new tracking object can be tracked in the next frame, and the generation indication state quantity and the disappearance indication state quantity of the tracking object can be initialized. If there is no detection object matching the tracking object, occlusion analysis can be performed on the tracking object, that is, it is judged whether the tracking object is occluded. If it is determined that the tracking object is occluded, the tracking object can be continuously tracked. If the tracking object is not occluded, the tracking object can be disappeared, that is, the tracking of the tracking object is stopped. Specifically, occlusion analysis can be performed according to the difference between the predicted distance value between the predicted position of the tracking object and the acquisition device and the actual predicted value corresponding to the predicted position in the input data. If the difference between the predicted distance value and the actual distance value is too large, it can be determined that the tracking object is occluded. If no object is detected at the predicted position, it can be determined that the tracking object is not tracked, or the tracking object has exceeded the acquisition range of the acquisition device, and there is no need to track the tracking object.

[0120] If the tracking object meets the output condition, such as the value of the generated indication state quantity exceeds the preset value, the tracking object can be output. If the tracking object meets the extinction condition, such as the value of the extinction indication state quantity exceeds the preset value, the tracking object can be extinct, that is, the tracking object is deleted from the tracking object set, that is, the tracking of the tracking object is stopped.

[0121] Therefore, in an embodiment of the present application, a detector can be used to detect objects in the input data, and based on the detection accuracy of the detector in different areas and / or different categories, the output or extinction of the tracked object can be analyzed to obtain accurate analysis results, thereby improving the output or extinction efficiency of the tracked object.

[0122] Furthermore, for ease of understanding, the target tracking method provided in the present application can be divided into multiple stages, such as detector performance encoding and target tracking, etc. The multiple stages are introduced below respectively.

[0123] Phase 1: Detector Performance Coding

[0124] The purpose of the detector performance coding is to count the performance of the detector according to the region or object category, determine the detector performance in different regions and different categories, so as to carry out targeted life cycle management of the tracked object. Specific indicators for measuring the performance of the detector may include recall rate, precision index or average precision, etc. This embodiment uses the recall rate and precision as examples for illustrative explanation. The recall rate and precision mentioned below may also be replaced by average precision or other parameters for measuring the performance of the detector.

[0125] Exemplarily, the process of encoding detector performance can be as follows Figure 6 shown.

[0126] First, perform grid division. Define a detection area of ​​interest as the detection range of the detector, and divide the detection range into multiple grids. The grids can be divided according to a two-dimensional plane or a three-dimensional space, which is equivalent to discretizing the detection space of the detector to facilitate the subsequent statistics of the performance of the detector in different areas. For example, if the method provided in the present application is deployed in a vehicle, a rectangular area can be divided from the central axis of the vehicle as the detection range. It can be understood that by performing grid division on the detection area of ​​the detector, a two-dimensional grid map is obtained, thereby realizing the spatial discretization of the detection area, so as to facilitate the subsequent statistics of the performance differences of the detection area in different areas.

[0127] Obtain data with ground truth annotations, which includes data collected by sensors, and perform annotations on the data collected by sensors, annotating the information of the ground truth objects included in each frame of the collected data, that is, the ground truth annotation, such as information about the sequence, position, or category of the ground truth objects.

[0128] Then use the data collected by the sensor as the input of the detector, and output the information of one or more detected objects in each frame of the collected data, such as position or category. The detector may include a pre-trained model for detecting or identifying targets in the input data, such as a target detection model, a classification model, etc.

[0129] Subsequently, calculate the accuracy based on the information of the detected objects output by the detector and the information of the annotated ground truth objects, such as calculating the recall rate and precision in different regions, so as to represent the detection accuracy of the detector through the recall rate and precision, and thus obtain the performance map of the detector. Of course, in addition to representing the detection accuracy of the detector through the recall rate and precision, other parameters can also be used to represent the detection accuracy of the detector, such as the average precision, which can be specifically adjusted according to the actual application scenario, and the present application does not limit this.

[0130] For example, the detection accuracy of the detector can be expressed as:

[0131] Recall rate:

[0132] Precision:

[0133] Among them, (u, v) represents the index value or coordinate value of the grid, etc., used to represent the position of the grid. For example, a coordinate system can be established in the detection area. u can be the abscissa of the center point of a certain grid in the detection area, v can be the ordinate of the center point of a certain grid in the detection area, or u can be the longitude of the center point of a certain grid in the detection area, v can be the latitude of the center point of a certain grid in the detection area, etc. TP represents the number of true positives, FN represents the number of false negatives, FP represents the number of false positives, Recall(u, v) and Precision(u, v) respectively represent the recall rate and precision of the detector in different regions. The recall rate can be used to evaluate the performance of the detector not to miss detections, that is, the higher the Recall value, the fewer missed detections. The precision is used to evaluate the performance of the detector not to make false detections, that is, the higher the Precision, the fewer false detections.

[0134] After obtaining the recall rate and precision of each grid, the recall rate and precision of each grid can be statistically analyzed, and finally a performance map can be obtained to facilitate querying the detection accuracy of each grid in the online part. Among them, the performance map can be saved in the memory so that the performance map can be extracted from the memory when performing the second stage, or the first stage can be performed before the second stage, so that the second stage can be executed based on the results of the first stage. The specific execution timing can be adjusted according to the actual application scenario. This application is only an example and is not a limitation.

[0135] In addition, in addition to statistically analyzing the overall detection accuracy of each grid, at a finer granularity, the detection degree of different classes within each grid by the detector can also be statistically analyzed. For example, usually the detection accuracy of the detector for different classes of objects may be different. When calculating the recall rate and precision, the different classes can be statistically analyzed separately, and the TP, FN, and FP of different classes can be calculated, so as to calculate the recall rate and precision of different classes within each grid based on the TP, FN, and FP of different classes. Thus, the performance of the detector can be divided at a finer granularity to obtain more accurate detection accuracy.

[0136] Among them, the first stage can be performed offline, that is, before performing target tracking, the performance of the detector can be analyzed and the analysis results can be saved in the memory. When performing target tracking, the performance map of the detector can be extracted from the saved data, so as to accurately track the target in combination with the performance of the detector. Of course, it can also be executed before the second stage, and can be specifically adjusted according to the actual application scenario. This application does not limit this.

[0137] Generally, for some common 3D point cloud object detection tasks, such as using deep neural networks such as PointRCNN and TANet to perform object detection tasks, only the overall performance of the algorithm is evaluated on the test set to obtain the overall recall rate and precision, without considering the performance differences of the detector caused by different features such as point cloud density and quantity for different orientations and different classes of targets. And this application statistically analyzes the detection accuracy of different regions and / or different classes, which is equivalent to obtaining a finer granularity of detection accuracy, so as to facilitate subsequent more accurate estimation of the state of the tracked object based on the finer detection accuracy, improve the accuracy of the estimation result, and indirectly improve the efficiency of outputting or eliminating the tracked object.

[0138] Stage Two: Target Tracking

[0139] Exemplarily, the process of target tracking can be as Figure 7 shown.

[0140] First, the detector outputs information about one or more detected objects in the current frame of the input data, that is, one or more in the set of detected objects, such as information about the category of the recognized detected object, its position in the current frame, etc.

[0141] Then, the tracking objects and the detected objects are matched to determine whether there is a detected object associated with the tracking object. A tracking object is any object in the set of tracking objects.

[0142] The set of tracking objects includes the objects recognized during the detection of the previous frame. For example, when processing the first frame, the objects detected in the first frame are added to the set of tracking objects. When processing the second frame, the tracking objects in the set of tracking objects are tracked against the objects recognized in the second frame, and the set of tracking objects is updated based on the new detected objects so as to continue tracking the set of tracking objects in the next frame, and so on.

[0143] Among them, there are various situations for the matching result between the tracking object and the detected object, such as there being a detected object associated with the tracking object, there being no detected object associated with the tracking object, or there being no tracking object associated with the detected object, etc. These will be described separately below.

[0144] Situation 1: There is a detected object associated with the tracking object

[0145] If there is a detected object associated with the tracking object, then the corresponding detection accuracy is queried in the performance map according to the position and category of the detected object, and a positive gain is applied to the target generation management model, while a negative gain is applied to the target extinction management model. Among them, the target generation management model is used to calculate the aforementioned first output indication state quantity, and the target extinction management model is used to calculate the aforementioned first extinction indication state quantity.

[0146] Situation 2: There is no tracking object associated with the detected object

[0147] Among them, if the set of detected objects includes redundant detected objects that do not match the tracking objects in the set of tracking objects, then these detected objects can be used to update the set of tracking objects, that is, these detected objects are added to the set of tracking objects to track these detected objects in the next frame.

[0148] Situation 3: There is no detected object associated with the tracking object

[0149] If there is no detected object associated with the tracking object, then the position of the tracking object in the current frame can be predicted, and then the corresponding detection accuracy is queried in the performance map according to the predicted position and the category of the tracking object.

[0150] Then, it is determined whether the tracked object is occluded. If it is confirmed that the tracked object is occluded, the positive gain can continue to be applied to the target generation management model, while the negative gain is applied to the target extinction management model.

[0151] If it is confirmed that the tracked object is not occluded, the negative gain can be applied to the target generation management model, while the positive gain is applied to the target extinction management model. That is, the possibility of outputting this tracked object is reduced, and the possibility of extinguishing this tracked object is increased.

[0152] Generally, during the process of tracking a target, it often happens that the tracked object is occluded due to the influence of other obstacles or other tracked objects. One of the significances of multi-target tracking compared to detecting a target is to be able to maintain the target output for a period of time even when the detection fails due to occlusion of the tracked object and missed detections occur, so as to provide important information for further behavior decision-making, such as providing important information for the behavior decision-making of an autonomous driving vehicle.

[0153] Therefore, occlusion analysis is introduced during the tracking process, and at the same time, the handling method when the target is occluded is considered in the tracking management module. For example, as Figure 8 shown, a 3D lidar scanner can be set on the roof of the vehicle itself. The lidar scanning range is 360 degrees. If the angular resolution is 0.5 degrees, the scanning range can be divided into 720 angular ranges, that is, 720 virtual rays. When performing occlusion analysis, it is necessary to determine whether there are other obstacles on these 720 virtual rays according to the obstacle point cloud of each frame, and the radial distance value from the origin of the coordinate system to the obstacle of the ray, which is initially default set to no obstacle and the radial distance value is 0.

[0154] Specifically, as Figure 8 shown, at the t-1 moment (i.e., the previous frame), the scanning surface of the tracked object is completely visible relative to the vehicle itself. At this time, according to the spatial position occupied by the tracked object, the ray distance value within the virtual ray range covering the tracked object can be calculated. According to the obstacle point cloud distribution at this moment, the ray distance value within the angular range corresponding to the tracked object can also be calculated. At the t-1 moment, due to no occlusion, these two distance values of each ray are basically equal. And at the t moment, due to the occlusion effect of another obstacle on the tracked object, at this time, most areas of the tracked object cannot be scanned by the lidar, Figure 8 the solid line shown in

[0155] Similarly, based on the spatial position occupied by the tracked object, the theoretical ray distance value (i.e., the predicted distance value) within the virtual ray covering the tracked object can be calculated. Additionally, the actual ray distance value (i.e., the actual distance value) can be calculated based on the obstacle point cloud distribution at this moment. Due to the influence of occlusion at this time, the theoretical ray distance value of the occluded part of the tracked object will be greater than the actual ray distance value. Therefore, the relationship between the predicted distance and the actual distance of the virtual ray is used as the judgment condition for whether the tracked object is occluded.

[0156] For example, if half or more of the virtual rays within the coverage of the virtual ray of the tracked object are occluded, it is considered that the tracked object is occluded.

[0157] In addition, if the tracked object is occluded for multiple consecutive frames, the tracked object can be eliminated. For example, if the tracked object is detected to be occluded for 10 consecutive frames, the tracked object can be eliminated to reduce the workload.

[0158] Therefore, in the embodiments of the present application, the occlusion analysis of the tracked object is increased. Even if the tracked object is occluded, it can be further processed according to the occluded situation, so that the output or elimination of the tracked object can be determined more accurately and efficiently.

[0159] Next, an exemplary introduction to the target generation management model and the target generation model involved in the above various situations is given.

[0160] First, a management mechanism is established for each tracked object, including a target generation management model and a target elimination management model, used to determine whether the tracked object should be generated or eliminated. These two models need to consider the occluded situation of the tracked object, that is, select appropriate gains for the occluded and unoccluded situations, which can solve the problem of loss of tracked objects caused by occlusion to a certain extent and achieve more continuous tracking of the tracked object. For example, if the tracked object is directly eliminated when it disappears, then when the tracked object is detected again, the state of the tracked object needs to be initialized and the tracking of the tracked object needs to be restarted, with low efficiency. However, in the present application, occlusion analysis is performed. Even if the tracked object is temporarily occluded, the tracking of the tracked object can be restored in a timely manner later, improving the efficiency of tracking the tracked object and avoiding low efficiency caused by the loss of the tracked object.

[0161] Specifically, for the sequence Θ = θ 1,t , θ 2,t , …, θ i,t , θ M,t of M tracked objects at time t, considering the historical tracking information of the tracked objects, the generation and elimination management process of each tracked object can be represented by a recursive Bayesian model, as Figure 9 shown as the recursive Bayesian schematic diagram of the tracking state of a certain tracked object.

[0162] Among them, θ t represents the specific state quantity (i.e., the first state quantity) of the tracked object concerned at time t. This state quantity can be understood as a latent variable, that is, it cannot be directly observed and can be indirectly inferred through the directly observable quantities O 1,t = [O1, O2…, O t , where O t = o t,1 , o t,2 , …, o t,N to indirectly infer this latent variable and estimate the state of the tracked object.

[0163] For example, the latent variable can include a variable x used to represent whether the tracked object can be generated i,t and a variable y used to represent whether the tracked object can disappear. i,t .

[0164] The following will introduce the specific situations of generation management and disappearance management respectively.

[0165] 1. Target generation management model

[0166] To determine whether to output the tracked object according to the observable quantity, it is necessary to calculate the posterior probability (i.e., the first posterior probability) about the latent variable x i,t . According to the recursive Bayesian posterior estimation formula, the posterior probability of whether the tracked object is generated at time t is:

[0167]

[0168] Among them, C is the normalization coefficient, k ∈ {0, 1}, and when k takes the value of 0, it means not to generate (i.e., not to output) the tracked object, and 1 means that the tracked object needs to be generated. is the likelihood estimation, and:

[0169]

[0170] Therefore, to determine whether to output the tracked object, the following ratio is calculated here:

[0171]

[0172] If the ratio r is greater than 1, it is considered that the tracked object truly exists, that is, the tracked object can be output for downstream tasks. For the convenience of calculation, the above formula is converted into the logarithmic function form here:

[0173]

[0174] At this time, the original formula changes from a multiplication formula to an accumulative iterative formula, greatly reducing the computational complexity. Therefore, the key is to calculate the observation likelihood ratio.

[0175] For this reason, the embodiment of the present application establishes an observation model for the management of target generation. For target generation (i.e., outputting a tracking object), the recall rate and detection accuracy of the detector have a greater impact on target generation. If the recall rate and accuracy of the detector used are higher, it means that if the detector detects an object, the probability that the object is output as a real tracking object is greater. And during the tracking process of the tracking object, the tracking object may appear in different orientations relative to the acquisition device, and there may also be differences in the performance of the detector. Therefore, the observation model established in this embodiment needs to combine the recall rate and accuracy of the detector.

[0176] Here, the detection accuracy queried from the detector performance map is used as an important observable for judging target generation, and we can obtain:

[0177]

[0178]

[0179]

[0180]

[0181] Among them, R(u, v) and P(u, v) respectively represent the recall rate and accuracy value queried from the detector performance map according to the region where the observable matching the tracking object is located. Thus, the observation likelihood ratio is obtained as:

[0182]

[0183]

[0184] represents the positive gain, usually a positive value, that is, increasing the confidence of target generation. In the foregoing Figure 7 when performing positive gain on the target generation management model, the foregoing is represents the negative gain. When performing negative gain on the target generation management model, the foregoing can be selected Usually a negative value, indicating a decrease in the confidence of generating a tracking formation.

[0185] Thus, the decision condition for the final generation of the tracking object can be obtained as:

[0186]

[0187] If the current frame is the first frame in which a tracking object is detected, then it can be set That is, the first frame in which a tracking object is detected does not output the tracking object by default to avoid false detection.

[0188] Therefore, in the embodiments of the present application, the detection accuracy of the detector is combined to determine whether to output the tracking object, and different gain directions are adaptively selected for different situations, so as to achieve more accurate tracking of the tracking object, reduce the duration of delayed output, and obtain more accurate output results more efficiently.

[0189] 2. Target extinction management model

[0190] Similar to the target generation management model, the purpose of target extinction is to maintain the number of internal tracking object sequences, timely delete misdetected tracking objects, avoid maintaining a tracking object sequence with an increasing number over time, and eliminating the updated tracking object sequence after management can also reduce the association uncertainty when associating tracking objects and detection objects.

[0191] When managing target extinction, in order to determine whether a certain tracking object should become extinct, it is necessary to calculate the posterior probability about the latent variable y i,t through the observed quantity, and the decision condition for tracking object extinction can also be established through recursive Bayesian inference:

[0192]

[0193] Similar to the aforementioned target generation management, an observation model can also be established here to calculate the likelihood ratio for whether a tracking object should become extinct. For the target extinction process, the recall rate and detection accuracy of the detector have a greater impact. If the recall rate and accuracy of the detector used are higher, it indicates that if the detector observes the tracking object, the probability of the target becoming extinct is smaller. Therefore, the observed likelihood ratio for judging target extinction can be obtained as:

[0194]

[0195]

[0196] represents the positive gain, usually a positive value, that is, increasing the target extinction confidence and accelerating the extinction of the tracking object. In the foregoing Figure 7 when performing positive gain on the target extinction management model, the foregoing is represents the negative gain. When performing negative gain on the target extinction management model, the foregoing can be selected

[0197] Therefore, in the embodiments of the present application, the detection accuracy corresponding to the tracking object can be extracted from the performance map of the detector. Thus, based on this accuracy, the confidence level for tracking the tracking object can be accurately calculated, improving the efficiency of outputting or eliminating the tracking object, efficiently outputting the tracking object to improve the efficiency of further processing based on the tracking object, or, invalid tracking objects can be eliminated in a timely manner, reducing the workload required for tracking invalid tracking objects and improving the working efficiency of the device. It can be understood that the present application can avoid manually setting the generation and elimination fixed thresholds through the Bayesian estimation method, improving the efficiency and accuracy of outputting or eliminating tracking objects. At the same time, it can also adapt to the problem of the tracking object being occluded, outputting the optimal sequence of tracking objects, and effectively solving the problem that it is difficult to balance the detector performance and the tracking output performance by setting various thresholds in the common tracking management solutions.

[0198] The process of the target tracking method provided by the present application is introduced above. Next, the device for executing the foregoing method will be introduced in detail.

[0199] First, refer to Figure 10 , the structural schematic diagram of a target tracking device provided by the present application is as follows.

[0200] The target tracking device includes:

[0201] A detection module 1001, configured to obtain at least one detected object in the current frame through a detector, where the current frame is any frame in the input data;

[0202] A tracking module 1002, configured to obtain a tracking object, where the tracking object includes an object detected by the detector in the previous frame of the current frame;

[0203] A matching module 1003, configured to match the tracking object with at least one detected object to obtain a matching result;

[0204] The tracking module 1002 is further configured to determine a first state quantity of the tracking object according to the matching result and the performance map of the detector, where the first state quantity is used to indicate whether to output the tracking object or whether to eliminate the tracking object, and the performance map includes the detection accuracy in multiple grids within the detection range of the detector.

[0205] In a possible implementation manner, the tracking module 1002 is specifically configured to: determine the position information of the tracking object in the current frame according to the matching result; query the detection accuracy corresponding to the tracking object in the performance map based on the position information of the tracking object in the current frame; calculate the first state quantity according to the detection accuracy.

[0206] In a possible implementation, the tracking module 1002 is specifically configured to: if there is no detected object matching the tracking object among at least one detected object, determine the position information of the tracking object in the current frame according to the motion state information of the tracking object; if the tracking object matches the first detected object among at least one detected object, use the position information of the first detected object as the position information of the tracking object in the current frame.

[0207] In a possible implementation, the tracking module 1002 is specifically configured to: obtain the predicted position of the tracking object in the current frame according to the motion state information of the tracking object; calculate the predicted distance value between the tracking object and the acquisition device that acquires the current frame based on the predicted position; obtain the actual distance value between the predicted position and the acquisition device in the current frame through input data; if the difference between the predicted distance value and the actual distance value is greater than the first threshold, use the predicted position as the position information of the tracking object in the current frame.

[0208] In a possible implementation, the first state quantity includes a first output indication state quantity and a first extinction indication state quantity. The first output indication state quantity is used to indicate whether to output the tracking object in the current frame, and the first extinction indication state quantity is used to indicate whether to extinguish the tracking object in the current frame. The tracking module 1002 is specifically configured to: obtain the direct observation quantity, where the direct observation quantity is used to represent the motion state of the tracking object; calculate the first posterior probability and the second posterior probability of the tracking object in the current frame according to the detection accuracy and the direct observation quantity. The first posterior probability is used to represent the probability of outputting the tracking object in the current frame, and the second posterior probability is used to represent the probability of extinguishing the tracking object in the current frame; obtain the first output indication state quantity based on the first posterior probability, and obtain the first extinction indication state quantity based on the second posterior probability.

[0209] In a possible implementation, the tracking module 1002 is specifically configured to: obtain the second output indication state quantity and the second extinction indication state quantity of the tracking object in the previous frame. The first state quantity includes the first output indication state quantity and the first extinction indication state quantity. The second output indication state quantity is used to indicate whether the tracking object was output in the previous frame, and the second extinction indication state quantity is used to indicate whether the tracking object was extinguished in the previous frame; fuse the first posterior probability and the second output indication state quantity to obtain the first output indication state quantity, and fuse the second posterior probability and the second extinction indication state quantity to obtain the first extinction indication state quantity.

[0210] In a possible implementation, the detection accuracy of each grid in the performance map includes the accuracies corresponding to multiple categories. The tracking module 1002 is specifically configured to: obtain the category of the tracking object; query the detection accuracy corresponding to the tracking object in the performance map based on the position information of the tracking object in the current frame and the category of the tracking object.

[0211] In a possible implementation, the tracking module 1002 is further configured to: if at least one detected object includes a second detected object that does not match the tracking object, use the second detected object as a new tracking object and track the new tracking object in the next frame of the current frame.

[0212] In a possible implementation, the target tracking device further includes an encoding module 1004, configured to divide the detection range of the detector into multiple grids before obtaining at least one detected object in the current frame through the detector; obtain ground truth data, where the ground truth data includes the acquisition data collected by the acquisition device and the information of the corresponding ground truth object; use the detector to detect the acquisition data to obtain the information of the predicted object; and calculate the detection accuracy corresponding to each grid in the multiple grids according to the predicted object and the ground truth object to obtain a performance map.

[0213] Please refer to Figure 11 , a schematic structural diagram of another target tracking device provided by this application is described as follows.

[0214] The target tracking device may include a processor 1101 and a memory 1102. The processor 1101 and the memory 1102 are interconnected by a line. Among them, program instructions and data are stored in the memory 1102.

[0215] The program instructions and data corresponding to the steps described above are stored in the memory 1102. Figures 4 - 9

[0216] The processor 1101 is configured to execute the method steps performed by the target tracking device shown in any of the foregoing Figures 4 - 9 embodiments.

[0217] Optionally, the target tracking device may further include a transceiver 1103, configured to receive or transmit data.

[0218] In an embodiment of this application, a computer-readable storage medium is further provided. A program for generating the vehicle driving speed is stored in the computer-readable storage medium. When it runs on a computer, it causes the computer to execute the steps in the method described in the foregoing Figures 4 - 9 embodiments.

[0219] Optionally, the foregoing Figure 11 target tracking device is a chip.

[0220] In an embodiment of this application, a target tracking device is further provided. The target tracking device may also be referred to as a digital processing chip or a chip. The chip includes a processing unit and a communication interface. The processing unit obtains program instructions through the communication interface, and the program instructions are executed by the processing unit. The processing unit is configured to execute the foregoing Figures 4 - 9The method steps executed by the target tracking device shown in any of the embodiments.

[0221] An embodiment of the present application further provides a digital processing chip. Circuits for implementing the foregoing processor 1101 or the functions of processor 1101 and one or more interfaces are integrated in the digital processing chip. When a memory is integrated in the digital processing chip, the digital processing chip can complete the method steps of any one or more of the foregoing embodiments. When a memory is not integrated in the digital processing chip, it can be connected to an external memory through a communication interface. The digital processing chip implements the actions executed by the target tracking device in the foregoing embodiment according to the program code stored in the external memory.

[0222] An embodiment of the present application also provides a computer program product, which, when running on a computer, causes the computer to execute the steps executed by the target tracking device in the method described in the foregoing Figures 4 - 9 embodiment shown.

[0223] The target tracking device provided in the embodiment of the present application may be a chip, and the chip includes: a processing unit and a communication unit. The processing unit may be a processor, for example, and the communication unit may be an input / output interface, a pin, a circuit, etc. The processing unit can execute the computer execution instructions stored in the storage unit to cause the chip in the server to execute the Figures 4 - 9 neural network training method described in the foregoing embodiment shown. Optionally, the storage unit is a storage unit inside the chip, such as a register, a cache, etc., and the storage unit may also be a storage unit outside the chip in the radio access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0224] Specifically, the aforementioned processing unit or processor may be a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0225] Exemplarily, please refer to Figure 12 , Figure 12 which is a schematic structural diagram of a chip provided by an embodiment of the present application. The chip may be embodied as a neural network processor NPU 120. The NPU 120 is mounted on a main CPU (Host CPU) as a coprocessor, and tasks are allocated by the Host CPU. The core part of the NPU is an arithmetic circuit 1203, and the arithmetic circuit 1203 is controlled by a controller 1204 to extract matrix data from a memory and perform a multiplication operation.

[0226] In some implementations, the arithmetic circuit 1203 internally includes a plurality of processing units (process engine, PE). In some implementations, the arithmetic circuit 1203 is a two-dimensional systolic array. The arithmetic circuit 1203 may also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1203 is a general matrix processor.

[0227] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory 1202 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory 1201 and performs a matrix operation with matrix B, and the partial result or the final result of the obtained matrix is stored in an accumulator 1208.

[0228] The unified memory 1206 is used to store input data and output data. The weight data directly passes through the direct memory access controller (DMAC) 1205, and the DMAC transfers it to the weight memory 1202. The input data is also transferred to the unified memory 1206 through the DMAC.

[0229] The bus interface unit (BIU) 1210 is used for the interaction between the AXI bus, the DMAC, and the instruction fetch buffer (IFB) 1209.

[0230] The bus interface unit 1210 (bus interface unit, BIU) is used for the instruction fetch buffer 1209 to obtain instructions from the external memory, and also for the storage unit access controller 1205 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0231] The DMAC is mainly used to transfer the input data in the external memory DDR to the unified memory 1206, transfer the weight data to the weight memory 1202, or transfer the input data to the input memory 1201.

[0232] The vector calculation unit 1207 includes multiple arithmetic processing units. When necessary, it further processes the output of the arithmetic circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for neural network non-convolution / full connection layer network calculations, such as batch normalization, pixel-level summation, upsampling of the feature plane, etc.

[0233] In some implementations, the vector calculation unit 1207 can store the processed output vector in the unified memory 1206. For example, the vector calculation unit 1207 can apply a linear function and / or a non-linear function to the output of the arithmetic circuit 1203, such as linearly interpolating the feature plane extracted by the convolutional layer, or for example, the vector of the accumulated value, to generate the activation value. In some implementations, the vector calculation unit 1207 generates normalized values, pixel-level summation values, or both. In some implementations, the processed output vector can be used as the activation input to the arithmetic circuit 1203, for example, for use in subsequent layers in the neural network.

[0234] The instruction fetch buffer 1209 connected to the controller 1204 is used to store the instructions used by the controller 1204;

[0235] The unified memory 1206, the input memory 1201, the weight memory 1202, and the fetch memory 1209 are all On-Chip memories. The external memory is private to the NPU hardware architecture.

[0236] Among them, the operations of each layer in the recurrent neural network can be executed by the operation circuit 1203 or the vector calculation unit 1207.

[0237] Among them, the processor mentioned anywhere above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the above-mentioned Figures 4 - 9 method.

[0238] In addition, it should be noted that the device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.

[0239] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits, or dedicated circuits. However, for this application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, a USB flash drive, a mobile hard disk, a read only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc of a computer, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of this application.

[0240] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0241] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0242] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims, and drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0243] Finally, it should be noted that the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all of them should be covered by the protection scope of the present application.

Claims

1. A target tracking method, characterized in that, Including: Obtaining at least one detected object in the current frame through a detector, where the current frame is any frame in the input data; Obtaining a tracked object, where the tracked object includes an object detected by the detector in the previous frame of the current frame; Matching the tracked object and the at least one detected object to obtain a matching result; Determining a first state quantity of the tracked object according to the matching result and a performance map of the detector, where the first state quantity is used to indicate whether to output the tracked object or whether to eliminate the tracked object, and the performance map includes detection accuracies within a plurality of grids in the detection range of the detector; The determining the first state quantity of the tracked object according to the matching result and the performance map of the detector includes: Determining the position information of the tracked object in the current frame according to the matching result; Querying the detection accuracy corresponding to the tracked object in the performance map based on the position information of the tracked object in the current frame; Calculating the first state quantity according to the detection accuracy.

2. The method according to claim 1, wherein The determining the position information of the tracked object in the current frame according to the matching result includes: If there is no detected object in the at least one detected object that matches the tracked object, determining the position information of the tracked object in the current frame according to the motion state information of the tracked object; If the tracked object matches a first detected object among the at least one detected object, using the position information of the first detected object as the position information of the tracked object in the current frame.

3. The method according to claim 2, wherein The determining the position information of the tracked object in the current frame according to the motion state information of the tracked object includes: Obtaining a predicted position of the tracked object in the current frame according to the motion state information of the tracked object; Calculating a predicted distance value between the tracked object in the current frame and the acquisition device that acquires the current frame according to the predicted position; Obtaining an actual distance value between the predicted position and the acquisition device in the current frame through the input data; If the difference between the predicted distance value and the actual distance value is greater than a first threshold, using the predicted position as the position information of the tracked object in the current frame.

4. The method according to any one of claims 1 to 3, characterized in that The first state quantity includes a first output indication state quantity and a first elimination indication state quantity, where the first output indication state quantity is used to indicate whether to output the tracked object in the current frame, and the first elimination indication state quantity is used to indicate whether to eliminate the tracked object in the current frame; The calculating the first state quantity according to the detection accuracy includes: Obtaining a direct observation quantity, where the direct observation quantity is used to represent the motion state of the tracked object; Calculating a first posterior probability and a second posterior probability of the tracked object in the current frame according to the detection accuracy and the direct observation quantity, where the first posterior probability is used to represent the probability of outputting the tracked object in the current frame, and the second posterior probability is used to represent the probability of eliminating the tracked object in the current frame; Obtain the first output indication state quantity based on the first posterior probability, and obtain the first extinction indication state quantity based on the second posterior probability.

5. The method according to claim 4, wherein The obtaining the first output indication state quantity based on the first posterior probability, and obtaining the first extinction indication state quantity based on the second posterior probability includes: Obtain the second output indication state quantity and the second extinction indication state quantity of the tracking object in the previous frame. The first state quantity includes the first output indication state quantity and the first extinction indication state quantity. The second output indication state quantity is used to indicate whether the tracking object is output in the previous frame, and the second extinction indication state quantity is used to indicate whether the tracking object is extinct in the previous frame; Fuse the first posterior probability and the second output indication state quantity to obtain the first output indication state quantity, and fuse the second posterior probability and the second extinction indication state quantity to obtain the first extinction indication state quantity.

6. The method according to any one of claims 1-3, characterized in that, The detection accuracy of each grid in the performance map includes the accuracies corresponding to multiple categories; The querying the detection accuracy corresponding to the tracking object in the performance map based on the position information of the tracking object in the current frame includes: Obtain the category of the tracking object; Based on the position information of the tracking object in the current frame and the category of the tracking object, query the detection accuracy corresponding to the tracking object in the performance map.

7. The method according to any one of claims 1 to 3, characterized in that, The method further includes: If the at least one detected object includes a second detected object that does not match the tracking object, then use the second detected object as a new tracking object, and track the new tracking object in the next frame of the current frame.

8. The method according to any one of claims 1 to 3, characterized in that Before obtaining the at least one detected object in the current frame through the detector, the method includes: Divide the detection range of the detector into multiple grids; Obtain ground truth data, where the ground truth data includes the acquisition data collected by the acquisition device and the information of the corresponding ground truth object; Use the detector to detect the acquisition data to obtain the information of the predicted object; According to the predicted object and the ground truth object, calculate the detection accuracy corresponding to each grid in the multiple grids to obtain the performance map.

9. A target tracking device, characterized in that, Includes: A detection module, configured to obtain at least one detected object in the current frame through a detector, where the current frame is any frame in the input data; A tracking module, configured to obtain a tracking object, where the tracking object includes an object detected by the detector in the previous frame of the current frame; A matching module, configured to match the tracking object and the at least one detected object to obtain a matching result; The tracking module is further configured to determine a first state quantity of the tracking object according to the matching result and the performance map of the detector. The first state quantity is used to indicate whether to output the tracking object or whether to extinct the tracking object. The performance map includes the detection accuracies in multiple grids within the detection range of the detector; The tracking module is specifically configured to: Determine the position information of the tracking object in the current frame according to the matching result; Query the detection accuracy corresponding to the tracked object in the performance map based on the position information of the tracked object in the current frame; Calculate the first state quantity according to the detection accuracy.

10. The device according to claim 9, characterized in that, The tracking module is specifically configured to: If there is no detected object matching the tracked object among the at least one detected object, determine the position information of the tracked object in the current frame according to the motion state information of the tracked object; If the tracked object matches the first detected object among the at least one detected object, use the position information of the first detected object as the position information of the tracked object in the current frame.

11. The device according to claim 10, wherein, The tracking module is specifically configured to: Obtain the predicted position of the tracked object in the current frame according to the motion state information of the tracked object; Calculate the predicted distance value between the tracked object and the acquisition device that acquires the current frame according to the predicted position; Obtain the actual distance value between the predicted position and the acquisition device in the current frame through the input data; If the difference between the predicted distance value and the actual distance value is greater than the first threshold, use the predicted position as the position information of the tracked object in the current frame.

12. The device according to any one of claims 9-11, characterized in that, The first state quantity includes a first output indication state quantity and a first extinction indication state quantity. The first output indication state quantity is used to indicate whether to output the tracked object in the current frame, and the first extinction indication state quantity is used to indicate whether to extinguish the tracked object in the current frame; The tracking module is specifically configured to: Obtain a direct observation quantity, where the direct observation quantity is used to represent the motion state of the tracked object; Calculate the first posterior probability and the second posterior probability of the tracked object in the current frame according to the detection accuracy and the direct observation quantity. The first posterior probability is used to represent the probability of outputting the tracked object in the current frame, and the second posterior probability is used to represent the probability of extinguishing the tracked object in the current frame; Obtain the first output indication state quantity based on the first posterior probability, and obtain the first extinction indication state quantity based on the second posterior probability.

13. The device according to claim 12, characterized in that The tracking module is specifically configured to: Obtain the second output indication state quantity and the second extinction indication state quantity of the tracked object in the previous frame. The first state quantity includes a first output indication state quantity and a first extinction indication state quantity. The second output indication state quantity is used to indicate whether to output the tracked object in the previous frame, and the second extinction indication state quantity is used to indicate whether to extinguish the tracked object in the previous frame; Fuse the first posterior probability and the second output indication state quantity to obtain the first output indication state quantity, and fuse the second posterior probability and the second extinction indication state quantity to obtain the first extinction indication state quantity.

14. The device according to any one of claims 9-11, characterized in that, The detection accuracy of each grid in the performance map includes accuracies corresponding to multiple categories; The tracking module is specifically configured to: Obtain the category of the tracked object; Query the detection accuracy corresponding to the tracked object in the performance map based on the position information of the tracked object in the current frame and the category of the tracked object.

15. The device according to any one of claims 9-11, characterized in that, The tracking module is further configured to: If the at least one detected object includes a second detected object that does not match the tracked object, use the second detected object as a new tracked object and track the new tracked object in the next frame of the current frame.

16. The device according to any one of claims 9-11, characterized in that, The apparatus further includes an encoding module, configured to: Before obtaining at least one detected object in the current frame through the detector, divide the detection range of the detector into a plurality of grids; Obtain ground truth data, where the ground truth data includes acquisition data collected by an acquisition device and information about corresponding ground truth objects; Use the detector to detect the acquisition data to obtain information about predicted objects; Calculate the detection accuracy corresponding to each grid in the plurality of grids according to the predicted object and the ground truth object to obtain the performance map.

17. A target tracking device, characterized in that, It includes a processor, the processor is coupled to a memory, and the memory stores a program, and when the program instructions stored in the memory are executed by the processor, the method according to any one of claims 1 to 8 is implemented.

18. A computer-readable storage medium includes a program, and when it is executed by a processing unit, it executes the method according to any one of claims 1 to 8.

19. A target tracking device, characterized in that, It includes a processing unit and a communication interface, the processing unit obtains program instructions through the communication interface, and when the program instructions are executed by the processing unit, the method according to any one of claims 1 to 8 is implemented.

20. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • A method and apparatus for target identification

    CN109472809A

  • Lane line detection method and device, electronic equipment and storage medium

    CN110232368A

  • Multi-target tracking method

    CN111127513A