Object Tracking Method, Electronic Device and Storage Medium
Through the first tracker and Kalman filter predicting candidate detection boxes, combined with the decision network of the graph attention mechanism, the path loss problem caused by occlusion in multi-object tracking is solved, and the accuracy and completeness of the tracking path is improved.
Patent Information
- Application Number
- CN202210232982.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-09
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-03-09
AI Technical Summary
Existing multi-object tracking algorithms tend to lose objects when objects are blocked, affecting the accuracy of path generation.
The initial tracking result is obtained by the first tracker, a candidate detection box in the target video frame is predicted using a Kalman filter, and a decision network based on the graph attention mechanism is used to determine whether the candidate detection box belongs to the motion path of the object.
Effectively repair fragmented paths in complex occlusion scenarios, improving object discrimination capabilities and tracking path accuracy and completeness.
Smart Images

Figure CN114758266B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technologies, and in particular, to an object tracking method, an electronic device, and a storage medium. Background Art
[0002] With the development of object detection technologies, object detection results are becoming more and more reliable, enabling extensive research on multi-object tracking algorithms based on object detection. Multi-object tracking is to simultaneously locate multiple objects of interest in a given video, maintain the identities of these objects, and record the paths of these objects.
[0003] The process of a multi-object tracking framework based on object detection is as follows: First, use an offline-trained detector to detect objects frame by frame, use a similarity matching method to associate the detected objects, and then continuously match the generated paths with the detection results to generate more reliable paths.
[0004] However, when an object is occluded, the online multi-object tracking algorithm cannot handle the negative impact brought by the detector, easily loses the object, and thus affects the accuracy of path generation. Summary of the Invention
[0005] In view of the above problems, embodiments of the present application are proposed to provide an object tracking method, an electronic device, and a storage medium that overcome the above problems or at least partially solve the above problems.
[0006] According to a first aspect of the embodiments of the present application, an object tracking method is provided, including:
[0007] Tracking the motion paths of each object in a video to be tracked through a first tracker to obtain initial tracking results of each object, where the motion path of the object is represented by a target detection box of the object in each video frame of the video to be tracked;
[0008] For each object, determining a target video frame in which the object is not detected according to the initial tracking result of the object;
[0009] Predicting a candidate detection box of the object in the target video frame through a Kalman filter;
[0010] Determining whether the candidate detection box belongs to the motion path of the object through a decision network based on a graph attention mechanism.
[0011] According to a second aspect of the embodiments of the present application, an object tracking device is provided, including:
[0012] A target tracking module, which is used to track the movement paths of each object in the video to be tracked through a first tracker, and obtain the initial tracking results of each object, where the movement path of the object is characterized by the target detection box of the object in each video frame of the video to be tracked;
[0013] A target video frame determination module, which is used to determine, for each object, the target video frames in which the object is not detected according to the initial tracking result of the object;
[0014] A candidate detection box prediction module, which is used to predict the candidate detection box of the object in the target video frame through a Kalman filter;
[0015] A decision module, which is used to determine whether the candidate detection box belongs to the movement path of the object through a decision network based on a graph attention mechanism.
[0016] According to the third aspect of the embodiments of the present application, an electronic device is provided, including: a processor, a memory, and a computer program stored on the memory and executable on the processor, where when the computer program is executed by the processor, the object tracking method described in the first aspect is implemented.
[0017] According to the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the object tracking method described in the first aspect is implemented.
[0018] According to the fifth aspect of the embodiments of the present application, a computer program product is provided, including a computer program or computer instructions, where when the computer program or computer instructions are executed by a processor, the object tracking method described in the first aspect is implemented.
[0019] The object tracking method, electronic device, and storage medium provided by the embodiments of the present application track the movement paths of each object in the video to be tracked through a first tracker, obtain the initial tracking results of each object, determine, for each object, the target video frames in which the object is not detected according to the initial tracking result of the object, predict the candidate detection box of the object in the target video frame through a Kalman filter, and determine whether the candidate detection box belongs to the movement path of the object through a decision network based on a graph attention mechanism. Since the candidate detection box of the object in the target video frame is predicted through a Kalman filter, and a decision network based on a graph attention mechanism is used to decide whether the candidate detection box belongs to the movement path of the object, fragmented paths can be effectively repaired, the discrimination ability of objects in complex occlusion scenarios is improved, and thus the accuracy and integrity of the tracking path are improved.
[0020] The above description is only an overview of the technical solution of this application. In order to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of this application more obvious and understandable, the specific embodiments of this application are specifically exemplified below. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of this application.
[0022] Figure 1 is a flowchart of the steps of an object tracking method provided by an embodiment of this application;
[0023] Figure 2 is a structural block diagram of an object tracking device provided by an embodiment of this application;
[0024] Figure 3 is a structural block diagram of an electronic device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The exemplary embodiments of this application will be described in more detail below with reference to the drawings. Although the exemplary embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that this application can be understood more thoroughly and the scope of this application can be fully conveyed to those skilled in the art.
[0026] In recent years, important progress has been made in the research of technologies such as computer vision, deep learning, machine learning, image processing, and image recognition based on artificial intelligence. Artificial Intelligence (AI) is a new science and technology that studies and develops theories, methods, technologies, and application systems for simulating and extending human intelligence. The discipline of artificial intelligence is a comprehensive discipline that involves many technical categories such as chips, big data, cloud computing, the Internet of Things, distributed storage, deep learning, machine learning, and neural networks. As an important branch of artificial intelligence, computer vision specifically enables machines to recognize the world. Computer vision technologies usually include face recognition, liveness detection, fingerprint recognition and anti-counterfeiting verification, biometric recognition, face detection, pedestrian detection, object detection, pedestrian recognition, image processing, image recognition, image semantic understanding, image retrieval, text recognition, video processing, video content recognition, 3D reconstruction, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), computational photography, robot navigation and positioning, etc. With the research and progress of artificial intelligence technology, this technology has been applied in many fields, such as security prevention and control, urban management, traffic management, building management, park management, face access, face attendance, logistics management, warehouse management, robots, intelligent marketing, computational photography, mobile phone imaging, cloud services, smart home, wearable devices, driverless, autonomous driving, intelligent healthcare, face payment, face unlocking, fingerprint unlocking, person-certificate verification, smart screen, smart TV, cameras, mobile Internet, webcasting, beauty, makeup, medical beauty, intelligent temperature measurement, etc. The embodiments of this application also relate to computer vision technology, specifically an object tracking method. In a multi-object tracking scenario, for a scenario where objects are frequently occluded, an accurate and complete tracking path can be effectively given. The specific solution is as follows:
[0027] Figure 1 is a step flowchart of a target tracking method provided by an embodiment of this application, as Figure 1 shown, the method may include:
[0028] Step 101, track the motion paths of each object in the video to be tracked through a first tracker, and obtain the initial tracking results of each of the objects, where the motion path of the object is represented by the target detection box of the object in each video frame of the video to be tracked.
[0029] Among them, the first tracker can be any multi-object tracker, for example, it can be a Re-ID tracker.
[0030] Input the video to be tracked into the first tracker. The first tracker tracks the movement paths of various objects in the video to be tracked, obtains the output result of the first tracker, and gets the initial tracking results of various objects in the video to be tracked. If an object is occluded by other objects in a video frame of the video to be tracked, the movement path of the object in the initial tracking results may not be tracked in this video frame. Embodiments of the present application can accurately predict the movement path of the object in this case.
[0031] Step 102, for each object, determine the target video frames in which the object is not detected according to the initial tracking results of the object.
[0032] For each object, the identifier (ID) corresponding to the target detection box of the object can be given in the initial tracking results. If the target detection box of the object is missing in a video frame, it is determined that the movement path of the object is not tracked in this video frame, and this video frame is determined as the target video frame, and subsequent path prediction is performed on this target video frame.
[0033] Step 103, predict the candidate detection box of the object in the target video frame through a Kalman filter.
[0034] Input the target detection boxes of the same object in the video frames before the target video frame in the video to be tracked into the Kalman filter, and predict the candidate detection box of the movement path of the object in the target video frame through the Kalman filter.
[0035] Step 104, determine whether the candidate detection box belongs to the movement path of the object through a decision network based on a graph attention mechanism.
[0036] Obtain the features of the candidate detection box, and obtain the features of the target detection box of the same object in the video frames before the target video frame, and input the features into the decision network based on the graph attention mechanism. The decision network processes the features to obtain output features, compares the output features with the features of the candidate detection box, and determines whether the candidate detection box belongs to the movement path of the object.
[0037] In one embodiment of the present application, determining whether the candidate detection box belongs to the motion path of the object by the decision network based on the graph attention mechanism includes: determining the spatio-temporal features of the candidate detection box according to the candidate detection box and the target detection box of the object in the adjacent video frames of the target video frame; wherein the spatio-temporal features include spatial features and temporal features; inputting the spatio-temporal features of the target detection boxes corresponding to the object in each video frame within the sliding time window into the decision network based on the graph attention mechanism to obtain the predicted features of the candidate detection box; and determining whether the candidate detection box belongs to the motion path of the object according to the spatio-temporal features and the predicted features of the candidate detection box.
[0038] Encode the occlusion situation of the candidate detection box by the target detection box of other objects as spatial features, and encode the comparison features between the candidate detection box and the target detection box of the same object in adjacent frames as temporal features, and splice the temporal features and spatial features to obtain the spatio-temporal features of the candidate detection box. Wherein, the adjacent frames may be the video frames with a preset number of frames before the target video frame.
[0039] Determine each video frame within the sliding time window before the target video frame, and determine the spatio-temporal features of the target detection boxes corresponding to the object in each video frame within the sliding time window according to the determination method of the spatio-temporal features, and input the spatio-temporal features of the target detection boxes corresponding to the object in each video frame within the sliding time window into the decision network based on the graph attention mechanism, and process each spatio-temporal feature through the decision network based on the graph attention mechanism to predict the features of the same object in the target video frame to obtain the predicted features of the candidate detection box. Compare the spatio-temporal features and the predicted features of the candidate detection box to determine whether the candidate detection box belongs to the motion path of the object. When the distance between the spatio-temporal features and the predicted features of the candidate detection box is small, it can be determined that the candidate detection box belongs to the motion path of the object. When the distance between the spatio-temporal encoded features and the predicted features of the candidate detection box is large, it can be determined that the candidate detection box does not belong to the motion path of the object.
[0040] By using the decision network to predict the features that the candidate detection box should have in the target video frame according to the spatio-temporal features of the target detection boxes corresponding to the object in each video frame within the sliding time window, the predicted features of the candidate detection box are obtained. Based on the spatio-temporal features and the predicted features of the candidate detection box, it is determined whether the candidate detection box belongs to the motion path of the object, which can accurately repair the missed detection caused by occlusion to optimize the tracking path, thereby improving the robustness of multi-object tracking in complex occlusion scenarios.
[0041] In one embodiment of the present application, determining whether the candidate detection box belongs to the motion path of the object according to the spatio-temporal feature and the prediction feature of the candidate detection box includes: determining the matching degree between the spatio-temporal feature of the candidate detection box and the prediction feature; if the matching degree is less than the matching degree threshold, determining that the candidate detection box does not belong to the motion path of the object; if the matching degree is greater than or equal to the matching degree threshold, determining that the candidate detection box belongs to the motion path of the object.
[0042] The matching degree between the spatio-temporal feature and the prediction feature of the candidate detection box can be characterized by the distance or similarity between the spatio-temporal feature of the candidate detection box and the prediction feature. Taking the distance between the spatio-temporal feature of the candidate detection box and the prediction feature as an example to characterize the matching degree, the greater the distance, the smaller the matching degree, and the smaller the distance, the greater the matching degree. Calculate the distance between the spatio-temporal feature of the candidate detection box and the prediction feature, and compare the distance with the distance threshold. If the distance between the spatio-temporal feature of the candidate detection box and the prediction feature is greater than the distance threshold, it means that the prediction feature is far from the spatio-temporal encoding feature of the candidate detection box, and it is determined that the candidate detection box does not belong to the motion path of the object; if the distance between the spatio-temporal feature of the candidate detection box and the prediction feature is less than or equal to the distance threshold, it means that the prediction feature is close to the spatio-temporal feature of the candidate detection box, and it is determined that the candidate detection box belongs to the motion path of the object. Among them, the distance can be, for example, Euclidean distance, Manhattan distance, etc.
[0043] By using the matching degree based on the spatio-temporal feature and the prediction feature of the candidate detection box to determine whether the candidate detection box belongs to the motion path of the object, the accuracy of the decision result can be improved, and further the accuracy of the final path tracking result can be improved.
[0044] In one embodiment of the present application, the spatial feature of the candidate detection box is determined through the following process: determining the confidence of the candidate detection box, and determining the occlusion ratio of the candidate detection box according to the initial tracking result; determining the spatial feature of the candidate detection box according to the confidence and the occlusion ratio.
[0045] It is determined that the confidence of the candidate detection box predicted by the Kalman filter is 0, and the confidence of the target detection box tracked in the initial tracking result can be directly obtained by the first tracker. According to the initial tracking results of each object, determine the target detection boxes of other objects in the current target video frame, and based on the position of the candidate detection box and the positions of the target detection boxes of other objects, determine the ratio of the candidate detection box being occluded by the target detection boxes of other objects to obtain the occlusion ratio of the candidate detection box. According to the confidence and the occlusion ratio, the spatial feature of the candidate detection box is Among them, represents the confidence of the candidate detection box, Indicates the occlusion ratio of the candidate detection box, where \(i\) represents that the object corresponding to the candidate detection box is the \(i\)-th object, and \(t\) represents that the target video frame is the \(t\)-th frame in the video to be tracked.
[0046] In an embodiment of the present application, the temporal feature of the candidate detection box is determined through the following process: determining the intersection over union (IoU) between the candidate detection box and the target detection box of the same object in the previous frame of the target video frame, determining the similarity between the candidate detection box and the target appearance feature of the target detection box of the same object in the previous frame of the target video frame, and determining the center point coordinate offset between the candidate detection box and the target detection box of the same object in the previous frame of the target video frame; determining the temporal feature of the candidate detection box according to the IoU, the similarity, and the center point coordinate offset.
[0047] Calculate the intersection between the candidate detection box and the target detection box of the same object in the previous frame of the target video frame, and calculate the union of the candidate detection box and the target detection box of the same object in the previous frame of the target video frame, and determine the ratio of the intersection to the union as the IoU; input the candidate detection box into the ReID model to obtain the target appearance (ReID) feature of the candidate detection box, and input the target detection box of the same object in the previous frame of the target video frame into the ReID model to obtain the target appearance feature of the target detection box in the previous frame, and calculate the similarity between the target appearance feature of the candidate detection box and the target appearance feature of the target detection box in the previous frame. The similarity can be, for example, cosine similarity; determine the center point coordinate of the candidate detection box, and determine the center point coordinate of the target detection box of the same object in the previous frame of the target video frame, and calculate the distance between the center point coordinate of the candidate detection box and the center point coordinate of the target detection box in the previous frame to obtain the center point coordinate offset between the candidate detection box and the detection box of the same object in the previous frame of the target video frame; determine the temporal feature of the candidate detection box according to the IoU, similarity, and center point coordinate offset as Where, represents the IoU between the candidate detection box and the target detection box of the same object in the previous frame of the target video frame, represents the target appearance feature of the target detection box in the previous frame, represents the target appearance feature of the candidate detection box, represents the similarity between the candidate detection box and the target appearance feature of the target detection box of the same object in the previous frame, represents the center point coordinate offset between the candidate detection box and the target detection box of the same object in the previous frame.
[0048] Concatenate the spatial feature and the temporal feature of the candidate detection box to obtain the spatio-temporal feature of the candidate detection box as
[0049] By determining the spatial features and temporal features in the above manner, relatively accurate spatio-temporal features can be obtained.
[0050] In an embodiment of the present application, inputting the spatio-temporal features of the target detection boxes corresponding to each video frame of the object within the sliding time window into the decision network based on the graph attention mechanism to obtain the predicted features of the candidate detection boxes includes:
[0051] Inputting the spatio-temporal features of the target detection boxes corresponding to each video frame of the object within the sliding time window into the decision network based on the graph attention mechanism, and determining the initial embedding vectors corresponding to the spatio-temporal features of each target detection box through the decision network;
[0052] Processing the structural features between the initial embedding vectors through the decision network to obtain the final embedding vectors;
[0053] According to the embedding vectors and the spatio-temporal features of the target detection boxes, determining attention coefficients through the decision network, and determining the aggregated representation features corresponding to the spatio-temporal features of the target detection boxes according to the attention coefficients;
[0054] Processing the final embedding vectors and the aggregated representation features through the decision network to obtain the predicted features of the candidate detection boxes.
[0055] The decision network based on the graph attention mechanism may include an embedding representation layer, a graph structure learning layer, a graph attention layer, and a fully connected layer.
[0056] Represent the spatio-temporal features corresponding to the target detection boxes in each video frame of the object within the sliding time window as the following input features:
[0057]
[0058] where W represents the width of the sliding time window, represents the spatio-temporal features corresponding to the target detection boxes in each video frame of the object within the sliding time window.
[0059] Input the input features into a decision network based on the graph attention mechanism. The embedding representation layer (Embedding Vector) in the decision network sets an initial embedding vector for the spatio-temporal features corresponding to the target detection boxes in each video frame of the object within the sliding time window. Continuously learn the structural features between the initial embedding vectors through the graph structure learning layer to obtain the final embedding vector. Concatenate the final embedding vector with the linear transformation of the spatio-temporal features corresponding to each target detection box to obtain a new embedding vector. Concatenate the new embedding vectors between each network node in the graph attention layer to obtain a concatenated vector. Perform a dot product calculation between the concatenated vector and the learnable weight vector in the decision network to obtain a calculation result. Apply the LeakyReLU activation function to the calculation result to obtain the initial attention score. Normalize the initial attention score using the softmax function to obtain the attention coefficient. Aggregate and calculate the spatio-temporal features corresponding to the target detection boxes in each video frame of the object within the sliding time window according to the attention coefficient to obtain the aggregated representation features of the spatio-temporal features corresponding to the target detection boxes in each video frame of the object within the sliding time window at each network node in each graph attention layer. The graph attention layer calculates the aggregated representation features of each network node. The aggregated representation feature of node m is expressed as follows:
[0060]
[0061] where, σ represents the activation function, W represents the weight matrix of the decision network, α mm represents the self-attention coefficient, and α mn represents the attention coefficient of the nth node to the mth node, and L(m) represents the nodes other than the mth node.
[0062] Multiply the aggregated representation feature element-wise with the final embedding vector to obtain a multiplied vector, and obtain the prediction feature of the candidate detection box after processing the multiplied vector through the fully connected layer. The prediction feature is expressed as follows:
[0063]
[0064] where, f0 represents the processing of the fully connected layer, v1, v2, …, v L represents the final embedding vector, and L is the number of network nodes in the graph attention layer.
[0065] Process the spatio-temporal features corresponding to the target detection boxes in each video frame of the object through the decision network based on the graph attention mechanism to obtain the prediction features corresponding to the candidate detection boxes, which can improve the accuracy of determining the prediction features, and further improve the accuracy of the target tracking path.
[0066] The object tracking method provided in this embodiment tracks the motion paths of each object in the video to be tracked through a first tracker, obtains the initial tracking results of each object, for each object, determines the target video frames in which the object is not detected according to the initial tracking results of the object, predicts the candidate detection boxes of the object in the target video frames through a Kalman filter, and determines whether the candidate detection boxes belong to the motion path of the object through a decision network based on a graph attention mechanism. Since the Kalman filter is used to predict the candidate detection boxes of the object's motion path in the target video frames, and the decision network based on the graph attention mechanism is used to decide whether the candidate detection boxes belong to the object's motion path, fragmented paths can be effectively repaired, the discrimination ability of the target in complex occlusion scenarios is improved, and thus the accuracy and integrity of the tracking path are improved.
[0067] Based on the above technical solution, the decision network of the graph attention mechanism is trained through the following process:
[0068] Obtain unlabeled video samples;
[0069] Use a second tracker to track the motion paths of each object in the video sample to obtain the first tracking results of each object, and use a third tracker to track the motion paths of each object in the video sample to obtain the second tracking results of each object;
[0070] For each object, according to the first tracking result and the second tracking result, determine the overlapping paths of the paths obtained by the second tracker and the third tracker for tracking the object;
[0071] Train the decision network based on the graph attention mechanism through the overlapping paths to obtain a trained decision network.
[0072] Among them, the second tracker and the third tracker are two different object trackers. For example, the second tracker can be a Re-ID tracker, and the third tracker can be an IoU tracker. The first tracker can be one of the second tracker and the third tracker, or other trackers.
[0073] Obtain unlabeled video samples, use the second tracker and the third tracker to track the motion paths of each object in the video sample respectively to obtain the first tracking results and the second tracking results of each object, for each object, according to the first tracking result and the second tracking result, determine the overlapping paths of the paths obtained by the second tracker and the third tracker for tracking the object, and use the overlapping paths to train the decision network based on the graph attention mechanism to obtain a trained decision network.
[0074] When training a decision network using overlapping paths, for each overlapping path, when making a decision on the detection boxes in the current frame of the overlapping path using the decision network based on the graph attention mechanism, determine the spatio-temporal features of the detection boxes in the current frame of the overlapping path, and obtain the spatio-temporal features of each video frame within the sliding time window before the current frame as the input features of the decision network. Input the input features into the decision network to obtain the predicted features corresponding to the detection boxes in the current frame. Determine the distance between the predicted features and the spatio-temporal features of the detection boxes in the current frame, and adjust the network parameters of the decision network based on this distance to train the decision network until the distance reaches the minimum convergence, and obtain the trained decision network.
[0075] By using the second tracker and the third tracker to track the motion paths of each object in the same video sample respectively, and screening the overlapping paths based on the two tracking results, it is possible to mine the pseudo-labels of unlabeled video samples based on the tracking consistency of the same object by different trackers, reducing the dependence on labeled data.
[0076] On the basis of the above technical solution, for each object, according to the first tracking result and the second tracking result, determine the overlapping paths of the paths obtained by the second tracker and the third tracker for tracking the object, including: for each object, determine the average path length of the first tracking result and the second tracking result, determine the number of overlapping detection boxes in the first tracking result and the second tracking result, determine the maximum number of path interruptions in the first tracking result and the second tracking result, and determine the maximum number of missed detections in the first tracking result and the second tracking result; according to the average path length, the number of detection boxes, the maximum number of path interruptions, and the maximum number of missed detections, determine the mutual likelihood confidence of the first tracking result and the second tracking result; determine the mutual likelihood confidence of the first tracking result and the second tracking result; if the mutual likelihood confidence is greater than or equal to the confidence threshold, then determine the path corresponding to the overlapping detection boxes as the overlapping path.
[0077] Perform object matching on the first tracking results of each object and the second tracking results of each object to determine the first tracking result and the second tracking result that belong to the same object. Of course, it is also possible not to determine the first tracking result and the second tracking result that belong to the same object, but directly calculate the mutual likelihood confidence of the two tracking results.
[0078] For the same object, determine the number of detection boxes in the first tracking result and the number of detection boxes in the second tracking result. Based on the number of detection boxes in the first tracking result and the number of detection boxes in the second tracking result, calculate the average number of detection boxes in the first tracking result and the second tracking result, and determine the average number of detection boxes in the first tracking result and the second tracking result as the average path length of the first tracking result and the second tracking result. For the same object, determine the number of detection boxes with overlapping positions in the first tracking result and the second tracking result.
[0079] Table 1 shows the number of detection boxes with overlapping positions in the first tracking result and the second tracking result. As shown in Table 1, the first tracking result includes the tracking results of 4 targets, and the second tracking result includes the tracking results of 5 targets. Among the first tracking result and the second tracking result, Object 02 and Object 34 are the same object, and the number of overlapping detection boxes is 20; Object 05 and Object 92 are the same object, and the number of overlapping detection boxes is 46; Object 11 and Object 13 are the same object, and the number of overlapping detection boxes is 31; Object 09 and Object 03 are the same object, and the number of overlapping detection boxes is 15.
[0080] Table 1
[0081]
[0082]
[0083] For the first tracking result or the second tracking result, if an object is tracked on the target detection box in a video frame, but not tracked on the target detection box in the subsequent frame, and then is tracked on the target detection box of the object again in the subsequent frame, it is determined that an interruption occurs in the first tracking result or the second tracking result. If the object's target detection box is not tracked for two consecutive frames, it is determined that two interruptions occur in the first tracking result or the second tracking result; if an object is tracked on the target detection box in a video frame, but the target detection box of the object is not tracked until the end of the video sample, it is determined that this video frame is the path end frame of the first tracking result or the second tracking result. According to the interruption situation of the path, determine the number of interruptions in the first tracking result and the number of interruptions in the second tracking result, compare the number of interruptions in the first tracking result and the number of interruptions in the second tracking result, and determine the maximum number of interruptions in the first tracking result and the second tracking result. For each object, based on the start video frame and the end video frame corresponding to the first tracking result and the second tracking result in the video sample, determine the total number of video frames including the object in the video sample, and determine the difference between the total number of video frames including the object and the number of video frames in which the object is tracked in the first tracking result as the missed detection times of the first tracking result, and determine the difference between the total number of video frames including the object and the number of video frames in which the object is tracked in the second tracking result as the missed detection times of the second tracking result. Compare the missed detection times of the first tracking result and the missed detection times of the second tracking result, and determine the maximum number of missed detections in the first tracking result and the second tracking result.
[0084] After determining the average path length of the first tracking result and the second tracking result, the number of overlapping detection boxes in the first tracking result and the second tracking result, the maximum number of interruptions in the first tracking result and the second tracking result, and the maximum number of missed detections in the first tracking result and the second tracking result, based on the average path length, the number of detection boxes, the maximum number of interruptions, and the maximum number of missed detections, determine the mutual likelihood confidence of the first tracking result and the second tracking result according to the following formula:
[0085]
[0086] where T i represents the first tracking result, which is the tracking path of the i-th object, and T j represents the second tracking result, which is the tracking path of the j-th object. MLC(T i , T j ) represents the mutual likelihood confidence of the first tracking result and the second tracking result, L represents the average path length, O represents the number of detection boxes, F represents the maximum number of interruptions, and M represents the maximum number of missed detections.
[0087] When the mutual likelihood confidence between the first tracking result and the second tracking result is greater than or equal to the confidence threshold, the paths corresponding to the overlapping detection boxes in the first tracking result and the second tracking result are determined as overlapping paths. By determining the first tracking result and the second tracking result with relatively high mutual likelihood confidence in the two tracking results, and using the overlapping paths in the first tracking result and the second tracking result to train the decision network, a self-supervised decision mechanism is provided, which generates overlapping paths with high confidence in a data-driven manner based on the consistency and continuity of the ideal paths, and uses them as samples for training the decision network, which can improve the judgment accuracy of the trained decision network.
[0088] It should be noted that, for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present application are not limited by the described action sequences, because according to the embodiments of the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present application.
[0089] Figure 2 It is a structural block diagram of an object tracking device provided by an embodiment of the present application, as Figure 2 shown. The object tracking device may include:
[0090] An object tracking module 201, configured to track the motion paths of each object in the video to be tracked through a first tracker, and obtain the initial tracking results of each object, where the motion path of the object is characterized by the target detection box of the object in each video frame of the video to be tracked;
[0091] A target video frame determination module 202, configured to, for each object, determine the target video frame in which the object is not detected according to the initial tracking result of the object;
[0092] A candidate detection box prediction module 203, configured to predict the candidate detection box of the object in the target video frame through a Kalman filter;
[0093] A decision module 204, configured to determine whether the candidate detection box belongs to the motion path of the object through a decision network based on a graph attention mechanism.
[0094] Optionally, the decision module includes:
[0095] A feature determination unit, configured to determine the spatio-temporal features of the candidate detection box according to the candidate detection box and the target detection boxes of the object in adjacent video frames of the target video frame; where the spatio-temporal features include spatial features and temporal features;
[0096] A prediction feature determination unit, configured to input the spatio-temporal features of the target detection boxes corresponding to each video frame of the object within a sliding time window into the decision network based on the graph attention mechanism, so as to obtain the prediction features of the candidate detection boxes;
[0097] A decision-making unit, configured to determine whether the candidate detection box belongs to the motion path of the object according to the spatio-temporal features and the prediction features of the candidate detection box.
[0098] Optionally, the decision-making unit is specifically configured to:
[0099] Determine the matching degree between the spatio-temporal features and the prediction features of the candidate detection box;
[0100] If the matching degree is less than the matching degree threshold, it is determined that the candidate detection box does not belong to the motion path of the object; if the matching degree is greater than or equal to the matching degree threshold, it is determined that the candidate detection box belongs to the motion path of the object.
[0101] Optionally, the feature determination unit includes:
[0102] A spatial feature determination subunit, configured to determine the confidence of the candidate detection box, and determine the occlusion ratio of the candidate detection box according to the initial tracking result; determine the spatial features of the candidate detection box according to the confidence and the occlusion ratio.
[0103] Optionally, the feature determination unit further includes:
[0104] A temporal feature determination subunit, configured to determine the intersection over union of the candidate detection box and the target detection box of the same object in the previous frame of the target video frame, determine the similarity of the target appearance features of the candidate detection box and the target detection box of the same object in the previous frame of the target video frame, and determine the center point coordinate offset of the candidate detection box and the detection box of the same object in the previous frame of the target video frame; determine the temporal features of the candidate detection box according to the intersection over union, similarity and the center point coordinate offset.
[0105] Optionally, the prediction feature determination unit is specifically configured to:
[0106] Input the spatio-temporal features of the target detection boxes corresponding to each video frame of the object within a sliding time window into the decision network based on the graph attention mechanism, and determine the initial embedding vectors corresponding to the spatio-temporal features of each target detection box through the decision network;
[0107] Process the structural features between the initial embedding vectors through the decision network to obtain the final embedding vectors;
[0108] Based on the final embedding vector and the spatio-temporal features of the target detection box, determine an attention coefficient through the decision network, and determine an aggregated representation feature corresponding to the spatio-temporal features of the target detection box according to the attention coefficient;
[0109] Process the final embedding vector and the aggregated representation feature through the decision network to obtain the prediction feature of the candidate detection box.
[0110] Optionally, the device further includes a decision network training module, and the decision network training module includes:
[0111] A video sample acquisition unit for acquiring unlabeled video samples;
[0112] Two tracker tracking units for using a second tracker to track the movement paths of each object in the video sample to obtain a first tracking result of each object, and using a third tracker to track the movement paths of each object in the video sample to obtain a second tracking result of each object;
[0113] An overlapping path determination unit for, for each object, determining an overlapping path of the paths obtained by the second tracker and the third tracker for tracking the object according to the first tracking result and the second tracking result;
[0114] A decision network training unit for training the decision network based on the graph attention mechanism through the overlapping path to obtain a trained decision network.
[0115] Optionally, the overlapping path determination unit includes:
[0116] An associated data determination subunit for, for each object, determining the average path length of the first tracking result and the second tracking result, determining the number of overlapping detection boxes in the first tracking result and the second tracking result, determining the maximum number of path interruptions in the first tracking result and the second tracking result, and determining the maximum number of missed detections in the first tracking result and the second tracking result;
[0117] A confidence determination subunit for determining the mutual likelihood confidence of the first tracking result and the second tracking result according to the average path length, the number of detection boxes, the maximum number of path interruptions, and the maximum number of missed detections;
[0118] An overlapping path determination subunit for, if the mutual likelihood confidence is greater than or equal to a confidence threshold, determining the path corresponding to the overlapping detection boxes as the overlapping path.
[0119] For the specific implementation processes of the functions corresponding to the various modules and units in the device provided in the embodiments of the present application, reference may be made to Figure 1 the method embodiments shown. The specific implementation processes of the functions corresponding to the various modules and units in the device part will not be elaborated herein.
[0120] The object tracking device provided in this embodiment tracks the movement paths of the objects in the video to be tracked through a first tracker, obtains the initial tracking results of the objects. For each object, the target video frames in which the object is not detected are determined according to the initial tracking results of the object. The candidate detection boxes of the object in the target video frames are predicted through a Kalman filter, and a decision network based on a graph attention mechanism is used to determine whether the candidate detection boxes belong to the movement paths of the objects. Since the Kalman filter is used to predict the candidate detection boxes of the object in the target video frames, and the decision network based on the graph attention mechanism is used to decide whether the candidate detection boxes belong to the movement paths of the objects, fragmented paths can be effectively repaired, and the discrimination ability for targets in complex occlusion scenes is improved, thereby improving the accuracy and integrity of the tracking paths.
[0121] For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For related parts, reference may be made to the partial description of the method embodiments.
[0122] Figure 3 FIG. is a schematic structural diagram of an electronic device provided in the embodiments of the present application. As Figure 3 shown, the electronic device 300 may include one or more processors 310 and one or more memories 320 connected to the processor 310. The electronic device 300 may further include an input interface 330 and an output interface 340 for communicating with another device or system. The program code executed by the processor 310 may be stored in the memory 320.
[0123] The processor 310 in the electronic device 300 calls the program code stored in the memory 320 to execute the object tracking method in the above embodiments.
[0124] According to an embodiment of the present application, there is also provided a computer-readable storage medium, which includes but is not limited to a disk memory, a CD-ROM, an optical memory, etc. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the object tracking method in the foregoing embodiments is implemented.
[0125] According to an embodiment of the present application, there is also provided a computer program product, including a computer program or computer instructions, and when the computer program or computer instructions are executed by a processor, the object tracking method described in the above embodiments is implemented.
[0126] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0127] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the embodiments of the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0128] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0129] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0130] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0131] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the embodiments of the present application.
[0132] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.
[0133] The above has introduced in detail a method for object tracking, an electronic device and a storage medium provided by the present application. Specific examples are used in this text to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. An object tracking method, characterized in that, Including: Tracking the motion paths of each object in the video to be tracked through a first tracker to obtain the initial tracking results of each object, where the motion path of the object is characterized by the target detection box of the object in each video frame of the video to be tracked; For each object, determining the target video frames in which the object is not detected according to the initial tracking result of the object; Predicting the candidate detection box of the object in the target video frame through a Kalman filter; Determining whether the candidate detection box belongs to the motion path of the object through a decision network based on a graph attention mechanism; The determining whether the candidate detection box belongs to the motion path of the object through a decision network based on a graph attention mechanism includes: Determining the spatio-temporal features of the candidate detection box according to the candidate detection box and the target detection boxes of the object in the adjacent video frames of the target video frame; where the spatio-temporal features include spatial features and temporal features; Inputting the spatio-temporal features of the target detection boxes corresponding to each video frame of the object in the sliding time window into the decision network based on a graph attention mechanism to obtain the predicted features of the candidate detection box; Determining whether the candidate detection box belongs to the motion path of the object according to the spatio-temporal features and the predicted features of the candidate detection box.
2. The method according to claim 1, wherein Determining whether the candidate detection box belongs to the motion path of the object according to the spatio-temporal features and the predicted features of the candidate detection box includes: Determining the matching degree between the spatio-temporal features and the predicted features of the candidate detection box; If the matching degree is less than the matching degree threshold, determining that the candidate detection box does not belong to the motion path of the object; if the matching degree is greater than or equal to the matching degree threshold, determining that the candidate detection box belongs to the motion path of the object.
3. The method according to claim 2, wherein Determining the spatial features of the candidate detection box through the following process: Determining the confidence of the candidate detection box and determining the occlusion ratio of the candidate detection box according to the initial tracking result; Determining the spatial features of the candidate detection box according to the confidence and the occlusion ratio.
4. The method according to claim 2, wherein Determining the temporal features of the candidate detection box through the following process: Determining the intersection over union of the candidate detection box and the target detection box of the same object in the previous frame of the target video frame, determining the similarity of the target appearance features of the candidate detection box and the target detection box of the same object in the previous frame of the target video frame, and determining the center point coordinate offset of the candidate detection box and the target detection box of the same object in the previous frame of the target video frame; Determining the temporal features of the candidate detection box according to the intersection over union, the similarity and the center point coordinate offset.
5. The method according to claim 1, wherein Inputting the spatio-temporal features of the target detection boxes corresponding to each video frame of the object in the sliding time window into the decision network based on a graph attention mechanism to obtain the predicted features of the candidate detection box includes: Input the spatio-temporal features of the target detection boxes corresponding to each video frame of the object within the sliding time window into the decision network based on the graph attention mechanism, and determine the initial embedding vectors corresponding to the spatio-temporal features of each target detection box through the decision network; Process the structural features between the initial embedding vectors through the decision network to obtain the final embedding vectors; According to the final embedding vectors and the spatio-temporal features of the target detection boxes, determine the attention coefficients through the decision network, and determine the aggregated representation features corresponding to the spatio-temporal features of the target detection boxes according to the attention coefficients; Process the final embedding vectors and the aggregated representation features through the decision network to obtain the prediction features of the candidate detection boxes.
6. The method according to any one of claims 1-5, characterized in that, Train the decision network of the graph attention mechanism through the following process: Obtain unlabeled video samples; Use the second tracker to track the motion paths of each object in the video sample to obtain the first tracking results of each object, and use the third tracker to track the motion paths of each object in the video sample to obtain the second tracking results of each object; For each object, according to the first tracking result and the second tracking result, determine the overlapping paths of the paths obtained by the second tracker and the third tracker for tracking the object; Train the decision network based on the graph attention mechanism through the overlapping paths to obtain the trained decision network.
7. The method according to claim 6, wherein For each object, according to the first tracking result and the second tracking result, determining the overlapping paths of the paths obtained by the second tracker and the third tracker for tracking the object includes: For each object, determine the average path length of the first tracking result and the second tracking result, determine the number of overlapping detection boxes in the first tracking result and the second tracking result, determine the maximum number of path interruptions in the first tracking result and the second tracking result, and determine the maximum number of missed detections in the first tracking result and the second tracking result; According to the average path length, the number of detection boxes, the maximum number of path interruptions, and the maximum number of missed detections, determine the mutual likelihood confidence of the first tracking result and the second tracking result; If the mutual likelihood confidence is greater than or equal to the confidence threshold, determine the paths corresponding to the overlapping detection boxes as the overlapping paths.
8. An electronic device, characterized in that, Includes: A processor, a memory, and a computer program stored on the memory and executable on the processor, where the computer program, when executed by the processor, implements the object tracking method according to any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, it implements the object tracking method according to any one of claims 1-7.
10. A computer program product, characterized in that, Includes a computer program or computer instructions, and when the computer program or computer instructions are executed by the processor, they implement the object tracking method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Urban vehicle path optimization method based on data-driven swarm intelligence calculation
CN112270047A
Target tracking detection method and device for obstacle shielding, equipment and medium
CN114092515A