Target rematching method based on motion trail in single target tracking

By combining reverse tracking and trajectory pool maintenance with the bipartite graph maximum matching algorithm, the drift problem caused by interference from similar targets in single-target tracking is solved, achieving efficient and real-time target re-matching and improving tracking accuracy.

CN121213612APending Publication Date: 2025-12-26CHINA OPTICS (HANGZHOU) INTELLIGENT OPTOELECTRONICS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511340444.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

In complex scenarios, single-target trackers are easily affected by targets with similar appearances, leading to target box drift and matching errors, which are difficult to solve effectively with existing technologies.

Method used

A strategy of reverse tracking, trajectory pool maintenance, and bipartite graph maximum matching is adopted. Target rematch is performed by combining the Hopcroft-Karp algorithm with penalized confidence, non-maximum suppression, and Kalman filter prediction to construct a two-dimensional motion model to correct target drift.

Benefits of technology

It significantly improves the accuracy of target rematching, reduces computational overhead, enables real-time tracking, and eliminates the need to retrain existing deep trackers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121213612A_ABST
    Figure CN121213612A_ABST
Patent Text Reader

Abstract

The invention discloses a target rematching method based on a motion trail in single target tracking, and relates to the technical field of intelligent visual monitoring and video analysis. The method comprises the following steps: firstly, screening a target candidate box, and performing multi-dimensional punishment and screening on a current frame tracking result to obtain the target candidate box; then, whether tracking ambiguity exists or not is judged in combination with a historical real trajectory; and then constructing a two-dimensional motion model based on Kalman filtering, and predicting the current position of the target. And finally, constructing a trajectory pool through reverse tracking, solving bipartite graph maximum matching by using a Hopcroft-Karp algorithm to realize target re-matching, and introducing an interference target trajectory pool dynamic updating mechanism to improve the matching accuracy. According to the method, tracking deviation caused by similar target interference can be effectively corrected, long-term stable tracking of an initial target by a tracker is ensured, the method is suitable for various single-target tracking scenes with similar target interference, and the long-term tracking robustness in complex scenes is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent visual surveillance and video analysis technology, and in particular to a target rematching method based on motion trajectory in single target tracking. Background Technology

[0002] Single Object Tracking (SOT) aims to continuously estimate the position and scale of a target in subsequent video frames after an initial target state is established. While depth-based trackers have achieved high accuracy in simple scenarios in recent years, complex scenarios such as UAV air-to-ground surveillance, dense traffic intersections, and warehousing logistics often present multiple targets with extremely similar appearances (e.g., vehicles or containers with identical models). When trackers are interfered with by similar targets, template drift or even target loss can easily occur. In such cases, it is necessary to trigger "re-identification / re-association" in the current frame to correct deviations and maintain long-term stable tracking. Chinese patent CN2022106955275 suppresses background through dual cross-correlation, but it does not distinguish between "highly similar non-targets," still prone to matching errors. Chinese patent CN2023110736472 uses a multi-level visual Transformer to extract features, but its ability to distinguish extremely similar targets is limited, and the model is large, requiring retraining for cross-platform portability and exhibiting poor real-time performance. In air-to-ground target tracking, target objects often experience bounding box drift and target matching errors due to factors such as the presence of multiple targets with extremely similar appearances in the surrounding area. In certain tracking scenarios, target matching errors can lead to catastrophic consequences. Therefore, developing a re-matching method for a specified target that experiences bounding box drift due to interference from similar objects during single-target tracking has become a pressing problem in this field. Summary of the Invention

[0003] To address the aforementioned problems, the present invention aims to propose a target re-matching method based on motion trajectory in single-target tracking. Without altering the original tracking network structure, this method employs a three-step strategy of "reverse tracking—trajectory pool maintenance—bipartite graph maximum matching" to quickly retrieve the disturbed target and significantly suppress drift.

[0004] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0005] A target re-matching method based on motion trajectory in single-target tracking includes the following steps:

[0006] (1) Filtering target candidate boxes: Obtain the target boxes to be selected and their corresponding classification confidence scores output by the single target tracking model in the current frame t. Perform a post-processing operation to penalize the classification confidence scores to obtain the penalized confidence scores. Filter out target boxes with scores lower than a set threshold, where the set threshold is a preset percentage of the highest score. Then remove redundant target boxes by non-maximum suppression (NMS) to obtain the target candidate boxes and their corresponding scores in the current frame t.

[0007] (2) Maintain a historical true trajectory of a specified length K: The historical true trajectory of the target contains the target position in the historical frame; if the number of target candidate boxes is 1, then directly enter the next frame for tracking; if the number of target candidate boxes is greater than 1, temporarily select the target box corresponding to the highest score as the result of the current frame, take the current frame t as the initial frame and the target box with the highest score as the initial box for reverse tracking K frames, obtain the reverse tracking trajectory, calculate the average intersection-union ratio (MIoU) between the reverse tracking trajectory and the historical true trajectory of the target, if the average intersection-union ratio is greater than the set threshold, then enter the next frame for tracking, otherwise execute the target rematch step;

[0008] (3) Constructing a target motion model and predicting the target position: Construct a two-dimensional motion model. The two-dimensional state vector of the two-dimensional motion model at time t contains eight state variables, including the target's upper left coordinate, width and height, and velocity and acceleration in the x and y directions. Set the state transition matrix, covariance matrix, observation matrix, process noise covariance matrix, and observation noise covariance matrix. Use Kalman filtering to initialize the state estimation and initial covariance matrix based on historical frame information. Obtain the current state and covariance matrix through the state transition equation. Calculate the Kalman gain. Use the final output position of the current frame t as the observation value to update the state estimation and covariance matrix and predict the current target position.

[0009] (4) Establish the interference target trajectory pool: For each target candidate box in step 1, perform reverse tracking for K frames, obtain the corresponding reverse tracking trajectory and form a reverse tracking trajectory pool; obtain the reverse tracking trajectory pool of the previous frame t-1, remove the trajectory that successfully matches the target's historical true trajectory in the previous frame, move the remaining trajectory forward 1 frame, remove the earliest frame trajectory, add it to the reverse tracking trajectory of the current frame, and obtain the interference target trajectory pool of the current frame t.

[0010] (5) Rematching the target: Combine the interference target trajectory pool with the target historical true trajectory in the current frame t to form a trajectory pool to be matched; take the reverse tracking trajectory pool and the trajectory pool to be matched as two vertex sets of a bipartite graph, calculate the average intersection-union ratio (MIoU) of any two trajectories in the two sets as the edge weight, and use the Hopcroft-Karp algorithm to solve the maximum matching of the bipartite graph; if a trajectory in the reverse tracking trajectory pool is successfully matched with the trajectory pool to be matched, the target candidate box corresponding to the trajectory is the rematching result of the current frame t; if the match is unsuccessful, determine whether the highest average intersection-union ratio of the trajectories in the reverse tracking trajectory pool and the trajectory pool to be matched is greater than a set threshold. If it is, take the corresponding candidate box as the result; otherwise, use the prediction value of the Kalman filter in step 3 as the result.

[0011] Furthermore, the penalty includes at least one of a cosine window penalty, a scale penalty, and a proportional penalty.

[0012] The reverse tracking uses the same feature extraction and cross-correlation matching network as the forward tracking, requiring no additional training.

[0013] The criterion for determining the maximum matching of the bipartite graph in step 5 is: when the average intersection-union ratio (MIoU) of the two trajectories is greater than a set matching threshold, the two trajectories are determined to be successfully matched. The set matching threshold ranges from 0.5 to 0.8.

[0014] Compared with the prior art, the beneficial effects of the present invention are:

[0015] This invention utilizes short-time backward tracking to generate a "backward trajectory" for each candidate box; it constructs a bipartite graph with the "backward trajectory pool" and the "historical true trajectory pool," using the average intersection-union ratio (MIoU) as the edge weights; and employs the Hopcroft-Karp algorithm to solve for the maximum match. If a match is successful, a rematch is performed; otherwise, the Kalman prediction position is used as a fallback output to help the tracker accurately rematch the initially locked target. This significantly improves rematch accuracy, addressing the target matching error problem caused by interference from similar targets in single-target tracking. It features low computational overhead, low single-frame latency, and real-time operation at edge environments; it is decoupled from existing depth trackers, allowing for plug-and-play functionality without retraining. Attached Figure Description

[0016] Figure 1 The overall flowchart of the integrated target re-matching single target tracking described in this invention;

[0017] Figure 2 Flowchart of the target rematching sub-process based on motion trajectory described in this invention

[0018] Figure 3 This invention illustrates the use of the Hopcroft-Karp algorithm to solve for the maximum matching in a bipartite graph. Detailed Implementation

[0019] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not intended to limit the present invention.

[0020] See Figure 1-3 A target re-matching method based on motion trajectory in single target tracking includes the following steps:

[0021] (1) Target candidate box filtering: Obtain the candidate target boxes and their corresponding classification confidence scores output by the single-target tracking model in the current frame t. Perform a post-processing operation to penalize the classification confidence scores to obtain the penalized confidence scores. Filter out target boxes with scores lower than a set threshold, where the set threshold is a preset percentage of the highest score. Then, remove redundant target boxes using non-maximum suppression (NMS) to obtain the target candidate boxes and their corresponding scores for the current frame t. For details, see [link to documentation]. Figure 1 Let the current tracking frame be frame t. After passing through the single-target tracking model, a series of target boxes to be selected are obtained. and corresponding classification confidence level A common post-processing method is to directly take the target box corresponding to the highest confidence as the tracking result of the current frame. However, this often causes the target box to drift due to interference from similar objects, making it impossible for the tracker to match the initially locked target.

[0022] To address the above issues, this solution first targets classification confidence. Penalties were applied, including post-processing operations such as cosine window penalty, scaling penalty, and proportional penalty, to initially alleviate the target matching error problem and obtain the confidence score after penalties.

[0023]

[0024] Then filter out the target boxes that do not score high enough. express The highest score in the range, α is a percentage between [0,1], considered to be above Only α percent of the bounding boxes are likely to become candidate bounding boxes.

[0025]

[0026] Finally, non-maximum suppression (NMS) is used to remove redundant target boxes, ultimately obtaining the target candidate boxes for the current frame t. and corresponding scores

[0027] (2) Maintain a historical true trajectory of a specified length K: The target historical true trajectory contains the target position in the historical frames; if the number of target candidate boxes is 1, directly proceed to the next frame for tracking; if the number of target candidate boxes is greater than 1, temporarily select the target box corresponding to the highest score as the result of the current frame, use the current frame t as the initial frame and the target box with the highest score as the initial box for reverse tracking K frames, obtain the reverse tracking trajectory, calculate the average intersection-union ratio (MIoU) between the reverse tracking trajectory and the target historical true trajectory, if the average MIoU is greater than a set threshold, proceed to the next frame for tracking, otherwise execute the target re-matching step; for details, see Figure 1 In the current frame t, a target historical true trajectory of length K is maintained simultaneously. loc i =(x i ,y i ,w i ,h i () refers to the target location in the historical frame.

[0028] If n = 1, B t If there is only one candidate box, the current frame is considered unambiguous, and tracking proceeds directly to the next frame. If n ≠ 1 (n > 1), the target box corresponding to the highest score is temporarily selected. As the result of the current frame, with the current frame t as the initial frame, and with Using the initial bounding box, perform reverse tracking for k frames to obtain the reverse tracking trajectory. calculate and The average intersection-over-union ratio (MIoU) is used. If the MIoU is greater than a set threshold, the tracking proceeds directly to the next frame; otherwise, a target re-matching method is used to correct the tracking results.

[0029]

[0030] (3) Constructing a target motion model and predicting the target position: Constructing a two-dimensional motion model, the two-dimensional state vector of the two-dimensional motion model at time t contains eight state variables, including the target's upper left coordinate, width and height, velocity and acceleration in the x and y directions. Set the state transition matrix, covariance matrix, observation matrix, process noise covariance matrix and observation noise covariance matrix; use Kalman filtering, initialize the state estimation and initial covariance matrix based on historical frame information, obtain the current state and covariance matrix through the state transition equation, calculate the Kalman gain, use the final output position of the current frame t as the observation value to update the state estimation and covariance matrix, and predict the current target position; specifically, the candidate target boxes output by the single target model are obtained based on appearance consistency, so Kalman filter prediction values ​​based on trajectory prediction are added as candidate boxes.

[0031] Construct a two-dimensional motion model, State tIt is a two-dimensional state vector at time t, using 8 state variables, where x t y t w t h t Let v be the top-left coordinate and width / height of the target at time t. x,t and v y,t These are the velocities along the x and y axes at time t, respectively, and a x,t and a y,t and are the accelerations along the x-axis and y-axis at time t, respectively. F is the state transition matrix, and P0 is the covariance matrix.

[0032]

[0033]

[0034] Measure t The final output position of the current frame t is used as the observation value to correct the Kalman filter tracking trajectory and update the covariance matrix. H is the observation matrix. In practical applications, due to measurement errors and other uncertainties, noise needs to be added to the motion model. Q is the process noise covariance matrix, and R is the observation noise covariance matrix; their specific values ​​are adjusted according to the actual situation.

[0035]

[0036]

[0037] After determining the motion model, Kalman filtering is used to predict the current target's state vector. The state estimate (State) is initialized, along with the initial covariance matrix (Co). The current state (State) is obtained using the state transition equation (F). t The covariance matrix Co is calculated, and the Kalman gain K is also calculated using the observed values ​​Measure. t Update State estimation t The covariance matrix Co. Historical frame information is used to complete the series of updates, which will not be elaborated further. In the current frame t, the Kalman filter predicts the target position as K. t =(x t ,y t ,w t ,h t ).

[0038] (4) Establish the interference target trajectory pool: For each target candidate box in step 1, perform reverse tracking for K frames, obtain the corresponding reverse tracking trajectory and form a reverse tracking trajectory pool; obtain the reverse tracking trajectory pool of the previous frame t-1, remove the trajectory that successfully matches the target's historical true trajectory in the previous frame, move the remaining trajectory forward 1 frame, remove the earliest frame trajectory, add it to the reverse tracking trajectory of the current frame, and obtain the interference target trajectory pool of the current frame t.

[0039] (5) Target rematch: Combine the interference target trajectory pool with the target's historical true trajectory in the current frame t to form a trajectory pool to be matched; using the reverse tracking trajectory pool and the trajectory pool to be matched as two vertex sets of a bipartite graph, calculate the average intersection-union ratio (MIoU) of any two trajectories in the two sets as edge weights, and use the Hopcroft-Karp algorithm to solve for the maximum matching of the bipartite graph; if a trajectory in the reverse tracking trajectory pool successfully matches the trajectory pool to be matched, the target candidate box corresponding to that trajectory is the rematch result of the current frame t; if no match is found, determine whether the highest average intersection-union ratio of the trajectories in the reverse tracking trajectory pool and the trajectory pool to be matched is greater than a set threshold. If so, take the corresponding candidate box as the result; otherwise, use the predicted value of the Kalman filter in step 3 as the result. Specifically, as shown in the figure... Figure 2 and Figure 3 As shown, for the target candidate box Each candidate box Reverse tracking k frames to obtain the corresponding reverse tracking trajectory Construct a reverse tracking trajectory pool

[0040] The reverse tracking trajectory pool of the previous frame t-1 in Remove the target's historical true trajectory from the previous frame. The successfully matched trajectory will All remaining trajectories Move forward 1 frame, i.e., remove. join in get Will This is called the interference target trajectory pool for the current frame t, and it is maintained and updated every frame. It also includes the historical true trajectories from the current frame t. Constructing the trajectory pool M t .in,

[0041] Reverse tracking trajectory pool T t back The arbitrary trajectory and the trajectory pool M to be matched t Any trajectory in T can form an edge, and the two endpoints of any edge are not in the same set. t back and Mt This can form a bipartite graph. The average intersection-union ratio (MIoU) between different trajectories in the two trajectory pools is calculated as the edge weights. The Hopcroft-Karp algorithm (an improved Hungarian algorithm) is then used to find the maximum matching in the bipartite graph. If T... t back There is a trajectory and If the match is successful, then The target candidate box in the current frame t is the result of the target rematch. Otherwise, determine... Mid-trajectory and If the highest MIoU is greater than a set threshold, the candidate box corresponding to this trajectory in the current frame t is used as the result; otherwise, the predicted value K from the Kalman filter is used. t =(x t ,y t ,w t ,h t As a result.

[0042] The penalty includes at least one of the following: cosine window penalty, scale penalty, and proportion penalty.

[0043] In one embodiment, the preset percentage is selected from a range of 50% to 80%, and the length K is selected from a range of 10 frames to 30 frames.

[0044] The reverse tracking uses the same feature extraction and cross-correlation matching network as the forward tracking, requiring no additional training.

[0045] The criterion for determining the maximum matching of the bipartite graph in step 5 is: when the average intersection-union ratio (MIoU) of the two trajectories is greater than a set matching threshold, the two trajectories are determined to be successfully matched. The set matching threshold ranges from 0.5 to 0.8.

Claims

1. A target re-matching method based on motion trajectory in single target tracking, characterized in that, Includes the following steps: (1) Filtering target candidate boxes: Obtain the target boxes to be selected and their corresponding classification confidence scores output by the single target tracking model in the current t-th frame. Perform post-processing operations to penalize the classification confidence scores to obtain the penalized confidence scores. Filter out target boxes with scores below a set threshold, where the set threshold is a preset percentage of the highest score, and then remove redundant target boxes using non-maximum suppression (NMS) to obtain the target candidate boxes and their corresponding scores for the current frame t. (2) Maintain a historical true trajectory of a specified length K: The historical true trajectory of the target contains the target position in the historical frame; if the number of target candidate boxes is 1, then directly enter the next frame for tracking; if the number of target candidate boxes is greater than 1, temporarily select the target box corresponding to the highest score as the result of the current frame, take the current frame t as the initial frame and the target box with the highest score as the initial box for reverse tracking K frames, obtain the reverse tracking trajectory, calculate the average intersection-union ratio (MIoU) between the reverse tracking trajectory and the historical true trajectory of the target, if the average intersection-union ratio is greater than the set threshold, then enter the next frame for tracking, otherwise execute the target rematching step; (3) Constructing a target motion model and predicting the target position: Constructing a two-dimensional motion model, wherein the two-dimensional state vector of the two-dimensional motion model at time t contains eight state variables, including the target's upper left coordinate, width and height, velocity and acceleration in the x and y directions, and setting the state transition matrix, covariance matrix, observation matrix, process noise covariance matrix and observation noise covariance matrix; Kalman filtering is used to initialize the state estimate and initial covariance matrix based on historical frame information. The current state and covariance matrix are obtained through the state transition equation. The Kalman gain is calculated. The final output position of the current frame t is used as the observation value to update the state estimate and covariance matrix and predict the current target position. (4) Establishing an interference target trajectory pool: For each target candidate box in step 1, perform reverse tracking for K frames, obtain the corresponding reverse tracking trajectory and form a reverse tracking trajectory pool; obtain the reverse tracking trajectory pool of the previous frame t-1, remove the trajectory that successfully matches the target's historical true trajectory in the previous frame, move the remaining trajectory forward 1 frame, remove the earliest frame trajectory, add the current frame's reverse tracking trajectory, and obtain the interference target trajectory pool of the current frame t. (5) Rematch the target: Combine the interference target trajectory pool with the target historical real trajectory in the current frame t to form a trajectory pool to be matched; Using the backtracking trajectory pool and the trajectory pool to be matched as two vertex sets of a bipartite graph, the average intersection-union ratio (MIoU) of any two trajectories in the two sets is calculated as the edge weight, and the Hopcroft-Karp algorithm is used to solve the maximum matching problem in the bipartite graph. If a trajectory in the backtracking trajectory pool is successfully matched with the trajectory pool to be matched, the target candidate box corresponding to that trajectory is the rematch result of the current frame t. If no match is found, determine whether the highest average intersection-union ratio of the reverse tracking trajectory pool and the trajectory pool to be matched is greater than the set threshold. If so, take the corresponding candidate box as the result; otherwise, use the predicted value of the Kalman filter in step 3 as the result.

2. The method according to claim 1, characterized in that, The penalty includes at least one of the following: cosine window penalty, scale penalty, and proportion penalty.

3. The method according to claim 1, characterized in that, The reverse tracking uses the same feature extraction and cross-correlation matching network as the forward tracking, requiring no additional training.

4. The method according to claim 1, characterized in that, The criterion for determining the maximum matching of the bipartite graph in step 5 is: when the average intersection-union ratio (MIoU) of the two trajectories is greater than a set matching threshold, the two trajectories are determined to be successfully matched. The set matching threshold ranges from 0.5 to 0.8.