Robust multi-target tracking method based on strong clue and weak clue

By introducing robust trajectory confidence modeling, pseudo-depth metrics, advanced observation center recovery and window denoising modules in multi-objective tracking technology, the problem of tracking performance degradation in complex motion scenarios is solved, and more robust and efficient multi-objective tracking is achieved.

CN120198459APending Publication Date: 2025-06-24UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510282582.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing multi-objective tracking techniques show limitations in complex motion scenarios, especially when dealing with nonlinear motion and occlusion, it is difficult to fully utilize strong and weak tracking cues, resulting in a degradation of tracking performance.

Method used

A robust multi-objective tracking method based on strong and weak cues is proposed. Through the robust trajectory confidence modeling module, pseudo-depth metric indicator, advanced observation center recovery module and window denoising module, a variety of clue information is integrated to improve the robustness and accuracy of tracking.

Benefits of technology

In complex motion scenarios, the robustness and accuracy of multi-objective tracking are significantly improved, ID switching and trajectory loss are reduced, and the system's real-time tracking capabilities are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198459A_ABST
    Figure CN120198459A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image processing, in particular to a robust multi-target tracking method based on a strong clue and a weak clue, which comprises the following steps of: S1, inputting a video stream, detecting a target of each frame of image in the input video stream, and recording coordinate information and detection confidence of points at the left upper corner and the right lower corner of a target detection frame; the invention provides a robust multi-target tracking method which better utilizes strong clues and weak clues in tracking, challenges such as non-linear motion and shielding in a complex motion scene are solved, specifically, a robust trajectory confidence modeling module realizes a more robust trajectory confidence modeling method, and the robustness of the trajectory confidence modeling is improved. Under the condition that target distinguishing becomes difficult due to the fact that crowded and sheltered objects only depend on strong clues, valuable clues are provided for distinguishing front and back objects by means of the confidence degree of modeling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and more specifically, to a robust multi-object tracking method based on strong and weak cues. Background Art

[0002] With the rapid development of artificial intelligence, multi-object tracking has become one of the hot research issues in the field of computer vision. As a basic visual perception task, multi-object tracking has a wide range of applications in the fields of national defense and military, aviation, transportation, etc. In an intelligent monitoring system, multi-object tracking can accurately and effectively track video targets, providing timely and effective information support for monitoring management personnel, thereby maintaining social security. In the scenario of autonomous driving, vehicles analyze the road conditions, track the vehicles and pedestrians within the field of view, and make motion planning and risk avoidance strategies based on their moving speeds and positions. The core of the multi-object tracking task lies in accurately locating the targets and maintaining the consistency of the target identities. However, the actual application scenarios are usually complex and crowded, the background changes frequently, and the target motions are complex, variable and irregular, which easily lead to target blurring, mutual occlusion, and frequent disappearance and reappearance of targets. These challenges pose great difficulties to the existing multi-object tracking technologies.

[0003] Currently, the mainstream paradigm of multi-object tracking is detection-based tracking, which consists of two core steps: detection and association. In the detection stage, the system identifies the targets in each frame; in the association stage, the detected targets are matched with the targets in the previous frame. For the successfully matched targets, their original identity IDs are retained; while for the newly emerged targets, new IDs are assigned. In this process, strong tracking cues such as position and appearance significantly enhance the target matching effect. However, the existing methods have deficiencies in utilizing weak tracking cues such as confidence and depth. In a crowded environment, situations such as severe occlusion, fast target movement, and similar appearances often lead to a decline in tracking performance.

[0004] In recent years, the Transformer model has achieved remarkable results in the field of computer vision, giving rise to a multi-object tracking paradigm based on Tracking-by-Attention. This method solves the problem of long-term association failure caused by target occlusion by establishing global spatio-temporal associations. However, the multi-object tracking model based on Transformer has the disadvantages of large number of parameters and long inference time. In addition, this method usually relies on a high-performance detector to obtain an ideal tracking effect, which further increases the computational burden of the system and is difficult to meet the requirements of real-time tracking.

[0005] As can be seen from the above analysis, existing multi-object tracking methods fail to fully utilize strong and weak tracking cues to improve tracking performance, especially showing certain limitations when dealing with complex motion scenarios. Therefore, how to provide a robust multi-object tracking method in complex scenarios has become an urgent problem for those skilled in the art to solve.

[0006] Therefore, we propose a robust multi-object tracking method based on strong cues and weak cues to solve the above problems. Summary of the Invention

[0007] To overcome the above defects of the prior art, embodiments of the present invention provide a robust multi-object tracking method based on strong cues and weak cues to solve the problems raised in the above background art.

[0008] To achieve the above object, the present invention provides the following technical solution: A robust multi-object tracking method based on strong cues and weak cues, comprising the following steps:

[0009] Step S1: Input a video stream, detect the objects in each frame of the input video stream, record the coordinate information of the points at the upper left corner and the lower right corner of the object detection box and the detection confidence, classify the objects into high-score detections, medium-score detections and low-score detections according to the comparison between the detection confidence and a set threshold, and discard the objects with low-score detections; when the object detection box on each frame matches the prediction box obtained by using historical information in the trajectory set, update the information of the corresponding matching trajectory in the trajectory set with the object detection box of the current frame;

[0010] Step S2: Use Kalman filtering to predict the positions of the existing trajectories in the trajectory set in the current frame, and at the same time use a robust trajectory confidence modeling module to predict the confidence of the trajectories in the current frame.

[0011] Step S3: Use the pseudo-depth metric PDIoU, confidence weak cues, motion direction weak cues and appearance strong cues to calculate the association cost between the trajectory set and the high-score detection objects, fuse them into an embedding vector and then perform matching through the Hungarian algorithm. The matched trajectories are sent to the trajectory manager for update, and the trajectories that fail to match enter the second-stage association;

[0012] Step S4: Use the pseudo-depth metric PDIoU, confidence, motion direction weak cues and appearance strong cues to calculate the association cost between the remaining unmatched trajectory set and the medium-score detection objects, perform matching through the Hungarian algorithm, and the matched trajectories are sent to the trajectory manager for update, which is the same as the operation in S3;

[0013] Step S5: Try to recover the trajectory association through the advanced observation center recovery module, calculate the association cost between the unmatched high-score detection backtracking and the unmatched trajectory backtracking, and the association cost between the unmatched high-score detection backtracking and the last observation box of the unmatched trajectory, and obtain the association cost by weighted averaging or multiplying the two, and perform matching through the Hungarian algorithm. The matched trajectories are sent to the trajectory manager for update;

[0014] Step S6: Send both the unmatched trajectory set and the matched trajectory set to the trajectory manager. For the unmatched medium-score detections, they are directly discarded. For the unmatched high-score detections, the window denoising module is optionally used to remove noise according to the needs of the scenario. The remaining unmatched high-score detections are used by the trajectory manager to create corresponding new trajectories;

[0015] Step S7: End the tracking of the current frame, track the next frame, and repeat steps S1 - S6.

[0016] In a preferred embodiment, in step S1, the general detector YOLOX and the common appearance detector SBS-50 are used to detect the image of the current video frame, and the coordinate information of the upper left and lower right points of the target detection box in the image, the detection confidence, and the corresponding appearance embedding vector are obtained.

[0017] In a preferred embodiment, in step S2, when the detection confidence is relatively high under no occlusion or slight occlusion, we use the Kalman filter for continuous state estimation and use exponential smoothing to estimate the trajectory confidence, as shown in the following formula:

[0018]

[0019] Where:

[0020] is the trajectory confidence to be modeled at the current time t;

[0021] is the trajectory confidence at time t - 1;

[0022] is the trajectory confidence predicted by the Kalman filter at time t;

[0023] α is the smoothing exponent.

[0024] In a preferred embodiment, in step S2, when occlusion occurs, the detection confidence is relatively low, and the Kalman filter cannot quickly capture the sudden change of the confidence. The first-order and second-order difference linear correction algorithms are used to linearly predict the trajectory confidence;

[0025] As follows:

[0026]

[0027] Wherein:

[0028] is the confidence of the trajectory at time t-2;

[0029] The updated confidence is based on the linear trend of the change in the previous confidence, and the difference between the last two trends is used to correct the linear change, that is, the following formula, which enables the confidence to adapt faster and more accurately during occlusion:

[0030]

[0031] In a preferred embodiment, in step S3: the association cost is composed of the sum of the confidence cost, the appearance cost, the motion direction cost, and the pseudo-depth metric cost. The confidence weak cue calculates the association cost between the trajectory set and the detection target by the absolute value of the difference in confidence between the two, that is, the following formula:

[0032]

[0033] The appearance strong cue calculates the association cost between the trajectory set and the detection target by extracting the appearance embedding vectors of the two by the general detector SBS-50, and then calculating the negative of the cosine similarity of the two embedding vectors as the association cost C a The motion direction is the cosine distance of the motion directions of the trajectory set and the detection target calculated by a general method as the association cost C v Both are calculated by general methods.

[0034] In a preferred embodiment, in step S3, the pseudo-depth metric PDIoU incorporates the height weak cue and depth weak cue information into the strong cue IoU calculation;

[0035] First, calculate the height intersection over union HIoU, as follows:

[0036] HIoU = y inter / y cover

[0037] Wherein:

[0038] y inter is the length of the overlap between the height of the trajectory box and the height of the detection box;

[0039] y cover is the height of the minimum bounding box of the trajectory box and the detection box;

[0040] The relative depth relationship between objects is approximated by measuring the distance from the bottom of the detection box to the bottom of the image, which is called pseudo-depth;

[0041] First, calculate the pseudo - depths (d1, d2) of the trajectory set and the detected object, as well as the absolute value δd of the difference between them;

[0042] As shown in the following formula:

[0043]

[0044] δd = |d1 - d2|

[0045] Where:

[0046] H is the height of the video frame;

[0047] and are respectively the y - axis coordinates of the lower - right corners of the trajectory box and the detection box;

[0048] Then, calculate the thresholds for pseudo - depth reward and penalty. Give a penalty for those exceeding the threshold and a reward for those not exceeding the threshold to obtain the pseudo - depth - related scores, as shown in the following two formulas:

[0049]

[0050] Where:

[0051] i and j are the i - th trajectory set and the j - th detected object;

[0052] k, α1, and α2 are all hyperparameters;

[0053] Finally, PDIoU can be defined as the following formula. Take the negative of PDIoU to obtain the pseudo - depth cost:

[0054] PDIoU = PD·HIoU·IoU

[0055] Add the above - mentioned association costs according to a certain weight to form an association - cost embedding vector, and use the general Hungarian algorithm for association matching. The successfully - matched trajectory - detection set, unmatched trajectory set, and unmatched high - score detections will be obtained.

[0056] In a preferred embodiment, in step S5, assume that at time T, there are m trajectory segments and n detections. Trace back the i - th predicted trajectory to the last moment before it was lost;

[0057] Denote it as T - ΔT i , and the corresponding trajectory is marked as OBox i ;

[0058] Use linear interpolation to obtain the virtual trajectory box PBox i at the intermediate time T - βΔT i , where β ∈ [0, 1] is a back - tracking factor affected by the target speed;

[0059] For the j-th detected trajectory, we apply the same method to the OBox j . Interpolation is performed at T - ΔT i to obtain the virtual trajectory box DBox at T - βΔT j ; j ;

[0060] We calculate the IoU between PBox i and DBox j , and the IoU between OBox i and DBox j respectively to obtain scores Score1 and Score2. The two scores are weighted averaged or multiplied to obtain the total association score, and the negative value of it is the association cost.

[0061] In a preferred embodiment, in step S6, the successfully matched detected trajectory updates the Kalman parameters and appearance, while the activated trajectory that fails to be successfully matched freezes the Kalman parameters and updates the Kalman parameters again after re-matching to prevent the accumulation of parameter errors during the period of unmatched detections;

[0062] Trajectories that have not been matched for a long time are considered to have left the tracking scene and are deleted by the trajectory manager;

[0063] For the unmatched high-score detections, if the tracking scene changes little and the processed video frames have exceeded a certain number of frames, the window denoising module is used to select a region in the video frame as the denoising region according to the scene,

[0064] If the current frame number exceeds the threshold τ and a new detection box appears within the filtering window, it may be a false detection or a failure to associate and is considered a new target, then the detection is discarded.

[0065] Since the object fails to be successfully associated with ID7, it is assigned a new ID (ID14);

[0066] In addition, due to the false detection of the detector, ID15 is wrongly generated;

[0067] However, through the window denoising module, these noises are effectively filtered;

[0068] Select a denoising region. If the current frame number exceeds the threshold τ and a new detection box appears within the filtering window, it may be a false detection or a failure to associate and is considered a new target, then the detection is discarded.

[0069] Technical effects and advantages of the present invention:

[0070] The present invention proposes a robust multi-object tracking method that makes better use of strong and weak cues in tracking, solving the challenges in complex motion scenarios, such as non-linear motion and occlusion. Specifically, the robust trajectory confidence modeling module implements a more robust method for modeling trajectory confidence. In situations where relying solely on strong cues in crowded and occluded scenarios makes it difficult to distinguish objects, the modeled confidence provides valuable cues to distinguish between front and rear objects; the advanced observation center recovery module creates virtual boxes through backtracking, thereby promoting trajectory recovery and alleviating the distance problem caused by the inter-frame interval. The pseudo-depth metric integrates two stable and informative cues, height and depth, into the IoU calculation, improving the spatial perception ability and tracking accuracy of the tracker; the optional window denoising module can remove noise interference in some stable scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 is the overall flowchart of the robust multi-object tracking method in the present invention;

[0072] Figure 2 is the schematic diagram of the calculation of the pseudo-depth metric in the present invention;

[0073] Figure 3 is the schematic diagram of the advanced observation center recovery module in the present invention;

[0074] Figure 4 is the schematic diagram of the window denoising module in the present invention;

[0075] Figure 5 is the performance comparison of the method proposed in the present invention with state-of-the-art trackers and other related trackers on the DanceTrack dataset;

[0076] Figure 6 is the performance comparison of the method proposed in the present invention with state-of-the-art trackers and other related trackers on the MOT20 dataset;

[0077] Figure 7 is the comparison chart of the tracking results of the method of the present invention and state-of-the-art trackers. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0078] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0079] Refer to Figure 1, A robust multi-object tracking method based on strong cues and weak cues, including the following steps:

[0080] S1. Input a video stream, detect the objects in each frame of the input video stream, record the coordinate information of the upper left and lower right points of the object detection box and the detection confidence, and classify the objects into high-score detections, medium-score detections, and low-score detections according to the comparison between the detection confidence and a set threshold, and discard the objects with low-score detections; when the object detection box on each frame matches the prediction box obtained using historical information in the trajectory set, update the information of the corresponding matching trajectory in the trajectory set using the object detection box of the current frame.

[0081] Specifically as follows:

[0082] Use the general detector YOLOX and the public appearance detector SBS-50 to detect the image of the current video frame, and obtain the coordinate information of the upper left and lower right points of the object detection box in the image, the detection confidence, and the corresponding appearance embedding vector.

[0083] S2. Use the Kalman filter to predict the position of the existing trajectories in the trajectory set at the current frame, and at the same time use the robust trajectory confidence modeling module to predict the confidence of the trajectories at the current frame, specifically as follows:

[0084] When the detection confidence is relatively high under no occlusion or slight occlusion, we use the Kalman filter for continuous state estimation and use exponential smoothing to estimate the trajectory confidence, as shown in the following formula:

[0085]

[0086] Where:

[0087] is the trajectory confidence to be modeled at the current time t;

[0088] is the trajectory confidence at time t - 1;

[0089] is the trajectory confidence predicted by the Kalman filter at time t;

[0090] α is the smoothing exponent.

[0091] When occlusion occurs, the detection confidence is relatively low, and the Kalman filter cannot quickly capture the sudden change in confidence. Use the first-order and second-order difference linear correction algorithm to linearly predict the trajectory confidence.

[0092] As follows:

[0093]

[0094] Where:

[0095] is the confidence of the trajectory at time t-2;

[0096] The updated confidence is based on the linear trend of the previous confidence change, and the difference between the last two trends is used to correct the linear change, that is, the following formula, which enables the confidence to adapt faster and more accurately during occlusion:

[0097]

[0098] S3. Calculate the cost of the association between the trajectory set and the high-score detection target using the pseudo-depth metric PDIoU, weak confidence cues, weak motion direction cues, and strong appearance cues. After fusing into the embedding vector, perform matching through the Hungarian algorithm. The matched trajectories are sent to the trajectory manager for update, and the unmatched trajectories enter the second-stage association;

[0099] Specifically as follows:

[0100] The association cost consists of the sum of the confidence cost, appearance cost, motion direction cost, and pseudo-depth metric cost. The weak confidence cue calculates the association cost between the trajectory set and the detection target by the absolute value of the difference between their confidences, that is, the following formula:

[0101]

[0102] The strong appearance cue calculates the association cost between the trajectory set and the detection target by extracting the appearance embedding vectors of the two using the general detector SBS-50, and then calculating the negative of the cosine similarity of the two embedding vectors as the association cost C a , while the motion direction calculates the cosine distance between the motion directions of the trajectory set and the detection target as the association cost C v , both of which are calculated by general methods.

[0103] The pseudo-depth metric PDIoU incorporates the height weak cue and depth weak cue information into the strong cue IoU calculation. For specific reference Figure 2 , the left part shows the application scenario of the pseudo-depth, which is obtained by calculating the distance from the bottom of the detection box to the bottom of the image, and the right part depicts the visualization details of the PDIoU.

[0104] First, calculate the height intersection over union HIoU, as follows:

[0105] HIoU = y inter / y cover

[0106] Where:

[0107] yi nter is the length of the overlap between the height of the trajectory box and the height of the detection box;

[0108] y cover is the height of the minimum bounding box of the trajectory box and the detection box;

[0109] Although the Height Intersection over Union (HIoU) effectively utilizes the stable feature of height, there is still some spatial information that is not fully utilized.

[0110] During the tracking process, we observed that foreground objects may occlude background targets, indicating a correlation between occlusion relationships and depth information. However, directly obtaining the true depth is challenging. Specifically, referring to Figure 2 the left figure, we approximate the relative depth relationship between objects by measuring the distance from the bottom of the detection box to the bottom of the image, which is called pseudo-depth. First, calculate the pseudo-depths (d1, d2) of the trajectory set and the detection target, as well as the absolute value δd of the difference between the two;

[0111] As follows:

[0112]

[0113] δd = |d1 - d2|

[0114] Where:

[0115] H is the height of the video frame;

[0116] and are the y-axis coordinates of the lower right corners of the trajectory box and the detection box respectively;

[0117] Then calculate the thresholds for pseudo-depth reward and penalty. Give a penalty for those exceeding the threshold and a reward for those not exceeding the threshold to obtain the pseudo-depth related scores. The following two equations are as follows:

[0118]

[0119] Where:

[0120] i and j are the i-th trajectory set and the j-th detection target;

[0121] k, α1, and α2 are all hyperparameters;

[0122] Finally, PDIoU can be defined as the following equation. Take the negative of PDIoU to obtain the pseudo-depth cost:

[0123] PDIoU = PD · HIoU · IoU

[0124] Add the above association costs according to a certain weight to form an association cost embedding vector. Use the general Hungarian algorithm for association matching to obtain the successfully matched trajectory detection set, the unmatched trajectory set, and the unmatched high-score detections.

[0125] S4. Calculate the association cost between the remaining unmatched trajectory set and the mid-division detection target using the pseudo-depth metric PDIoU, confidence, weak motion direction cues, and strong appearance cues, and perform matching through the Hungarian algorithm. The matched trajectories are sent to the trajectory manager for update, similar to the operation in S3;

[0126] S5. Try to recover the trajectory association through the advanced observation center recovery module. Calculate the association cost between the unmatched high-score detections after backtracking and the unmatched trajectories after backtracking, as well as the association cost between the unmatched high-score detections after backtracking and the last observation box of the unmatched trajectories. Weight the average or multiply the two to obtain the association cost, and perform matching through the Hungarian algorithm. The matched trajectories are sent to the trajectory manager for update, as follows:

[0127] When some frames span multiple time steps, a large distance may be generated, making the association difficult.

[0128] To solve this problem, we appropriately reduce the time interval between the unmatched detection boxes and the previous observation frames of the unmatched trajectories. In addition, the same processing is performed on the unmatched trajectories and their respective previous observation frames.

[0129] Specifically refer to Figure 3 As shown, assume that at time T, there are m trajectory segments and n detections. Backtrack the i-th predicted trajectory to the last moment before it is lost;

[0130] Denoted as T - ΔT i , and the corresponding trajectory is marked as OBox i .

[0131] Use linear interpolation to obtain the virtual trajectory box PBox at the intermediate time T - βΔT i , where β ∈ [0, 1] is a backtracking factor affected by the target speed. i For the j-th detected trajectory, we apply the same method to OBox

[0132] . Interpolate at T - ΔT j to obtain the virtual trajectory box DBox at T - βΔT i , j and j .

[0133] We calculate the IoU between PBox i and DBox j , as well as the IoU between OBox i and DBox jThe IoU between them is used to obtain scores Score1 and Score2 respectively. The two scores are weighted averaged or multiplied to obtain the total association score, and after taking the negative value, it becomes the association cost.

[0134] S6. Send both the unmatched track set and the matched track set into the track manager. For the unmatched medium-score detections, they are directly discarded. For the unmatched high-score detections, a window denoising module is optionally used to remove noise according to the needs of the scenario. The remaining unmatched high-score detections are used by the track manager to create corresponding new tracks, as follows:

[0135] The track manager mainly consists of functions such as Kalman filter update, appearance update, robust track confidence modeling, track deletion, and new track creation.

[0136] For the tracks that successfully match the detections, the Kalman parameters and appearance are updated. For the active tracks that do not successfully match the detections, the Kalman parameters are frozen, and the Kalman parameters are updated again after re-matching to prevent the accumulation of parameter errors during the period of unmatched detections. Tracks that have not been matched for a long time are considered to have left the tracking scenario and are deleted by the track manager. For the unmatched high-score detections, if the tracking scenario changes little and the processed video frames have exceeded a certain number of frames, the window denoising module is used. A region in the video frame is selected as the denoising region according to the scenario. If the current frame number exceeds the threshold τ and a new detection box appears within the filtering window, it may be a false detection or a failure to associate and is considered a new target, then the detection is discarded.

[0137] Specifically refer to Figure 4 As shown, since the object fails to successfully associate with ID7, it is assigned a new ID (ID14).

[0138] In addition, due to the false detection of the detector, ID15 is wrongly generated.

[0139] However, through the window denoising module, these noises are effectively filtered. A denoising region is selected. If the current frame number exceeds the threshold τ and a new detection box appears within the filtering window, it may be a false detection or a failure to associate and is considered a new target, then the detection is discarded.

[0140] S7. End the tracking of the current frame, track the next frame, and repeat S1 - S6.

[0141] In Embodiment 1, refer to Figure 5 , to evaluate the robustness of the tracker we proposed against non-linear motion and occlusion, the present invention reports the performance comparison of the method proposed by the present invention with state-of-the-art trackers on DanceTrack. The methods using the detection results of the same general detector are classified below the dashed line.

[0142] The method proposed in the present invention ranks first among all trackers that do not use additional data for training, achieving the highest performance in almost all major tracking metrics. For methods using learnable matchers, although the results are also good, they have higher complexity compared to mainstream heuristic matchers and cannot meet the requirements of real-time online processing.

[0143] Its association metrics HOTA and IDF1 increased by 1.0 and 3.4 respectively compared to the state-of-the-art tracker Hybrid-SORT. These results highlight that effectively utilizing weak cues (such as pseudo-depth, height, and confidence) and strong cues (such as position and appearance) can significantly reduce ID switches and trajectory loss problems, demonstrating the good association performance of our model in complex motion scenarios.

[0144] Meanwhile, referring to Figure 6 , the method proposed in the present invention also ranks first in terms of the most important metric HOTA for tracking performance on the MOT20 dataset, reflecting the excellent generalization of the present invention.

[0145] Figure 7 Visual comparisons between the present invention and state-of-the-art trackers on the DanceTrack dataset are shown. In these challenging scenarios, the state-of-the-art trackers suffer from ID switches due to missed detections. In contrast, our method improves the robustness and effectiveness of tracking performance in complex motion scenarios such as occlusion and non-linear motion by comprehensively utilizing a robust trajectory confidence modeling module, an advanced observation center recovery module, a pseudo-depth metric, and a window denoising module, successfully tracking all targets.

[0146] Distinctive technical features of the present invention:

[0147] Robust trajectory confidence modeling module: It conducts more robust modeling and utilization of weak confidence cues, combines exponential smoothing denoising with first-order and second-order difference linear correction algorithms, and quickly adapts to occlusion scenarios through the historical confidence change trend. Existing technologies only update the trajectory state through fixed thresholds (such as detection confidence) and trajectory embedding scores, without involving denoising dynamic modeling of confidence and multi-order linear correction in occlusion scenarios.

[0148] Pseudo-depth metric: It fuses height and depth weak cues into the IoU strong cue, enhancing the spatial perception ability of the tracker. Existing technologies use traditional IoU-ReID fusion methods, only combining the IoU distance matrix with appearance features (such as cosine similarity), without utilizing height or pseudo-depth information.

[0149] Advanced Observation Center Recovery Module: By backtracking the trajectory through linear interpolation to reduce the position change of the trajectory caused by long-term occlusion, making it easier for the lost trajectory to be re-associated. The existing technology uses the observation amplification method, by expanding the width and height of the detection box to increase the IoU matching probability, and does not involve the backtracking mechanism of virtual box interpolation.

[0150] Window Denoising Module: In a stable scene, filter the unmatched high-score detections through a regional window, dynamically discard false detections. The existing technology's trajectory management is only based on the active / inactive state and the number of matches, and does not introduce prior knowledge to denoise the scene.

[0151] Overall Algorithm Structure Multi-Cue Phased Association: In the data association stage, hierarchically process high-score detections and medium-score detections, and fuse strong cues (appearance, position) and weak cues (confidence, motion direction, pseudo-depth, height) to form a multi-dimensional association cost matrix. The existing technology uses two-stage matching (active / inactive trajectories), only relying on IoU and appearance features (IoU-ReID), and does not introduce the difference in motion direction or confidence as the association basis.

[0152] Finally: The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A robust multi-target tracking method based on strong cues and weak cues, characterized in that; The following steps are involved: Step S1: Input a video stream, detect the target of each frame in the input video stream, record the coordinate information and detection confidence of the points in the upper left corner and lower right corner of the target detection frame, and classify the targets into high-score detection, medium-score detection and low-score detection according to the comparison between the detection confidence and the set threshold, and discard the targets with low-score detection; when the target detection frame on each frame matches the predicted frame obtained by using historical information in the trajectory set, use the target detection frame of the current frame to update the information of the corresponding matching trajectory in the trajectory set; Step S2: Use Kalman filtering to predict the position of the existing trajectory in the trajectory set in the current frame, and use the robust trajectory confidence modeling module to predict the confidence of the trajectory in the current frame; Step S3: Use the pseudo-depth metric PDIoU, confidence weak clues, motion direction weak clues, and appearance strong clues to calculate the cost of associating the trajectory set with the high-score detection target. After fusion into an embedded vector, match it through the Hungarian algorithm. The matched trajectories are sent to the trajectory manager for update, and the unmatched trajectories enter the second stage of association. Step S4: Use the pseudo-depth metric PDIoU, confidence, weak clues of motion direction, and strong clues of appearance to calculate the association cost of the remaining unmatched track set and the middle-point detection target, match them through the Hungarian algorithm, and send the matched tracks to the track manager for update, which is consistent with the operation of S3; Step S5: Try to restore the trajectory association through the advanced observation center recovery module, calculate the association cost after the failed high-score detection backtracking and the failed trajectory backtracking, and the association cost after the failed high-score detection backtracking and the last observation frame of the failed trajectory, and weighted average or multiply the two to get the association cost, match through the Hungarian algorithm, and send the matched trajectory to the trajectory manager for update; Step S6: Send both the unmatched trajectory set and the matched trajectory set to the trajectory manager, and directly discard the unmatched medium-score detections. For the unmatched high-score detections, the window denoising module can be used to remove noise according to the needs of the scene. The trajectory manager creates corresponding new trajectories for the remaining unmatched high-score detections. Step S7: End tracking of the current frame, track the next frame, and repeat steps S1 to S6.

2. A robust multi-target tracking method based on strong cues and weak cues according to claim 1, characterized in that: In step S1, the general detector YOLOX and the public appearance detector SBS-50 are used to detect the image of the current video frame to obtain the coordinate information of the points in the upper left corner and the lower right corner of the target detection box in the image, the detection confidence and the corresponding appearance embedding vector.

3. The robust multi-target tracking method based on strong cues and weak cues according to claim 1, characterized in that: In step S2, the detection confidence is high when there is no occlusion or slight occlusion. We use the Kalman filter for continuous state estimation and use exponential smoothing to estimate the trajectory confidence, as shown in the following formula: in: The confidence level of the trajectory that needs to be modeled at the current time t; is the trajectory confidence at time t-1; is the trajectory confidence predicted by the Kalman filter at time t; α is the smoothing exponent.

4. The robust multi-target tracking method based on strong cues and weak cues according to claim 1, characterized in that: In step S2, when occlusion occurs, the detection confidence is low, and the Kalman filter cannot quickly capture the sudden change of confidence. The first-order and second-order differential linear correction algorithms are used to linearly predict the trajectory confidence; As follows: in: is the confidence of the trajectory at time t-2; The updated confidence is based on the linear trend of the previous confidence change, using the difference between the last two trends to correct the linear change, which allows the confidence to adapt faster and more accurately during occlusion:

5. The robust multi-target tracking method based on strong cues and weak cues according to claim 1, characterized in that: In step S3: the association cost is composed of the sum of the confidence cost, appearance cost, motion direction cost and pseudo-depth metric cost. The confidence weak clue calculation trajectory set and the detection target association cost are calculated by the absolute value of the confidence difference between the two, that is, the following formula: The association cost between the trajectory set and the detection target is calculated by using the general detector SBS-50 to extract the appearance embedding vectors of the two, and then calculating the negative of the cosine similarity of the two embedding vectors as the association cost C. a , and the motion direction is calculated by a general method, the cosine distance between the trajectory set and the motion direction of the detected target as the association cost C v , both calculated by the general method.

6. The robust multi-target tracking method based on strong cues and weak cues according to claim 1, characterized in that: In step S3, the pseudo depth metric PDIoU integrates the height weak clue and depth weak clue information into the strong clue IoU calculation; First, calculate the high intersection-over-union ratio HIoU as follows: HIoU=y inter / and cover in: y inter It is the length of the overlap between the height of the trajectory frame and the height of the detection frame; y cover is the height of the minimum bounding box of the trajectory box and the detection box; The relative depth relationship between objects is approximated by measuring the distance from the bottom of the detection box to the bottom of the image, which is called pseudo depth; First, calculate the pseudo depth (d1, d2) of the trajectory set and the detected target and the absolute value of the difference between the two δd; As follows: δd=|d1-d2| in: H is the height of the video frame; and They are the y-axis coordinates of the lower right corner of the trajectory box and the detection box respectively; Then calculate the thresholds of pseudo-depth rewards and penalties, give penalties for those exceeding the threshold, and give rewards for those that do not exceed the threshold to get the pseudo-depth related scores, as shown in the following two formulas: in: i and j are the i-th trajectory set and the j-th detection target; k, α1, and α2 are all hyperparameters; Finally, PDIoU can be defined as the following formula, and the pseudo depth cost is obtained by taking the negative number of PDIoU: PDIoU=PD·HIoU·IoU The above association costs are added according to certain weights to form an association cost embedding vector. The general Hungarian algorithm is used for association matching, and a successfully matched trajectory detection set, an unmatched trajectory set, and an unmatched high-score detection set will be obtained.

7. The robust multi-target tracking method based on strong cues and weak cues according to claim 1, characterized in that: In step S5, assuming that at time T, there are m trajectory segments and n detections, the i-th predicted trajectory is traced back to the last moment before it is lost; Recorded as T-ΔT i , the corresponding trajectory is marked as OBox i ; Use linear interpolation to obtain the value of T-βΔT at the intermediate time i Virtual trajectory box PBox i , where β∈[0,1] is a backtracking factor affected by the target speed; For the jth detected tracklet, we apply the same method with OBox j .In T-ΔT i Interpolation is performed at T-βΔT j Virtual trajectory box DBox j ; We calculate PBox i With DBox j The IoU between them, and OBox i With DBox j The IoU between them is calculated to get scores Score1 and Score2 respectively. The total association score can be obtained by weighted average or multiplication of the two scores. The negative number is the association cost.

8. The robust multi-target tracking method based on strong cues and weak cues according to claim 1, characterized in that: In step S6, the Kalman parameters and appearance are updated for the trajectory that successfully matches the detection, while the Kalman parameters are frozen for the activated trajectory that fails to successfully match the detection, and the Kalman parameters are updated again after matching again to prevent the accumulation of parameter errors during the unmatched detection period; Tracks that have not been matched for a long time are considered to have left the tracking scene and are deleted by the track manager; For unmatched high-score detections, if the tracking scene does not change much and the number of processed video frames exceeds a certain number, the window denoising module is used to select an area in the video frame as the denoising area according to the scene. If the current frame number exceeds the threshold τ and a new detection box appears in the filter window, it may be a false detection or it may fail to be associated and is considered to be a new target, so the detection is discarded. Since the object failed to be successfully associated with ID7, it was assigned a new ID (ID14); In addition, ID15 was generated incorrectly due to the misdetection of the detector; However, these noises are effectively filtered through the window denoising module; Select a denoising area. If the current frame number exceeds the threshold τ and a new detection box appears in the filter window, it may be a false detection or it may be considered a new target due to failure to associate. In this case, the detection is discarded.