Online multi-target tracking method based on prediction residual driving

By predicting multiple motion models in parallel in a unified state space and utilizing model transfer and uncertainty modeling driven by prediction residuals, the robustness and stability issues of online multi-target tracking in complex scenarios are solved, achieving efficient perception and response to target motion and reducing the risk of false association.

CN122135049APending Publication Date: 2026-06-02TAIYUAN UNIVERSITY OF SCIENCE AND TECHNOLOGY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TAIYUAN UNIVERSITY OF SCIENCE AND TECHNOLOGY
Filing Date
2026-04-07
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing online multi-target tracking methods struggle to effectively perceive target motion uncertainties and maneuvering behaviors in complex scenarios, leading to trajectory interruptions or identity switching. Furthermore, the prediction residuals are not explicitly modeled as key indicators reflecting the reliability of target motion, affecting the reliability of data association.

Method used

Multiple motion models are predicted in parallel in a unified state space. By adaptively adjusting the model transition matrix driven by the prediction residual and modeling motion uncertainty, and combining the geometric overlap of the target and the consistency of the motion direction, a comprehensive association cost is constructed to achieve the synergistic optimization of data association and motion modeling.

Benefits of technology

It improves the robustness and stability of online multi-target tracking in complex scenarios, reduces computational complexity, enhances the ability to perceive and respond to target maneuvering behavior, and reduces the risk of false association.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135049A_ABST
    Figure CN122135049A_ABST
Patent Text Reader

Abstract

This invention discloses a prediction residual-driven multi-target tracking method, belonging to the field of computer vision technology. First, target detection is performed on the current frame. Based on the updated model transition matrix and model probabilities from the previous frame, multiple motion models are used to predict existing target trajectories in parallel within a unified state space, and the prediction results are weighted and fused to obtain the target prediction state. Motion constraints are constructed using the motion uncertainty index from the previous frame, and together with geometric overlap and motion direction consistency, they form an association cost, realizing data association between the trajectory prediction state and the detection boxes. Kalman filtering is performed on the successfully associated detection boxes to update the data, and the prediction residual is calculated. Based on the prediction residual, the model transition matrix is ​​adaptively adjusted, and an equivalent prediction residual is constructed to quantify the target motion uncertainty index for data association in subsequent frames. This method improves the robustness and stability of multi-target tracking in complex scenes without relying on appearance features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision technology, specifically relating to an online multi-target tracking method based on prediction residuals, and more particularly to a multi-target tracking method based on adaptive motion modeling and data association driven by prediction residuals, which is suitable for real-time target tracking applications in complex scenarios. Background Technology

[0002] Multi-target tracking technology is a key component of intelligent video analytics, widely used in scenarios such as video surveillance, autonomous driving, intelligent transportation, and crowd behavior analysis. Online multi-target tracking methods typically employ a "detection-tracking" paradigm, where the target detection result is first obtained in each frame, and then motion modeling and data association are used to maintain the target's identity over time.

[0003] In existing online multi-target tracking methods, target motion modeling is typically based on simplified linear motion assumptions, such as using a uniform motion model to predict the target state. These methods exhibit good real-time performance and stability when the target motion is stable and the scene is relatively simple. However, in real-world applications, target motion often exhibits significant non-stationarity and uncertainty. For example, targets may accelerate, decelerate, suddenly turn, or experience short-term anomalies in their motion state due to occlusion, interaction, or other factors. In such cases, a single motion model struggles to accurately describe the target's true motion behavior, easily leading to significant prediction bias, thus reducing the reliability of data association and potentially causing trajectory interruptions or target identity switching.

[0004] To enhance the ability to model complex motion behaviors, some methods have introduced multi-model motion modeling mechanisms, which maintain multiple motion models in parallel to characterize different types of target motion states. However, existing multi-model methods typically employ fixed model transfer strategies, exhibiting significant response lag when the target's motion pattern changes. This makes it difficult to adapt to the target's maneuvering behavior in a timely manner, thus limiting the practical effectiveness of multi-model methods in complex scenarios.

[0005] On the other hand, in the data association stage, existing online multi-target tracking methods typically introduce constraints based on the target's historical motion information to utilize the continuity and inertial characteristics of the target's motion. However, such motion constraints often implicitly assume that the target's motion is approximately linear. When the target is in a maneuvering state, or affected by factors such as short-term occlusion or dense interaction, motion constraints based on fixed empirical forms can easily amplify prediction errors, leading to an unreasonable increase in association costs and thus causing misassociation problems.

[0006] Furthermore, in existing technologies, prediction residuals are typically treated merely as intermediate computational costs in the filtering update process. They are not explicitly modeled as key indicators reflecting the reliability or uncertainty of target motion, nor are they uniformly incorporated into the two core stages of motion modeling and data association for collaborative utilization. This, to some extent, limits the system's ability to perceive the uncertainty and maneuvering behavior of target motion.

[0007] Therefore, how to effectively utilize prediction residual information without significantly increasing computational complexity, enhance the ability to perceive the uncertainty of target motion and maneuvering intentions, and integrate it into the motion modeling and data association process to improve the robustness of online multi-target tracking in complex scenarios has become a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0008] To address the aforementioned technical problems, this invention proposes an online multi-target tracking method based on predictive residuals. This invention enhances the perception and response capabilities to target maneuvering behavior through adaptive collaboration between predictive residual-driven motion modeling and data association. Without relying on appearance features, it significantly improves the robustness and stability of multi-target tracking in complex scenarios.

[0009] The technical solution protected by this invention is: an online multi-target tracking method based on prediction residuals, characterized by comprising the following steps:

[0010] S1. Video Input and Object Detection: Process the input video sequence frame by frame and use an object detection algorithm to obtain the set of object detection boxes for the current frame. ;

[0011] S2. Multi-model parallel prediction and motion state estimation: based on the model transition matrix updated from the previous frame. and model probability In a unified state space, at least two different motion models are used to predict the trajectory of an existing target in parallel, and the prediction results are then used to... Weighted fusion yields the predicted state of the target in the current frame. The calculation formula is as follows:

[0012]

[0013] in, Indicates the number of motion models;

[0014] S3. Based on motion uncertainty trajectory data association: Extract the target motion uncertainty index from the previous frame. During the data association stage, utilize Constructing motion constraint terms based on motion uncertainty And combined with the cost of the target geometric overlap relationship and the cost of consistency in motion direction Constructing a comprehensive related cost Complete the prediction state of the current frame. With the set of detection boxes The data association between them, and the overall association cost, are expressed as follows:

[0015]

[0016] in, The weights of the motion constraint terms;

[0017] S4. Observation-based filter update: Utilizing data to correlate successfully detected bounding boxes. Kalman filtering is performed on each motion model to update the posterior model probability of each motion model in the current frame. ;

[0018] S5. Prediction Residual Calculation: Based on the trajectory prediction status of each motion model and the successfully associated detection boxes, calculate the prediction residuals of each motion model in the current frame. The calculation formula is as follows:

[0019]

[0020] in, For the observation matrix, To predict measurement covariance;

[0021] S6. Residual-driven adaptive update of the model transition matrix: Based on the predicted residuals, the model transition matrix in the multi-model framework is updated. Perform online adaptive adjustments for use in model prediction of the next video frame;

[0022] S7. Motion uncertainty modeling based on prediction residuals: Constructing equivalent prediction residuals based on the prediction residuals of each motion model. Specifically, it is expressed as:

[0023]

[0024] in, For the first The prediction residuals of each model, For the first The posterior model probability of each model is calculated, and then the target motion uncertainty index is calculated based on the equivalent prediction residual. This information is saved and used in the data association stage of the next video frame to construct motion constraint terms based on motion uncertainty. The calculation formula is as follows:

[0025]

[0026] S8. Trajectory Lifecycle Management: Manages the trajectory set based on the association results of the current frame data. Perform lifecycle management, including trajectory initialization, trajectory maintenance, and trajectory termination.

[0027] Furthermore, step S2 specifically involves: the motion model including a uniform motion model and a uniformly accelerated motion model; the unified state space being represented by the target state vector as follows:

[0028]

[0029] in, Indicates the center location of the target. and These represent the scale and aspect ratio, respectively. For velocity components, This represents the acceleration component.

[0030] Furthermore, step S3 specifically involves: during the data association phase, only when the trajectory prediction state... The intersection-union ratio between the predicted bounding box and the detected bounding box obtained by mapping Meet the preset threshold conditions Only then are motion constraint terms constructed from motion uncertainty indices introduced. It participates in the association cost calculation; when the Intersection over Union (IoU) does not meet the threshold condition, the motion constraint term is not introduced. The motion constraint term is expressed as:

[0031]

[0032] in, This indicates an indicator function.

[0033] Furthermore, step S5 specifically involves: the predicted residual Based on the successfully associated detection boxes, the residuals between the predicted boxes obtained from the target trajectory prediction state mapping and the detection boxes are normalized using measurement covariance to obtain the prediction residuals in Mahalanobis distance form used to quantify motion consistency. The predicted residual The calculation is based on the two-dimensional center position component of the target. This is to reduce the noise impact introduced by scale changes or detection jitter.

[0034] Furthermore, step S6 specifically involves: predicting the residuals... Mapped to continuous model confidence weights And based on the model's credibility weights Model transition matrix in a multi-model framework Perform online adaptive adjustments. Represented as:

[0035]

[0036] in, This is the scale parameter.

[0037] Furthermore, the model credibility weights This is a non-probabilistic adjustment factor constructed based on the predicted residuals. Its value range is limited to a preset interval, and it is used to adjust the model transition matrix. Retention probability of each motion model and transition probability The specific formula is as follows:

[0038]

[0039]

[0040] in, This represents the lower bound of credibility.

[0041] Furthermore, step S7 specifically involves: analyzing the prediction residuals of each motion model... According to the corresponding model probability We perform weighted fusion to construct a continuous form of equivalent prediction residuals. ; and the equivalent prediction residual An uncertainty index for the target motion is constructed using a nonlinear mapping. It is used to characterize the uncertainty of the target's motion.

[0042] Furthermore, the uncertainty of target motion Motion constraint terms used to construct the association cost When the motion uncertainty index When increasing, the motion constraint term is increased. The degree of contribution of the motion uncertainty index to the association cost is determined to reduce the dependence of the association process on the consistency of motion direction; when the motion uncertainty index When reduced, the influence of the motion constraint term on the association cost is weakened, thereby enhancing the constraint of the association process on the consistency of historical motion.

[0043] Another technical solution to be protected by the present invention is a computer device, including a memory and a processor, wherein the memory stores a computer program, characterized in that the processor executes the computer program to implement the method steps of any one of claims 1 to 8.

[0044] Another technical solution to be protected by the present invention is: a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is used to implement the method of any one of claims 1 to 8.

[0045] Compared with the prior art, the present invention has the following advantages: The present invention constructs multiple motion models in a unified state space to predict existing target trajectories in parallel, and performs weighted fusion of prediction results based on the credibility of historical models before data association; after completing the data association between the prediction boxes and detection boxes obtained by trajectory prediction state mapping, the prediction residuals corresponding to each motion model are further used to characterize the uncertainty of target motion, and this uncertainty information is used for subsequent motion modeling and data association costs, thereby improving the robustness and stability of online multi-target tracking in complex scenarios without relying on target appearance features and without significantly increasing computational complexity. Attached Figure Description

[0046] Figure 1 This is an overall flowchart of the present invention.

[0047] Figure 2 Schematic diagram of multi-model parallel prediction and fusion structure.

[0048] Figure 3 Visualization examples of the tracking results of the method of the present invention on the DanceTrack dataset, where (a) frame 50, (b) frame 60, (c) frame 80, and (d) frame 90. Detailed Implementation

[0049] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0050] like Figure 1 As shown, the online multi-target tracking method based on prediction residuals proposed in this invention mainly includes the following steps:

[0051] Step S1, Video Input and Object Detection: First, the input video sequence is read frame by frame, and an object detection algorithm is executed on each frame to obtain a set of object detection boxes in the current frame. .

[0052] Specifically, object detection can be implemented using a deep learning-based object detection model, such as the anchorless bounding box YOLOX; this model is used to identify targets in the scene from the input image and output the corresponding bounding boxes and their confidence information. It should be noted that the object detection module is only used to provide observation input for the subsequent multi-target tracking process, and this invention does not limit the specific detection algorithm used.

[0053] For each frame of the image, the object detector outputs several detection results, each of which can be represented as an observation vector. ;in This indicates the two-dimensional position of the target center in the image coordinate system. The scale representing the target. This represents the detection confidence level of the target. All detection results constitute the target detection box set for the current frame. This data is then used as input for subsequent multi-target tracking algorithms.

[0054] Step S2, Multi-model Parallel Prediction and Motion State Estimation: Step 2 corresponds to Figure 2 The multi-model parallel prediction and fusion structure is shown. To unify the state representation under different motion models and ensure effective state interaction and probability fusion between different models, this embodiment uses a unified state space to model the target state. We construct a unified 9-dimensional state vector:

[0055]

[0056] in, Indicates the center location of the target. and These represent the scale and aspect ratio, respectively. For velocity components, This refers to the acceleration component. Through the aforementioned unified state modeling method, the uniform motion model (CV) and the uniformly accelerated motion model (CA) can be described in the same state space, thus avoiding the conversion problem between different state dimensions and providing a unified foundation for subsequent model interaction and probabilistic fusion. Based on this state space, this embodiment maintains multiple motion models simultaneously for the target trajectory in each frame and completes parallel filtering updates of multiple models through an interactive multi-model framework. The specific process includes the following sub-steps.

[0057] S2-1, Model Interaction and State Fusion: At time... Each motion model Each corresponds to a set of state estimates Covariance matrix and model probability Based on the model transition matrix from the previous time step. Calculate the mixture probability between each motion model. As shown in formula (2), the states of different models are weighted and fused to obtain the mixed initial states corresponding to each model. With the initial covariance matrix As shown in formulas (3) and (4):

[0058]

[0059]

[0060]

[0061] in Indicates the number of motion models. Represents the elements of the model transition matrix. Indicates from the model Transfer to model The mixed probability.

[0062] It should be noted that: in the first During the frame prediction stage, the model transition matrix The system uses the transition matrix obtained from the previous frame based on the adaptive update of the prediction residual; when processing the first frame, the system uses a preset fixed transition matrix.

[0063] S2-2, Parallel Prediction Using Multiple Motion Models: After obtaining the mixed initial state, the target state is predicted using each motion model separately. This embodiment employs two motion models: uniform velocity motion (CV) and uniform acceleration motion (CA). The predicted state for each model is obtained through a Kalman filter prediction step. and covariance matrix As shown in formulas (5) and (6):

[0064]

[0065]

[0066] in: This is the state transition matrix for the corresponding motion model; Let be the process noise covariance matrix.

[0067] This step allows us to obtain a set of predicted states under different motion models.

[0068] S2-3, Fusion Prediction State Generation: To be used in the subsequent data association process, the prediction results of each motion model need to be fused together. In this embodiment, the prediction results of each model are weighted and fused based on the model probability of the previous time step to obtain the prediction state of the target in the current frame. As shown in formula (7):

[0069]

[0070] This fusion prediction state is used to represent the comprehensive motion prediction of the target at the current moment and serves as an important basis for data association between the trajectory and the detection box.

[0071] Step S3: Correlation of trajectory data based on motion uncertainty: After obtaining the target prediction state fused in step S2, it is necessary to correlate the detection results of the current frame with the existing trajectory to determine the correspondence between the trajectory and the observation.

[0072] Specifically, for the first For a video image frame, let the set of detection boxes for the current frame be:

[0073]

[0074] in: This indicates the number of targets detected in the current frame; each detection box... Represented as:

[0075]

[0076] in: Indicates the center location of the target; This indicates the width and height of the detection frame; This indicates the confidence level of the detection.

[0077] For each existing trajectory Using the trajectory prediction state obtained in step S2 Construct predicted bounding boxes and build a comprehensive association cost with candidate bounding boxes in the detection box set. .

[0078] In this embodiment, the overall association cost is considered. Mainly due to the cost of target geometric overlap Cost of consistency with direction of motion and motion constraints based on motion uncertainty Together they constitute.

[0079] It should be noted that: in the data association process, the motion constraint cost is included in the association cost. Motion uncertainty index calculated based on the previous frame It is constructed and participates in the calculation of the comprehensive correlation cost. When processing the first frame, since the prediction residual has not yet been calculated, no motion uncertainty constraints are applied.

[0080] S3-1, Target Geometric Overlap Cost: First, calculate the Intersection over Union (IoU) between the predicted bounding box and the detected bounding box, as shown in Formula (8):

[0081]

[0082] in: Indicates the first Predicted bounding boxes for each trajectory; Indicates the first Each detection bounding box.

[0083] Based on the above overlap relationship, construct the target geometric overlap cost. As shown in formula (9):

[0084]

[0085] S3-2 Motion Direction Consistency Cost: To further utilize the target's historical motion information, this embodiment also constructs a motion direction consistency constraint term. Let the target's motion direction vector in the historical frame be... The current detected direction vector relative to the historical trajectory is Consistency cost of movement direction Represented as formula (10):

[0086]

[0087] in: , and The coordinates of the center point observed at two different times are two-dimensional coordinates.

[0088] S3-3, Threshold-guided motion uncertainty constraints: Based on the motion uncertainty index calculated in the previous frame... Subsequently, in the data association stage, this embodiment constructs motion constraint terms based on motion uncertainty.

[0089] Considering that motion constraints are only meaningful when candidate matches have basic geometric rationality, this embodiment introduces the intersection-over-union (IoU) threshold between the predicted box and the detected box as a prerequisite.

[0090] Let the trajectory With detection The intersection-union ratio is ,when Only then are constraints imposed on the associated costs. The motion constraint term is shown in formula (11):

[0091]

[0092] in: Indicates an indicator function, only when Exceeding the threshold The compensation will only take effect at that time.

[0093] Through the above mechanism, when the target motion prediction is stable, the motion uncertainty is small and the impact of motion constraints on the association results is weak. However, when the target motion is unstable, by adjusting the contribution of the motion constraint terms, the dependence of the association process on the consistency of a single historical motion direction can be reduced, and the matching robustness in complex motion situations can be improved.

[0094] Combining the aforementioned cost items, the final comprehensive associated cost is... Define formula (12):

[0095]

[0096] in: The weights are the weights of the motion constraint terms.

[0097] After obtaining the correlation cost matrix, S3-4 uses the Hungarian Algorithm to solve for the optimal matching relationship, thereby obtaining the trajectory set. With detection set The matching results between them.

[0098] Step S4, Observation-Based Filtering Update: After completing multi-model prediction, the system retains the predicted state and covariance matrix of each motion model for subsequent updates. When subsequent steps complete the data association between the trajectory and observations, if a trajectory successfully matches a detection box, the detection box is used to perform Kalman filtering updates on each motion model, thereby obtaining the updated state estimate. and Kalman gain As shown in formulas (13) and (14):

[0099]

[0100]

[0101] in: The observation matrix; To observe the noise covariance matrix; This is the currently matched detection box.

[0102] Simultaneously, the observation likelihood function is calculated based on the matched observations of each motion model, and the model probability is updated using Bayesian rules, as shown in formula (15):

[0103]

[0104] in Representation Model Likelihood probability for the current observation.

[0105] S5. Prediction Residual Calculation: After completing step S4 (filter update), to characterize the motion consistency between the prediction results of each motion model and the detection observations, this embodiment further calculates the prediction residuals of each motion model. Specifically, for the trajectory... In motion model Predicted state under According to the observation matrix The predicted state is mapped to the measurement space and successfully associated with the detection box in the data. By comparing the results, the predicted residuals can be obtained. As in formula (16):

[0106]

[0107] in The center position of the detection box indicating successful data association; This is the observation matrix. This embodiment only uses the target's two-dimensional center position. It is used in residual calculation to reduce the noise impact caused by changes in the detection frame size or detection jitter.

[0108] Subsequently, to eliminate the influence of uncertainties in different states on the residual scale, the residual vector is normalized for covariance. The predictive measurement covariance is defined by formula (17):

[0109]

[0110] Representation Model The predicted covariance matrix; This represents the measurement noise covariance matrix.

[0111] Based on this, the prediction residuals of each motion model are calculated using Mahalanobis distance, as shown in formula (18):

[0112]

[0113] income The prediction residuals are in Mahalanobis distance form, used to quantify the degree of consistency between the predicted target state and the detected observations. When When the value is small, it indicates that the current motion model's prediction is largely consistent with the observed data; when... A larger value indicates a significant discrepancy between the current model prediction and actual observation.

[0114] In this invention, the predicted residual This will be further used in subsequent steps for adaptive adjustment of the model transition matrix and construction of the target motion uncertainty index.

[0115] S6. Residual-driven adaptive update of the model transition matrix: To improve the system's responsiveness to changes in the target motion pattern, this embodiment further utilizes the prediction residuals to adaptively adjust the model transition matrix online. First, the prediction residuals are mapped to model confidence weights, as shown in formula (19):

[0116]

[0117] Then, the model transition matrix is ​​dynamically adjusted using this credibility weight. retention probability As shown in formulas (20) and (21):

[0118]

[0119]

[0120] in: For scale parameters; This represents the lower bound of credibility.

[0121] Through the residual-driven adaptive model transfer matrix mechanism described above, this invention can automatically adjust the switching behavior between different motion models based on real-time detection consistency, thereby improving the system's response to target maneuvering motion and reducing the model switching lag problem caused by traditional fixed transfer matrices.

[0122] Step S7: Motion uncertainty modeling based on prediction residuals: In complex scenes, the target may accelerate, turn, or be affected by target interactions, which will significantly increase the prediction error of the motion model. To characterize this motion uncertainty, this embodiment further utilizes prediction residuals to construct a motion uncertainty index. After the trajectory and observation are matched and the filtering update results are obtained, this embodiment further utilizes the prediction residuals of each motion model to evaluate the target motion uncertainty, and uses this uncertainty information in the data association process of subsequent frames to improve the robustness of the algorithm under complex motion conditions.

[0123] S7-1. Calculation of Equivalent Prediction Residuals: To obtain a unified measure of motion uncertainty, this embodiment performs weighted fusion of the residuals of each model to construct equivalent prediction residuals. Let the first The Mahalanobis distance generated by each motion model is: Its corresponding posterior model probability is Then the equivalent residual is defined by formula (22):

[0124]

[0125] in: The number of motion models; Indicates the first The model probability corresponding to each model.

[0126] S7-2, Quantification of Motion Uncertainty: Obtaining Equivalent Prediction Residuals Subsequently, this embodiment further constructs a motion uncertainty index to describe the reliability of the target's motion state. The motion uncertainty index is defined as shown in formula (23):

[0127]

[0128] When the target motion is stable and the prediction is accurate, the equivalent residual is small, corresponding to a high motion similarity, and the motion uncertainty is low. When the target accelerates, turns, or is affected by complex interactions, the prediction residual increases significantly, and the corresponding motion uncertainty index increases accordingly. Through the above modeling, the prediction residual is explicitly enhanced into a continuous quantity characterizing the target motion credibility and maneuvering intention, providing a direct and stable basis for the adaptive adjustment of subsequent association costs.

[0129] Step S8, Trajectory Lifecycle Management: After completing the data association in Step S4, the trajectory set needs to be updated and managed to maintain the stable existence of valid trajectories in the system and to properly handle newly appearing and disappearing targets. This embodiment constructs a trajectory lifecycle management mechanism through three stages: trajectory initialization, trajectory maintenance, and trajectory termination.

[0130] S8-1, Trajectory Initialization

[0131] After data association is completed, some unmatched detection targets still exist in the current frame. These unmatched detection sets... If the detection confidence of a target in the data is higher than a preset threshold, a new target trajectory is created for it.

[0132] Specifically, for unmatched detection The initial trajectory state is:

[0133]

[0134] in: Initialized from the center position of the detection frame; Initialize the detection frame size and aspect ratio; velocity term. With acceleration term Initialized to zero; covariance matrix Set the initial uncertainty as preset.

[0135] At the same time, a unique trajectory number is assigned to the new trajectory, and the trajectory survival time and match count are initialized.

[0136] S8-2, Track Maintenance

[0137] For the set of trajectories that successfully match the observation in the current frame The trajectory status is updated using the filtering update result in step S4, and the trajectory loss counter is reset to zero.

[0138] For the set of unmatched detection trajectories This embodiment allows the trajectory to continue existing within a certain time window and maintains its trajectory state through motion prediction. Specifically, the trajectory loss counter is incremented by 1, and when the counter is less than the maximum allowed number of lost frames... At that time, the trajectory remains in the system and continues to participate in data association in the next frame.

[0139] This mechanism can maintain trajectory continuity when the target is briefly occluded or detection fails.

[0140] S8-3, Track Termination

[0141] When the undifferentiated trajectory set If a certain trajectory fails to match a detection for multiple consecutive frames, and its loss counter exceeds the maximum allowed threshold... If the target is deemed to have left the current scene or tracking has failed, the trajectory is deleted from the system.

[0142] The above strategies can avoid retaining invalid trajectories for a long time, thereby reducing the computational burden on the system and improving the overall tracking stability.

[0143] In addition, embodiments of the present invention also provide an electronic device, which includes: a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the aforementioned online multi-target tracking method based on prediction residual driving.

[0144] Specifically, the processor may be a CPU (Central Processing Unit), an ASIC (Application-Specific Integrated Circuit), or one or more integrated circuits configured to implement embodiments of the present invention; the memory is used to store programs that can run on the processor, and the memory may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk storage; the program may include program code, which includes computer-executable instructions; and the communication interface is used for communication between the storage and the processor.

[0145] This invention also provides a computer-readable storage medium storing computer instructions for causing the computer to execute the online multi-target tracking method based on prediction residuals in the foregoing method embodiments.

[0146] The method of the present invention has been described in detail above. In order to verify the effectiveness of the method of the present invention, the method of the present invention will be compared with the classical algorithm below.

[0147] Table 1 compares the performance metrics of the proposed method with classic multi-object tracking algorithms such as SORT, ByteTrack, and OCSort on the Dancetrack test set. HOTA, AssA, MOTA, and IDF1 represent tracking performance. Comparative experiments show that the proposed method can effectively improve tracking progress and robustness in different scenarios while ensuring online tracking.

[0148]

[0149] The visualization structure of this invention on the dancetrack dataset is as follows: Figure 3 As shown in the figure. Frames 50, 60, 80 and 90 of the video sequence are selected as typical frames for visualization, corresponding to Figures (a) to (d) respectively.

[0150] As shown in Figure (a), targets numbered 1, 5 and 8 gradually approach each other in spatial position in frame 50, and show obvious motion changes in subsequent frames, while there is a potential occlusion relationship.

[0151] As shown in Figure (b), in frame 60, target 5 and target 1 are occluded, causing target 1 to be briefly missing in the detection results.

[0152] As shown in Figure (c), in frame 80, target 8 and target 5 are occluded, causing target 5 to be briefly lost during this phase.

[0153] During the time period shown in Figures (b) to (c), although target 1 experienced occlusion and detection loss, the method of the present invention was still able to restore its original identity when it reappeared.

[0154] As shown in Figure (d), in frame 90, target 5 re-enters the detection area, and the identities of target 1, target 5 and target 8 remain consistent, with no identity switching.

[0155] The above results show that, even when targets undergo motion changes and mutual occlusion, the method of the present invention can maintain stable trajectory associations and achieve accurate identity recovery when the target reappears, thereby improving the association stability and tracking continuity in the multi-target tracking process.

[0156] Through the above steps, this invention realizes an online multi-target tracking method driven by predictive residuals. It is understood that the above specific description of this invention is only for illustrative purposes and is not intended to limit the technical solutions described in the embodiments of this invention. Those skilled in the art should understand that modifications or equivalent substitutions can still be made to this invention to achieve the same technical effects; as long as the usage requirements are met, they are all within the protection scope of this invention.

Claims

1. An online multi-target tracking method based on prediction residuals, characterized in that, Includes the following steps: S1. Video Input and Object Detection: The input video sequence is processed frame by frame, and the set of object detection boxes for the current frame is obtained using an object detection algorithm. ; S2. Multi-model parallel prediction and motion state estimation: based on the model transition matrix updated from the previous frame. and model probability In a unified state space, at least two different motion models are used to predict the trajectory of an existing target in parallel, and the prediction results are then used to... Weighted fusion yields the predicted state of the target in the current frame. The calculation formula is as follows: ; in, Indicates the number of motion models; S3. Based on motion uncertainty trajectory data association: Extract the target motion uncertainty index from the previous frame. During the data association stage, utilize Constructing motion constraint terms based on motion uncertainty And combined with the cost of the target geometric overlap relationship and the cost of consistency in motion direction Constructing a comprehensive related cost Complete the prediction state of the current frame. With the set of detection boxes The data association between them, and the overall association cost, are expressed as follows: ; in, The weights of the motion constraint terms; S4. Observation-based filter update: Utilizing data to correlate successfully detected bounding boxes. Kalman filtering is performed on each motion model to update the posterior model probability of each motion model in the current frame. ; S5. Prediction Residual Calculation: Based on the trajectory prediction status of each motion model and the successfully associated detection boxes, calculate the prediction residuals of each motion model in the current frame. The calculation formula is as follows: ; in, For the observation matrix, To predict measurement covariance; S6. Residual-driven adaptive update of the model transition matrix: Based on the predicted residuals, the model transition matrix in the multi-model framework is updated. Perform online adaptive adjustments for use in model prediction of the next video frame; S7. Motion uncertainty modeling based on prediction residuals: Constructing equivalent prediction residuals based on the prediction residuals of each motion model. Specifically, it is expressed as: ; in, For the first The prediction residuals of each model, For the first The posterior model probability of each model is calculated, and then the target motion uncertainty index is calculated based on the equivalent prediction residual. This information is saved and used in the data association stage of the next video frame to construct motion constraint terms based on motion uncertainty. The calculation formula is as follows: ; S8. Trajectory Lifecycle Management: Manages the trajectory set based on the association results of the current frame data. Perform lifecycle management, including trajectory initialization, trajectory maintenance, and trajectory termination.

2. The online multi-target tracking method based on prediction residual driving according to claim 1, characterized in that, Specifically, step S2 involves: the motion model includes a uniform motion model and a uniformly accelerated motion model; the unified state space is represented by the target state vector as follows: ; in, Indicates the center location of the target. and These represent the scale and aspect ratio, respectively. For velocity components, This represents the acceleration component.

3. The online multi-target tracking method based on prediction residual driving according to claim 2, characterized in that, Specifically, step S3 involves: during the data association phase, only when the trajectory prediction state is... The intersection-union ratio between the predicted bounding box and the detected bounding box obtained by mapping Meet the preset threshold conditions Only then are motion constraint terms constructed from motion uncertainty indices introduced. It participates in the association cost calculation; when the Intersection over Union (IoU) does not meet the threshold condition, the motion constraint term is not introduced. The motion constraint term is expressed as: ; in, Indicates an indicator function.

4. The online multi-target tracking method based on prediction residual driving according to claim 3, characterized in that, Specifically, step S5 involves: the predicted residual Based on the successfully associated detection boxes, the residuals between the predicted boxes obtained from the target trajectory prediction state mapping and the detection boxes are normalized using measurement covariance to obtain the prediction residuals in Mahalanobis distance form used to quantify motion consistency. The predicted residual The calculation is based on the two-dimensional center position component of the target. This is to reduce the noise impact introduced by scale changes or detection jitter.

5. The online multi-target tracking method based on prediction residual driving according to claim 4, characterized in that, Specifically, step S6 involves: predicting the residual... Mapped to continuous model confidence weights And based on the model's credibility weights Model transition matrix in a multi-model framework Perform online adaptive adjustments. Represented as: ; in, This is the scale parameter.

6. The online multi-target tracking method based on prediction residual driving according to claim 5, characterized in that, The model credibility weight This is a non-probabilistic adjustment factor constructed based on the predicted residuals. Its value range is limited to a preset interval, and it is used to adjust the model transition matrix. Retention probability of each motion model and transition probability The specific formula is as follows: ; ; in, This represents the lower bound of credibility.

7. The online multi-target tracking method based on prediction residual driving according to claim 6, characterized in that, Specifically, step S7 involves: analyzing the prediction residuals of each motion model... According to the corresponding model probability We perform weighted fusion to construct a continuous form of equivalent prediction residuals. ; and the equivalent prediction residual An uncertainty index for target motion is constructed using a nonlinear mapping. It is used to characterize the uncertainty of the target's motion.

8. The online multi-target tracking method based on prediction residual driving according to claim 7, characterized in that: Target motion uncertainty Motion constraint terms used to construct the association cost When the motion uncertainty index When increasing, by increasing the motion constraint term The degree of contribution of the motion uncertainty index to the association cost is determined to reduce the dependence of the association process on the consistency of motion direction; when the motion uncertainty index When reduced, the influence of the motion constraint term on the association cost is weakened, thereby enhancing the constraint of the association process on the consistency of historical motion.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method steps of any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the method of any one of claims 1 to 8.