Multi-target tracking method and device based on multi-similarity matrix fusion

By adopting the multi-similarity matrix fusion method in the multi-objective tracking algorithm, combining SIFT and WRN feature extractor and Gaussian smooth interpolation method, the ID switching problem of unstable target appearance features in complex scenarios is solved, and higher tracking accuracy and robustness are achieved.

CN120013991APending Publication Date: 2025-05-16ZHEJIANG UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510077525.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing multi-objective tracking algorithm is difficult to accurately track target appearance characteristics caused by factors such as blur, small size, and occlusion in complex scenarios, resulting in ID switching problems.

Method used

Using a method based on multi-similarity matrix fusion, the target's eigenvector is obtained through SIFT and WRN feature extractors, the Mahayana distance and cosine distance are calculated, the similarity matrix is ​​constructed and the weight fusion is performed. Combined with the jiyali algorithm and Gaussian smooth interpolation method, the target's precise tracking is achieved.

Benefits of technology

In complex scenarios such as small target size, occlusion, and blur, tracking accuracy and robustness are improved, ID switching problems are reduced, and the performance of multi-target tracking is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013991A_ABST
    Figure CN120013991A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-target tracking method and device based on multi-similarity matrix fusion, and the method comprises the steps: inputting an input image into a target detection model, obtaining a detection frame of each target in the image, predicting the motion track of a previous frame of target through a Kalman filter, and obtaining a prediction frame; inputting the detection box and the prediction box into appearance model WRN and SIFT feature extractors to obtain corresponding SIFT features and WRN features; calculating a mahalanobis distance and a cosine distance between a detection frame target and a prediction frame target, constructing a corresponding WRN feature similarity matrix and a corresponding SIFT feature similarity matrix, and performing fusion according to a certain weight to obtain a final cost matrix; matching the targets through a Hungary algorithm to obtain a first matching result, and carrying out IOU matching on unmatched targets and unmatched detection in the first matching result to obtain tracks of the targets; and obtaining a final trajectory of the target through a connection model irrelevant to the appearance and a Gaussian smooth interpolation method. The method shows good performance in complex scenes such as small target size, shielding and blurring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of multi-target tracking, and relates to a multi-target tracking method and device based on multi-similarity matrix fusion. Background Art

[0002] Computer vision is the process of simulating biological vision through computers and related equipment, and studies how to enable machines to have the ability to "see". As an important branch of computer vision, multi-target tracking is not only of great significance for improving the understanding and analysis capabilities of video content, but also plays an important role in ensuring public safety, optimizing traffic management, and improving human-computer interaction experience. The goal of multi-target tracking is to accurately identify, locate and track multiple targets of interest, such as pedestrians and vehicles, from complex video streams. By continuously tracking the motion trajectories of these targets, rich dynamic information can be obtained, such as the position, speed, acceleration, etc. of the target, thereby achieving accurate analysis and prediction of the target behavior.

[0003] With the widespread application of deep learning in various computer vision tasks, multi-target tracking algorithms based on deep learning and detection can be divided into detection-based tracking and joint detection tracking. The detection-based tracking mode has the main advantages of modular design and flexibility, separating the target detection and target tracking tasks so that each task can be optimized independently. Especially when utilizing deep learning technology, the detection-based tracking paradigm can effectively utilize advanced detection algorithms to achieve accurate target positioning and perform well in dealing with complex scenes involving object occlusion, deformation or scale changes. In contrast, the joint detection and tracking paradigm emphasizes global optimization and computational efficiency, combining detection and tracking tasks together, sharing feature extraction networks, and forming a unified optimization framework. Its goal is to find the best solution in a global environment, thereby reducing the computational burden associated with independent detection and tracking processes. Shared feature calculations not only reduce redundant calculations, but also increase algorithm speed and improve overall tracking efficiency.

[0004] Although the above studies have improved the accuracy of multi-target tracking in different ways, the appearance features of the target will become unstable during the target tracking process due to factors such as blur, small size and occlusion of the target, resulting in ID switching problems in the tracking task. Therefore, extracting more accurate target appearance features is crucial for calculating the similarity measurement between targets. However, the existing algorithms use simple convolutional features to extract target appearance feature information, which does not fully describe the target appearance features and cannot meet the needs of actual application scenarios. Summary of the invention

[0005] The present invention aims to solve the above problems in the prior art and proposes a multi-target tracking method and device based on multi-similarity matrix fusion.

[0006] To achieve the above object, the technical solution adopted by the present invention is:

[0007] A first aspect of the present invention proposes a multi-target tracking method based on multi-similarity matrix fusion, comprising:

[0008] S1. Input the input image into the target detection model to obtain the detection frame of each target in the image, and predict the motion trajectory of the target in the previous frame through the Kalman filter to obtain the predicted frame.

[0009] S2. Input the detection box and the prediction box into the appearance model WRN and SIFT feature extractor to obtain the corresponding SIFT features and WRN features.

[0010] S3. Calculate the Mahalanobis distance and cosine distance between the detection box and the prediction box target, construct the corresponding WRN feature similarity matrix and SIFT feature similarity matrix, and fuse them according to certain weights to obtain the final cost matrix.

[0011] S4. Match the target through the Hungarian algorithm to obtain the first matching result, and perform IOU matching on the unmatched targets and unmatched detections in the first matching result to obtain the trajectory of the target.

[0012] S5. The final trajectory of the target is obtained through the appearance-independent connection model and Gaussian smoothing interpolation method.

[0013] The second aspect of the present invention relates to a multi-target tracking device based on multiple similarity matrix fusion, comprising a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the multi-target tracking method based on multiple similarity matrix fusion of the present invention.

[0014] A third aspect of the present invention relates to a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the multi-target tracking method based on multi-similarity matrix fusion of the present invention.

[0015] Aiming at the unstable appearance features of the target caused by blurry targets, small target size, and occlusion, which easily cause the ID switching problem in the tracking task, the present invention introduces the fusion of SIFT and WRN similarity matrices to enrich the description of the target's appearance information. In addition, the present invention also uses a connection model that is independent of appearance to associate short trajectories into complete trajectories, and fills the trajectory gaps caused by missing detection through Gaussian smoothing interpolation. It shows good performance in complex scenes such as small target size, occlusion, blur, etc., has better robustness and tracking accuracy, and can be better applied in the field of target tracking.

[0016] The advantages of the present invention are: it shows good performance in complex scenes such as small target size, occlusion, blur, etc., has better robustness and tracking accuracy, and can be better applied in the field of target tracking. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a flow chart of the method of the present invention;

[0018] Figure 2 is a flow chart of step S2 in the method of the present invention;

[0019] Figure 3 is a flow chart of step S3 in the method of the present invention;

[0020] Figure 4 Schematic diagram of the process of step S3 in the present method;

[0021] Figure 5 Schematic diagram of the process of step S4 in the method of the present invention;

[0022] Figure 6 It is a schematic diagram of the structure of the connection model of the present invention that is independent of appearance. DETAILED DESCRIPTION

[0023] The following is a detailed description of an example of the present invention in conjunction with the accompanying drawings: This example is implemented on the premise of the technical solution of the present invention, and a detailed implementation method and specific operation process are given.

[0024] Example 1

[0025] like Figure 1 As shown, this embodiment provides a multi-target tracking method based on multi-similarity matrix fusion, comprising the following steps:

[0026] S1. Input the input image into the trained target detection model to obtain the detection frame of each target in the image. Use the Kalman filter to predict the motion trajectory of the target in the previous frame to obtain the predicted frame.

[0027] S2. Input the detection box and the prediction box into the appearance model WRN and SIFT feature extractor, and extract the SIFT feature vector of the detection box, the SIFT feature vector of the prediction box, the WRN feature vector of the detection box, and the WRN feature vector of the prediction box respectively.

[0028] Among them, the appearance model WRN has 2 convolutional layers, 6 residual blocks and 1 Dense layer. Each convolutional layer uses a 3×3 convolution kernel. After the convolutional layer, the network performs a pooling operation. Next is a six residual block structure, each of which contains multiple convolutional layers, and the input and output are added through residual connections. After extracting the 128-dimensional feature vector, the features are mapped using L2 normalization. The SIFT feature extractor extracts the SIFT feature vectors of the targets in the detection box and the prediction box through the process of scale space extreme value detection, key point location, key point direction determination, and key point description, and then flattens them into a one-dimensional vector and performs BatchNorm normalization on the SIFT feature vector.

[0029] S3, such as Figure 3 , Figure 4 As shown in the figure, the Mahalanobis distance and cosine distance between the detection box and the prediction box target are calculated, the corresponding WRN feature similarity matrix and SIFT feature similarity matrix are constructed, and the final cost matrix is ​​obtained by fusion according to a certain weight. Specifically:

[0030] The Mahalanobis distance between the detection box and the prediction box, the cosine distance between the detection box SIFT feature and the prediction box feature, and the cosine distance between the detection box WRN feature and the prediction box WRN feature are calculated, and then the WRN feature similarity matrix and the SIFT feature similarity matrix are constructed based on these three distances.

[0031] The Mahalanobis distance calculation formula is as follows:

[0032]

[0033] Where i is the i-th track, j is the j-th detection, and y i , S i are the predicted position and predicted covariance matrix of a target in a certain frame, d j It is the detection box information of the next frame.

[0034] The formula for calculating cosine distance is as follows:

[0035]

[0036] Among them, r j is the appearance feature vector of the jth detection, are all appearance feature vectors of the i-th detection.

[0037] After obtaining the two similarity matrices, they are fused according to the weights to obtain the final cost matrix.

[0038] c i,j =λc (1) (i,j)+(1-λ)c (2)(i,j) (3)

[0039] Among them, λ is the weight coefficient, c (1) (i,j) is the WRN feature similarity matrix, c (2) (i,j) is the WRN feature similarity matrix.

[0040] S4, such as Figure 5 As shown, after the cost matrix is ​​obtained, the Hungarian algorithm is used to increase the number of frames that have not been successfully matched from 0 to A. max The trajectory is matched cyclically, and the frames with the smallest number of unsuccessful matches are given priority for optimal matching. The successfully matched trajectory and detection need to be within the threshold of the gate matrix, and the number of cycles is greater than A. max When , the loop matching ends, and the unmatched trajectories and detections are returned to obtain the matching results. The unmatched trajectories and unmatched detections in the matching results are matched again by IOU. Undetermined trajectories and trajectories with a lifespan greater than the set maximum are deleted. The two successfully matched trajectories are filtered through Kalman filtering, and the unmatched detections are created as new trajectories, which are updated to the trajectory of the next frame. After the matching is completed, the final trajectory of each target is obtained.

[0041] S5: After obtaining a series of tracks of the target, use the appearance-independent connection model to associate short tracks into complete tracks. At the same time, Gaussian smoothing interpolation is used to fill in track gaps caused by missing detections. Specifically:

[0042] Figure 6 The two-branch framework of the AFLink model is presented, which provides an efficient and accurate method for trajectory association. i and T j As input, the trajectory From the latest N = 30 frames f k and the corresponding position (x k ,y k ), for those trajectories with a length of less than 30 frames, the model uses zero padding to ensure the consistency of the input and facilitate subsequent processing. Next, the model uses a 7×1 convolution kernel to perform a convolution operation on the input trajectory along the time dimension in the timing module to extract time-related features. Subsequently, the model integrates information from different feature dimensions (i.e., f, x, y) through a 1×3 convolution operation in the fusion module. After processing by the fusion module, the two feature maps obtained are converted into feature vectors after pooling and compression, and then the two feature vectors are concatenated. These feature vectors contain rich spatiotemporal information. Finally, the model uses a multi-layer perceptron to predict the confidence score of the trajectory association.

[0043] The timing module consists of 4 convolutional layers with kernel size 7×1 and output channels of {32, 64, 128, 256}. Each convolution is followed by a BN layer and a ReLU activation layer. The fusion module includes a 1×3 convolutional layer, a BN layer, and a ReLU activation layer. The classifier is a multilayer perceptron with two fully connected layers and a ReLU layer inserted in between.

[0044] During training, the association process is formulated as a binary classification task. Then, it is optimized using the binary cross entropy loss:

[0045]

[0046] where x n ∈[0,1] is the association probability of the predicted sample pair, y n ∈{0,1} is the sample label value.

[0047] In the final reasoning, a temporal distance threshold of 30 frames and a spatial distance threshold of 75 pixels are used to filter out unreasonable association pairs. If the prediction score of an association pair is greater than 0.95, the trajectory is considered to be associated.

[0048] The core idea of ​​the Gaussian smoothing interpolation method is to use the Gaussian function to smooth the trajectory data and interpolate based on the smoothed data. First, the Gaussian weighted average of the trajectory data is calculated. This process takes into account the spatial and temporal relationship between the data points, as well as their weights. The weight is usually determined based on the distance between the data point and the point to be interpolated or other related factors. The closer the distance is, the higher the weight is. Through Gaussian weighting, Gaussian smoothing interpolation can better capture the local characteristics and changing trends of the trajectory. After obtaining the smoothed trajectory data, Gaussian smoothing interpolation uses this data to estimate the value of the missing position. During the interpolation process, not only the values ​​of the known data points are considered, but also the spatial and temporal relationship between them, as well as the overall trend of the trajectory. This enables Gaussian smoothing interpolation to generate interpolation results that are closer to the true trajectory.

[0049] Gaussian smoothing interpolation models the i-th trajectory as follows:

[0050] p t =f (i) (t)+∈ (5)

[0051] Where t is the frame in the frame set F, p t is the position coordinate of the tth frame, p t ∈P, ∈ is Gaussian noise, obeying N(0,σ 2 ).

[0052] Given a trajectory of length L By fitting the function f (i) (t) to solve the nonlinear motion modeling problem. Assume that it obeys the Gaussian process f (i) (t)∈GP(0,k(·,·)), where is the radial basis kernel function. According to the properties of the Gaussian process, given a new frame set F * In the case of smoothing position P * A prediction was made.

[0053] P * =K(F * ,F)(K(F,F)+σ 2 I) -1 P (6)

[0054] where K(·,·) is the covariance function based on k(·,·).

[0055] In addition, the hyperparameter λ controls the smoothness of the trajectory, and its size is related to the trajectory length l.

[0056] λ=τ*log(τ 3 / l) (7)

[0057] The value of τ is set to 10.

[0058] Example 2

[0059] The present embodiment relates to a multi-target tracking device based on multiple similarity matrix fusion, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the multi-target tracking method based on multiple similarity matrix fusion of Example 1.

[0060] Example 3

[0061] This embodiment relates to a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the multi-target tracking method based on multi-similarity matrix fusion of Embodiment 1 is implemented.

[0062] The contents described in the embodiments of this specification are merely an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms described in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.

Claims

1. A multi-target tracking method based on multi-similarity matrix fusion, comprising the following steps: S1. Input the input image into the target detection model to obtain the detection frame of each target in the image, and predict the motion trajectory of the target in the previous frame through the Kalman filter to obtain the predicted frame; S2, input the detection box and the prediction box into the appearance model WRN and SIFT feature extractor to obtain the corresponding SIFT features and WRN features; S3, calculate the Mahalanobis distance and cosine distance between the detection box and the prediction box target, construct the corresponding WRN feature similarity matrix and SIFT feature similarity matrix, and fuse them according to certain weights to obtain the final cost matrix; S4, match the target through the Hungarian algorithm to obtain the first matching result, and perform IOU matching again on the unmatched targets and unmatched detections in the first matching result to obtain the trajectory of the target; S5. The final trajectory of the target is obtained through the appearance-independent connection model and Gaussian smoothing interpolation method.

2. A multi-target tracking method based on multi-similarity matrix fusion as claimed in claim 1, characterized in that: The appearance model WRN described in step S2 has 2 convolutional layers, 6 residual blocks and 1 Dense layer; each convolutional layer uses a 3×3 convolution kernel; after the convolutional layer, the network performs a pooling operation; followed by six residual block structures, each residual block contains multiple convolutional layers, and the input and output are added through residual connections; after extracting the 128-dimensional feature vector, the feature is mapped using L2 normalization; the SIFT feature extractor extracts the SIFT feature vectors of the targets in the detection box and the prediction box through the process of scale space extreme value detection, key point positioning, key point direction determination, and key point description, and then flattens them into a one-dimensional vector and performs BatchNorm normalization on the SIFT feature vector.

3. A multi-target tracking method based on multi-similarity matrix fusion as claimed in claim 1, characterized in that: Step S3 specifically includes: Specifically: Calculate the Mahalanobis distance between the detection box and the prediction box, the cosine distance between the detection box SIFT feature and the prediction box feature, and the cosine distance between the detection box WRN feature and the prediction box WRN feature, and then construct the WRN feature similarity matrix and SIFT feature similarity matrix based on these three distances; The Mahalanobis distance calculation formula is as follows: Where i is the i-th track, j is the j-th detection, and y i , S i are the predicted position and predicted covariance matrix of a target in a certain frame, d j The detection box information of the next frame; The formula for calculating cosine distance is as follows: Among them, r j is the appearance feature vector of the jth detection, is the appearance feature vector of all the i-th detections; After obtaining the two similarity matrices, they are fused according to the weights to obtain the final cost matrix; c i,j =λc (1) (i,j)+(1-λ)c (2) (i,j) (3) Among them, λ is the weight coefficient, c (1) (i,j) is the WRN feature similarity matrix, c (2) (i,j) is the WRN feature similarity matrix.

4. A multi-target tracking method based on multi-similarity matrix fusion as claimed in claim 1, characterized in that: Step S4 specifically includes: after the cost matrix is ​​obtained, the Hungarian algorithm is used to increment from 0 to A according to the number of frames that have not been successfully matched. max The trajectory is matched cyclically, and the frames with the smallest number of unsuccessful matches are given priority for optimal matching; the successfully matched trajectory and detection need to be within the threshold of the gate matrix, and the number of cycles is greater than A max When , the loop matching ends, and the unmatched trajectories and detections are returned to obtain the matching results; the unmatched trajectories and unmatched detections in the matching results are matched again by IOU; the uncertain trajectories and trajectories with a lifespan greater than the set maximum are deleted; the two successfully matched trajectories are filtered through Kalman filtering, and the unmatched detections are created as new trajectories, which are updated to the trajectories of the next frame; after the matching is completed, the final trajectory of each target is obtained.

5. A multi-target tracking method based on multi-similarity matrix fusion as claimed in claim 1, characterized in that: Step S5 specifically includes: The AFLink model is based on the trajectory T i and T j As input, the trajectory From the latest N = 30 frames f k and the corresponding position (x k ,y k ), for those trajectories with a length of less than 30 frames, the model uses zero padding to ensure the consistency of the input and facilitate subsequent processing; then, the model uses a 7×1 convolution kernel in the timing module to perform convolution operations on the input trajectory along the time dimension to extract time-related features; then, the model uses a 1×3 convolution operation in the fusion module to integrate information from different feature dimensions (f, x, y); after being processed by the fusion module, the two feature maps obtained are converted into feature vectors after pooling and compression, and then the two feature vectors are concatenated. These feature vectors contain rich spatiotemporal information; finally, the model uses a multi-layer perceptron to predict the confidence score associated with the trajectory; The timing module consists of 4 convolutional layers with kernel size of 7×1 and output channels of {32, 64, 128, 256}; each convolution is followed by a BN layer and a ReLU activation layer; the fusion module includes a 1×3 convolutional layer, a BN layer and a ReLU activation layer; the classifier is a multi-layer perceptron with two fully connected layers and a ReLU layer inserted in the middle; During training, the association process is formulated as a binary classification task; then, it is optimized using the binary cross entropy loss: where x n ∈[0,1] is the association probability of the predicted sample pair, y n ∈{0,1} is the sample label value; In the final reasoning, a temporal distance threshold of 30 frames and a spatial distance threshold of 75 pixels are used to filter out unreasonable association pairs; if the prediction score of an association pair is greater than 0.95, the trajectory is considered to be relevant; The Gaussian smoothing interpolation method includes: firstly, calculating the Gaussian weighted average of the trajectory data, which takes into account the spatial and temporal relationship between the data points and their weights; the weights are usually determined based on the distance between the data point and the point to be interpolated or other related factors, and the closer the point is, the higher the weight is; through Gaussian weighting, Gaussian smoothing interpolation can better capture the local characteristics and changing trends of the trajectory; after obtaining the smoothed trajectory data, Gaussian smoothing interpolation uses these data to estimate the values ​​of the missing positions; in the interpolation process, not only the values ​​of the known data points are considered, but also the spatial and temporal relationship between them, as well as the overall trend of the trajectory; this enables Gaussian smoothing interpolation to generate interpolation results that are closer to the real trajectory; Gaussian smoothing interpolation models the i-th trajectory as follows: p t =f (i) (t)+∈ (5) Where t is the frame in the frame set F, p t is the position coordinate of the tth frame, p t ∈P, ∈ is Gaussian noise, obeying N(0,σ 2 ); Given a trajectory of length L By fitting the function f (i) (t) to solve the nonlinear motion modeling problem; assuming that it obeys a Gaussian process f (i) (t)∈GP(0,k(·,·)), where is the radial basis kernel function; according to the properties of the Gaussian process, given a new frame set F * In the case of smoothing position P * Predictions were made; P * =K(F * ,F)(K(F,F)+σ 2 I) -1 P (6) where K(·,·) is the covariance function based on k(·,·); In addition, the hyperparameter λ controls the smoothness of the trajectory, and its size is related to the trajectory length l; λ=τ*log(τ 3 / l) (7) The value of τ is set to 10.

6. A multi-target tracking device based on multi-similarity matrix fusion, characterized in that: It comprises a memory and one or more processors, wherein the memory stores executable codes, and when the one or more processors execute the executable codes, they are used to implement the multi-target tracking method based on multi-similarity matrix fusion according to any one of claims 1 to 5.

7. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the multi-target tracking method based on multi-similarity matrix fusion described in any one of claims 1 to 5 is implemented.

Citation Information

Cited By

  • Multi-target tracking detection method, device and equipment and readable storage medium

    CN120726100A