A decouplable end-to-end multi-object tracking method and device

By building an adaptive information interaction module in the end-to-end multi-objective tracking model, decoupling of detection tasks and associated tasks is solved, and the lack of data adaptation of existing models is improved, and the learning efficiency and adaptability of the model are improved.

CN118710877BActive Publication Date: 2025-06-10HEBEI FORESTRY ECOLOGICAL CONSTR INVESTMENT CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410752932.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-12
Publication Date
2025-06-10
Estimated Expiration
2044-06-12

AI Technical Summary

Technical Problem

The existing end-to-end multi-objective tracking model cannot be adaptively adjusted based on data, resulting in coupling and conflict between detection tasks and associated tasks, increasing the difficulty of model optimization and deployment costs.

Method used

By building a decoupled end-to-end multi-objective tracking model for adaptive information interaction, the multi-head cross-attention module, the multi-head self-attention module and the feature-sampling point alignment module are used to realize the adaptive interaction and decoupling of detection tasks and associated tasks.

Benefits of technology

The conflict between detection tasks and associated tasks is reduced, the learning efficiency of the model and adaptability to different scenarios is improved, and the deployment cost and the need for hyperparameter configuration is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118710877B_ABST
    Figure CN118710877B_ABST
Patent Text Reader

Abstract

The present invention discloses a decoupled end-to-end multi-object tracking method and device, belonging to the technical field of computer vision. The end-to-end multi-object tracking model constructed by the present invention includes an adaptive information interaction module. In the adaptive information interaction module, the detection-tracking information transfer module transfers the information of the tracked object from the detection embedding vector to the tracking embedding vector. The detection-tracking information suppression module adaptively suppresses the detection embedding vector information. The feature-sampling point alignment module aligns the updated detection embedding vector and detection sampling point coordinates, and the tracking embedding vector and tracking sampling point coordinates. Thus, the decoupled end-to-end multi-object tracking model based on adaptive information interaction realizes the decoupling of the detection task and the association task in the traditional end-to-end multi-object tracking model. Through the adaptive information interaction method, the conflict between the detection task and the association task is reduced, and the learning efficiency of the end-to-end multi-object tracking model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and more specifically, to a decoupled end-to-end multi-object tracking method and device. Background Art

[0002] In recent years, computer vision has been more and more widely applied in actual scenarios. As a comprehensive task of classification, localization, and object re-identification, multi-object tracking plays an important role in aspects such as autonomous driving, security monitoring, and video understanding. With the development of deep learning, multi-object tracking models have gradually transitioned from two-stage detection-based tracking architectures to one-stage end-to-end tracking architectures, making the model structure more concise and the data utilization method more efficient. The one-stage end-to-end tracking framework mainly focuses on the efficient fusion of detection information and association information. It is necessary that new objects can be detected in time and existing objects can be stably tracked. However, serious coupling phenomena occur in the process of processing these two types of information, resulting in conflicts between the detection task and the association task, thus greatly increasing the optimization difficulty of the model. In very few one-stage decoupled end-to-end multi-object tracking methods, detection information and association information are processed through a rule-based matching algorithm. Therefore, hyperparameters need to be set manually, and the model cannot be fully adaptively adjusted according to the data, making it difficult to handle complex scenarios such as target aggregation and occlusion, and increasing the deployment cost of the model and restricting the application scope.

[0003] After retrieval, a Chinese patent application, application number 202210239608.4, publication date August 16, 2022, discloses a multi-object tracking method, device, electronic device, storage medium, and product. The method includes: obtaining a video to be measured; inputting the video to be measured into an end-to-end multi-object tracking model to obtain multi-object information included in the video to be measured output by the end-to-end multi-object tracking model; wherein the multi-object information includes the detection box where the object is located and the object identity information, the end-to-end multi-object model is trained based on a video sample data set, and the features extracted during the training process of the end-to-end multi-object model are enhanced features based on historical trajectory features extracted according to the video sample data set. This method integrates an object detection branch, a feature branch, and an identity association branch through an end-to-end multi-object tracking model, solves the defect of inaccurate multi-object tracking results in the prior art, and improves the multi-object detection accuracy. However, the model used in this method has a high construction cost, and the model cannot be adaptively adjusted according to the data characteristics, and does not have strong practicability and wide applicability. Summary of the Invention

[0004] 1. Technical Problems to be Solved

[0005] In view of the problems in the existing technology that the end-to-end multi-object tracking model cannot be adaptively adjusted according to data, resulting in coupling and conflicts between the detection task and the association task, the present invention provides a decoupled end-to-end multi-object tracking method and device, which realizes the adaptive interaction of the detection task and the association task information in the traditional decoupled end-to-end multi-object tracking model through a decoupled end-to-end multi-object tracking model with adaptive information interaction, reduces the conflict between the detection task and the association task, improves the learning efficiency of the end-to-end multi-object tracking model, and enhances the adaptability of the end-to-end multi-object tracking model to different scenarios.

[0006] 2. Technical solution

[0007] The object of the present invention is achieved through the following technical solutions.

[0008] A decoupled end-to-end multi-object tracking method includes the following steps:

[0009] Construct an end-to-end multi-object tracking model, which includes an object detector, an adaptive information interaction module, an associator, and a tracking head network;

[0010] Obtain a video to be measured, and input the video to be measured into the object detector to detect the encoder embedding vector, feature map position encoding vector, detection embedding vector, and detection result of each frame of image frame by frame;

[0011] For the first frame of image in the video to be measured, input the encoder embedding vector, feature map position encoding vector, detection embedding vector, and detection result of the first frame of image into the associator to obtain an association embedding vector, and input the association embedding vector into the tracking head network to obtain a tracking result;

[0012] Starting from the second frame of image in the video to be measured, interact the tracking result of the previous frame of image with the detection embedding vector and detection result of the current frame of image through the adaptive information interaction module, and input the interacted tracking result of the previous frame of image, the detection embedding vector and detection result of the current frame of image, as well as the encoder embedding vector and feature map position encoding vector of the current frame of image into the associator to obtain an association embedding vector, and input the association embedding vector into the tracking head network to obtain the tracking result of the current frame of image.

[0013] Further, set a classification confidence threshold. In the first frame of image, mark the tracking results with a classification confidence higher than the classification confidence threshold as newly emerged objects, add them to the tracking sequence, use the association embedding vector of the newly emerged objects as the tracking embedding vector, and use the corresponding positioning information as the tracking sampling point coordinates.

[0014] Further, the adaptive information interaction module includes a detection-tracking information transfer module, a detection-tracking information suppression module, and a feature-sampling point alignment module;

[0015] The detection - tracking information transfer module includes a multi - head cross - attention module and a feed - forward connection network; the multi - head cross - attention module calculates the similarity between the tracking embedding vector and the detection embedding vector, performs weighted summation on the detection embedding vector based on the similarity, applies the residual connection to the tracking embedding vector, and updates the tracking embedding vector through the feed - forward connection network. The process is expressed as:

[0016]

[0017] Te 2 = FFN TM (Norm TM (Te 1 + Oe 1 ))

[0018] Among them, represents the first tracking embedding vector with position information, Te 1 represents the first tracking embedding vector, PE(Tref1 1 ) represents the position encoding of the first tracking sampling point coordinates, represents the first detection embedding vector with position information, De 1 represents the first detection embedding vector, PE(Dref 1 ) represents the position encoding of the first detection sampling point coordinates, Oe 1 represents the result obtained by weighted summation of the first detection embedding vector based on the similarity, CrossAttn represents the multi - head cross - attention module, Q represents the query vector, K represents the key vector, V represents the value vector, Te 2 represents the second tracking embedding vector, FFN TM represents the feed - forward connection network in the detection - tracking information transfer module, Norm TM represents the normalization operation in the detection - tracking information transfer module.

[0019] Furthermore, the detection - tracking information suppression module includes a multi - head self - attention module and a feed - forward connection network; the detection - tracking information suppression module adaptively suppresses the similar information in the tracking embedding vector and the detection embedding vector. The process is expressed as:

[0020] Me 1 =(Te 1 , De 1 )

[0021]

[0022] ( -, De 2 ) = FFN SM (Norm SM (Me1 +Oe 2 ))

[0023] Among them, Me 1 represents the first mixed embedding vector, represents the tracking embedding vector with position information and marker information, represents the mixed embedding vector with position information and marker information, Oe 2 represents the result obtained by adaptively suppressing the similar information in the first tracking embedding vector and the first detection embedding vector. SelfAttn represents the multi-head self-attention module, (-, De 2 ) represents taking only the updated detection embedding vector as the second detection embedding vector, FFN SM represents the forward connection network in the detection-tracking information suppression module, Norm SM represents the normalization operation in the detection-tracking information suppression module.

[0024] Furthermore, the feature-sampling point alignment module includes a dynamic anchor box cross-attention module, a forward connection network, and a position update sub-network. The feature-sampling point alignment module aligns the interacted tracking embedding vector and detection embedding vector, and its process is expressed as:

[0025]

[0026] Me 2 =(Te 2 , De 2 )

[0027]

[0028] (Te 3 , De 3 ) = FFN PCAM (Norm PCAM (Me 2 +Oe 3 ))

[0029] (Tref 2 , Dref 2 ) = MLP PUN (Oe 3 )+(Tref 1 , Dref 1 )

[0030] Among them, represents the second tracking embedding vector with position information, represents the second detection embedding vector with position information, De 2 represents the second detection embedding vector, Denote the hybrid embedding vector with position information, Me 2 Denote the second hybrid embedding vector, Oe 3 Denote the output result of the detection cross-attention module, DABCrossAten represents the dynamic anchor box cross-attention module, Te 3 Denote the third tracking embedding vector, De 3 Denote the third detection embedding vector, FFN PCAM Denote the forward connection network in the feature-sampling point alignment module, Norm PCAM Denote the normalization operation in the feature-sampling point alignment module, Tref 2 Denote the second tracking sampling point coordinates, Dref 2 Denote the second detection sampling point coordinates, MLP PUN Denote the position update sub-network.

[0031] Furthermore, starting from the second frame image in the video to be measured, mark the tracking results with classification confidence higher than the classification confidence threshold as positive predictions, mark the positive predictions corresponding to the detection embedding vectors as newly emerged objects, add them to the tracking sequence, use the associated embedding vectors corresponding to all positive predictions as the tracking embedding vectors, and use the tracking results corresponding to all positive predictions as the tracking sampling point coordinates; the positive prediction discrimination method is:

[0032]

[0033] Among them, PosPre represents positive prediction, i represents the prediction index number, p i Denote the maximum class probability value in the i-th prediction, b i Denote the position information in the i-th prediction, s i Denote the classification confidence in the i-th prediction, σ represents the classification confidence threshold, Denote the maximum class probability value among all predictions, Denote the position information among all predictions;

[0034] Mark the tracking results with classification confidence not higher than the classification confidence threshold as negative predictions, mark the newly emerged objects as disappeared states, and cancel the newly emerged objects that have disappeared for more than M frames from the tracking sequence; the negative prediction discrimination method is:

[0035]

[0036] Among them, NegPre represents negative prediction.

[0037] Further, in the training stage, the video to be tested is sampled at intervals of τ, and the continuous T frames obtained by sampling are used as a video clip. Taking the video clip as a unit, the classification ground truth label and the localization ground truth label of the video clip are labeled. Local bipartite graph matching is performed according to the tracking results output by the tracking head network and the labeled classification ground truth label and localization ground truth label. The loss function is calculated according to the matching results, and the end-to-end multi-object tracking model is trained by the gradient descent method;

[0038] In the inference stage, end-to-end multi-object tracking is performed frame by frame on the video to be tested, and tracking sequence management is carried out. The positive predictions corresponding to the detection embedding vectors are marked as newly emerging objects and added to the tracking sequence; the negative predictions corresponding to the tracking embedding vectors are marked as disappeared states, and the newly emerging objects that have disappeared continuously for more than M frames are removed from the tracking sequence.

[0039] A decoupled end-to-end multi-object tracking device includes:

[0040] A construction module that constructs an end-to-end multi-object tracking model, and the end-to-end multi-object tracking model includes an object detector, an adaptive information interaction module, an associator, and a tracking head network;

[0041] An input module that obtains the video to be tested and inputs the video to be tested into the object detector to detect the encoder embedding vector, the feature map position encoding vector, the detection embedding vector, and the detection result of each frame of image frame by frame;

[0042] A detection module, for the first frame image in the video to be tested, inputs the encoder embedding vector, the feature map position encoding vector, the detection embedding vector, and the detection result of the first frame image into the associator to obtain an associated embedding vector, and inputs the associated embedding vector into the tracking head network to obtain a tracking result; starting from the second frame image in the video to be tested, the adaptive information interaction module is used to interact the tracking result of the previous frame image with the detection embedding vector and the detection result of the current frame image, and inputs the interacted tracking result of the previous frame image, the detection embedding vector and the detection result of the current frame image, as well as the encoder embedding vector and the feature map position encoding vector of the current frame image into the associator to obtain an associated embedding vector, and inputs the associated embedding vector into the tracking head network to obtain the tracking result of the current frame image.

[0043] A computer device includes a memory and a processor, and a computer program that can run on the processor is stored on the memory. When the processor executes the computer program, the above-mentioned method is implemented.

[0044] A computer-readable storage medium stores a computer program, and when the computer program is run by a processor, the above-mentioned method is executed.

[0045] 3. Beneficial effects

[0046] Compared with the prior art, the advantages of the present invention are as follows:

[0047] (1) The decoupled end-to-end multi-object tracking method and device of the present invention realizes the decoupling of the detection task and the association task in the traditional end-to-end multi-object tracking model by constructing a decoupled end-to-end multi-object tracking model with adaptive information interaction, effectively reducing the conflict between the detection task and the association task, and significantly improving the learning efficiency of the end-to-end multi-object tracking model.

[0048] (2) The decoupled end-to-end multi-object tracking method and device of the present invention can realize the adaptive and efficient fusion of the detection task and the association information while decoupling the end-to-end multi-object tracking by constructing a decoupled end-to-end multi-object tracking model with adaptive information interaction, effectively avoiding the manual configuration of hyperparameters based on matching methods, and strengthening the efficient learning and in-depth mining of scene data.

[0049] (3) The decoupled end-to-end multi-object tracking method and device of the present invention can simply and directly utilize the advanced high-performance transformer detector by constructing a decoupled end-to-end multi-object tracking model with adaptive information interaction, without retraining, which can reduce the development cost of the end-to-end multi-object tracking model and has strong practicability and wide applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is the structure diagram of the end-to-end multi-object tracking model in the embodiment of the present invention;

[0051] Figure 2 It is the structure diagram of the adaptive information interaction module in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] The present invention will be described in detail below with reference to the accompanying drawings of the specification and specific embodiments.

[0053] Embodiment

[0054] A decoupled end-to-end multi-object tracking method provided in this embodiment realizes multi-object tracking by constructing a decoupled end-to-end multi-object tracking model. As Figure 1 shown, a decoupled end-to-end multi-object tracking model provided in this embodiment includes an object detector, an adaptive information interaction module, an associator, and a tracking head network.

[0055] Specifically in this embodiment, for the target detector, an independent Transformer-based target detector is pre-trained, or a powerful basic vision detection model is obtained as the target detector. In this embodiment, the end-to-end multi-object tracking model can be compatible with two types of target detectors. One is a conventional Transformer-based target detector, which needs to be pre-trained in the target scenario to obtain good detection performance. The other is a Transformer-based basic vision detection model trained on a large amount of data. Since this type of vision detection model has strong generalization ability and already has good detection performance in the target scenario, it does not need to be trained. It should be noted that both of these two types of target detectors are composed of a backbone network, an encoder, and a decoder, and each decoder network is externally connected to a tracking head network. In this embodiment, the DAB-DETR detector based on Transformer or the basic vision detection model GroundingDINO can be used as the target detector.

[0056] Further, the video to be measured is obtained, and the video to be measured is input into the target detector for frame-by-frame detection to obtain the detection information of each frame of image. The detection information includes the encoder embedding vector, the feature map position encoding vector, the detection embedding vector, and the detection result of each frame of image.

[0057] Specifically, the encoder embedding vector (Encoder Embedding, Ee) is obtained from the output of the encoder of the target detector.

[0058] The feature map position encoding vector (Positional Encoding, Pe) is a position encoding vector obtained according to the size of the feature map. In this embodiment, the feature map position encoding is performed based on the sine and cosine functions. It is assumed that the size of the feature map is H×W and the number of channels is c, where H represents the height of the feature map, W represents the width of the feature map, and for the feature point with coordinates (pos x , pos y ), the position encoding is represented as:

[0059]

[0060] where pos represents the feature point, pos x represents the abscissa of the feature point, pos y represents the ordinate of the feature point, i and j respectively represent the indexes of the position encoding in the horizontal and vertical dimensions, and the value ranges of i and j are both [0, …, d model / 2], d model represents the dimension after encoding each feature point, d model = c / 2, The sine encoding vector representing the abscissa of the feature point, The cosine encoding vector representing the abscissa of the feature point, The sine encoding vector representing the ordinate of the feature point, The cosine encoding vector representing the ordinate of the feature point, and Pe represents the feature map position encoding vector.

[0061] The detection embedding vector (De) is obtained from the decoder output of the object detector.

[0062] The detection result includes classification information and localization information. Specifically, the classification information includes the class information and classification confidence of the newly emerged object predicted. The localization information is the coordinates of the detection sampling points (Detection reference points, Dref). In this embodiment, the localization information includes the localization coordinates and size of the rectangular bounding box of the newly emerged object predicted, which is represented as (x, y, w, h) in this embodiment. Among them, (x, y) are the abscissa and ordinate of the center point of the rectangular bounding box, and (w, h) are the width and length of the rectangular bounding box.

[0063] For the adaptive information interaction module, as Figure 2 shown, in this embodiment, the adaptive information interaction module includes a detection-tracking information transfer module (Transfer Module, TM), a detection-tracking information suppression module (SuppressionModule, SM), and a feature-sampling point alignment module (Content and Position Alignment Module, CPAM).

[0064] The detection-tracking information transfer module TM is composed of a multi-head cross-attention module CrossAttn and a feed-forward network (Feed Forward Network, FFN). In this embodiment, the detection-tracking information transfer module TM transfers the information of the tracked object from the detection embedding vector De to the tracking embedding vector (Tracking Embedding, Te), realizing the update of the detection embedding vector De and the tracking embedding vector Te.

[0065] The detection-tracking information suppression module SM is composed of a multi-head self-attention module SelfAttn and a feed-forward network FFN. In this embodiment, the detection-tracking information suppression module SM adaptively suppresses the information of the detection embedding vector De, avoiding the same object being simultaneously responsible for by the detection embedding vector De and the tracking embedding vector Te, thereby alleviating the conflict between the detection task and the association task.

[0066] The feature-sampling point alignment module PCAM consists of a dynamic anchor box cross-attention module DABCrossAttn, a forward connection network FFN, and a position update sub-network (Position Update Network, PUN). In this embodiment, the position update sub-network PUN consists of a multi-layer linear mapping layer, a non-linear mapping layer, and a normalization operation module. In this embodiment, the feature-sampling point alignment module PCAM aligns the updated detection embedding vector De and the detection sampling point coordinates Dref, the tracking embedding vector Te, and the tracking sampling point coordinates (Tracking reference points, Tref) to avoid the generation gap of information in the time dimension.

[0067] In this embodiment, for the forward connection network FFN, given an input vector x, the forward connection network FFN is defined as:

[0068] FFN = Norm(x + Linear2(ReLU(Linear1(x))))

[0069] Where, Linear1 and Linear2 represent two linear mapping layers with the same structure but different parameters, ReLU represents a non-linear activation operation, and Norm represents a normalization operation.

[0070] For the associator, in this embodiment, the associator (Association Module, AM) consists of a layer of network with the same structure as the decoder network in the target detector.

[0071] For the tracking head network, in this embodiment, the tracking head network (Track Head, TH) includes a classification sub-network and a regression sub-network. It should be noted that in this embodiment, both the classification sub-network and the regression sub-network consist of a multi-layer linear mapping layer, a non-linear mapping layer, and a normalization operation module. Specifically, given that the size of the input vector x is (d 1 , d 2 ), then the classification sub-network is expressed as:

[0072] MLP cls = Linear1(ReLU(Linear2(ReLU(Linear3(x))))))

[0073] Where, MLP cls represents the classification sub-network, Linear1 represents the parameter matrix of 1×d 2 , Linear2 represents the parameter matrix of d 2 ×d 2 , Linear3 represents the parameter matrix of d 2 ×d 2The parameter matrix, where ReLU represents the standard non - linear mapping operation.

[0074] Set the size of the input vector x to be (d 1 , d 2 ), then the regression sub - network is represented as:

[0075] MLP reg = Linear1(ReLU(Linear2(ReLU(Linear3(x)))))

[0076] Among them, MLP reg represents the regression sub - network, Linear1 represents the parameter matrix of 4×d 2 , Linear2 represents the parameter matrix of d 2 ×d 2 , Linar3 represents the parameter matrix of d 2 ×d 2 .

[0077] Set the size of the input vector x to be (d 1 , d 2 ), then the forward - connected network is represented as:

[0078] FFN = Norm(x + Linear2(ReLU(Linear1(x))))

[0079] Among them, Linear1 represents the parameter matrix of d 2 ×d 2 , Linear2 represents the parameter matrix of d 2 ×d 2 .

[0080] Thus, an end - to - end multi - object tracking model with adaptive information interaction and decoupling provided in this embodiment can fully learn the data characteristics, completely adaptively adjust the end - to - end multi - object tracking model according to the data, reduce the deployment cost of the end - to - end multi - object tracking model, and expand the application scope of the end - to - end multi - object tracking model.

[0081] An end - to - end multi - object tracking method with decoupling provided in this embodiment realizes multi - object tracking based on the end - to - end multi - object tracking model with decoupling. Specifically, in this embodiment, first, the video to be measured is obtained, and the video to be measured is input into the object detector to detect the encoder embedding vector Ee, the feature map position encoding vector Pe, the detection embedding vector De and the detection results frame by frame. The detection results include classification information and localization information, where the localization information is the detection sampling point coordinates Dref.

[0082] Further, for the first frame image in the video to be measured, the encoder embedding vector Ee, the feature map position encoding vector Pe, the detection embedding vector De, and the detection result of the first frame image are input into the associator AM to obtain the association embedding vector (Ae), and the association embedding vector Ae is input into the tracking head network TH to obtain the tracking result.

[0083] Specifically, for the first frame image, the encoder embedding vector Ee, the feature map position encoding vector Pe, the detection embedding vector De, and the detection sampling point coordinates Dref of the first frame image are input into the associator AM to obtain the association embedding vector Ae, and the association embedding vector Ae is input into the classification sub-network and the regression sub-network of the tracking head network TH to obtain the tracking result. Among them, represents the category of the predicted newly emerged object, represents the classification confidence of the predicted newly emerged object, represents the position information of the predicted newly emerged object, and its process is expressed as:

[0084] Ae = AM(tgt = De, anchor_box = Dref, memory = Ee, pos = Pe)

[0085]

[0086] Among them, Ae represents the association embedding vector, AM represents the associator, tgt represents the input result of the content query vector, anchor_box represents the input result of the learnable anchor box, memory represents the input result of the encoder embedding vector, and pos represents the input result of the feature map position encoding vector.

[0087] In the tracking result, the discrimination method for the newly emerged object is:

[0088]

[0089] Among them, NewObject represents the newly emerged object, i represents the predicted index number, p i represents the maximum category probability value in the i-th prediction, b i represents the position information in the i-th prediction, s i represents the classification confidence in the i-th prediction, σ represents the classification confidence threshold, represents the maximum category probability value in all predictions, represents the position information in all predictions.

[0090] In this embodiment, a classification confidence threshold σ is set. In the first frame image, the tracking results with a classification confidence higher than the classification confidence threshold σ are marked as newly emerged objects and added to the tracking sequence. The associated embedding vector Ae of the newly emerged objects is used as the tracking embedding vector Te, and the corresponding positioning information is used as the tracking sampling point coordinates Tref. In this embodiment, the classification confidence threshold σ is set to 0.5. Specifically, in this embodiment, the tracking embedding vector Te is obtained as follows:

[0091] Te = {e i |s i > σ, e i ∈ Ae}

[0092] where e i represents the i-th vector in the associated embedding vector Ae.

[0093] The tracking sampling point coordinates Tref are obtained as follows:

[0094]

[0095] where ref i represents the position information in the i-th prediction.

[0096] Starting from the second frame image in the video to be tested, the tracking results of the previous frame image are interacted with the detection embedding vector De and the detection results of the current frame image through the adaptive information interaction module. Specifically, starting from the second frame image, the tracking embedding vector Te, the tracking sampling point coordinates Tref of the previous frame image, the detection embedding vector De, and the detection sampling point coordinates Dref of the current frame image are input into the adaptive information interaction module to perform adaptive information interaction on the tracking embedding vector Te, the tracking sampling point coordinates Tref, the detection embedding vector De, and the detection sampling point coordinates Dref.

[0097] Specifically, in the adaptive information interaction module, the detection-tracking information transfer module TM includes a multi-head cross-attention module CrossAttn and a feed-forward network FFN. The multi-head cross-attention module CrossAttn calculates the similarity between the tracking embedding vector Te and the detection embedding vector De, performs weighted summation on the detection embedding vector De based on the similarity, and uses the residual connection to act on the tracking embedding vector Te. The tracking embedding vector Te is updated through the feed-forward network FFN, and the process is expressed as:

[0098]

[0099] Te 2 = FFN TM (Norm TM (Te 1 + Oe 1))

[0100] Among them, represents the first tracking embedding vector with position information, Te 1 represents the first tracking embedding vector, PE(Tref 1 ) represents position encoding for the coordinates of the first tracking sampling point, represents the first detection embedding vector with position information, De 1 represents the first detection embedding vector, PE(Dref 1 ) represents position encoding for the coordinates of the first detection sampling point, Oe 1 represents the result obtained by weighted summation of the first detection embedding vector based on similarity. CrossAttn represents the multi-head cross-attention module, Q represents the query vector, K represents the key vector, and V represents the value vector, Te 2 represents the second tracking embedding vector, FFN TM represents the feed-forward connection network in the detection-tracking information transfer module, Norm TM represents the normalization operation in the detection-tracking information transfer module.

[0101] The detection-tracking information suppression module SM includes the multi-head self-attention module SelfAttn and the feed-forward connection network FFN. The tracking embedding vector Te and the detection embedding vector De are directly concatenated to obtain the mixed embedding vector (Mixed Embedding, Me). The tracking embedding vector with position information and the detection embedding vector with position information are directly concatenated to obtain the mixed embedding vector with position information Since the tracking embedding vectors with position information and the detection embedding vectors with position information of similar objects are relatively close, in order to guide the end-to-end multi-object tracking model to distinguish, the last dimension of the tracking embedding vector with position information is set to 1 to obtain the first tracking embedding vector with position and marking information, marked as Thus, the detection-tracking information suppression module SM adaptively suppresses the similar information in the tracking embedding vector Te and the detection embedding vector De, and its process is expressed as:

[0102] Me 1 =(Te 1 , De 1 )

[0103]

[0104] (-, De 2 ) = FFNSM (Norm SM (Me 1 +Oe 2 ))

[0105] Among them, Me 1 represents the first mixed embedding vector, which is directly concatenated by the first tracking embedding vector and the first detection embedding vector, represents the tracking embedding vector with position information and marking information, represents the mixed embedding vector with position information and marking information, which is directly concatenated by the tracking embedding vector with position information and marking information and the first detection embedding vector with position information, Oe 2 represents the result obtained by adaptively suppressing the similar information in the first tracking embedding vector and the first detection embedding vector. SelfAttn represents the multi-head self-attention module, (-, De 2 ) represents taking only the updated detection embedding vector as the second detection embedding vector, FFN SM represents the forward connection network in the detection-tracking information suppression module, Norm SM represents the normalization operation in the detection-tracking information suppression module.

[0106] Thus, in the detection-tracking information suppression module SM, only the updated detection embedding vector De is obtained, realizing the suppression of the similar information between the detection embedding vector De and the tracking embedding vector Te, and further avoiding an object being responsible for both the detection embedding vector De and the tracking embedding vector Te, alleviating the conflict between the detection task and the association task.

[0107] The feature-sampling point alignment module CPAM includes the dynamic anchor box cross-attention module DABCrossAttn, the forward connection network FFN, and the position update sub-network MLP PUN . The feature-sampling point alignment module CPAM aligns the updated tracking embedding vector Te and the detection embedding vector De, and its process is expressed as:

[0108]

[0109] Me 2 =(Te 2 ,De 2 )

[0110]

[0111] (Te 3 ,De 3 )=FFN PCAM (Norm PCAM (Me 2 +Oe3 ))

[0112] (Tref 2 , Dref 2 ) = MLP PUN (Oe 3 ) + (Tref 1 , Dref 1 )

[0113] Among them, represents the second tracking embedding vector with position information, represents the second detection embedding vector with position information, De 2 represents the second detection embedding vector, represents the mixed embedding vector with position information, Me 2 represents the second mixed embedding vector, Oe 3 represents the output result of the detection cross-attention module, DABCrossAttn represents the dynamic anchor box cross-attention module, Te 3 represents the third tracking embedding vector, De 3 represents the third detection embedding vector, FFN PCAM represents the feed-forward connection network in the feature-sampling point alignment module, Norm PCAM represents the normalization operation in the feature-sampling point alignment module, Tref 2 represents the second tracking sampling point coordinates, Dref 2 represents the second detection sampling point coordinates, MLP PUN represents the position update sub-network.

[0114] Thus, in this embodiment, the detection information of the current frame image and the tracking information of the previous frame image are adaptively interacted through the adaptive information interaction module, so as to ensure the correct association of the existing objects and the accurate detection of the newly emerging objects, and improve the tracking accuracy of the decoupled end-to-end multi-object tracking model.

[0115] Furthermore, the tracking result of the updated previous frame image, the detection embedding vector De of the current frame image, the detection result, and the encoder embedding vector Ee and the feature map position encoding vector Pe of the current frame image are input into the associator AM to obtain the associated embedding vector Ae, and the associated embedding vector Ae is input into the tracking head network TH to obtain the tracking result of the current frame image.

[0116] Specifically, the updated tracking embedding vector Te, the tracking sampling point coordinates Tref, the detection embedding vector De, the detection sampling point coordinates Dref, as well as the encoder embedding vector Ee and the feature map position encoding vector Pe of the current frame image are input into the associator AM to obtain the associated embedding vector Ae. The associated embedding vector Ae is input into the tracking head network TH to obtain the tracking result of the current frame image. In this embodiment, for other frame images in the video to be measured, the above steps are repeated through the end-to-end multi-object tracking model for frame-by-frame detection to obtain the final tracking result. Specifically in this embodiment, the third tracking embedding vector Te 3 and the third detection embedding vector De 3 are directly connected to obtain the third hybrid embedding vector Me 3 , and the second tracking sampling point coordinates Tref 2 and the second detection sampling point coordinates Dref 2 are directly connected to obtain the hybrid sampling point Mref, and the process is expressed as:

[0117] Mref = (Tref 2 , Dref 2 )

[0118] Me 3 = (Te 3 , De 3 )

[0119] Ae = AM(tgt = Me 3 , anchor_box = Mref, memory = Ee, pos = Pe)

[0120]

[0121] In this embodiment, is the tracking result.

[0122] Furthermore, starting from the second frame image in the video to be measured, the tracking results with classification confidence higher than the classification confidence threshold σ are marked as positive predictions. The positive predictions corresponding to the detection embedding vector De are marked as newly emerging objects and added to the tracking sequence. The associated embedding vectors Ae corresponding to all positive predictions are used as the tracking embedding vector Te, and the tracking results corresponding to all positive predictions are used as the tracking sampling point coordinates Tref for tracking the next frame image. The positive prediction discrimination method is:

[0123]

[0124] where PosPre represents positive prediction, i represents the prediction index number, p i represents the maximum class probability value in the i-th prediction, b i represents the position information in the i-th prediction, si represents the classification confidence in the i-th prediction, σ represents the classification confidence threshold, represents the maximum class probability value among all predictions, Represents the position information in all predictions.

[0125] The tracking results whose classification confidence is not higher than the classification confidence threshold σ are marked as negative predictions, and the newly appeared objects are marked as disappeared. The newly appeared objects that disappear continuously for more than M frames are cancelled from the tracking sequence; the negative prediction judgment method is:

[0126]

[0127] Among them, NegPre represents negative prediction.

[0128] Therefore, the present embodiment provides a decoupled end-to-end multi-target tracking method. In the training phase, the video to be tested is sampled at an interval of τ, and the sampled continuous T frames are taken as a video segment. The classification truth labels and positioning truth labels of the video segments are annotated in units of video segments. Local bipartite graph matching is performed based on the tracking results output by the tracking head network TH and the annotated classification truth labels and positioning truth labels. The loss function is calculated based on the matching results, and the end-to-end multi-target tracking model is trained by the gradient descent method. In the present embodiment, local bipartite graph matching refers to the use of bipartite graph matching for the prediction results of newly appeared objects and detection embedding vectors De, and the use of the matching relationship at the previous moment for the prediction results of already appeared objects and tracking embedding vectors De. For bipartite graph matching, the truth labels involved in the matching are expressed as y=(c, s, b), and the prediction results of the end-to-end multi-target tracking model are expressed as c represents category information, b represents location information, and s represents classification confidence. The true value label y involved in the matching and the prediction result of the end-to-end multi-target tracking model are The bipartite graph matching between is:

[0129]

[0130] in, Represents the index set from the true value label i to all predicted values ​​such that The index that reaches the minimum value on all true value labels, arg min represents the value of the variable when the minimum value is reached, represents the mapping variable from the true value label to the predicted value, N represents the number of predictions, i represents the true value label index, and the number of predictions N must be greater than the number of true value labels. In this embodiment, N=300. In order to make the predictions and true value labels correspond one to one, in this embodiment, Perform a padding operation on the true value label, ω NDenotes the set of indices from the true label i to all predicted values, y i Denotes the i-th true label, Denotes the prediction corresponding to the i-th true label, Denotes the loss function when the mapping variable from the true label to the predicted value is mapped to ω(i), c i Denotes the category, Denotes the foreground category, Denotes the predicted probability value b of category c when the mapping variable from the true label to the predicted value is mapped to ω(i) i i Denotes the position information in the i-th prediction, Denotes the regression loss function value corresponding to the i-th true label when the mapping variable from the true label to the predicted value is mapped to ω(i), Denotes the classification loss function value of the foreground category, Denotes the regression loss function of the foreground category.

[0131] In local bipartite graph matching, the above bipartite graph matching strategy is applied to the matching between newly emerging objects in the current frame image and the predictions of the detection embedding vector De, and the result is denoted as The matching between the objects that have appeared at the previous moment and the tracking embedding vector follows the previous matching relationship, and the result is denoted as The final matching result is the union of the two parts of the matching relationship, which is expressed as:

[0132]

[0133] Calculate the loss function according to the matching result. The loss function of the j-th frame image is expressed as:

[0134]

[0135] Among them, Loss j Denotes the loss function of the j-th frame image, Denotes the classification loss function, λ cls Denotes the weight factor. In this embodiment, λ cls = 1, is Calculated by the GIOU algorithm.

[0136] It should be noted that in this embodiment, the loss function is calculated in units of video segments. Then, the loss function of a video segment with a length of T is expressed as:

[0137]

[0138] ​Among them, Loss represents the loss function of a video clip with length T. In this embodiment, the gradients of the trainable parameters in the end-to-end multi-object tracking model are calculated based on the loss function Loss of a video clip with length T, and the end-to-end multi-object tracking model is trained using the gradient descent method. It should be noted that in this embodiment, T is randomly sampled from integers between [1, 10]. Thus, in this embodiment, T = 5 is selected.

[0139] In the inference stage, the end-to-end multi-object tracking is performed frame by frame on the video to be measured, and the tracking sequence management is carried out. The positive predictions corresponding to the detection embedding vectors are marked as newly emerging objects and added to the tracking sequence; the negative predictions corresponding to the tracking embedding vectors are marked as disappearing states, and the newly emerging objects that have disappeared continuously for more than M frames are removed from the tracking sequence.

[0140] Thus, a decoupled end-to-end multi-object tracking method provided in this embodiment realizes the decoupling of the detection task and the association task in the traditional end-to-end multi-object tracking model by constructing a decoupled end-to-end multi-object tracking model with adaptive information interaction, effectively reducing the conflict between the detection task and the association task, and significantly improving the learning efficiency of the end-to-end multi-object tracking model. In addition, the decoupled end-to-end multi-object tracking model with adaptive information interaction constructed in this embodiment can achieve the adaptive and efficient fusion of the detection task and the association information while decoupling the end-to-end multi-object tracking, effectively avoiding the manual configuration of hyperparameters based on matching methods, and strengthening the efficient learning and in-depth mining of scene data. At the same time, the decoupled end-to-end multi-object tracking model with adaptive information interaction can simply and directly utilize the advanced high-performance transformer detector without retraining, which can reduce the development cost of the end-to-end multi-object tracking model and has strong practicability and wide applicability.

[0141] This embodiment also provides a decouplable end-to-end multi-object tracking device, including a construction module, an input module, and a detection module. In the construction module, an end-to-end multi-object tracking model is constructed. The end-to-end multi-object tracking model includes an object detector, an adaptive information interaction module, an associator, and a tracking head network. In the input module, the video to be measured is acquired, and the video to be measured is input into the object detector to detect the encoder embedding vector, feature map position encoding vector, detection embedding vector, and detection result of each frame of image frame by frame. In the detection module, for the first frame of image in the video to be measured, the encoder embedding vector, feature map position encoding vector, detection embedding vector, and detection result of the first frame of image are input into the associator to obtain an association embedding vector, and the association embedding vector is input into the tracking head network to obtain a tracking result; starting from the second frame of image in the video to be measured, through the adaptive information interaction module, the tracking result of the previous frame of image and the detection embedding vector and detection result of the current frame of image are interacted, and the interacted tracking result of the previous frame of image, the detection embedding vector and detection result of the current frame of image, and the encoder embedding vector and feature map position encoding vector of the current frame of image are input into the associator to obtain an association embedding vector, and the association embedding vector is input into the tracking head network to obtain the tracking result of the current frame of image. The decouplable end-to-end multi-object tracking device provided in this embodiment can implement any of the above-mentioned decouplable end-to-end multi-object tracking methods, and the specific working process of the decouplable end-to-end multi-object tracking device can refer to the corresponding process in the embodiment of the decouplable end-to-end multi-object tracking method. The methods and devices provided in this embodiment can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for example, the division of a certain module is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed connections or communication connections with each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can also be electrical, mechanical, or other forms of connections.

[0142] This embodiment also provides a computer device. A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned decouplable end-to-end multi-object tracking method is implemented.

[0143] This embodiment also provides a computer-readable storage medium. A computer-readable storage medium has a computer program stored thereon, and when the computer program is run by a processor, it executes the decoupled end-to-end multi-object tracking method described in this embodiment. Among them, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, device, or component; the program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0144] The present invention and its implementation manners are schematically described above. This description is not restrictive. Without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Any reference numeral in the claims should not limit the claimed claim. Therefore, if those of ordinary skill in the art are inspired by it and, without departing from the purpose of this creation, design similar structural manners and embodiments to this technical solution without creative efforts, they shall fall within the protection scope of the present invention. In addition, the word "including" does not exclude other elements or steps, and the word "a" before an element does not exclude including "a plurality of" such elements. The plurality of elements stated in the product claims can also be implemented by one element through software or hardware. The words such as first, second, etc. are used to indicate names and do not indicate any specific order.

Claims

1. A decoupled end-to-end multi-target tracking method, comprising the following steps: Construct an end-to-end multi-target tracking model, which includes a target detector, an adaptive information interaction module, an associator, and a tracking head network; Obtain the video to be tested, input the video to be tested into the target detector for frame-by-frame detection to obtain the encoder embedding vector, feature map position encoding vector, detection embedding vector and detection result of each frame image; For the first frame image in the video to be tested, the encoder embedding vector, the feature map position encoding vector, the detection embedding vector and the detection result of the first frame image are input into the associator to obtain the associated embedding vector, and the associated embedding vector is input into the tracking head network to obtain the tracking result; Starting from the second frame image in the video to be tested, the tracking result of the previous frame image and the detection embedding vector and detection result of the current frame image are interacted through the adaptive information interaction module, and the tracking result of the previous frame image and the detection embedding vector and detection result of the current frame image after interaction as well as the encoder embedding vector and feature map position encoding vector of the current frame image are input into the associator to obtain the associated embedding vector, and the associated embedding vector is input into the tracking head network to obtain the tracking result of the current frame image.

2. The decoupled end-to-end multi-target tracking method according to claim 1, characterized in that: Set the classification confidence threshold. In the first frame image, mark the tracking results with classification confidence higher than the classification confidence threshold as new objects, add them to the tracking sequence, use the associated embedding vector of the new objects as the tracking embedding vector, and use the corresponding positioning information as the tracking sampling point coordinates.

3. The decoupled end-to-end multi-target tracking method according to claim 2, characterized in that: The adaptive information interaction module includes a detection-tracking information transfer module, a detection-tracking information suppression module, and a feature-sampling point alignment module; The detection-tracking information transfer module includes a multi-head cross-attention module and a forward connection network; the multi-head cross-attention module calculates the similarity between the tracking embedding vector and the detection embedding vector, performs a weighted summation of the detection embedding vector based on the similarity, uses the residual connection to act on the tracking embedding vector, and updates the tracking embedding vector through the forward connection network. The process is expressed as: Te2=FFN TM (Norm TM (Te1+Oe1)) in, represents the first tracking embedding vector with position information, Te1 represents the first tracking embedding vector, PE(Tref1) represents the position encoding of the coordinates of the first tracking sampling point, represents the first detection embedding vector with position information, De1 represents the first detection embedding vector, PE(Dref1) represents the position encoding of the first detection sampling point coordinates, Oe1 represents the result of weighted summation of the first detection embedding vector based on similarity, CrossAttn represents the multi-head cross attention module, Q represents the query vector, K represents the key vector, V represents the value vector, Te2 represents the second tracking embedding vector, FFN TM represents the forward connection network in the detection-tracking information transfer module, Norm TM Represents the standardized operations in the detection-tracking information transfer module.

4. The decoupled end-to-end multi-target tracking method according to claim 3, characterized in that: The detection-tracking information suppression module includes a multi-head self-attention module and a forward connection network; the detection-tracking information suppression module adaptively suppresses similar information in the tracking embedding vector and the detection embedding vector. The process is expressed as: Me1=(Te1,De1) (-,De2)=FFNs M (Norm SM (Me1+Oe2)) Among them, Me1 represents the first mixed embedding vector, represents the tracking embedding vector with position information and tag information, represents a mixed embedding vector with position information and tag information, Oe2 represents the result obtained after adaptively suppressing similar information in the first tracking embedding vector and the first detection embedding vector, SelfAttn represents a multi-head self-attention module, (-, De2) represents taking only the updated detection embedding vector as the second detection embedding vector, and FFN SM represents the forward connection network in the detection-tracking information suppression module, Norm SM Represents the normalization operation in the detection-tracking information suppression module.

5. The decoupled end-to-end multi-target tracking method according to claim 4, characterized in that: The feature-sampling point alignment module includes a dynamic anchor box cross attention module, a forward connection network, and a position update subnetwork. The feature-sampling point alignment module aligns the tracking embedding vector and the detection embedding vector after interaction. The process is expressed as follows: Me2=(Te2,De2) <h2 style=";text-align:left;direction:ltr">(Te3, De3)=FFN<h2 style=";text-align:left;direction:ltr"> PCAM <h2 style=";text-align:left;direction:ltr"> (Norm<h2 style=";text-align:left;direction:ltr"> PCAM <h2 style=";text-align:left;direction:ltr"> (Me2+Oe3)) (Tref2,Dref2)=MLP PUN (Oe3)+(Tref1,DreF1) in, represents the second tracking embedding vector with position information, represents the second detection embedding vector with position information, De2 represents the second detection embedding vector, represents a mixed embedding vector with position information, Me2 represents the second mixed embedding vector, Oe3 represents the output of the detection cross attention module, DABCrossAttn represents the dynamic anchor box cross attention module, Te3 represents the third tracking embedding vector, De3 represents the third detection embedding vector, FFNP CAM Represents the forward connection network in the feature-sample point alignment module, Norm PCAM represents the standardization operation in the feature-sampling point alignment module, Tref2 represents the coordinates of the second tracking sampling point, Dref2 represents the coordinates of the second detection sampling point, MLP PUN Represents the location update subnetwork.

6. The decoupled end-to-end multi-target tracking method according to claim 5, characterized in that: Starting from the second frame image in the video to be tested, the tracking results with classification confidence higher than the classification confidence threshold are marked as positive predictions, and the positive predictions corresponding to the detection embedding vectors are marked as newly appeared objects. The tracking sequence is added, and the associated embedding vectors corresponding to all positive predictions are used as tracking embedding vectors, and the tracking results corresponding to all positive predictions are used as tracking sampling point coordinates. The positive prediction discrimination method is: Among them, PosPre represents positive prediction, i represents the prediction index number, and p i represents the maximum category probability value in the i-th prediction, b i represents the position information in the i-th prediction, s i represents the classification confidence in the i-th prediction, σ represents the classification confidence threshold, represents the maximum class probability value among all predictions, Represents the position information in all predictions; The tracking results whose classification confidence is not higher than the classification confidence threshold are marked as negative predictions, and the newly appeared objects are marked as disappeared. The newly appeared objects that disappear continuously for more than M frames are cancelled from the tracking sequence; the negative prediction judgment method is: Among them, NegPre represents negative prediction.

7. The decoupled end-to-end multi-target tracking method according to claim 6, characterized in that: In the training phase, the video to be tested is sampled at an interval of τ, and the continuous T frames obtained by sampling are taken as a video segment. The classification truth label and positioning truth label of the video segment are annotated in units of video segments. Local bipartite graph matching is performed based on the tracking results output by the tracking head network and the annotated classification truth label and positioning truth label. The loss function is calculated based on the matching results, and the end-to-end multi-target tracking model is trained by the gradient descent method. In the inference phase, end-to-end multi-target tracking is performed frame by frame on the video to be tested, and tracking sequence management is performed. The positive prediction corresponding to the detection embedding vector is marked as a newly appeared object and added to the tracking sequence. The negative prediction corresponding to the tracking embedding vector is marked as disappeared, and the newly appeared objects that disappear continuously for more than M frames are deregistered from the tracking sequence.

8. A decoupled end-to-end multi-target tracking device, characterized in that: include: Building modules, building an end-to-end multi-target tracking model, which includes a target detector, an adaptive information interaction module, an associator, and a tracking head network; The input module obtains the video to be tested, inputs the video to be tested into the target detector for frame-by-frame detection to obtain the encoder embedding vector, feature map position encoding vector, detection embedding vector and detection result of each frame image; The detection module, for the first frame image in the video to be tested, inputs the encoder embedding vector, feature map position encoding vector, detection embedding vector and detection result of the first frame image into the associator to obtain the associated embedding vector, and inputs the associated embedding vector into the tracking head network to obtain the tracking result; starting from the second frame image in the video to be tested, the tracking result of the previous frame image and the detection embedding vector and detection result of the current frame image are interacted through the adaptive information interaction module, and the tracking result of the previous frame image after the interaction and the detection embedding vector and detection result of the current frame image, as well as the encoder embedding vector and feature map position encoding vector of the current frame image are input into the associator to obtain the associated embedding vector, and the associated embedding vector is input into the tracking head network to obtain the tracking result of the current frame image.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is executed.

Citation Information

Patent Citations

  • Multi-target tracking method and device, electronic equipment, storage medium and product

    CN114913201A

  • Multi-target tracking method and system

    CN116703962A

  • Target detection and tracking method and device, equipment and storage medium

    CN118115539A