Multi-target tracking method based on feature decoupling and feature reinforcement learning
By using the dual-branch feature fusion network and attention mechanism in the multi-objective tracking method for feature decoupling and enhancement, the problem of insufficient feature caused by task competition within the model is solved, and higher tracking accuracy and real-time performance are achieved.
Patent Information
- Application Number
- CN202510028237.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-08
AI Technical Summary
In the existing multi-objective tracking method, competition between the internal detection tasks and the re-identification tasks in the model leads to insufficient features, making it difficult to achieve accurate identification in scenarios with severe occlusion and background chaos, resulting in target loss and identity switching problems.
A multi-objective tracking method based on feature decoupling and feature enhancement is proposed, and a dual-branch feature fusion network and attention mechanism is adopted. Through feature decoupling and differentiation and task-oriented enhancement, task competition is reduced, and a new lost target refining strategy is introduced to improve the robustness and real-timeness of the model.
It effectively alleviates the competition between detection tasks and re-identification tasks, improves the task orientation of features, reduces the occurrence of target loss and identity switching, and improves the accuracy and real-timeness of multi-target tracking.
Smart Images

Figure CN119992411A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and artificial intelligence, and in particular relates to a single-stage multi-target tracking method based on feature decoupling and feature enhancement. Background Art
[0002] Multi-object tracking (MOT) aims to accurately estimate the positions and identities of multiple objects in a video sequence. It is one of the most fundamental and challenging tasks in computer vision, and has a wide range of applications in intelligent surveillance, autonomous driving, and biomedicine. In the past decade, due to the advancement of object detection technology, researchers have generally adopted the tracking-by-detection (TBD) paradigm to solve this task. Traditional TBD methods usually use two independent models: a detector model to generate the bounding box of the object, and an embedding-based tracking model to extract the object features and establish object associations between frames based on the features and motion information. Although these TBD methods can achieve good tracking results, the separation of the two networks in the TBD method leads to problems such as the inability to share features between models and insufficient feature utilization, which in turn results in high computational costs and difficulty in achieving real-time tracking.
[0003] In recent years, joint detection and tracking (JDT) methods have achieved a good balance between inference speed and tracking accuracy by adopting an end-to-end strategy, thereby alleviating some problems of tracking-by-detection (TBD) methods. TransTrack establishes a new JDT framework based on the query-key mechanism of Transformer. It tracks the target in the current frame by querying the relevant features of the previous frame in the sequence. Transcenter further proposes a dense pixel-level multi-scale query method based on the center point heat map to enhance the ability to detect targets in crowded scenes. FairMOT finds that there is an inherent conflict between the goals of detection and tracking tasks: detection aims to maximize inter-class differences, while tracking aims to maximize intra-class differences. It alleviates this problem by constructing two parallel branches to achieve target detection and feature extraction. RelationTrack further proposes re-identification (reID) embedding learning under global information to distinguish the features required by the two tasks. However, it still does not completely solve the competition problem between detection and recognition tasks within the model. In complex scenes with severe occlusion and cluttered background, the features extracted by task competition are not enough to accurately identify the target, resulting in problems such as target loss and identity switching. Summary of the invention
[0004] In order to overcome the shortcomings of the existing technology, the present invention provides a lightweight and effective multi-target tracking method based on feature decoupling and feature enhancement learning. In the method, a new dual-branch feature enhancement learning network and a new lost target re-finding strategy are proposed. Through the comprehensive application of this network and strategy, a more powerful tool is brought to the field of multi-target tracking.
[0005] The technical solution adopted by the present invention to solve its technical problem is:
[0006] A multi-target tracking method based on feature decoupling and feature enhancement learning includes the following steps:
[0007] Step 1: Obtain image features and perform feature decoupling and enhancement. The process is as follows:
[0008] S11, inputting a pedestrian multi-target tracking sequence, where each frame in the sequence contains multiple pedestrian tracking targets;
[0009] S12, input the pedestrian image data into the backbone network and obtain the 4-layer feature map, then perform feature decoupling and differentiation according to the 4-layer feature map, and then introduce the modified attention mechanism to enhance the task-orientedness of the decoupled and differentiated image features;
[0010] S13, two parts of image features after feature decoupling and feature enhancement;
[0011] Step 2: Target tracking. The process is as follows:
[0012] S21, taking the two parts of image features output by S13 as input;
[0013] S22, after the input features pass through a specific classification head, box_size, center_off, heatmap and reid data are obtained, and target association is completed based on these data based on the joint detection and tracking paradigm;
[0014] S23, target association: The multi-target tracking model outputs data according to S22, and retains the coordinate position of each tracked target in a specific frame sequence, while considering the re-identification features of each target for target association and re-recovery;
[0015] S24, comprehensive loss: The multi-target tracking model will comprehensively consider the target detection loss and target association loss to optimize the overall performance and balance the competition among tasks within the model;
[0016] S25. After the balance loss calculation, the multi-target tracking model has good robustness and its lightweight size can maintain the real-time requirements of tracking.
[0017] Furthermore, in S12, the process of feature decoupling is:
[0018] Firstly, DLA-34 is used as the backbone network. Considering the different semantic levels of features required for detection and re-identification tasks, a dual-branch feature fusion network is proposed. The network uses two different feature aggregation methods to obtain the feature representation required for detection and re-identification.
[0019] For the detection part, the reconstructed IDA-UP structure is used, which is based on operating the upsampled three-layer deep feature maps and fusing them with the shallow feature maps, and is called a top-down feature fusion strategy;
[0020] For the re-identification part, a bottom-up feature aggregation strategy is proposed. represents the feature map input to the network, where N represents the number of layers with feature maps of different resolutions extracted by the backbone network, Fi represents the feature map of the i-th layer, and then the bottom-up feature fusion strategy is expressed as:
[0021]
[0022] in, Represents the final fused features of each layer, represents the features obtained after fusion of the previous layer, UpSample(·) represents the upsampling operation consisting of deformable convolution and deconvolution, Conv 1×1 (·) represents a convolutional layer of size 1*1, which is used to change the number of feature channels, and σ(·) represents the Sigmoid activation layer.
[0023] Furthermore, in S12, an attention mechanism is introduced for feature enhancement, and a spatial feature enhancement module SFEM is proposed. SFEM is a superposition of a recursive spatial attention mechanism with a depth of two layers. Considering the feature fusion strategy of the constructed top-down detection part, the highest-level feature map obtained by the backbone is passed into the SFEM module. The fusion formula of each layer is:
[0024]
[0025] where Q,K,V∈R C×H×W is the input feature F C×H×W The matrix obtained by three learnable weight vectors, Q u represents any unit eigenvector in the Q matrix, ε i,u ,φ i,u represents the i-th value in the set of feature vectors in the same row or column as u in the K and V matrices, respectively. Softmax(·) is used to change the number of feature channels;
[0026] An identity recognition feature enhancement module (IEM) adapted to pixel-level fine-grained tasks is proposed. IEM completely folds features in one direction while maintaining high resolution of features in its orthogonal direction. Then, the dynamic range of attention is increased through Softmax() and Sigmoid() operations on the bottleneck vector, thereby enhancing the feature vector. The channel attention formula of IEM is:
[0027] A c (X) = F SG [W z ((σ1(W v (X))×F SM (σ2(W q (X))))))]
[0028] where X∈R C×H×W , W q ,W v ,W z They are convolutional layers of size 1X1, σ1, σ2 are two vector resetting factors, F SM (·) is the SoftMax operator, W v ,W z The number of internal channels is C / 2, A c (X)∈R C×1×1 , output Z c =A c (X) c X∈R C×H×W ;
[0029] The spatial attention formula of IEM is:
[0030] A s (X) = F SG [σ3(F SM (σ1(F GP (W q (X))))×σ2(W v (X)))]
[0031] where X∈R C×H×W , W q ,W v They are convolutional layers of size 1X1, σ1, σ2, σ3 are three vector reset factors, F SM (·) is the SoftMax operator, output Z p =A p (X) p X∈R C×H×W ;
[0032] The above two branches are added in parallel to get the final result:
[0033] PSAP(X)=Z c +Z p .
[0034] In S23, the multi-target tracking model adopts a tracking strategy based on IOU and appearance embedding distance. First, the appearance embedding information is used to compensate for the drift of the Kalman filter to obtain the final predicted target center position, and the average motion speed (V i ,V j ), considering the average number of frames processed per second t and the number of frames k that the target has lost, the Kalman filter is used to select the prediction center (C i ,C y ) The appearance embedding vectors in the range are calculated and compared with the appearance embedding vector E of the lost target. l The distance between them is calculated, and the formula is used to determine whether they belong to the same trajectory:
[0035]
[0036] Where D min Indicates the prediction center (C i ,C y ) in the N×N range near E l The cosine distance between r Indicates the matching threshold of trajectory recovery. If To represent the appearance embedding vector at pixel (i, j), then D min The calculation formula is as follows:
[0037]
[0038] Among them, F cosdistance (·) indicates calculating the cosine distance between two vectors. If the target object of the lost track is finally found, the found center point is set as the center point of the track in the current frame, so that its size is consistent with the size of the tracked target in the previous frame.
[0039] The technical concept of the present invention is: in order to alleviate the problem of internal task competition in the multi-target tracking method of joint detection and re-identification, the present invention proposes a new feature decoupling network, which integrates top-down and bottom-up feature extraction strategies, and differentiates the features required for the internal detection task and re-identification task of the model; by introducing and modifying a specific attention mechanism network, the specified dimensionality of the differentiated features is enhanced, thereby making the acquired features more task-oriented; through the proposed new tracking strategy, the model can maintain robustness in complex scenes with severe occlusion and cluttered background, and fully alleviates the problems of easy target loss and easy identity switching.
[0040] The beneficial effects of the present invention are mainly manifested in:
[0041] (1) A new re-identification feature construction branch is proposed. Compared with the existing technology, this method has significant advantages in obtaining re-identification features and reducing identity switching and misidentification problems. The traditional feature decoupling method still divides the unified image features into two parts required for the detection task and the re-identification task. However, this method reconstructs the low-dimensional and high-dimensional feature branches of the two tasks according to the feature map obtained by the backbone network and obtains the corresponding features, which greatly alleviates the problem of task competition within the model and the impact of the shared neural network on the feature effect.
[0042] (2) The attention mechanism suitable for detection\re-identification tasks is introduced into the multi-target tracking method for joint detection and re-identification. Since the method of the present invention adjusts and modifies the specific attention mechanism in terms of dimensions, the feature enhancement of the attention mechanism performs better on the overall task, thereby improving the overall tracking performance of the model.
[0043] (3) A new tracking strategy is proposed. This tracking strategy can combine the appearance distance to help recover the detected target lost due to temporary occlusion. At the same time, it can help the Kalman filter better locate the tracking target, update the parameters of the Kalman filter and reduce the cumulative error of the Kalman filter, so that the model can achieve better tracking effect.
[0044] (4) It provides more flexible use. The feature decoupling reconstruction branch, attention mechanism and tracking method proposed in this method can be easily inserted on the basis of various models to achieve a plug-and-play effect. Extensive experiments conducted on multiple comparative experiments and data sets show that the method of the present invention can enhance the incremental learning performance of existing models without adding additional costs. Ablation studies further verify the effectiveness of each component in the method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a principle block diagram of a multi-target tracking method based on feature decoupling and feature enhancement learning. DETAILED DESCRIPTION
[0046] The present invention will be further described below in conjunction with the accompanying drawings.
[0047] Reference Figure 1 , a multi-target tracking method based on feature decoupling and feature enhancement learning, comprising the following steps:
[0048] Step 1: Obtain image features and perform feature decoupling and enhancement. The process is as follows:
[0049] S11, inputting a pedestrian multi-target tracking sequence, where each frame in the sequence contains multiple pedestrian tracking targets;
[0050] S12, input the pedestrian image data into the backbone network and obtain the 4-layer feature map, then perform feature decoupling and differentiation according to the 4-layer feature map, and then introduce the modified attention mechanism to enhance the task-orientedness of the decoupled and differentiated image features;
[0051] The goal of feature decoupling is to provide different task branches with the features of the required dimensions in the same model, thereby alleviating the performance problems caused by feature ambiguity due to task competition. The process is as follows:
[0052] Firstly, DLA-34 is adopted as the backbone network to achieve a good balance between detection speed and accuracy. Considering the different semantic levels of features required for detection and re-identification tasks, a dual-branch feature fusion network is proposed, which uses two different feature aggregation methods to obtain the feature representation required for detection and re-identification.
[0053] For the detection part, the reconstructed IDA-UP structure is used, which is based on operating on the upsampled three-layer deep feature maps and fusing them with the shallow feature maps, and is called a top-down feature fusion strategy;
[0054] For the re-identification part, it requires shallower features to distinguish different samples within each category; a bottom-up feature aggregation strategy is proposed, represents the feature map input to the network, where N represents the number of layers with feature maps of different resolutions extracted by the backbone network, Fi represents the feature map of the i-th layer, and then the bottom-up feature fusion strategy is expressed as:
[0055]
[0056] in, Represents the final fused features of each layer, represents the features obtained after fusion of the previous layer, UpSample(·) represents the upsampling operation consisting of deformable convolution and deconvolution, Conv 1×1 (·) represents a convolutional layer of size 1*1, which is used to change the number of feature channels, and σ(·) represents the Sigmoid activation layer.
[0057] The goal of feature enhancement is to enhance the image features obtained by decoupling and differentiation in a task-oriented manner, so as to obtain more discriminative detection features and re-identification features. Therefore, the attention mechanism is introduced for feature enhancement.
[0058] In order to strengthen the network's attention to foreground information and suppress background noise, a spatial feature enhancement module SFEM is proposed to improve the robustness of feature encoding to complex environments and target sizes and increase its global representation. SFEM is a superposition of a two-layer recursive spatial attention mechanism, which is simpler and more efficient than other attention mechanisms; considering the feature fusion strategy of the top-down detection part, the highest-level feature map obtained by the backbone is passed into the SFEM module, and the fusion formula for each layer is:
[0059]
[0060] where Q,K,V∈R C×H×W is the input feature F C×H×W The matrix obtained by three learnable weight vectors, Q u represents any unit eigenvector in the Q matrix, ε i,u ,φ i,u represents the i-th value in the set of feature vectors in the same row or column as u in the K and V matrices, respectively. Softmax(·) is used to change the number of feature channels;
[0061] Different from the detection task, the re-identification task requires obtaining higher-level features of the object to distinguish each instance within the object class. An identity recognition feature enhancement module (IEM) adapted to pixel-level fine-grained tasks is proposed. IEM completely folds features in one direction while maintaining high resolution of features in its orthogonal direction. Then, the Softmax() and Sigmoid() operations of the bottleneck vector (the smallest feature vector in the attention block) are used to increase the dynamic range of attention, thereby enhancing the feature vector. The channel attention formula of IEM is:
[0062] A c (X) = F SG [W z ((σ1(W v (X))×F SM (σ2(W q (X))))))]
[0063] where X∈R C×H×W , W q ,W v ,W z They are convolutional layers of size 1X1, σ1, σ2 are two vector resetting factors, F SM (·) is the SoftMax operator, W v ,W z The number of internal channels is C / 2, A c (X)∈R C×1×1 , output Z c =Ac (X) c X∈R C×H×W ;
[0064] The spatial attention formula of IEM is:
[0065] A s (X) = F SG [σ3(F SM (σ1(F GP (W q (X))))×σ2(W v (X)))]
[0066] where X∈R C×H×W , W q ,W v They are convolutional layers of size 1X1, σ1, σ2, σ3 are three vector reset factors, F SM (·) is the SoftMax operator, output Z p =A p (X) p X∈R C×H×W ;
[0067] The above two branches are added in parallel to get the final result:
[0068] PSAP(X)=Z c +Z p .
[0069] S13, two parts of image features after feature decoupling and feature enhancement;
[0070] Step 2: Target tracking. The process is as follows:
[0071] S21, taking the two parts of image features output by S13 as input;
[0072] S22, after the input features pass through a specific classification head, box_size, center_off, heatmap and reid data are obtained, and target association is completed based on these data based on the joint detection and tracking paradigm;
[0073] S23, target association: The multi-target tracking model outputs data according to S22, and retains the coordinate position of each tracked target in a specific frame sequence, while considering the re-identification features of each target for target association and re-recovery;
[0074] The multi-target tracking model adopts a tracking strategy based on IOU and appearance embedding distance. It follows the secondary matching principle of ByteTrack's high-resolution frame and low-resolution frame. After the secondary matching, it relies on the precise re-identification features obtained by the feature decoupling enhancement framework to re-check and retrieve the lost targets.
[0075] When the target is lost due to temporary blockage or other reasons, the appearance embedding information is first used to compensate for the drift of the Kalman filter to obtain the final predicted target center position, and the average motion speed (V) of the target center during the retention time (30 frames) is calculated. i ,V j ), considering the average number of frames processed per second t and the number of frames k that the target has lost, the Kalman filter is used to select the prediction center (C i ,C y ) The appearance embedding vectors in the range are calculated and compared with the appearance embedding vector E of the lost target. l The distance between them is calculated, and the formula is used to determine whether they belong to the same trajectory:
[0076]
[0077] Where D min Indicates the prediction center (C i ,C y ) in the N×N range near E l The cosine distance between r Indicates the matching threshold of trajectory recovery. If To represent the appearance embedding vector at pixel (i, j), then D min The calculation formula is as follows:
[0078]
[0079] Among them, F cosdistance (·) indicates calculating the cosine distance between two vectors. If the target object of the lost track is finally found, the found center point is set as the center point of the track in the current frame, so that its size is consistent with the size of the tracked target in the previous frame.
[0080] S24, comprehensive loss: The multi-target tracking model will comprehensively consider the target detection loss and target association loss to optimize the overall performance and balance the competition among tasks within the model;
[0081] S25. After the balance loss calculation, the multi-target tracking model has good robustness and its lightweight size can maintain the real-time requirements of tracking.
[0082] In this embodiment, in order to evaluate the method proposed in the present invention, DLA-34 is used as the backbone network to extract features. The size of the input image is 1088×608, and the size of the feature map is 272×152. The model is trained for 30 cycles on an NVIDIA GeForce RTX 3090 GPU, and the Adam optimizer is used to adjust the parameters. The initial learning rate is set to 1×10-5. The batch size is set to 12.
[0083] The same dataset as in FairMOT is used for training, including ETH, CityPerson, CalTech, CUHK-SYSU, PRW, MOT17, and CrowdHuman, with a total of 20,086 images for training. The tracker is evaluated on two tracking benchmark test sets: MOT17 and MOT20.
[0084] Tables 1 and 2 show the ablation study conducted on the MOT17 validation set and the experimental results on different datasets.
[0085]
[0086] Table 1
[0087]
[0088] Table 2
[0089] By introducing feature decoupling and enhancement and tracking methods into FairMOT, the competition between detection and re-ID tasks is alleviated. The performance of the tracker has achieved some improvements in MOTA (+0.6), IDF1 (+1.9) and TPR (+0.3). With the introduction of spatial feature enhancement (SFE) and identity feature enhancement (IFE) modules, the performance of the tracker has been further improved in MOTA (+1.3), IDF1 (+2.4) and TPR (+0.8). This shows that the features of detection and re-ID tasks have been enhanced in a targeted manner, showing stronger task specificity. When combined with the re-ID (RF) strategy, the model framework shows the highest performance improvement, surpassing the baseline in MOTA (+2.6), IDF1 (+4.3) and TPR (+1.0), while reducing identity switches (IDs) by 187 times, which shows that the RF strategy effectively enhances the identity recognition ability. In addition, the three modules can complement each other and jointly promote the performance improvement of detection and tracking tasks.
[0090] Table 3 is the experimental results obtained on the MOTChallenge;
[0091]
[0092] Table 3
[0093] The results show that this method can be plug-and-play with existing methods, improves the accuracy of existing models to varying degrees, and can significantly improve the performance of the model.
[0094] The contents described in the embodiments of this specification are merely enumerations of implementation forms of the inventive concept and are for illustrative purposes only. The protection scope of the present invention should not be considered to be limited to the specific forms described in this embodiment, and the protection scope of the present invention also extends to equivalent technical means that can be thought of by ordinary technicians in this field based on the inventive concept.
Claims
1. A multi-target tracking method based on feature decoupling and feature enhancement learning, characterized in that: The method comprises the following steps: Step 1: Obtain image features and perform feature decoupling and enhancement. The process is as follows: S11, inputting a pedestrian multi-target tracking sequence, where each frame in the sequence contains multiple pedestrian tracking targets; S12, input the pedestrian image data into the backbone network and obtain the 4-layer feature map, then perform feature decoupling and differentiation according to the 4-layer feature map, and then introduce the modified attention mechanism to enhance the task-orientedness of the decoupled and differentiated image features; S13, two parts of image features after feature decoupling and feature enhancement; Step 2: Target tracking. The process is as follows: S21, taking the two parts of image features output by S13 as input; S22, after the input features pass through a specific classification head, box_size, center_off, heatmap and reid data are obtained, and target association is completed based on these data based on the joint detection and tracking paradigm; S23, target association: The multi-target tracking model outputs data according to S22, and retains the coordinate position of each tracked target in a specific frame sequence, while considering the re-identification features of each target for target association and re-recovery; S24, comprehensive loss: The multi-target tracking model will comprehensively consider the target detection loss and target association loss to optimize the overall performance and balance the competition among tasks within the model; S25. After the balance loss calculation, the multi-target tracking model has good robustness and its lightweight size can maintain the real-time requirements of tracking.
2. The multi-target tracking method based on feature decoupling and feature enhancement learning as claimed in claim 1, characterized in that: In S12, the process of feature decoupling is: Firstly, DLA-34 is used as the backbone network. Considering the different semantic levels of features required for detection and re-identification tasks, a dual-branch feature fusion network is proposed. The network uses two different feature aggregation methods to obtain the feature representation required for detection and re-identification. For the detection part, the reconstructed IDA-UP structure is used, which is based on operating the upsampled three-layer deep feature maps and fusing them with the shallow feature maps, and is called a top-down feature fusion strategy; For the re-identification part, a bottom-up feature aggregation strategy is proposed. represents the feature map input to the network, where N represents the number of layers with feature maps of different resolutions extracted by the backbone network, Fi represents the feature map of the i-th layer, and then the bottom-up feature fusion strategy is expressed as: in, Represents the final fused features of each layer, represents the features obtained after fusion of the previous layer, UpSample(·) represents the upsampling operation consisting of deformable convolution and deconvolution, Conv 1×1 (·) represents a convolutional layer of size 1*1, which is used to change the number of feature channels, and σ(·) represents the Sigmoid activation layer.
3. The multi-target tracking method based on feature decoupling and feature enhancement learning as claimed in claim 2, characterized in that: In S12, an attention mechanism is introduced for feature enhancement, and a spatial feature enhancement module SFEM is proposed. SFEM is a superposition of a recursive spatial attention mechanism with a depth of two layers. Considering the feature fusion strategy of the constructed top-down detection part, the highest-level feature map obtained by the backbone is passed into the SFEM module. The fusion formula of each layer is: where Q,K,V∈R C×H×W is the input feature F C×H×W The matrix obtained by three learnable weight vectors, Q u represents any unit eigenvector in the Q matrix, ε i,u ,φ i,u represents the i-th value in the set of feature vectors in the same row or column as u in the K and V matrices, respectively. Softmax(·) is used to change the number of feature channels; An identity recognition feature enhancement module (IEM) adapted to pixel-level fine-grained tasks is proposed. IEM completely folds features in one direction while maintaining high resolution of features in its orthogonal direction. Then, the dynamic range of attention is increased through Softmax() and Sigmoid() operations on the bottleneck vector, thereby enhancing the feature vector. The channel attention formula of IEM is: A c (X)=F SG [W z ((σ1(W v (X))×F SM (σ2(W q (X)))))] where X∈R C×H×W , W q ,W v ,W z They are convolutional layers of size 1X1, σ1, σ2 are two vector resetting factors, F SM (·) is the SoftMax operator, W v ,W z The number of internal channels is C / 2, A c (X)∈R C×1×1 , output Z c =A c (X) c X∈R C×H×W ; The spatial attention formula of IEM is: A s (X)=F SG [σ3(F SM (σ1(F GP (W q (X))))×σ2(W v (X)))] where X∈R C×H×W , W q ,W v They are convolutional layers of size 1X1, σ1, σ2, σ3 are three vector reset factors, F SM (·) is the SoftMax operator, output Z p =A p (X) p X∈R C×H×W ; The above two branches are added in parallel to get the final result: PSAP(X)=Z c +Z p 。 4. The multi-target tracking method based on feature decoupling and feature enhancement learning as claimed in claim 2 or 3, characterized in that: In S23, the multi-target tracking model adopts a tracking strategy based on IOU and appearance embedding distance. First, the appearance embedding information is used to compensate for the drift of the Kalman filter to obtain the final predicted target center position, and the average motion speed (V i ,V j ), considering the average number of frames processed per second t and the number of frames k that the target has lost, the Kalman filter is used to select the prediction center (C i ,C y ) The appearance embedding vectors in the range are calculated and compared with the appearance embedding vector E of the lost target. l The distance between them is calculated and the formula is used to determine whether they belong to the same trajectory: Where D min Indicates the prediction center (C i ,C y ) in the N×N range near E l The cosine distance between r Indicates the matching threshold of trajectory recovery. If To represent the appearance embedding vector at pixel (i, j), then D min The calculation formula is as follows: Among them, F cosdistance (·) indicates calculating the cosine distance between two vectors. If the target object of the lost track is finally found, the found center point is set as the center point of the track in the current frame, so that its size is consistent with the size of the tracked target in the previous frame.
Citation Information
Patent Citations
Joint detection and tracking method, device and equipment based on multi-dimensional attention mechanism
CN114663812A
Anchor-free real-time multi-target tracking method based on joint detection and re-identification
CN117437260A
Multi-target tracking method based on enhancement of target appearance feature saliency
CN118840394A