Pedestrian multi-target tracking method based on attention mechanism
Through the pedestrian multi-objective tracking method based on attention mechanism, the problems of insufficient feature extraction and task conflict in the occlusion scenario are solved, and higher tracking accuracy and effectiveness are achieved. Especially in the occlusion scenario, the accuracy of feature extraction and trajectory association is significantly improved.
Patent Information
- Application Number
- CN202510461429.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-18
AI Technical Summary
The existing pedestrian multi-objective tracking method has insufficient feature extraction in the occluded scenario, conflicts between the object detection task and the Re-ID task, and unreasonable data correlation strategy design, resulting in poor tracking performance.
The pedestrian multi-objective tracking method based on attention mechanism is adopted, and the multi-scale pedestrian feature map is weighted and fusion through the target perception module. The feature decoupling module decouples the object detection and Re-ID task features, and optimizes the target trajectory association with the quadratic association algorithm.
The accuracy and effectiveness of pedestrian multi-target tracking are improved, especially in occlusion scenarios, which significantly improve feature extraction capabilities and trajectory correlation accuracy, reduce the loss of low-score detection boxes, and avoid target trajectory interruptions.
Smart Images

Figure CN120339333A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of Internet big data and computer vision, and particularly relates to a multi-object tracking method for pedestrians based on an attention mechanism. Background Art
[0002] Multi-object tracking is an important research direction in the field of computer vision, which not only has profound academic value but also shows broad prospects in practical applications. The application scope of this task is extensive. It can not only track specific targets (such as pedestrians, vehicles or animals), but also handle the tracking problem of multi-class targets. Among many research objects, pedestrians have become the focus of attention in the academic and industrial fields due to the frequent changes in their postures and appearances during movement, and their complexity in crowded and occluded scenarios. Compared with targets with less posture and appearance changes (such as vehicles), pedestrian tracking is more challenging. Pedestrian multi-object tracking technology has been widely applied in many life scenarios such as intelligent monitoring systems, autonomous driving, and human-computer interaction.
[0003] In recent years, with the development of deep learning, multi-object tracking algorithms based on deep learning can be mainly divided into two categories: two-stage methods and one-stage methods. Two-stage methods handle object detection and re-identification (Re-ID) as two independent sub-tasks: first, use a detection algorithm to generate object bounding boxes, then extract appearance features and generate appearance embedding vectors through a Re-ID model, and finally combine position information to achieve object association. Although two-stage methods perform excellently in tracking effects, the separation of modules also brings problems such as large computational overhead, high training cost, and reduced coupling between modules. One-stage methods integrate the detection task and the Re-ID task into one network. By sharing backbone features, they simultaneously output the position information of the detection box and the appearance embedding vector of the target, thereby reducing the model's computational amount, simplifying the training and tracking processes, and greatly improving the accuracy of multi-object tracking. However, one-stage methods still face many challenges:
[0004] First, in occluded scenarios, pedestrian targets may be partially or completely occluded, resulting in a significant weakening of the distinction between the target and the background. It is difficult for the model to effectively capture key pedestrian feature information, making it difficult for the tracking algorithm to continuously and accurately associate the positions and states of the target in different frames, thus affecting the overall performance of the model.
[0005] Secondly, compared with two-stage methods, one-stage methods often perform poorly in terms of performance. Two-stage methods handle object detection and Re-ID tasks separately, optimizing the object location and appearance features respectively, which enables them to focus more on the characteristics of their respective tasks. This separation approach allows two-stage methods to have more room for optimization in object detection and appearance matching, effectively improving the tracking accuracy. In contrast, one-stage methods have obvious limitations. Since there are significant differences in feature requirements between the object detection task and the Re-ID task. Specifically, the object detection task pays more attention to the similarity between objects of the same category, while the Re-ID task aims to maximize the differences between different objects. When these two tasks share features, conflicts in feature learning are likely to occur, which will directly affect the tracking performance of the algorithm.
[0006] Finally, one-stage methods have the following deficiencies in the association strategy: 1) When performing association matching at each layer, usually only a single feature is relied on to generate the similarity matrix. This method is difficult to comprehensively capture the complex correlations between objects, resulting in limited matching accuracy. 2) By setting a detection score threshold to filter out valid objects, although this is a commonly used strategy, detection results with low detection scores are often regarded as unreliable information and directly discarded. However, low-detection-score objects are usually caused by factors such as partial occlusion, object deformation, or complex background interference. Simply eliminating these objects may lead to the omission of real objects, thus compromising the coherence of the object trajectory.
[0007] In summary, the problems of poor performance in pedestrian multi-object tracking methods caused by insufficient extraction of key features of pedestrians in occlusion scenarios, conflicts between object detection tasks and Re-ID tasks, and unreasonable design of data association strategies need to be solved urgently. Summary of the Invention
[0008] Aiming at the above deficiencies of the prior art, the technical problem to be solved by the present invention is: how to provide a pedestrian multi-object tracking method based on the attention mechanism, which can systematically solve problems such as occlusion handling, task conflicts, and data association strategies in pedestrian multi-object tracking through the collaborative design of an object perception module, a feature decoupling module, and a secondary association algorithm, thereby improving the accuracy and effectiveness of pedestrian multi-object tracking.
[0009] To solve the above technical problems, the present invention adopts the following technical solutions:
[0010] A pedestrian multi-object tracking method based on the attention mechanism, comprising:
[0011] S1: Obtain the current frame image;
[0012] S2: Input the current frame image into the trained object tracking model, and obtain the tracking results of each pedestrian object based on the object detection results and appearance embedding vectors;
[0013] The processing steps of the object tracking model include:
[0014] S201: Input the current frame image into the backbone network for feature extraction to obtain pedestrian feature maps at multiple scales;
[0015] S202: Input the pedestrian feature maps at multiple scales into the object perception module based on the spatial attention mechanism for weighted fusion to obtain a fused feature map;
[0016] S203: Input the fused feature map into the feature decoupling module based on the channel attention mechanism for feature decoupling to obtain a first feature map for object detection and a second feature map for Re-ID;
[0017] S204: Input the first feature map and the second feature map into the object detection branch and the Re-ID branch respectively for object detection and Re-ID to obtain object detection results and appearance embedding vectors;
[0018] S205: Based on the object detection results and appearance embedding vectors, perform association matching between the detection boxes and trajectories of pedestrian objects through the quadratic association algorithm to obtain the tracking results of each pedestrian object;
[0019] S3: Take the tracking results of each pedestrian object as the multi-object tracking results of pedestrians in the image to be detected.
[0020] Preferably, in step S201, the backbone network is the DLA34 backbone network. The DLA34 backbone network is used to perform feature extraction on the image to be detected to obtain pedestrian feature maps with multiple resolutions, that is, pedestrian feature maps at multiple scales.
[0021] Preferably, in step S202, the processing steps of the object perception module include:
[0022] S2021: Align the regions of the pedestrian feature maps at each scale to obtain aligned feature maps at each scale;
[0023] S2022: Input the aligned feature maps at each scale into the corresponding spatial attention mechanism modules respectively to obtain spatial attention response maps at each scale;
[0024] S2023: After multiplying each pixel of the spatial attention response maps at each scale with the aligned feature maps at the corresponding scale, and then adding each pixel with the aligned feature maps at the corresponding scale, obtain spatial attention enhanced feature maps at each scale;
[0025] S2024: Assign corresponding weights to the spatial attention enhanced feature maps of each scale;
[0026] S2025: Perform weighted fusion on the spatial attention enhanced feature maps of each scale based on the assigned weights, and use 1×1 convolution to perform channel mapping on the fusion result to obtain the final fusion feature map.
[0027] Preferably, in step S2021, the pedestrian feature maps of multiple scales include pedestrian feature maps with resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32;
[0028] The processing steps of region alignment include:
[0029] For the pedestrian feature map of 1 / 4, input it into the deformable convolution module to adjust the spatial distribution;
[0030] For the pedestrian feature map of 1 / 8, after inputting it into the deformable convolution module to adjust the spatial distribution, upsample it to the 1 / 4 size;
[0031] For the pedestrian feature map of 1 / 16, after inputting it into the deformable convolution module to adjust the spatial distribution, upsample it to the 1 / 4 size;
[0032] For the pedestrian feature map of 1 / 32, after inputting it into the deformable convolution module to adjust the spatial distribution, upsample it to the 1 / 4 size.
[0033] Preferably, in step S2022, the processing steps of the spatial attention mechanism module include:
[0034] S20221: Perform average pooling and max pooling operations on the input pedestrian feature map in the channel direction respectively to obtain two single-channel feature maps;
[0035] S20222: Use a 7×7 convolutional layer to converge the two single-channel feature maps into a single-channel feature map;
[0036] S20223: Perform Sigmoid normalization on the converged single-channel feature map to obtain the spatial attention response map.
[0037] Preferably, in step S203, the processing steps of the feature decoupling module include:
[0038] S2031: Input the fusion feature map into the channel attention module to obtain the channel attention weight map;
[0039] S2032: Multiply the fusion feature map and the channel attention weight map element by element to obtain the channel attention enhanced feature map;
[0040] S2033: After inputting the channel attention enhanced feature map into the first convolutional layer, perform element-wise addition with the fusion feature map to obtain the first feature map for object detection;
[0041] S2034: After inputting the channel attention enhanced feature map into the second convolutional layer, perform element-wise addition with the fusion feature map to obtain the second feature map for Re-ID.
[0042] Preferably, in step S2031, the processing steps of the channel attention module include:
[0043] S20311: Perform global average pooling and global max pooling operations on the fusion feature map respectively to obtain two dimensionality-reduced feature maps;
[0044] S20312: Input the two dimensionality-reduced feature maps into a shared network containing a one-dimensional convolutional layer and a fully-connected layer for feature vector encoding respectively to obtain two feature vectors;
[0045] S20313: Perform element-wise addition on the two feature vectors and then perform Sigmoid normalization operation to obtain the channel attention weight map.
[0046] Preferably, in step S204, the object detection results output by the object detection branch include the spatial positions, sizes, and detection scores of the detection frames of each pedestrian object;
[0047] The output of the Re-ID branch includes the appearance embedding vectors of the detection frames of each pedestrian object.
[0048] Preferably, in step S205, the processing steps of the quadratic association algorithm include:
[0049] S2051: Obtain the object detection results and appearance embedding vectors in the current frame image;
[0050] S2052: Predict the prediction frames of all trajectories in the current frame image through the Kalman filter;
[0051] S2053: According to the detection score thresholds ds high and ds low classify the detection frames into high-score detections D high and low-score detections D low : Detection frames with detection scores higher than ds high are classified into D high , and scores higher than ds low less than ds high are classified into D low ;
[0052] S2054: Associate and match the detection frames in the high-score detections D high with all trajectories:
[0053] 1) Calculate the GIoU distance between the detection box and the predicted box of the trajectory;
[0054] 2) Calculate the cosine distance between the appearance embedding vector of the detection box and the appearance embedding vector of the target corresponding to the trajectory;
[0055] 3) Weight the GIoU distance and the cosine distance of the detection box to obtain the fusion matrix EG;
[0056] The formula is expressed as:
[0057] EG = λcosine+(1 - λ)GIoU;
[0058] Where: cosine represents the cosine distance; GIoU represents the GIoU distance; λ is set to 0.8;
[0059] 4) Perform data association based on the fusion matrix EG through the Hungarian algorithm;
[0060] 5) Apply auxiliary data association based on the GIoU distance;
[0061] 6) Store the unmatched detection boxes and unmatched trajectories in the remaining high-score detections D high and the remaining trajectories T remain and the remaining trajectories T remain ;
[0062] S2055: Associate and match the detection boxes in the low-score detections D low with the remaining trajectories T remain :
[0063] 1) Calculate the GIoU distance between the detection box and the predicted box of the trajectory;
[0064] 2) Perform data association based on the GIoU distance through the Hungarian algorithm;
[0065] 3) Store the unmatched trajectories in the remaining trajectories T remain in the remaining trajectories T re-remain ;
[0066] 4) Delete the unmatched detection boxes in the low-score detections D low as the background;
[0067] S2056: Initialize the detection boxes in the remaining high-score detections D remain whose detection scores exceed the trajectory initialization threshold ts as new trajectories.
[0068] Preferably, in step S2051, if a certain trajectory is not matched for more than N frames of images continuously, then delete the trajectory.
[0069] Compared with the prior art, the multi-object pedestrian tracking method based on the attention mechanism in the present invention has the following beneficial effects:
[0070] In the present invention, the base target perception module performs weighted fusion on the pedestrian feature maps of multiple scales. The target perception module mines effective target information through spatial attention and deformable convolution, and assigns adaptive weights to the input feature layers of different scales, enabling the network to dynamically adjust the proportion of the multi-scale fusion features, aggregating the discriminative features of the pedestrian targets in the current frame, alleviating the problem that it is difficult for the model to capture key features due to occlusion, enabling the model to adaptively focus on the key features of pedestrians, enhancing the discriminability of pedestrian features, significantly improving the feature extraction ability of the model in occlusion scenarios, and thus improving the accuracy of multi-object pedestrian tracking.
[0071] In the present invention, the feature decoupling module decouples the fused feature map to obtain a first feature map for object detection and a second feature map for Re-ID. The shared features are decoupled into task-specific feature representations through the channel attention mechanism, providing targeted feature inputs for the object detection and pedestrian Re-ID tasks respectively, reducing the conflict between tasks. The first feature map focuses on low-level semantic information such as the position and scale of pedestrians, aiming to improve the regression accuracy of the detection box. The second feature map is used to encode high-level semantic information such as the appearance and pose of pedestrians, aiming to enhance the uniqueness of the appearance embedding vector. Through the decoupling design, the mutual interference between the object detection and Re-ID tasks during feature sharing is avoided, thereby improving the effectiveness of multi-object pedestrian tracking.
[0072] In the present invention, the first feature map and the second feature map are respectively input into the object detection branch and the Re-ID branch for object detection and Re-ID, and then the detection box and trajectory of the pedestrian target are associated and matched through the quadratic association algorithm based on the object detection result and the appearance embedding vector of Re-ID, obtaining a tracking detection image containing the detection box of the pedestrian target and the corresponding trajectory information. First, by introducing the quadratic association strategy, low-score detection boxes are included in the matching process, effectively reducing the loss of low-score detection boxes. Secondly, in the first matching of the high-score detection results with all trajectories, the appearance feature and the motion feature are fused for measurement, where the motion feature similarity is improved from IoU to the generalized intersection over union GIoU; at the same time, an auxiliary data association strategy based on GIoU is introduced to make full use of the high-score detections. Finally, in the second matching of the low-score detection results with the remaining unmatched target trajectories, only GIoU is used as the similarity metric. In summary, the present invention solves the ID switching problem caused by target occlusion through the quadratic association algorithm, and can optimize the target trajectory association accuracy, thereby further improving the accuracy of multi-object pedestrian tracking. Description of the Drawings
[0073] To make the objectives, technical solutions, and advantages of the invention clearer, the following will further describe the present invention in detail with reference to the accompanying drawings, where:
[0074] Figure 1 It is a logic block diagram of a multi-object pedestrian tracking method based on an attention mechanism.
[0075] Figure 2 (a) is the network structure diagram of the target perception module; Figure 2 (b) is the network structure diagram of the spatial attention mechanism module.
[0076] Figure 3 (a) is the network structure diagram of the feature decoupling module; Figure 3 (b) is the network structure diagram of the channel attention mechanism module.
[0077] Figure 4 It is a comparison of the tracking effects of the model of the present invention and FairMOT
[0078] Figure 5 It is the visualization effects on the 250th, 300th, and 350th frames of the MOT17 test set. Specific implementation manners
[0079] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Usually, the components of the embodiments of the present invention described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0080] It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the invention is usually placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance. In addition, terms such as "horizontal" and "vertical" do not mean that the components are required to be absolutely horizontal or hanging, but can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but can be slightly inclined. In the description of the present invention, it should also be noted that unless otherwise clearly specified and limited, the terms "set", "install", "connected", "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0081] The following is a more detailed description through specific embodiments:
[0082] Embodiment:
[0083] A multi-object pedestrian tracking method based on an attention mechanism is disclosed in this embodiment.
[0084] As Figure 1 shown, the multi-object pedestrian tracking method based on an attention mechanism includes:
[0085] S1: Obtain the image to be detected;
[0086] S2: Input the image to be detected into the trained target tracking model, and obtain the tracking results of each pedestrian target based on the target detection results and Re-ID (ReID) features. The finally output is a tracking detection image containing the pedestrian target detection box and the corresponding trajectory information.
[0087] The processing steps of the target tracking model include:
[0088] S201: Input the image to be detected into the backbone network for feature extraction to obtain pedestrian feature maps at multiple scales;
[0089] S202: Input the pedestrian feature maps at multiple scales into the target perception module based on the Spatial Attention Module (SAM) mechanism for weighted fusion to obtain a fused feature map;
[0090] S203: Input the fused feature map into the feature decoupling module based on the Channel Attention Module (CAM) mechanism for feature decoupling to obtain a first feature map for object detection and a second feature map for Re-ID;
[0091] S204: Input the first feature map and the second feature map into the object detection branch and the Re-ID branch respectively for object detection and Re-ID to obtain an object detection result and an appearance embedding vector;
[0092] S205: Perform association matching between the detection boxes and trajectories of pedestrian targets based on the object detection result and the appearance embedding vector through the quadratic association algorithm to obtain the tracking results of each pedestrian target;
[0093] S3: Take the tracking results of each pedestrian target as the multi-object tracking results of pedestrians in the image to be detected.
[0094] In the present invention, the target perception module is used to perform weighted fusion on pedestrian feature maps at multiple scales. The target perception module mines effective target information through spatial attention and deformable convolution, and assigns adaptive weights to input feature layers at different scales, enabling the network to dynamically adjust the proportion of multi-scale fusion features, aggregate the discriminative features of pedestrian targets in the current frame, alleviate the problem that it is difficult for the model to capture key features due to occlusion, enable the model to adaptively focus on the key features of pedestrians, enhance the discriminability of pedestrian features, significantly improve the feature extraction ability of the model in occlusion scenarios, and thus improve the accuracy of multi-object pedestrian tracking.
[0095] In the present invention, the feature decoupling module is used to perform feature decoupling on the fused feature map to obtain a first feature map for object detection and a second feature map for Re-ID. The shared features are decoupled into task-specific feature representations through the channel attention mechanism, providing targeted feature inputs for object detection and pedestrian Re-ID tasks respectively, reducing conflicts between tasks. The first feature map focuses on low-level semantic information such as pedestrian position and scale, aiming to improve the regression accuracy of detection boxes. The second feature map is used to encode high-level semantic information such as pedestrian appearance and pose, aiming to enhance the uniqueness of the appearance embedding vector. Through the decoupling design, the mutual interference between object detection and Re-ID tasks during feature sharing is avoided, thereby improving the effectiveness of multi-object pedestrian tracking.
[0096] The present invention inputs the first feature map and the second feature map into the target detection branch and the Re-ID branch respectively for target detection and Re-ID, and then performs the association matching between the detection box and the trajectory of the pedestrian target based on the target detection result and the appearance embedding vector of Re-ID through a quadratic association algorithm, so as to obtain a tracking detection image containing the pedestrian target detection box and the corresponding trajectory information. First, by introducing a quadratic association strategy, low-score detection boxes are incorporated into the matching process, effectively reducing the loss of low-score detection boxes (low-confidence detection results are often regarded as unreliable information and directly discarded, but such targets are usually caused by factors such as partial occlusion, target deformation, or complex background interference. Simple elimination may lead to the omission of real targets and damage the trajectory coherence). Secondly, in the first matching of high-score detection results with all trajectories, the appearance feature and the motion feature are fused for measurement, where the motion feature similarity is improved from IoU to Generalized Intersection Over Union (GIoU); at the same time, an auxiliary data association strategy based on GIoU is introduced to make full use of high-score detections. Finally, in the secondary matching of low-score detection results with the remaining unmatched target trajectories, only GIoU is used as the similarity index (because low-score detection boxes often have severe occlusion or motion blur, and the credibility of their appearance features is relatively low). In summary, the present invention solves the ID switching problem caused by target occlusion through a quadratic association algorithm, can optimize the target trajectory association accuracy, and thus further improves the accuracy of multi-target pedestrian tracking.
[0097] To better introduce the technical solution of the present invention, this embodiment is described through the following several parts.
[0098] I. Backbone Network
[0099] The backbone network is the DLA34 backbone network, which is a lightweight and efficient deep learning model. It adopts a tree structure and hierarchical aggregation technology to iteratively fuse the feature information of each layer of the network. It combines the dense connection of DenseNet and the spatial aggregation advantage of the feature pyramid, enhancing the feature extraction and representation ability of the model, and is especially suitable for multi-scale target detection and recognition tasks. In multi-target pedestrian tracking, DLA34 can output the detection box and trajectory information of pedestrian targets, helping to achieve accurate and efficient tracking.
[0100] The backbone network is the DLA34 backbone network. The DLA34 backbone network is used to extract features from the image to be detected, and pedestrian feature maps with multiple resolutions (1 / 4, 1 / 8, 1 / 16, 1 / 32) are obtained, that is, pedestrian feature maps of multiple scales.
[0101] II. Target Perception Module
[0102] Combined withFigure 2 As shown in Figure 2 , the processing steps of the target perception module include:
[0103] S2021: Perform regional alignment on the pedestrian feature maps of each scale to obtain the aligned feature maps of each scale;
[0104] S2022: Input the aligned feature maps of each scale into the corresponding spatial attention mechanism module respectively to obtain the spatial attention response maps of each scale;
[0105] S2023: After multiplying each pixel of the spatial attention response maps of each scale with the aligned feature maps of the corresponding scale, and then adding each pixel with the aligned feature maps of the corresponding scale, obtain the spatial attention enhanced feature maps of each scale;
[0106] S2024: Assign corresponding weights to the spatial attention enhanced feature maps of each scale;
[0107] Common feature fusion methods make the upper-layer feature map have the same resolution as the lower-layer feature map through upsampling, and then use channel concatenation (Concatenation, Concat) or pixel-level addition (Addition, Add) to achieve fusion. However, these two basic methods of Concat and Add have obvious limitations: they ignore the importance differences of features at different scales and default that all feature maps have the same importance. In fact, features at different scales contribute differently to the overall task, and direct concatenation or addition may lead to information interference or loss. Therefore, to better highlight the contributions of features at different scales, the present invention performs a weighting operation on the 4 spatially attention-enhanced feature maps after spatial attention enhancement before feature fusion, assigns learnable weight parameters W i (i = 1, 2, 3, 4) to the spatially attention-enhanced feature maps of each scale, normalizes the weight parameters through the Softmax function to ensure that the sum of the weights is 1, and then uses a 1×1 convolutional layer to complete channel mapping to provide features for subsequent tasks.
[0108] S2025: Perform weighted fusion on the spatial attention enhanced feature maps of each scale based on the assigned weights, and use a 1×1 convolution to perform channel mapping on the fusion result to obtain the final fused feature map.
[0109] The present invention enhances and fuses features by multiplying each point of the spatial attention response map with the pedestrian (aligned) feature map of the corresponding scale, thereby enhancing the saliency of target information and suppressing the interference of background noise. To ensure that information is not lost, the target perception module introduces a residual structure to combine the pedestrian feature map with the feature map processed by spatial attention.
[0110] 1. Regional alignment
[0111] In this embodiment, the pedestrian feature maps of multiple scales include pedestrian feature maps with resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32;
[0112] The processing steps of region alignment include:
[0113] For the pedestrian feature map of 1 / 4, input it into the deformable convolution module to adjust the spatial distribution;
[0114] For the pedestrian feature map of 1 / 8, after inputting it into the deformable convolution module to adjust the spatial distribution, upsample it to the 1 / 4 size;
[0115] For the pedestrian feature map of 1 / 16, after inputting it into the deformable convolution module to adjust the spatial distribution, upsample it to the 1 / 4 size;
[0116] For the pedestrian feature map of 1 / 32, after inputting it into the deformable convolution module to adjust the spatial distribution, upsample it to the 1 / 4 size.
[0117] 2. Spatial attention mechanism module
[0118] The processing steps of the spatial attention mechanism module include:
[0119] S20221: Perform average pooling (AvgPool) and max pooling (MaxPool) operations on the input pedestrian feature map in the channel direction respectively to obtain two single-channel feature maps with a size of 1×H×W;
[0120] S20222: Use a 7×7 convolutional layer to converge the two single-channel feature maps into a single-channel feature map;
[0121] S20223: Perform Sigmoid normalization on the converged single-channel feature map to obtain a spatial attention response map. In this spatial attention response map, the target information will show a high response intensity (close to 1) at the adaptive resolution, while the response of the background or irrelevant information is suppressed (close to 0).
[0122] III. Feature decoupling module
[0123] Combined with Figure 3 as shown, the processing steps of the feature decoupling module include:
[0124] S2031: Input the fused feature map into the channel attention module to obtain a channel attention weight map;
[0125] S2032: Multiply the fused feature map and the channel attention weight map element by element to obtain a channel attention enhanced feature map; The channel attention weight value is multiplied element by element with the original feature map to generate a feature expression that matches the specific task.
[0126] S2033: After inputting the channel attention enhanced feature map into the first convolutional layer (of the upper branch), perform element-wise addition with the fused feature map to obtain the first feature map for object detection;
[0127] S2034: After inputting the channel attention enhanced feature map into the second convolutional layer (of the lower branch), perform element-wise addition with the fused feature map to obtain the second feature map for Re-ID.
[0128] Channel attention module
[0129] The processing steps of the channel attention module include:
[0130] S20311: Perform global average pooling and global max pooling operations on the fused feature map respectively to obtain two dimensionality-reduced feature maps;
[0131] S20312: Input the two dimensionality-reduced feature maps into a shared network containing a one-dimensional convolution and a fully-connected layer for feature vector encoding respectively to obtain two feature vectors;
[0132] S20313: Perform element-wise addition on the two feature vectors and then perform Sigmoid normalization operation to obtain the channel attention weight map.
[0133] Figure 3 The upper branch in it is used to learn the characteristic features of the detection task, and the lower branch is used to learn the characteristic features of the Re-ID task, thus effectively alleviating the optimization conflict problem in multi-task learning. At the same time, a strategy combining global average pooling and global max pooling is adopted to compress the spatial dimension information, and finally a feature description vector at the channel level is obtained. The advantage of this dual-pooling strategy is that GAP can effectively extract the overall feature distribution, while GMP focuses on highlighting the significant feature regions. The synergistic effect of the two realizes the dual capture of global features and local significant information. In order to enhance the feature expression ability, a residual connection mechanism is also introduced in the feature decoupling module to fuse the initial feature and the attention-weighted feature to form the final task-oriented feature representation. It should be noted that when generating feature maps applicable to different tasks, this decoupling mechanism always maintains the consistency of the output feature and the input feature in the spatial dimension to ensure that the feature size remains unchanged.
[0134] IV. Object detection branch and Re-ID branch
[0135] The main task of the object detection branch is to accurately predict the spatial position and size of each object to ensure that each object can be accurately located between different frames. The Re-ID branch is responsible for extracting the appearance features of each object from the image and generating corresponding embedding vectors. These embedding vectors are the unique identifiers of each object, helping the network to distinguish different objects and maintain their consistency.
[0136] Specifically, the object detection results output by the object detection branch include the spatial positions, sizes, and detection scores (confidence levels) of the detection boxes of each pedestrian object;
[0137] The output of the Re-ID branch includes the appearance embedding vectors of the detection boxes of each pedestrian object.
[0138] V. Secondary Association Algorithm
[0139] As shown in Table 1, the processing steps of the secondary association algorithm include:
[0140] S2051: Obtain the object detection results and appearance embedding vectors (of each detection box) in the current frame image, as well as the trajectory information of all pedestrian objects.
[0141] S2052: Predict the prediction boxes of all trajectories in the current frame image (next state) through the Kalman filter;
[0142] S2053: According to the detection score threshold ds high and ds low classify the detection boxes into high-score detections D high and low-score detections D low : Detection boxes with detection scores higher than ds high are classified into D high , and those with scores higher than ds low less than ds high are classified into D low ; Detection boxes with scores lower than ds low are directly deleted;
[0143] S2054: Associate and match the detection boxes in the high-score detections D high with all trajectories:
[0144] 1) Calculate the GIoU distance between the detection box and the prediction box of the trajectory;
[0145] 2) Calculate the cosine distance between the appearance embedding vector of the detection box and the appearance embedding vector of the object corresponding to the trajectory;
[0146] 3) Weight the GIoU distance and cosine distance of the detection box to obtain the fusion matrix EG;
[0147] The formula is expressed as:
[0148] EG = λcosine + (1 - λ)GIoU;
[0149] In the formula: cosine represents the cosine distance; GIoU represents the GIoU distance; λ is set to 0.8;
[0150] 4) Perform data association based on the fusion matrix EG through the Hungarian algorithm, and associate and match the detection boxes and trajectories.
[0151] 5) Apply auxiliary data association based on the GIoU distance, which also associates and matches the detection boxes and trajectories.
[0152] 6) Store the unmatched detection boxes and unmatched trajectories in the remaining high-score detections D high and the remaining trajectories T remain ; remain
[0153] S2055: Associate and match the detection boxes in the low-score detections D low with the remaining trajectories T remain :
[0154] 1) Calculate the GIoU distance between the detection box and the predicted box of the trajectory.
[0155] 2) Perform data association based on the GIoU distance through the Hungarian algorithm, and associate and match the detection boxes and trajectories.
[0156] 3) Store the unmatched trajectories in the remaining trajectories T remain ; re-remain
[0157] 4) Delete the unmatched detection boxes in the low-score detections D low as the background.
[0158] S2056: Initialize the detection boxes with detection scores exceeding the trajectory initialization threshold ts in the remaining high-score detections D remain as new trajectories.
[0159] If a certain trajectory is not matched for more than N (30) consecutive frames, delete the trajectory.
[0160] Table 1 Pseudo-code for the tracking process of the secondary association algorithm
[0161]
[0162]
[0163] Data association is a key link affecting the performance of multi-target tracking. In the target association stage of the present invention, the detection boxes are optimized as follows: First, all detection boxes are retained and divided into two groups, high-confidence boxes and low-confidence boxes, according to the confidence level (detection score). Subsequently, for detection boxes with different confidence levels, a differentiated association strategy is designed to make full use of the detection information. Similarity matching mainly depends on the position information and appearance features of the targets. The target states exhibit different characteristics at different stages: For the targets in high-score detection boxes, their states are usually relatively stable, and high-quality association can be achieved by weighted fusion of position and appearance information; while for the targets in low-score detection boxes, due to the influence of factors such as occlusion and blur, the appearance information is often unreliable, so appearance similarity is not used for matching. In contrast, there may be partial occlusion for the targets in high-score detection boxes, and introducing appearance information can effectively improve the anti-occlusion performance, thus further enhancing the robustness of the matching.
[0164] VI. Experimental Explanation
[0165] To better illustrate the advantages of the technical solution of the present invention, the following experiments are disclosed in this embodiment.
[0166] The platform and resource configuration of this experiment are as follows: The operating system uses the Ubuntu 20.04 version, and the computing hardware includes 1 NVIDIA GeForce RTX 4090 graphics card (24G video memory). The experiment is based on the PyTorch deep learning framework, and the training and testing of the multi-target tracking model are completed in the environment of Python3.8 and CUDA 11.3.
[0167] 1. Comparative Experiment
[0168] To verify the superiority of the model proposed by the present invention, this experiment conducts a comparative experiment on the MOT16 and MOT17 datasets and makes a comparative analysis with the current advanced algorithms. The experimental results are shown in Tables 2 and 3.
[0169] On the MOT16 dataset, the method of the present invention achieves the best results in both the IDF1 and HOTA metrics, and other metrics also show strong competitiveness.
[0170] In the MOT17 benchmark test, it achieved 62.9% HOTA, 74.8% MOTA, and 77.4% IDF1, showing a significant improvement compared to the baseline model FairMOT: with the same training data, MOTA increased by 1.1% (from 73.7% to 74.8%), IDF1 increased by 5.1% (from 72.3% to 77.4%), and HOTA increased by 3.6% (from 59.3% to 62.9%). Compared with CSTrack, the model of the present invention leads by 3.6% and 4.8% in HOTA and IDF1 respectively. The above results verify the significant performance improvement of the method of the present invention in multi-object tracking tasks, especially showing a competitive advantage in the IDF1 and HOTA metrics.
[0171] Table 2 Experimental results of the MOT16 test set
[0172]
[0173] Table 3 Experimental results of the MOT17 test set
[0174]
[0175]
[0176] 2. Experimental visualization analysis
[0177] Figure 4 Shows the comparison of the tracking results between the original FairMOT model and the model of the present invention (ATTrack) for frames 928, 932, and 945 in video sequence 3 of the MOT17 test set. It can be observed from Figure 4 that the FairMOT model had an obvious missed detection phenomenon at frame 933: two male targets (ID 400 and ID 412) wearing black clothes in the upper left corner of the screen were not successfully detected; at frame 945, although the target ID 400 was successfully detected, ID 412 was missed again due to severe occlusion, resulting in the interruption of its trajectory. In contrast, the model proposed in the present invention showed stronger robustness throughout the tracking process and was able to maintain the ID consistency of these two targets all the time. Even in the case of severe occlusion, the model of the present invention could still successfully detect the targets, effectively avoiding problems such as trajectory interruption or ID swapping, demonstrating its superior performance in complex scenarios.
[0178] Figure 5 Shows the visualization effects of the method of the present invention on frames 250, 300, and 350 of the MOT17 test set. The experimental results show that even under complex conditions such as target scale change, occlusion, or camera movement, the method of the present invention can still run stably and demonstrates excellent tracking performance.
[0179] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Those of ordinary skill in the art should understand that any modifications or equivalent replacements made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions should be covered by the scope of the claims of the present invention.
Claims
1. A multi-object pedestrian tracking method based on the attention mechanism, characterized in that Including: S1: Obtain the current frame image; S2: Input the current frame image into the trained object tracking model, and obtain the tracking results of each pedestrian object based on the object detection results and appearance embedding vectors; The processing steps of the object tracking model include: S201: Input the current frame image into the backbone network for feature extraction to obtain pedestrian feature maps of multiple scales; S202: Input the pedestrian feature maps of multiple scales into the object perception module based on the spatial attention mechanism for weighted fusion to obtain a fused feature map; S203: Input the fused feature map into the feature decoupling module based on the channel attention mechanism for feature decoupling to obtain a first feature map for object detection and a second feature map for Re-ID; S204: Input the first feature map and the second feature map into the object detection branch and the Re-ID branch respectively for object detection and Re-ID to obtain object detection results and appearance embedding vectors; S205: Based on the object detection results and appearance embedding vectors, perform association matching between the detection boxes and trajectories of pedestrian objects through the quadratic association algorithm to obtain the tracking results of each pedestrian object; S3: Take the tracking results of each pedestrian object as the multi-object tracking results of pedestrians in the image to be detected.
2. The multi-object pedestrian tracking method based on the attention mechanism according to claim 1, wherein: In step S201, the backbone network is the DLA34 backbone network. The DLA34 backbone network is used to perform feature extraction on the image to be detected to obtain pedestrian feature maps of multiple resolutions, that is, pedestrian feature maps of multiple scales.
3. The multi-object pedestrian tracking method based on the attention mechanism according to claim 1, characterized in that: In step S202, the processing steps of the object perception module include: S2021: Align the regions of the pedestrian feature maps of each scale to obtain the aligned feature maps of each scale; S2022: Input the aligned feature maps of each scale into the corresponding spatial attention mechanism module respectively to obtain the spatial attention response maps of each scale; S2023: After multiplying each pixel of the spatial attention response map of each scale with the aligned feature map of the corresponding scale, and then adding each pixel with the aligned feature map of the corresponding scale, obtain the spatial attention enhanced feature maps of each scale; S2024: Assign corresponding weights to the spatial attention enhanced feature maps of each scale; S2025: Based on the assigned weights, perform weighted fusion on the spatial attention enhanced feature maps of each scale, and use 1×1 convolution to perform channel mapping on the fusion result to obtain the final fused feature map.
4. The multi-object pedestrian tracking method based on the attention mechanism according to claim 3, characterized in that: In step S2021, the pedestrian feature maps of multiple scales include pedestrian feature maps with resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32; The processing steps of region alignment include: For the pedestrian feature map of 1 / 4, input it into the deformable convolution module to adjust the spatial distribution; For the pedestrian feature map of 1 / 8, after inputting it into the deformable convolution module to adjust the spatial distribution, upsample it to the 1 / 4 size; For the pedestrian feature map of 1 / 16, after inputting it into the deformable convolution module to adjust the spatial distribution, upsample it to the 1 / 4 size; For the pedestrian feature map of 1 / 32, after inputting it into the deformable convolution module to adjust the spatial distribution, upsample it to the 1 / 4 size.
5. The multi-object pedestrian tracking method based on the attention mechanism according to claim 3, characterized in that: In step S2022, the processing steps of the spatial attention mechanism module include: S20221: Perform average pooling and max pooling operations on the input pedestrian feature map in the channel direction to obtain two single-channel feature maps; S20222: Use a 7×7 convolutional layer to converge the two single-channel feature maps into a single-channel feature map; S20223: Perform Sigmoid normalization on the converged single-channel feature map to obtain a spatial attention response map.
6. The multi-object pedestrian tracking method based on the attention mechanism according to claim 1, characterized in that: In step S203, the processing steps of the feature decoupling module include: S2031: Input the fused feature map into the channel attention module to obtain a channel attention weight map; S2032: Multiply the fused feature map and the channel attention weight map element-wise to obtain a channel attention enhanced feature map; S2033: After inputting the channel attention enhanced feature map into the first convolutional layer, add it to the fused feature map element-wise to obtain the first feature map for object detection; S2034: After inputting the channel attention enhanced feature map into the second convolutional layer, add it to the fused feature map element-wise to obtain the second feature map for Re-ID.
7. The multi-object pedestrian tracking method based on attention mechanism according to claim 6, characterized in that: In step S2031, the processing steps of the channel attention module include: S20311: Perform global average pooling and global max pooling operations on the fused feature map respectively to obtain two dimensionality-reduced feature maps; S20312: Input the two dimensionality-reduced feature maps into a shared network containing a one-dimensional convolution and a fully-connected layer for feature vector encoding respectively to obtain two feature vectors; S20313: Add the two feature vectors element-wise and then perform Sigmoid normalization operation to obtain a channel attention weight map.
8. The multi-object pedestrian tracking method based on the attention mechanism according to claim 1, characterized in that: In step S204, the object detection results output by the object detection branch include the spatial positions, sizes, and detection scores of the detection boxes of each pedestrian object; The output of the Re-ID branch includes the appearance embedding vectors of the detection boxes of each pedestrian object.
9. The multi-object pedestrian tracking method based on the attention mechanism according to claim 8, characterized in that: In step S205, the processing steps of the quadratic association algorithm include: S2051: Obtain the object detection results and appearance embedding vectors in the current frame image; S2052: Predict the prediction boxes of all trajectories in the current frame image through a Kalman filter; S2053: According to the detection score threshold ds high and ds low classify the detection boxes into high-score detections D high and low-score detections D low : Detection boxes with detection scores higher than ds high are classified as D high , with scores higher than ds low less than ds high are classified as D low ; S2054: Associate and match the detection boxes in high-score detection D high with all the trajectories: 1) Calculate the GIoU distance between the detection box and the prediction box of the trajectory; 2) Calculate the cosine distance between the appearance embedding vector of the detection box and the appearance embedding vector of the object corresponding to the trajectory; 3) Weight the GIoU distance and cosine distance of the detection box to obtain a fusion matrix EG; The formula is expressed as: EG = λcosine+(1 - λ)GIoU; In the formula: cosine represents the cosine distance; GIoU represents the GIoU distance; λ is set to 0.8; 4) Perform data association based on the fusion matrix EG through the Hungarian algorithm; 5) Apply auxiliary data association based on the GIoU distance; 6) Store the unmatched detection boxes and unmatched trajectories in the high-score detection D high that are not matched and store them in the remaining high-score detection D remain and the remaining trajectories T remain ; S2055: Associate the detection box in D with low score detection and the remaining trajectory T low for correlation matching: remain 1) Calculate the GIoU distance between the detection box and the prediction box of the trajectory; 2) Perform data association based on the GIoU distance through the Hungarian algorithm; 3) Store the unmatched trajectories in the remaining trajectory T remain in the remaining trajectory T re-remain ; 4) Delete the unmatched detection boxes in the low-score detection D low as background; S2056: Initialize the detection boxes with detection scores exceeding the trajectory initialization threshold ts in the remaining high-score detections D as new trajectories. remain Initialize the detection boxes with detection scores exceeding the trajectory initialization threshold ts in the remaining high-score detections D as new trajectories.
10. The multi-object pedestrian tracking method based on the attention mechanism according to claim 9, characterized in that: In step S2051, if a certain trajectory is not matched for more than N consecutive frame images, delete the trajectory.
Citation Information
Cited By
Target detection-tracking integrated method and system based on AFPN and LSK-Net
CN121661565A