Pedestrian Multi-Object Tracking Calculation Method Based on Spatio-Temporal Interaction Attention Mechanism
By introducing a spatiotemporal and spatial interaction attention mechanism in multi-objective tracking, the feature conflict problem caused by the differences in the characteristics of detection and Re-ID tasks is solved, and the accuracy and robustness of pedestrian multi-objective tracking in occlusion and strong light are improved.
Patent Information
- Application Number
- CN202210506694.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-07
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-05-07
AI Technical Summary
In the existing multi-objective tracking computing model, detection tasks and Re-ID tasks have characteristics conflicts due to different characteristics, which reduce performance, and lack tracking accuracy and robustness in occlusion and strong lighting.
A pedestrian multi-objective tracking calculation method based on the spatiotemporal interaction attention mechanism is introduced. Through global optimization and cross-channel feature interaction, more accurate Re-ID features are extracted, and the collaboration capabilities of detection and Re-ID subtasks are improved.
The accuracy and robustness of pedestrian multi-objective tracking in occlusion and strong light areas is significantly improved, and the collaboration capabilities of detection and Re-ID subtasks are enhanced.
Smart Images

Figure CN114998780B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to video scene analysis technology, and particularly to a pedestrian multi-object tracking calculation method based on a spatio-temporal interaction attention mechanism. Background Art
[0002] Multi-object tracking plays a crucial role in the scene understanding task of video analysis; its purpose is to estimate the trajectories of objects and associate the trajectories of target objects with the detection results of each frame in an online or offline manner; multi-object tracking is one of the basic challenges in computer vision and is widely applied in various application and research fields such as video surveillance, traffic control, autonomous driving, and human-computer interaction, having important theoretical research significance and application value.
[0003] Currently, in multi-object tracking calculation models, the detection task and the Re-ID task are two completely different tasks, and they require different characteristics; generally speaking, Re-ID characteristics require more low-level characteristics to distinguish different instances of the same class, while the detection characteristics need to be similar for different instances; if shared characteristics are used, it will lead to feature conflicts and thus reduce performance; if only simple convolution is used to extract Re-ID features, more attention is paid to the detailed feature information therein, thus losing the feature interaction of various effective information; therefore, introducing a spatio-temporal interaction attention mechanism pays more attention to global information and can more accurately extract Re-ID features by jointly using multi-channel features, which is expected to improve the accuracy and robustness of multi-object tracking under occlusion and strong illumination conditions. Summary of the Invention
[0004] Aiming at the problems involved in the above background art, the present invention provides a pedestrian multi-object tracking calculation method based on a spatio-temporal interaction attention mechanism, which improves the accuracy and robustness of pedestrian multi-object tracking under occlusion and strong illumination conditions through global optimization of detailed information and interaction of cross-channel features.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A pedestrian multi-object tracking calculation method based on a spatio-temporal interaction attention mechanism, and the steps are as follows:
[0006] 1) Input two consecutive frames of images in the image sequence;
[0007] 2) Input the first frame of the image sequence into a multi-layer fusion feature extraction network to obtain a fusion feature map;
[0008] 3) Input the extracted fusion feature map of the first frame into a detection branch and a Re-ID branch simultaneously, which are respectively used for detecting objects and extracting Re-ID features;
[0009] 4) In the detection branch, the input fusion feature map Generate the heatmap, object center offset, and bounding box size through 3×3 convolution and 1×1 convolution respectively, where C is the number of channels of the fused feature map, and H and W are the height and width of the fused feature map respectively;
[0010] The heatmap is responsible for estimating the position of the object center, with a size of
[0011] The object center offset is used to more accurately locate the object, with a size of
[0012] The size of the bounding box is responsible for estimating the height and width of the target box at each position, with a size of
[0013] 5) In the Re-ID branch, use the fused feature map of the first frame as the input of the spatio-temporal interaction attention mechanism, and successively pass through the channel interaction attention mechanism and the spatial attention mechanism;
[0014] 6) Input into the channel interaction attention mechanism, and capture more effective Re-ID features by the interactive fusion of effective information between different channels. The calculation method is as follows:
[0015]
[0016] In formula (1), first obtain the aggregated feature through global average pooling AvgPool. Secondly, use grouped convolution d k×k to determine the coverage range of the interaction, so as to obtain the channel weights, and then use the σ activation function Sigmoid to constrain the weight values to between (0, 1). Finally, represents element-wise multiplication; where the relationship between the size of the grouped convolution k and the channel C is as follows:
[0017] C = φ(k) (2)
[0018] According to the principle of analogy in formula (2), there is a proportional relationship between the one-dimensional convolution size k and the number of channels C. Therefore, there may be a mapping φ between k and C; since the number of channels C is usually set to a power of 2, φ(k) can be transformed into:
[0019] C = φ(k) = 2 (γ*k-b) (3)
[0020] In formula (3), the simplest linear mapping φ(k) = γ * k - b is used, where γ and b are custom constants; therefore, the adaptive one-dimensional convolution size can be calculated according to the number of channels C:
[0021]
[0022] In formula (4), |t odd represents the odd number closest to t; subsequently, after mapping ψ, the high-dimensional channels have longer-range interactions, while the low-dimensional channels have shorter-range interactions by using non-linear mapping;
[0023] 7) The feature map obtained through the channel interaction attention mechanism is input into the spatial attention mechanism to obtain the information parts at different spatial positions of the image, and the calculation method is as follows:
[0024]
[0025] where AvgPool(·) and MaxPool(·) respectively represent average pooling and max pooling operations, and two pooling operations are used to aggregate the channel information of a feature map to generate two two-dimensional maps and f 7×7 represents a convolution operation with a 7×7 convolution kernel, σ is the activation function Sigmoid, represents element-wise multiplication;
[0026] 8) For the output of the spatio-temporal interaction attention mechanism use 1×1 convolution to obtain the person re-identification feature so as to extract the 128-dimensional features of each object centered at (x, y);
[0027] 9) Use Kalman filtering to perform data association and prediction on the bounding box obtained in step 4) and the Re-ID feature obtained in step 8), and obtain and save the initial trajectory of the first frame and the predicted trajectory position of the second frame;
[0028] 10) Input the second-frame image into the multi-layer fusion feature extraction network, repeat steps 3) to 8), obtain the bounding box and Re-ID feature of the second frame, and use Kalman filtering to perform data association on the predicted trajectory position of the second frame and the bounding box of the second frame;
[0029] 11) Use the Hungarian algorithm to match the association results to obtain the final tracking result.
[0030] The present invention introduces a new cross-channel collaboration network in the Re-ID subtask of multi-object tracking to optimize and enhance high-level features, alleviate the competition problem of sharing features between detection and Re-ID tasks within a single network, improve the collaboration ability of detection and Re-ID subtasks in multi-object tracking, and have higher accuracy and stronger stability for pedestrian occlusion and multi-object tracking calculations in strong light areas; by using the interactive fusion of effective information between different channels, optimize and enhance high-level features, obtain more robust global features, and significantly improve the accuracy and robustness of multi-object pedestrian tracking in occlusion and strong light areas. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is the 62nd frame image in the MOT17-03-DPM image sequence of the present invention example;
[0032] Figure 2 is the 63rd frame image in the MOT17-03-DPM image sequence of the present invention example;
[0033] Figure 3 is the structural diagram of the multi-layer fusion feature extraction network of the present invention example;
[0034] Figure 4 is the structural diagram of the multi-object pedestrian tracking calculation network based on the spatio-temporal interaction attention mechanism of the present invention example;
[0035] Figure 5 is the calculation graph of the spatio-temporal interaction attention mechanism in pedestrian re-identification of the present invention example;
[0036] Figure 6 is the calculation graph of the interactive channel attention in the spatio-temporal interaction attention mechanism of the present invention example;
[0037] Figure 7 is the calculation graph of the spatial attention in the spatio-temporal interaction attention mechanism of the present invention example;
[0038] Figure 8 is the tracking result of the MOT17-03-DPM image sequence obtained by the calculation of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0039] Please refer to Figures 1 - 8 , the present invention provides a multi-object pedestrian tracking calculation method based on the spatio-temporal interaction attention mechanism, and uses the MOT17-03-DPM sequence image for experimental illustration:
[0040] 1) Input Figure 1 and Figure 2 are two consecutive frame images of the MOT17-03-DPM image sequence in the MOT Challenge dataset; where: Figure 1 is the first frame image,Figure 2 is the second frame image with a resolution of 1920×1080;
[0041] 2) As shown in Figure 3 input the preprocessed data into the multi-layer fusion feature extraction network to obtain a fused feature map with a 4-fold downsampling, and the resolution of the feature map is 272×152; Figure 1
[0042] 3) As shown in Figure 4 input the fused feature map of the first frame into the detection branch and the Re-ID branch simultaneously for object detection and Re-ID feature extraction respectively;
[0043] 4) In the detection branch, the input fused feature map generates a heat map, an object center offset, and the size of the bounding box through 3×3 convolution and 1×1 convolution respectively:
[0044] The heat map is responsible for estimating the position of the object center, with a size of
[0045] The purpose of the object center offset is to more accurately locate the object, with a size of
[0046] The size of the bounding box is responsible for estimating the height and width of the target box at each position, with a size of
[0047] In the above, C is the number of channels of the fused feature map, and H and W are the height and width of the fused feature map respectively;
[0048] 5) As shown in Figure 5 in the Re-ID branch, use the fused feature map of the first frame as the input of the spatio-temporal interaction attention mechanism, and successively pass through the channel interaction attention mechanism and the spatial attention mechanism;
[0049] 6) As shown in Figure 6 input into the channel interaction attention mechanism, and capture more effective Re-ID features through the interactive fusion of effective information between different channels. The calculation method is as follows:
[0050]
[0051] In Equation (1), first obtain the aggregated feature by global average pooling AvgPool, and then use the adaptive group convolution d to determine the coverage range of the interaction (i.e., the one-dimensional convolution size k) to obtain the channel weights, and then use the σ activation function Sigmoid to constrain the weight values to the range of (0,1). Finally k×k Denotes element-wise multiplication; where the size of the group convolution k has the following relationship with the number of channels C:
[0052] C = φ(k) (2)
[0053] According to the principle of analogy in Equation (2), there is a proportional relationship between the size k of the one-dimensional convolution and the number of channels C. Therefore, there may be a mapping φ between k and C; since the number of channels C is usually set to a power of 2, φ(k) can be transformed into:
[0054] C = φ(k) = 2 (γ*k-b) (3)
[0055] In Equation (3), the simplest linear mapping φ(k) = γ * k - b is used, where γ and b are 2 and 1 in this experiment; therefore, the adaptive one-dimensional convolution size can be calculated according to the number of channels C:
[0056]
[0057] In Equation (4), |t odd Denotes the odd number closest to t; subsequently, after the mapping ψ, the high-dimensional channels have a longer-range interaction, while the low-dimensional channels have a shorter-range interaction by using a non-linear mapping.
[0058] 7) As Figure 7 shown, the feature map obtained through the channel interaction attention mechanism is input into the spatial attention mechanism to obtain the information part at different spatial positions of the image. The calculation method is as follows:
[0059]
[0060] In the formula, AvgPool(·) and MaxPool(·) represent the average pooling and max pooling operations respectively. Two pooling operations are used to aggregate the channel information of a feature map to generate two two-dimensional maps and f 7×7 represents the convolution operation with a 7×7 convolution kernel, σ is the activation function Sigmoid, denotes element-wise multiplication;
[0061] 8) For the output of the spatio-temporal interaction attention mechanism Use a 1×1 convolution to obtain the pedestrian re-identification feature Thus, extract the 128-dimensional features of each object centered at (x, y);
[0062] 9) Use Kalman filtering to perform data association and prediction on the bounding box obtained in step 4) and the Re-ID feature obtained in step 8), and obtain and save the initial trajectory of the first frame and the predicted trajectory position of the second frame;
[0063] 10) Figure 2 Input the multi-layer fusion feature extraction network, repeat steps 3) - 8) to obtain the bounding box and Re-ID features of the second frame, and use Kalman filtering to perform data association between the predicted trajectory position of the second frame and the bounding box of the second frame;
[0064] 11) Use the Hungarian algorithm to match the association results, so as to obtain the tracking results of the 63rd frame of the MOT17-03-DPM image sequence.
[0065] As Figure 8 shown, the method of the present invention has high accuracy and effectiveness for tracking targets under occlusion and strong illumination conditions, and has wide practicability in video surveillance, traffic control, autonomous driving, and human-computer interaction, etc. The fast and lightweight characteristics make it more reliable in practical applications.
[0066] The pedestrian multi-object tracking calculation method based on the spatio-temporal interaction attention mechanism of the present invention first inputs the first frame image of two consecutive frame images in the image sequence into the deep aggregation feature extraction network for feature extraction; secondly, inputs the fusion feature map of the first frame into the detection branch and the Re-ID branch at the same time, where the position detection of the tracking object is completed in the detection branch, and the pedestrian appearance features are extracted through the spatio-temporal interaction attention mechanism in the Re-ID branch; then uses Kalman filtering to associate and predict the detection results with the Re-ID features, and obtains and saves them as the initial trajectory and the predicted box of the second frame; subsequently, input the second frame and repeat the above operations of the first frame, and use Kalman filtering to associate the predicted box of the second frame with the bounding box; finally, use the Hungarian algorithm to match the association results, so as to obtain the final tracking results; the pedestrian multi-object tracking calculation method based on the spatio-temporal interaction attention mechanism of the present invention captures more effective parts in different channels and different spatial positions, and at the same time uses the interaction and fusion of effective information between different channels to optimize and enhance the high-level features, obtains more robust global features, and significantly improves the accuracy and robustness of pedestrian multi-object tracking in occluded and strongly illuminated areas.
[0067] The pedestrian multi-object tracking calculation method based on the spatio-temporal interaction attention mechanism of the present invention improves the accuracy and robustness of pedestrian multi-object tracking under occlusion and strong illumination conditions through the global optimization of detail information and the interaction of cross-channel features.
[0068] The pedestrian multi-object tracking calculation method based on the spatio-temporal interaction attention mechanism of the present invention introduces a new cross-channel collaboration network in the Re-ID sub-task of multi-object tracking to optimize and enhance high-level features, so as to alleviate the competition problem of sharing features between detection and Re-ID tasks within a single network, improve the collaboration ability of detection and Re-ID sub-tasks in multi-object tracking, and has higher accuracy and stronger stability for multi-object tracking calculation in pedestrian occlusion and strong light areas.
Claims
1. A pedestrian multi-object tracking calculation method based on a spatio-temporal interaction attention mechanism, the steps are as follows: 1) Input two consecutive frames of images in the image sequence; 2) Input the first frame of the image sequence into a multi-layer fusion feature extraction network to obtain a fusion feature map; 3) Input the extracted fusion feature map of the first frame into the detection branch and the Re-ID branch simultaneously, which are used to detect objects and extract Re-ID features respectively; 4) In the detection branch, the input fused feature map generates the heatmap, object center offset, and bounding box size through 3×3 convolution and 1×1 convolution respectively: The heatmap is responsible for estimating the position of the object center and has a size of The purpose of the object center offset is to locate the object more precisely, with a size of The size of the bounding box is responsible for estimating the height and width of the target box at each position, and the size is Where, C is the number of channels of the fusion feature map, and H and W are the height and width of the fusion feature map respectively; 5) In the Re-ID branch, use the fusion feature map of the first frame as the input of the spatio-temporal interaction attention mechanism, and successively pass through the channel interaction attention mechanism and the spatial attention mechanism; 6) Input into the channel interaction attention mechanism, and capture more effective Re-ID features by the interactive fusion of effective information between different channels. The calculation method is as follows: In formula (1), first, aggregate features are obtained through global average pooling AvgPool secondly, group convolution d k×k is used to determine the coverage range of the interaction, thereby obtaining channel weights, and then the σ activation function Sigmoid is used to constrain the weight values to the range of (0, 1). Finally represents element-wise multiplication; the relationship between the size k of the one-dimensional convolution kernel and the channel C is as follows: C = φ(k) (2) According to the principle of analogy in Equation (2), there is a proportional relationship between the size k of the one-dimensional convolution kernel and the number of channels C. Therefore, there may be a mapping φ between k and C; since the number of channels C is usually set to a power of 2, φ(k) can be transformed into: C = φ(k) = 2 (γ*k-b) (3) In Equation (3), the simplest linear mapping φ(k) = γ * k - b is used, where γ and b are custom constants; therefore, the adaptive one-dimensional convolution size can be calculated according to the number of channels C: In formula (4), |t odd represents the odd number closest to t; subsequently, after mapping ψ, the high-dimensional channels have longer-range interactions, while the low-dimensional channels have shorter-range interactions by using non-linear mapping; 7) Feature map obtained through the channel interaction attention mechanism Input it into the spatial attention mechanism to obtain the information parts at different spatial positions of the image. The calculation method is as follows: Where AvgPool(·) and MaxPool(·) represent average pooling and max pooling operations respectively, and two pooling operations are used to aggregate the channel information of a feature map to generate two two-dimensional maps and f 7×7 represents a convolution operation with a 7×7 convolution kernel, σ is the activation function Sigmoid, represents element-wise multiplication; 8) Output of the spatio-temporal interaction attention mechanism Use 1×1 convolution to obtain person re-identification features Thereby extracting 128-dimensional features of each object centered at (x, y); 9) Use the heat map, object center offset, and size of the bounding box generated in step 4) and the Re-ID features obtained in step 8) for data association and prediction using the Kalman filter, and obtain and save the initial trajectory of the first frame and the predicted trajectory position of the second frame; 10) Input the second frame of the image into the multi-layer fusion feature extraction network, repeat steps 3) to 8), obtain the bounding box and Re-ID features of the second frame, and use the Kalman filter to perform data association between the predicted trajectory position of the second frame and the bounding box of the second frame; 11) Use the Hungarian algorithm to match the association results to obtain the final tracking result.
Citation Information
Patent Citations
Unmanned aerial vehicle video multi-target tracking method based on attention feature fusion
CN113807187A
FairMOT multi-class tracking method based on improved attention mechanism
CN114241053A