Dense Pedestrian Multi-Object Tracking Method Based on Feature Fusion, Computer Device, and Storage Medium

By adopting feature fusion method in dense pedestrian multi-object tracking, combining basic feature extraction, re-identification feature extraction and displacement prediction, the detection and tracking separation problems of pedestrian multi-object tracking in dense scenarios are solved, and efficient and real-time trajectory formation is achieved.

CN116311353BActive Publication Date: 2025-05-30HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310087699.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-09
Publication Date
2025-05-30
Estimated Expiration
2043-02-09

AI Technical Summary

Technical Problem

When pedestrian multi-target tracking in dense scenarios, the problem of separation of detection and tracking leads to inconsistent model optimization direction, and the calculation cost of re-identification model is high, which limits the real-time nature of the algorithm.

Method used

The intensive pedestrian multi-objective tracking method based on feature fusion is adopted to form the final trajectory through basic feature extraction, re-identification feature extraction, displacement prediction and feature weighted fusion, which enhances the connection between detection and tracking tasks.

Benefits of technology

It improves the accuracy and real-timeness of intensive pedestrian multi-target tracking, and is suitable for multi-task tracking in dense target scenarios, improving the algorithm's generalization ability in dense scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116311353B_ABST
    Figure CN116311353B_ABST
Patent Text Reader

Abstract

A dense multi-object pedestrian tracking method based on feature fusion, a computer device, and a storage medium belong to the field of computer vision tracking technology, and solve the problem that there is no existing method for tracking pedestrians in dense scenes. The method of the present invention includes: First, a new target center point modeling method is designed to facilitate more accurate positioning of the target center point position; Second, a lightweight re-identification feature extraction network is proposed, and a similarity comparison method based on the essential matrix is used to obtain the target inter-frame displacement prediction; Then, a feature enhancement network based on a hybrid attention mechanism is designed to fuse the inter-frame information in the time dimension and the static information in the space dimension, enhancing the connection between the detection task and the tracking task; Finally, the detection result and the target displacement are integrated through a quadratic data association method to obtain the final trajectory. The present invention is applicable to multi-pedestrian tracking in dense target scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision tracking technology, and particularly to dense pedestrian multi-object tracking. Background Art

[0002] The research field of pedestrian multi-object tracking is currently divided into two paradigms: one is the paradigm of separating object detection and tracking; the other is the paradigm of jointly detecting and tracking objects. In recent years, the detection-based tracking method has been the mainstream method in the field of multi-object tracking. The detection-based tracking method advocates that: first, use the existing detection model to generate the detection results for each frame; then use an additional object re-identification model to extract the appearance features of each detection result or use a motion model to directly predict the inter-frame motion state of the object; finally, use a correlation matching algorithm to complete the data association step and obtain the complete tracking trajectory. The method of jointly detecting and tracking has rapidly emerged in this field due to its structural advancement. Its re-examination of the relationship between the detection model and the tracking model has extremely high practical value for jointly optimizing the two. It integrates the originally completely separated detection model and tracking model into the same framework by transforming or inserting part of the existing detection model into the tracking model.

[0003] Although the detection-based tracking method has been the mainstream method in the field of multi-object tracking, it has two main drawbacks: 1) The separation of the detection and tracking parts is not conducive to the joint optimization of the model, and there are often situations where the optimization directions of the two parts of the model are inconsistent, ultimately resulting in the overall model not being able to obtain the optimal result globally; 2) In order to provide an optimization basis for the data association step, the re-identification models adopted by these methods are often independent and require high computational costs, which greatly limits the real-time performance of the multi-object tracking algorithm. Compared with the detection-based multi-object tracking paradigm, the multi-object tracking algorithm of the joint detection and tracking paradigm has good prospects in both theoretical research and practical applications due to its advanced structural form and advantages in tracking speed.

[0004] In the multi-object tracking task, pedestrian targets are usually the center of attention in the video scene, which makes detecting and tracking them a basic problem that needs to be studied in the field of computer vision. In addition, compared with other visual targets, pedestrians, as typical non-rigid targets, are ideal samples for studying multi-object tracking problems. However, the complexity of this task increases with the increase in the number of pedestrians to be tracked, and it is still an open research field. As the situation of large-scale dense pedestrians becomes more and more common, due to the sudden increase in target density, the model not only faces challenges in object detection, but also the occurrence of identity conversion situations becomes more frequent during the tracking trajectory generation process. The vast majority of existing methods do not specifically focus on the pedestrian tracking problem in dense scenes. Therefore, when migrating these methods to such scenes, good generalization performance is often not obtained. Summary of the Invention

[0005] The object of the present invention is to solve the problem that there is no existing method for tracking pedestrians in dense scenarios, and a multi-object tracking method for dense pedestrians based on feature fusion, a computer device, and a storage medium are provided.

[0006] The present invention is realized through the following technical solutions. On the one hand, the present invention provides a multi-object tracking method for dense pedestrians based on feature fusion, and the method includes:

[0007] Step 1: Perform basic feature extraction processing on adjacent input video frames to obtain the basic features of each frame;

[0008] Step 2: Based on the basic features of adjacent frames obtained in Step 1, use a re-identification feature extraction network to extract re-identification features to obtain re-identification features of adjacent frames; according to the re-identification features of adjacent frames, use a cost matrix module to obtain an inter-frame displacement prediction matrix of the same target;

[0009] Step 3: Obtain the target detection information of the current frame according to the basic features of adjacent frames, specifically including:

[0010] Step 3.1: Subtract the corresponding elements of the basic features of adjacent frames obtained in Step 1 to obtain an inter-frame difference feature;

[0011] Step 3.2: Integrate the inter-frame difference feature and the displacement prediction matrix obtained in Step 2 according to dimensions as the input of a deformable convolution offset extraction unit, so as to obtain the offset prediction required by the deformable convolution network;

[0012] Step 3.3: Weight the features of the earlier frame in adjacent frames with the predicted heat map as the input of the deformable convolution network DCN, and the deformation displacement of the convolution kernel is determined by the offset prediction above, thereby obtaining new features of the previous frame different from the basic features;

[0013] Step 3.4: Weight and fuse the new features of the previous frame and the basic features of the current frame to obtain new features of the current frame, and use the new features of the current frame for classification and regression to obtain the target detection information of the current frame;

[0014] Step 4: Form a final trajectory according to the target detection information of the current frame, specifically including:

[0015] Step 4.1: The result of enhancing the adjacent frame features obtained through Step 3 passes through the classification and regression branches to obtain the category and position information of the target;

[0016] Step 4.2: Associate the identities of the same targets between frames according to the category and location information of the target, as well as the re-identification features and inter-frame displacement prediction obtained in Step 2.

[0017] Step 4.3: Form the final trajectory through the linear assignment algorithm.

[0018] Further, in Step 1, the DLA-34 feature extraction network structure is used to perform basic feature extraction processing on the input adjacent video frames. The method for obtaining the target center point includes:

[0019] Adopt a center point constraint, and the effective radius r in the case of the target center key point center As shown in the following formula:

[0020]

[0021] where W is the width of the input image, H is the height of the input image, and IoU threshold is the intersection over union threshold.

[0022] Further, in Step 2, the re-identification feature extraction network includes three types of network modules, namely the convolutional layer conv, the batch normalization layer BN, and the non-linear activation layer SiLU.

[0023] Except for using 1×1 convolutional kernels in the first convolutional layer and the last convolutional layer, the remaining convolutional layers all use 3×3 convolutional kernels.

[0024] Further, according to the re-identification features of the adjacent frames, using the essential matrix module to obtain the inter-frame displacement prediction of the same target specifically includes:

[0025] Perform a correlation operation on the current frame part E t in the multi-frame re-identification embedding model extracted and the previous frame part E t-τ to obtain the essential matrix;

[0026] After obtaining the essential matrix of the inter-frame similarity metric, predict the motion direction and motion displacement of the target between frames;

[0027] Multiply the horizontal and vertical displacement templates M i,j and V i,j by the horizontal difference probability representation and the vertical difference probability representation respectively, to obtain the displacement change amount of the current frame relative to the previous frame.

[0028] Further, the deformable convolution offset extraction unit is a convolutional neural network based on a hybrid attention mechanism, specifically including: a convolutional layer conv, a batch normalization layer BN, non-linear activation layers ReLU and SiLU, a max pooling layer, an average pooling layer, a fully connected layer FC, a basic residual block, a spatial attention mechanism network, and a channel attention mechanism network;

[0029] The basic residual block is used for further feature extraction.

[0030] Further, the method of weighting the features of the earlier frame among adjacent frames with the predicted heatmap specifically includes:

[0031] For the feature map of the previous frame, it is not directly operated on in the composition of the deformable convolution input. Instead, its basic feature map is multiplied element by element with its heatmap. The formula is as follows:

[0032]

[0033] where, F p t-τ represents the basic feature map extracted by the backbone network layer of the (t - τ)-th frame, represents the heatmap result obtained by predicting through the detection model for the (t - τ)-th frame, only for the classification of pedestrians, represents the result of and being superimposed channel by channel and pixel by pixel. ⊙ represents the Hadamard product of matrices, and p = 1, 2,..., 64 represents the index values of each channel.

[0034] Further, the method of weighting and fusing the new features of the previous frame with the basic features of the current frame to obtain the new features of the current frame specifically includes:

[0035] Adding the integrated features of the previous frame and the basic features of the current frame through an adaptive weight matrix. The formula is as follows:

[0036]

[0037] where, represents the adaptive matrix of the current frame, w t-τ represents the adaptive matrix of the previous frame, and it satisfies the relationship T represents the number of previous frames used, ⊙ represents the Hadamard product of matrices, and the adaptive weight matrix is obtained by two sets of convolutional layers and the softmax function.

[0038] Further, step 4.2 specifically includes:

[0039] Step 4.2.1, Initialize multiple trajectory queues, which are divided into three categories: the tracked trajectory queue Ttracked , the track queue T of unmatched tracks in adjacent frames lost , the track queue T of ended tracks removed ; through two thresholds thresh low and thresh high , the detection results of the current frame are divided into two categories: high-confidence detection results and low-confidence detection results;

[0040] Step 4.2.2, perform the first data association, specifically including: for the cost matrix C IoU Use the Jonker-Volgenant linear assignment algorithm to obtain the set S of matching index pairs m , the set S of unmatched tracks um-track , the set S of unmatched detection results um-det ; for the set S of index pairs that can be matched m , which contains a tracked track element and a detection result element of the current frame; if the matched track belongs to the tracked track queue T tracked , directly add the detection result of the current frame to this track to become the subsequent tracked track; otherwise, the detected result will match the track in the track queue T of unmatched tracks in adjacent frames lost , then reactivate the unmatched track;

[0041] Step 4.2.3, perform the second data association, specifically including: for the low-confidence detection results, use exactly the same processing method as in the first association to obtain the tracked tracks and reactivated tracks;

[0042] Mark the tracks that still cannot be matched after the second data association as unmatched tracks in adjacent frames and summarize them into the corresponding queue T lost ; for the detection results unmatched in the first data association Calculate the position similarity with the unactivated tracks, and use the Jonker-Volgenant linear assignment algorithm to obtain the matching pair indexes and end the unactivated tracks that are not matched; for the detection results that still exist in the high-confidence detection results but are not matched, generate them as the starting points of new tracks; update the track status and check whether there are tracks in the track queue T of unmatched tracks in adjacent frames lost that exceed the association length threshold, and end these tracks.

[0043] In a second aspect, the present invention provides a computer device, including a memory and a processor, where a computer program is stored in the memory, and when the processor runs the computer program stored in the memory, it executes the steps of a multi-object tracking method for dense pedestrians based on feature fusion as described above.

[0044] In a third aspect, the present invention provides a computer-readable storage medium storing multiple computer instructions for causing a computer to execute a multi-object tracking method for dense pedestrians based on feature fusion as described above.

[0045] Advantages of the present invention:

[0046] In the present invention, an algorithm for dense pedestrian detection and tracking based on a joint detection and tracking paradigm is proposed. The algorithm uses re-identification features to construct a cost matrix to predict the inter-frame displacement of objects. And a hybrid attention mechanism is adopted to achieve inter-frame feature fusion, and displacement information is used for detection, enhancing the connection between the detection task and the tracking task. There is a great improvement in the migration from a conventional target density scenario to a dense target scenario. The visualization results can be seen in Figure 7 .

[0047] In the present invention, first, a new method for modeling the target center point is designed, which is beneficial to more accurately locate the position of the target center point;

[0048] Secondly, a lightweight re-identification feature extraction network is proposed, and a similarity comparison method based on the cost matrix is used to obtain the prediction of the inter-frame displacement of the target;

[0049] Then, a feature enhancement network based on a hybrid attention mechanism is designed to fuse the inter-frame information in the time dimension and the static information in the space dimension, enhancing the connection between the detection task and the tracking task;

[0050] Finally, the detection results and the target displacement are integrated through a secondary data association method to obtain the final trajectory.

[0051] The present invention is applicable to multi-pedestrian tracking in a dense target scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] To more clearly illustrate the technical solutions of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0053] Figure 1 is a schematic diagram of the method network flow of the present invention;

[0054] Figure 2 are four different constraint cases of the effective radius;

[0055] Figure 3 is a schematic diagram of the specific network structure corresponding to step 2 in the embodiment of the present invention;

[0056] Figure 4Schematic diagram of the results of a lightweight re-identification feature extraction module of the present invention;

[0057] Figure 5 Schematic diagram of the process of step three in the embodiment of the present invention;

[0058] Figure 6 Network structure diagram of the deformable convolution offset extraction unit of the present invention;

[0059] Figure 7 Visualization result of the present invention. Detailed implementation manners

[0060] The following details the implementation manners of the present invention. Examples of the implementation manners are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The implementation manners described below by referring to the drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.

[0061] Embodiment 1. A dense pedestrian multi-object tracking method based on feature fusion, the method comprising:

[0062] Step 1: Perform basic feature extraction processing on adjacent input video frames to obtain the basic features of each frame;

[0063] Step 2: Based on the adjacent frame basic features obtained in Step 1, use a re-identification feature extraction network to perform re-identification feature extraction to obtain adjacent frame re-identification features; according to the adjacent frame re-identification features, use the present quantity matrix module to obtain an inter-frame displacement prediction matrix of the same target;

[0064] Step 3: Obtain the target detection information of the current frame according to the adjacent frame basic features, specifically including:

[0065] Step 3.1: Subtract the corresponding positions of the adjacent frame basic features obtained in Step 1 element by element to obtain an inter-frame difference feature;

[0066] Step 3.2: Integrate the inter-frame difference feature and the displacement prediction matrix obtained in Step 2 according to dimensions as the input of the deformable convolution offset extraction unit, so as to obtain the offset prediction required by the deformable convolution network;

[0067] Step 3.3: Weight the features of the earlier frame in the adjacent frames (i.e., the previous frame features) with the predicted heat map as the input of the deformable convolution network DCN. The deformation displacement of the convolution kernel is determined by the offset prediction above, and thus new features of the previous frame different from the basic features are obtained;

[0068] It should be noted that the classification and regression part after obtaining the basic features will obtain a predicted heat map.

[0069] Step 3.4: Weightedly fuse the new features of the previous frame and the basic features of the current frame to obtain the new features of the current frame, and use the new features of the current frame for classification and regression to obtain the target detection information of the current frame;

[0070] Step 4: Form a final trajectory according to the target detection information of the current frame, specifically including:

[0071] Step 4.1: The result of enhanced adjacent frame features obtained through Step 3 passes through the classification and regression branches to obtain the category and position information of the target;

[0072] Step 4.2: According to the category and position information of the target, the re-identification features of adjacent frames obtained in Step 2, and the inter-frame displacement prediction, associate the identities of the same target between frames;

[0073] Step 4.3: Form a final trajectory through the linear assignment algorithm.

[0074] In this embodiment, first, accurately locate the position of the target center point;

[0075] Secondly, a lightweight re-identification feature extraction network is proposed, and the target inter-frame displacement prediction is obtained by using the similarity comparison method based on the essential matrix;

[0076] Then, a feature enhancement network based on a hybrid attention mechanism is designed to fuse the inter-frame information in the time dimension and the static information in the space dimension, enhancing the connection between the detection task and the tracking task;

[0077] Finally, the detection result and the target displacement are integrated through the method of quadratic data association to obtain the final trajectory.

[0078] Embodiment 2: This embodiment further limits a dense pedestrian multi-target tracking method based on feature fusion described in Embodiment 1. In this embodiment, Step 1 is further limited, specifically including:

[0079] Step 1 uses the DLA-34 feature extraction network structure to perform basic feature extraction processing on the input adjacent video frames. The method for obtaining the target center point includes:

[0080] Adopt a center point constraint, and the effective radius r in the case of the target center key point center As shown in the following formula:

[0081]

[0082] Among them, W is the width of the input image, H is the height of the input image, and IoU threshold is the intersection over union threshold.

[0083] In this embodiment, the basic feature extraction is used for the positioning and classification of the target center point. Therefore, this part not only depends on the network structure design, but also needs to model the representation of the target center point. The method proposed in this embodiment in this part is a more novel method for modeling the target center point, which is beneficial to more accurately locate the position of the target center point.

[0084] Embodiment 3 is a further limitation of a dense pedestrian multi-object tracking method based on feature fusion described in Embodiment 1. In this embodiment, the re-identification feature extraction network in step 2 is further limited, specifically including:

[0085] In step 2, the re-identification feature extraction network includes three types of network modules, namely the convolutional layer conv, the batch normalization layer BN, and the non-linear activation layer SiLU;

[0086] Except for the first convolutional layer and the last convolutional layer using 1×1 convolutional kernels, the remaining convolutional layers all use 3×3 convolutional kernels.

[0087] It should be noted that the module cascade structure here follows the order of convolutional layer - batch normalization layer - non-linear activation layer and is stacked. This cyclic module is recommended to be controlled within 4 groups.

[0088] This embodiment designs a lightweight re-identification feature extraction module to be connected after the backbone network layer, converting the original inter-class appearance features into high-dimensional intra-class discriminative features.

[0089] Embodiment 4 is a further limitation of a dense pedestrian multi-object tracking method based on feature fusion described in Embodiment 1. In this embodiment, the obtaining of the inter-frame displacement prediction of the same target by using the cost matrix module according to the adjacent frame re-identification features is further limited, specifically including:

[0090] The current frame part E in the extracted multi-frame re-identification embedding model t is correlated with the previous frame part E t-τ to obtain the cost matrix;

[0091] After obtaining the cost matrix of the inter-frame similarity metric, the ultimate goal is to predict the movement direction and movement displacement of the target between frames;

[0092] Multiply the horizontal and vertical displacement templates M i,j and V i,j by the horizontal difference probability representation and the vertical difference probability representation respectively to obtain the displacement change amount of the current frame relative to the previous frame.

[0093] In this embodiment, multiple methods can be adopted for the calculation of the cost matrix. In this embodiment, the vector inner product distance, the vector cosine distance with feature channel normalization, and the vector Euclidean distance (vector L2 norm distance) are respectively adopted to calculate the cost matrix, so as to obtain the inter-frame embedded feature similarity.

[0094] Embodiment 5 is a further limitation on a dense pedestrian multi-object tracking method described in Embodiment 1. In this embodiment, the deformable convolution offset extraction unit is further limited, specifically including:

[0095] The deformable convolution offset extraction unit is a convolutional neural network based on a hybrid attention mechanism, specifically including: a convolutional layer conv, a batch normalization layer BN, non-linear activation layers ReLU and SiLU, a max pooling layer, an average pooling layer, a fully connected layer FC, a basic residual block, a spatial attention mechanism network, and a channel attention mechanism network;

[0096] The basic residual block is used for further feature extraction.

[0097] In this embodiment, the inter-frame difference feature and the displacement prediction matrix output in Step 2 are integrated according to the dimension as the input of the deformable convolution offset extraction unit, so as to obtain the offset prediction required by the deformable convolution network.

[0098] Embodiment 6 is a further limitation on a dense pedestrian multi-object tracking method described in Embodiment 1. In this embodiment, the weighting of the feature of the earlier frame in adjacent frames with the predicted heat map is further limited, specifically including:

[0099] For the feature map of the previous frame, it is not directly operated on in the composition of the deformable convolution input, but its basic feature map is multiplied by each element of its heat map one by one, and the formula is as follows:

[0100]

[0101] Among them, represents the basic feature map extracted by the backbone network layer of the (t - τ)-th frame, represents the heat map result obtained by predicting through the detection model for the (t - τ)-th frame, only for the classification of pedestrians, represents the result of and being superimposed channel by channel and pixel by pixel. ⊙ represents the Hadamard product of matrices, and p = 1, 2,..., 64 represents the index values of each channel.

[0102] In this embodiment, after obtaining the convolutional kernel element offset, it is used as part of the input of the deformable convolutional network (DCN), and the fused feature of the base feature map and the predicted heat map of the previous frame is used as the convolutional object input of the DCN. For the feature map of the previous frame, it is not directly operated on in the input composition of the deformable convolution. Instead, each element of its base feature map is multiplied by the corresponding element of its heat map. The operation method for the (t - τ)-th frame is equivalent to weighting the predicted result of the previous frame for the target to the base feature map in the form of a two-dimensional Gaussian distribution, which can be understood as an attention mechanism for the target.

[0103] Embodiment 7. This embodiment further limits a dense pedestrian multi-object tracking method based on feature fusion described in Embodiment 1. In this embodiment, further limitations are made on the operation of weighting and fusing the new feature of the previous frame and the base feature of the current frame to obtain the new feature of the current frame, which specifically includes:

[0104] Adding the integrated feature of the previous frame and the base feature of the current frame through an adaptive weight matrix, and its formula is as follows:

[0105]

[0106] Among them, represents the adaptive matrix of the current frame, and w t-τ represents the adaptive matrix of the previous frame, and they satisfy the relationship T represents the number of previous frames used, and ⊙ represents the Hadamard product of matrices. The adaptive weight matrix is obtained by two convolutional layers and a softmax function.

[0107] This embodiment belongs to the feature enhancement part, which fuses the inter-frame information in the time dimension and the static information in the space dimension, enhancing the connection between the detection task and the tracking task.

[0108] Embodiment 8. This embodiment further limits a dense pedestrian multi-object tracking method based on feature fusion described in Embodiment 1. In this embodiment, further limitations are made on step 4.2, which specifically includes:

[0109] Step 4.2.1: Initialize multiple trajectory queues, which are divided into three categories: the tracked trajectory queue T tracked , the trajectory queue T lost of un-matched trajectories in adjacent frames, and the ended trajectory queue T removed ; Divide the detection results of the current frame into two categories through two thresholds thresh low and thresh high : high-confidence detection results and low-confidence detection results;

[0110] Step 4.2.2: Perform the first data association, specifically including: For the cost matrix C IoU Use the Jonker-Volgenant linear assignment algorithm to obtain a set S of matching index pairs m , an unmatched track set S um-track , and an unmatched detection result set S um-det ; For the set S of index pairs that can be matched m , which contains a tracked track element and a current frame detection result element. If the matched track belongs to the tracked track queue T tracked , directly add the current frame detection result to this track to become the subsequent tracked track. Otherwise, the detected result will match a track in the unmatched track queue T lost of adjacent frames, and then reactivate this unmatched track;

[0111] Step 4.2.3: Perform the second data association, specifically including: For low-confidence detection results, use exactly the same processing method as in the first association to obtain tracked tracks and reactivated tracks;

[0112] Mark the tracks that still cannot be matched after the second data association as unmatched tracks in adjacent frames and summarize them into the corresponding queue T lost ; For the detection results unmatched in the first data association Calculate the position similarity with the unactivated tracks, and use the Jonker-Volgenant linear assignment algorithm to obtain the matching pair indexes and end the unactivated tracks that are unmatched; For the detection results that still exist in the high-confidence detection results but are not matched, generate them as the starting points of new tracks; Update the track status and check whether there are tracks in the unmatched track queue T lost of adjacent frames that exceed the association length threshold, and end these tracks.

[0113] In this embodiment, comprehensively considering the category and position information of the target, the re-identification features of adjacent frames obtained in Step 2, and the predicted inter-frame displacement output, associate the identities of the same targets between frames, that is, integrate the detection results with the target displacement through the method of secondary data association to obtain the final track, so as to improve the accuracy of multi-target tracking of dense pedestrians.

[0114] Example:

[0115] The algorithm network flow of this embodiment is as Figure 1As shown in the figure, it can be composed of 4 main parts, namely the basic feature extraction part, the inter-frame displacement prediction part, the feature enhancement part, and the data association part. In this network, the input consists of multiple frames of video images, and the output is the trajectory information of the target of interest. In addition, there is partial input-output interaction between the basic feature extraction and the feature enhancement part, and the inter-frame displacement prediction part also needs to output some content to the data association part. Therefore, Figure 1 dashed arrows are used to indicate this meaning.

[0116] Step 1: Perform basic feature extraction processing on adjacent input video frames, and output the basic features of each frame, where the interval between adjacent frames does not exceed 5 frames.

[0117] In this embodiment, the same feature extraction network structure DLA-34 as in [1] is adopted, but the method for obtaining the target center point designed in this embodiment is used. It should be noted that basic feature extraction is used for the positioning and classification of the target center point. Therefore, this part not only depends on the network structure design, but also needs to model the representation of the target center point. The method proposed in this embodiment in this part is a more novel method for modeling the target center point. The following is a comparison with the method in [1], and the specific content of this method will be elaborated.

[0118] The method for modeling the target center point can be summarized as: mapping the input image The width of the input image is W, the height is H, and the number of channels is 3, representing an RGB image, into a heat map of key points where R represents the scaling ratio of the heat map relative to the original image size, and C represents the number of categories involved in the target. During the training process, after the target center point is mapped onto the heat map, it is a probability representation following a Gaussian distribution.

[0119] Therefore, in order to ensure that the coordinate of the target center point on the heat map does not deviate too much from the coordinate of the target center point in the annotation set, certain constraint conditions need to be added after the mapping to limit the scattered positions of the target center point set on the heat map. In a two-dimensional space, for the positional correlation between target sets, the intersection over union is usually used for measurement. For the above reasons, for such a mapping, the model often requires continuous smooth characteristics, assigning higher weight coefficients to the parts closer to the exact position of the target center point, and lower weight coefficients to the parts farther away from the target center point. Therefore, here the ground truth information is transformed into a probability representation using a two-dimensional Gaussian kernel function and mapped into a ground truth heat map Y∈[0,1] (W / R)×(H / R)×C . The functional form is shown in Equation (1-1):

[0120]

[0121] where, is the position coordinate of the target center point on the heat map, is the true coordinate of the target center point in the annotation set, and σ k is the standard deviation of the Gaussian kernel function for mapping. According to the relevant properties of the variable interval estimation of the two-dimensional Gaussian distribution, the confidence interval of the abscissa x is taken as (x - 3σ k , x + 3σ k ), and the confidence interval of the ordinate y is taken as (y - 3σ k , y + 3σ k ), which can ensure that the confidence level of the internal samples reaches 99.7%.

[0122] Based on this, in this embodiment, 3σ k is defined as the effective radius of the distribution of the target center point on the heat map under the two-dimensional Gaussian mapping. Based on this, a constraint on the effective radius smaller than the intersection over union threshold will be generated, as shown in Equation (1-2):

[0123]

[0124] Among them, the intersection over union is obtained from the bounding box determined by the target center point on the heat map and the target bounding box in the annotation set. S inter represents the area of the intersection part of the two, and S union represents the area of the union part of the two.

[0125] This embodiment proposes a heat map key point generation method different from that in [1]. In [1], this constraint relationship is established in three cases, as shown in Figure 2 (a), 2(b), and 2(c) respectively. Based on this, the effective radius is obtained, and the specific content is as shown in Equation (1-3):

[0126]

[0127] For the above three cases, this embodiment designs a generation method to simplify these three cases into one case. In Figure 2 (d), this embodiment simplifies the case of two corner point constraints into a case of one center point constraint. Based on this, the following formula can be obtained:

[0128] S 1 = (W - rsinθ)·(H - rcosθ)(1-4)

[0129] S 2 = W·H - S 1 (1-5)

[0130]

[0131] Among them, in Figure 2 (d), S 1It represents the area size of the intersection between the bounding box determined by the target center point on the heat map and the target bounding box in the annotation set, S 2 It represents the difference obtained by subtracting the intersection area S1 from the area of the target bounding box in the annotation set, and the intersection over union threshold is IoU threshold If and only if At this time, the equal sign situation can be obtained in Equation (1-6), so the effective radius r of the target center key point can be obtained center As shown in Equation (1-7):

[0132]

[0133] During the process of mapping the ground truth of the annotation set to the heat map plane, in this embodiment, an implicit constraint on the target size is added to it, and the prior of the target size ratio is added to the training process in advance.

[0134] Step 2: Based on the adjacent frame basic features output in Step 1, use the re-identification feature extraction network to extract re-identification features. The obtained adjacent frame re-identification features (i.e., re-identification embedded features) are used as the input of the volume matrix module to output the inter-frame displacement prediction of the same target. The specific network structure is as Figure 3 shown.

[0135] During the tracking process, since the target will be given a new identity label due to being occluded or having a drastic appearance change, etc., if it is directly used as the starting point of a new trajectory, a large number of trajectory fragments and identity conversion phenomena will occur. Using the re-identification embedded features not only benefits the distinction between the same type of targets, but also establishes a feature bank for the targets, which can provide a basis for trajectory continuation when the occluded target reappears. At the same time, since the feature extraction performed by the backbone network is mainly used for subsequent inter-class discrimination, that is, separating pedestrian targets from the background, while the re-identification network mainly performs intra-class discrimination of the targets to distinguish different individuals under the same type of targets. The re-identification feature extraction method proposed in this embodiment is different from the traditional method of obtaining local features for analysis and comparison. It constructs a high-dimensional embedded model to describe the differences between different individuals under the same category.

[0136] In this embodiment, a lightweight re-identification feature extraction module is designed to be connected after the backbone network layer, converting the original inter-class appearance features into high-dimensional intra-class discrimination features. This network consists of 3 types of network modules, namely the convolutional layer conv, the batch normalization layer BN, and the non-linear activation layer SiLU. Figure 4 (a) is a conventional implementation structure Figure 4(b) is the implementation structure of this embodiment. Here, the module cascade structure follows the sequential superposition of the convolutional layer - batch normalization layer - non-linear activation layer. This loop module is recommended to be controlled within 4 groups. In addition, except for the first layer convolutional layer and the last layer convolutional layer using 1×1 convolutional kernels, the remaining convolutional layers all use 3×3 convolutional kernels. Its mapping can be expressed by the following formula:

[0137] E t =σ(F t )(2-1)

[0138] Among them, (W, H) represents the resolution of the input image after affine transformation, F t represents the features obtained after the t-th frame image is extracted by the backbone network layer, E t represents the re-identification embedded feature of the t-th frame image, and σ(·) represents Figure 3 the mapping corresponding to the re-identification embedded model extraction network in

[0139] After that, the current frame part E t in the extracted multi-frame re-identification embedded model is t-τ correlated with the previous frame part E

[0140]

[0141]

[0142]

[0143] Among them, C i,j,k,l represents the cost matrix, represents the embedded multi-dimensional matrix of the t-th frame image, and (i, j) respectively represent the abscissa index and ordinate index of the matrix element, represents the embedded multi-dimensional matrix of the (t - τ)-th frame image, and (k, l) respectively represent the abscissa index and ordinate index of the matrix element, and (·) Τ represents the matrix transpose operation. Equation (2-2) corresponds to the method of calculating the cost matrix using the vector inner product; Equation (2-3) corresponds to the method of calculating the cost matrix using the vector cosine distance with feature channel normalization, where Norm L2 (·) represents the L2 norm calculation in the feature channel direction; Equation (2-4) corresponds to the method of calculating the cost matrix using the vector Euclidean distance, where (·) 2Denotes the square operation at the matrix element level, not the multiplication operation of the matrix itself.

[0144] After obtaining the essential matrix of the inter-frame similarity metric, the ultimate goal is to predict the motion direction and displacement of the target between frames. This part can be divided into three steps: 1) Perform max pooling on the essential matrix in the height and width directions to find the maximum horizontal difference of each pixel point in the current frame relative to the previous frame and the maximum vertical difference where and represent the similarity degree between the position (i, j) in the t-th frame image and all pixel positions in the (t - τ)-th frame image. For example: represents the similarity degree between the target appearing at the position (i, j) in the t-th frame image and all pixel positions in the column where the position (*, l) is located in the (t - τ)-th frame image; 2) For the matrix C W of the maximum horizontal difference and the matrix C H of the maximum vertical difference after pooling, use the softmax function for normalization to map the original correlation values to probabilities represented by [0, 1]; 3) After obtaining the similarity probability representations of each point in the current frame with respect to multiple previous frames, it is also necessary to convert these probability values into actual inter-frame displacement information. According to the position relationship of different pixel positions in the current frame relative to the previous frame, the following horizontal and vertical displacement templates can be designed, and their calculation methods are shown in Equation (2-5). Here, taking 1 / 8 of the input image size as an example, when the input image resolution is 512×512, the size of the obtained feature map is 64×64:

[0145]

[0146] where M i,j and V i,j represent the horizontal and vertical displacement templates respectively.

[0147] Finally, multiply the horizontal and vertical displacement templates M i,j and V i,j by the horizontal difference probability representation and the vertical difference probability representation respectively, to obtain the displacement change amount of the current frame relative to the previous frame. These probability representations of displacement changes can be used as the association basis in the data association step and as the position attention information for feature fusion. The above process can be represented by Equation (2-6):

[0148]

[0149] where O i,jIt represents the displacement prediction matrix of the target at the image coordinates (i, j) of the t-th frame relative to all positions on the image of the (t - τ)-th frame in the horizontal and vertical directions.

[0150] Step 3: Subtract the corresponding positions of the adjacent-frame basic features output in Step 1 element by element to obtain the inter-frame difference features; integrate the inter-frame difference features and the displacement prediction matrix output in Step 2 according to the dimensions as the input of the deformable convolution offset extraction unit, so as to obtain the offset prediction required by the deformable convolution network; weight the features of the earlier frame in the adjacent frames with the predicted heatmap as the input of the deformable convolution network DCN, and the deformation displacement of the convolution kernel is determined by the offset prediction above, thereby obtaining the new features of the previous frame different from the basic features; weight and fuse the new features of the previous frame and the basic features of the current frame to obtain the new features of the current frame, and use this feature for classification and regression to obtain the object detection information of the current frame. The specific process of this step is as Figure 5 shown.

[0151] In a 3×3 convolution kernel, there are a total of 9 convolution elements. Therefore, in deformable convolution, it is necessary to determine 8 horizontal offsets of elements and 8 vertical offsets of elements. Based on this, it is necessary to use the integrated inter-frame difference features and the displacement prediction matrix output in Step 2 as the input and map them to 16 offset outputs.

[0152] In this embodiment, a convolutional neural network based on a hybrid attention mechanism is designed to complete the training and learning of this mapping process. The specific content of this network structure is as Figure 6 shown. The input of this network is the integrated spliced features, and the output is the convolution kernel offset. This network structure includes multiple layers such as a convolutional layer conv, a batch normalization layer BN, a non-linear activation layer ReLU and SiLU, a max pooling layer, an average pooling layer, and a fully connected layer FC, and uses a basic residual block for further feature extraction. Figure 6 The part marked by the dashed box in

[0153] is the part where this network structure adopts the hybrid attention mechanism. Among them, the red dashed box represents the spatial attention mechanism network structure, and the blue dashed box represents the channel attention mechanism network structure.

[0154]

[0155] where, denotes the basic feature map extracted by the backbone network layer for the (t - τ)-th frame, denotes the heatmap result obtained by predicting the (t - τ)-th frame through the detection model. Under the problem of this embodiment, it only targets the classification of pedestrians, denotes the result of and being superimposed channel by channel and pixel by pixel. In formula (3 - 1), ⊙ represents the Hadamard product of matrices, and p = 1, 2,..., 64 represents the index values of each channel.

[0156] In formula (3 - 1), the operation method for the (t - τ)-th frame is equivalent to weighting the basic feature map with the prediction results of previous frames for the target in the form of a two-dimensional Gaussian distribution, which can be understood as an attention mechanism for the target.

[0157] The feature enhancement part adds the integrated features of previous frames and the basic features of the current frame through an adaptive weight matrix, and the specific form can be expressed by formula (3 - 2):

[0158]

[0159] where, denotes the integrated features of previous frames, denotes the basic features of the current frame, denotes the new features of the current frame, denotes the adaptive weight matrix of the current frame, w t-τ denotes the adaptive weight matrix of previous frames, which satisfies the relationship T represents the number of previous frames used, and ⊙ represents the Hadamard product of matrices. Among them, the adaptive weight matrix is obtained by two groups of convolutional layers and the softmax function.

[0160] Step 4: The result of adjacent frame feature enhancement obtained in Step 3 passes through the classification and regression branches to obtain the category and location information of the target; comprehensively considering the category and location information of the target, the adjacent frame re-identification features obtained in Step 2, and the predicted inter-frame displacement output, associate the identities of the same target between frames (i.e., data association); form the final trajectory through the linear assignment algorithm.

[0161] The specific data association method designed in this embodiment is as follows:

[0162] First, initialize multiple trajectory queues, mainly divided into three categories: the tracked trajectory queue T tracked , the trajectory queue T lost of un-matched trajectories in adjacent frames, and the trajectory queue T removed of ended (removed) trajectories; through two thresholds thresh low and thresh highThe current frame detection results are divided into two categories: high-confidence detection results and low-confidence detection results. Among them, the high-confidence detection results are used for the first data association, and the low-confidence detection results are used for the second data association; the track queue T of un-matched tracks in adjacent frames lost The identity identifiers contained in tracked are compared with those in the tracked track queue T tracked ; for the detection results det tracked in the tracked track queue T t -τ , the Kalman filtering method is adopted to generate the detection results predicted by the filter Calculate the detection results predicted by the Kalman filter and the detection results det output by the detection model t to form a position information similarity matrix C with the intersection over union as the cost IoU ;

[0163] The first data association: For this cost matrix C IoU the Jonker-Volgenant linear assignment algorithm is used. This algorithm is the shortest augmented path algorithm for dense and sparse linear assignment problems and can obtain a set S of matching index pairs m , an un-matched track set S um-track , and an un-matched detection result set S um-det . For the set S of index pairs that can be matched m , it contains a tracked track element and a current frame detection result element. If the matched track belongs to the tracked track queue T tracked , the current frame detection result is directly added to this track to become the subsequent tracked track. Otherwise, the track matched by this detection result will be the track in the track queue T of un-matched tracks in adjacent frames lost , and this un-matched track is re-activated;

[0164] The second data association: For the low-confidence detection results, exactly the same processing method as in the first association is adopted to obtain the tracked tracks and the re-activated tracks. The tracks that still cannot be matched after the second data association are marked as un-matched tracks in adjacent frames and summarized into the corresponding queue T lost ; for the un-matched detection results in the first data association Calculate the position similarity with the unactivated trajectories, and use the Jonker-Volgenant linear assignment algorithm to obtain the matching pair indices, and end the unmatched unactivated trajectories; for the detection results that still exist in the high-confidence detection results but have not been matched, use them as the starting points for generating new trajectories; update the trajectory status, and check whether there are trajectories in the unmatched trajectory queue T lost that exceed the association length threshold, and end these trajectories.

[0165] There is a significant improvement in the migration from the scenario of regular target density to the dense target scenario. The visualization results can be seen in Figure 7 . The part circled in yellow indicates the better result area of the algorithm of the present invention than the similar algorithm [1], and the part circled in red indicates the better result area of the latter. It can be found that in various scenarios and different video sequences, the algorithm of the present invention is superior to the similar algorithms in the dense pedestrian scenario.

[0166] The algorithm of the present invention can be implemented and deployed on most standard multi-object tracking datasets, and can be directly connected to the video stream for target tracking processing.

[0167] To prevent overfitting and cause the network detection effect to fail, the present invention can use a dataset with rich targets to pre-train the network to obtain a faster optimal value acquisition.

[0168] This embodiment uses a variant of the DLA-34 network as the backbone layer of the overall network. In addition, the algorithm of the present invention can also use other backbone networks for basic feature extraction without hindering the subsequent steps, and pre-train on the COCO dataset to initialize the backbone network model.

[0169] This embodiment uses the Adam optimizer to train the network, iterates 70 epochs, and starts training with a learning rate of 3.25e-5. The learning rate decays to 3.25e-6 at the 60th epoch.

[0170] The batchsize size set in this embodiment is 8. At the same time, some standard data augmentation strategies are used, including flipping, scale change, and color transformation. The input image size is reshaped to 960*544, and the feature map resolution at the regression branch position is 240*136.

[0171] This embodiment consumed approximately 12 hours in the training phase, on two RTX3090 graphics cards.

[0172] Algorithm [1]: ZHOU X, KOLTUN V, P.Tracking objects as points[C].European Conference on Computer Vision,2020:474-490。

Claims

1. A dense multi-object tracking method for pedestrians based on feature fusion, characterized in that, the method includes: Step 1: Perform basic feature extraction processing on adjacent input video frames to obtain the basic features of each frame; Step 2: Based on the adjacent frame basic features obtained in Step 1, use the re-identification feature extraction network to extract re-identification features, and obtain adjacent frame re-identification features; according to the adjacent frame re-identification features, use the essential matrix module to obtain the inter-frame displacement prediction matrix of the same target; Step 3: According to the adjacent frame basic features, obtain the target detection information of the current frame, specifically including: Step 3.1: Subtract the corresponding elements of the adjacent frame basic features obtained in Step 1 to obtain the inter-frame difference features; Step 3.2: Integrate the inter-frame difference features and the displacement prediction matrix obtained in Step 2 according to the dimension, and use it as the input of the deformable convolution offset extraction unit, so as to obtain the offset prediction required by the deformable convolution network; Step 3.3: Weight the features of the earlier frame in the adjacent frames with the predicted heat map, and use it as the input of the deformable convolution network DCN. The deformation displacement of the convolution kernel is determined by the offset prediction above, so as to obtain the new features of the previous frame different from the basic features; Step 3.4: Weight and fuse the new features of the previous frame and the basic features of the current frame to obtain the new features of the current frame, and use the new features of the current frame for classification and regression to obtain the target detection information of the current frame; Step 4: According to the target detection information of the current frame, form the final trajectory, specifically including: Step 4.1: The result of the enhanced adjacent frame features obtained through Step 3 passes through the classification and regression branches to obtain the category and position information of the target; Step 4.2: According to the category and position information of the target, the adjacent frame re-identification features and the inter-frame displacement prediction obtained in Step 2, associate the same targets between frames; Step 4.3: Form the final trajectory through the linear assignment algorithm.

2. A dense multi-object tracking method for pedestrians based on feature fusion according to claim 1, characterized in that, Step 1 uses the DLA-34 feature extraction network structure to perform basic feature extraction processing on adjacent input video frames, and the method for obtaining the target center point includes: Adopt a central point constraint, the effective radius r in the case of the target center key point center As shown in the following formula: Among them, W is the width of the input image, H is the height of the input image, and IoU threshold is the intersection over union threshold.

3. A dense multi-object tracking method for pedestrians based on feature fusion according to claim 1, characterized in that, In Step 2, the re-identification feature extraction network includes 3 types of network modules, namely the convolutional layer conv, the batch normalization layer BN, and the non-linear activation layer SiLU; Except for the first layer convolutional layer and the last layer convolutional layer using 1×1 convolutional kernels, the remaining convolutional layers all use 3×3 convolutional kernels.

4. A dense multi-object tracking method for pedestrians based on feature fusion according to claim 1, characterized in that, The obtaining of the inter-frame displacement prediction of the same target by using the essential matrix module according to the adjacent frame re-identification features specifically includes: The current frame part E in the extracted multi-frame re-identification embedding model t is correlated with the previous frame part E t-τ to perform a correlation operation and obtain a volume matrix; After obtaining the essential matrix of the inter-frame similarity metric, predict the motion direction and motion displacement of the target between frames; Multiply the horizontal and vertical displacement templates M i,j and V i,j by the horizontal difference probability representation and the vertical difference probability representation respectively, and the displacement change amount of the current frame relative to the previous frame can be obtained.

5. A dense pedestrian multi-object tracking method based on feature fusion according to claim 1, wherein, the deformable convolution offset extraction unit is a convolutional neural network based on a hybrid attention mechanism, specifically including: a convolutional layer conv, a batch normalization layer BN, non-linear activation layers ReLU and SiLU, a max pooling layer, an average pooling layer, a fully connected layer FC, a basic residual block, a spatial attention mechanism network, and a channel attention mechanism network; the basic residual block is used for further feature extraction.

6. A dense pedestrian multi-object tracking method based on feature fusion according to claim 1, wherein, the weighting of the features of the earlier frame in adjacent frames with the predicted heatmap specifically includes: For the feature map of the previous frame, instead of directly operating on it in the deformable convolution input composition, its basic feature map is multiplied element by element with its heatmap, and the formula is as follows: Among them, represents the basic feature map extracted by the backbone network layer for the (t - τ)-th frame, represents the heatmap result obtained by predicting the (t - τ)-th frame through the detection model, only for the classification of pedestrians, represents the result of and being superimposed channel by channel and pixel by pixel. ⊙ represents the Hadamard product of matrices, and p = 1, 2,..., 64 represents the index values of each channel.

7. A dense pedestrian multi-object tracking method based on feature fusion according to claim 1, wherein, the weighting and fusion of the new features of the previous frame and the basic features of the current frame to obtain the new features of the current frame specifically includes: The integrated features of the previous frame and the basic features of the current frame are added through an adaptive weight matrix, and the formula is as follows: Among them, represents the adaptive matrix of the current frame, w t-τ represents the adaptive matrix of the previous frame, which satisfies the relationship T represents the number of previous frames used, ⊙ represents the Hadamard product of matrices, and the adaptive weight matrix is obtained by two sets of convolutional layers and the softmax function.

8. A dense pedestrian multi-object tracking method based on feature fusion according to claim 1, wherein, Step 4.2 specifically includes: Step 4.2.1, Initialize multiple trajectory queues, which are divided into three categories: the tracked trajectory queue T tracked , the trajectory queue T of the adjacent frame that has not been matched lost , the trajectory queue T that has ended removed ; Divide the detection results of the current frame into two categories through two thresholds thresh low and thresh high : high-confidence detection results and low-confidence detection results; Step 4.2.2: Perform the first data association, specifically including: for the cost matrix C IoU Use the Jonker-Volgenant linear assignment algorithm to obtain the set S of matching index pairs m , the set S of unmatched trajectories um-track , and the set S of unmatched detection results um-det ; for the set S of index pairs that can be matched m , which contains a tracked trajectory element and a current frame detection result element; if the matched trajectory belongs to the tracked trajectory queue T tracked , directly add the current frame detection result to this trajectory to form a consecutive tracked trajectory; otherwise, the detected result will match a trajectory in the queue T of unmatched trajectories in adjacent frames lost , then reactivate this unmatched trajectory Step 4.2.3, perform the second data association, specifically including: for the low-confidence detection results, use exactly the same processing method as in the first association to obtain the tracked trajectories and the reactivated trajectories; Trajectories that still cannot be matched after the second data association are marked as trajectories that are not matched in adjacent frames, and they are summarized and put into the corresponding queue T lost ; For the detection results that are not matched in the first data association Calculate the position similarity between them and the unactivated trajectories, and use the Jonker-Volgenant linear assignment algorithm to obtain the matching pair indexes, and end the unactivated trajectories that are not matched; For the detection results that still exist in the high-confidence detection results but have not been matched, generate them as the starting points of new trajectories; Update the trajectory status, and check the queue T lost of the trajectories that are not matched in adjacent frames to see if there are trajectories that exceed the association length threshold, and end these trajectories.

9. A computer device, including a memory and a processor, and a computer program is stored in the memory, wherein, when the processor runs the computer program stored in the memory, it executes the steps of the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, wherein, multiple computer instructions are stored in the computer-readable storage medium, and the multiple computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • End-to-end multi-target detection and tracking combined method based on target association learning

    CN113139620A

  • Pedestrian multi-target tracking method based on multivariate difference fusion

    CN113221787A