An online updating strategy pedestrian single-target tracking method fusing pedestrian characteristics

By integrating pedestrian characteristics and online update strategies, a single-target pedestrian tracking method is developed. This method utilizes ResNet50 feature extraction and model predictors to address the issues of insufficient real-time performance and accuracy in existing technologies. It achieves stable tracking in complex scenarios and improves both tracking accuracy and speed.

CN114067240BActive Publication Date: 2025-10-28TIANJIN UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111294661.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-03
Publication Date
2025-10-28
Estimated Expiration
2041-11-03

AI Technical Summary

Technical Problem

Existing single-target pedestrian tracking technologies are insufficient in terms of real-time performance and accuracy, making it difficult to maintain stable tracking in complex scenarios, especially when there are occlusions or changes in lighting conditions, which can easily lead to target loss.

Method used

An online update strategy incorporating pedestrian characteristics is adopted. The classification filter is trained and optimized online using a ResNet50 feature extraction network and a model predictor. Combining classification and regression tasks, data augmentation and response maps are used to determine the tracking status, and the update strategy is adjusted to adapt to pedestrian motion characteristics.

Benefits of technology

It achieves high-precision and real-time single-target pedestrian tracking, can accurately track pedestrians in complex scenes, reduces occlusion and drift phenomena, and has practical value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114067240B_ABST
    Figure CN114067240B_ABST
Patent Text Reader

Abstract

A pedestrian single-target tracking method incorporating online update strategies based on pedestrian characteristics decomposes the pedestrian target tracking problem in subsequent frames into classification and regression tasks, given only the target state of the initial frame. The classification task aims to classify image regions into foreground and background using a classification filter to predict the coarse location of the target in the image. The regression task estimates the target state by combining the coarse localization obtained from the classification task with candidate bounding boxes, typically represented by bounding boxes. The method further refines the coarse location of the target by incorporating the inherent characteristics of pedestrian movement and defines different tracking states for different scenes based on the complexity of the current scene. Different online update strategies are applied to the classification filter to increase the discriminative power of the classifier, thereby improving the tracking performance of single pedestrian targets and achieving a higher tracking success rate, thus possessing practical value.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to the fields of pattern recognition, image processing, and computer vision, and specifically to a single-target pedestrian tracking method that integrates pedestrian characteristics with an online update strategy. [Background Technology]

[0002] Visual tracking technology is an important topic in the field of computer vision (a branch of artificial intelligence), with significant research value and has received widespread attention in recent years. Pedestrian tracking is key to achieving intelligent analysis such as pedestrian analysis.

[0003] In existing technologies, methods for single-target pedestrian tracking mainly fall into two categories:

[0004] 1. Using a pedestrian recognition algorithm to detect each video frame individually and then connecting the pedestrian target boxes to form the target trajectory, this scheme can achieve long-term target tracking. However, the effectiveness of the detection network is inversely proportional to the detection time. The more complex the detection algorithm, the better the system can extract image features, resulting in better detection performance. However, due to the deeper network and more system parameters, the system detection time increases, making real-time pedestrian target tracking impossible. Conversely, the weaker the network's expressive power, the lower its corresponding detection accuracy, making the algorithm prone to losing track of pedestrians. Neither approach makes it easy to apply the algorithm to real-world scenarios. Therefore, to apply this scheme to real-world scenarios, it is necessary to increase the system's computing power and upgrade the system configuration. The advantage of this scheme is its ability to extract deep semantic features of pedestrian targets, resulting in strong recognition capabilities. However, its disadvantages lie in frame-by-frame detection, which does not utilize video context information, thus failing to improve the system's detection frame rate. Furthermore, in real-time video target detection, motion blur and other issues can lead to local tracking failures, reducing the system's tracking efficiency.

[0005] 2. Using a tracking algorithm, the pedestrian target is manually outlined in the first frame, or a recognition algorithm is used to detect the pedestrian target, and then the tracking algorithm is used to track the pedestrian. This method can achieve short-term target tracking. The advantage of using a target tracking algorithm is that the tracking algorithm has a simple structure and can achieve real-time tracking. However, its disadvantage is that during the tracking process, the pedestrian target is prone to deformation, occlusion, or changes in lighting, which can easily lead to the target being lost. Furthermore, once tracking fails, the target cannot be retrieved, thus causing the algorithm to fail.

[0006] Therefore, this application analyzes pedestrian, single-target, and short-term tracking, and proposes corresponding solutions to address the problems encountered by existing trackers in single-target pedestrian tracking. Given only the target state in the initial frame, the tracking problem is decomposed into classification and regression tasks based on the tracking algorithm. The classification task aims to provide a coarse, stable location of the target in the image by classifying image regions into foreground and background. The regression task estimates the target state, typically represented by bounding boxes, thereby enabling the tracker to perform single-target pedestrian tracking in subsequent frames of a video sequence. [Summary of the invention]

[0007] The purpose of this invention is to propose an online update strategy for single-target pedestrian tracking that integrates pedestrian characteristics. This method can overcome the shortcomings of existing technologies and is a method with high accuracy and fast real-time tracking speed for single-target pedestrians, and it has certain practical value.

[0008] The technical solution of this invention: A method for single-target pedestrian tracking with an online update strategy that integrates pedestrian characteristics, characterized by comprising the following steps:

[0009] Step 1: Classification task during the tracking process:

[0010] (1.1) The features of the reference frame and each frame in the video sequence are extracted by the feature extraction network. That is, the first frame of the video sequence is selected as the reference frame, and the target to be tracked is selected by manually marking the target bounding box. In each subsequent frame of the video sequence, i.e. the test frame, the target to be tracked selected in the reference frame is identified and detected, so as to realize the tracking of the target in the video sequence, and thus realize the single target tracking process of pedestrians.

[0011] The feature extraction network is ResNet50, which consists of four residual blocks connected in series. The four residual blocks are named Block1, Block2, Block3, and Block4. Each of the four residual blocks contains 50 convolutional operations, which is a well-known technique. Two convolutional layers are connected after ResNet50 to form a backbone network for feature extraction, which is used to extract features from the current reference frame or test frame image.

[0012] (1.2) Using the features of each frame extracted from the reference frame selected in step (1.1) and its manually annotated target bounding box, a classification filter is trained online by the model predictor to distinguish the foreground and background of the target to be tracked in subsequent frames and to predict the target position.

[0013] The model predictor consists of an initializer module and an optimizer module. The initializer module can effectively provide an initial estimate of the classification filter using only the appearance of the target to be tracked. The optimizer module is used to optimize the initially estimated classification filter init_filter to obtain the optimized classification filter new_filter, which performs target foreground and background classification on subsequent frames of the tracked video sequence and predicts the center position of the target to be tracked for coarse localization.

[0014] The classification filter new_filter is obtained through online training of the model predictor using information from the reference frame. This process consists of the following steps:

[0015] (1.2.1) Using the reference frame information of the currently tracked video sequence, including the reference frame image information and the manually specified target annotation bounding box, flip, mirror, blur and rotate data augmentation operations are performed on the reference frame respectively, and images with the respective operation effects are obtained. These images are combined into a set as the image set after data augmentation, and the corresponding annotation bounding box after data augmentation is obtained.

[0016] (1.2.2) The feature extraction network in step (1.1) is used to extract features from the image set processed in step (1.2.1) to obtain a set of training sample features, which are then fed into the model initialization module for initial estimation of the classification filter. The obtained initial estimated classification filter, init_filter, is as follows: Figure 2 As shown; the initial estimation of the classification filter in the model initialization module refers to extracting features within the labeled bounding box by performing PrRoi Pooling on the features of the training samples. The resulting features are the target features, and the output is the initially estimated classification filter init_filter.

[0017] (1.2.3) The initial estimated classification filter init_filter obtained in step (1.2.2) and the feature extraction network in step (1.1) are used to extract features from the image set processed in step (1.2.1) to obtain a set of training sample features, which are then fed into the optimizer module. After optimization, the optimized classification filter new_filter is obtained and used for target foreground and background classification in subsequent test frames. The specific implementation process can be described by the following steps:

[0018] ① Calculate the reference frame response graph s ref :

[0019] The initial estimated classification filter init_filter from step (1.2.2) is convolved with the training sample features to obtain the response map s. refThe response diagram of the reference frame is shown in equation (1):

[0020] s ref =x ref *f init (1)

[0021] Where x ref The features representing the reference frame are the same as the features of the training samples; * represents convolution calculation; f init The initial estimated classification filter is init_filter;

[0022] ② Calculate the reference frame response graph s ref and reference frame labels The difference between them is r(s,c):

[0023] The response graph s of the reference frame calculated according to step (1.2.2) ref Its representation is a 19×19 two-dimensional matrix; by scaling the bounding box of the pedestrian target annotation in the reference frame onto the 19×19 two-dimensional matrix, the position c of the annotated pedestrian target in the 19×19 two-dimensional matrix is ​​obtained;

[0024]

[0025]

[0026] In formulas (2) and (3), t represents each position of a 19×19 two-dimensional matrix, and c represents the target position; The spatial coefficients, ρ, are obtained during training. k The distance calculation function is determined by formulas (4-1) and (4-2):

[0027]

[0028]

[0029] Among them, y hn The true label represents the current frame response map; m c This represents the mask parameters, used to determine the region where the current target is located in a 19×19 two-dimensional matrix, within the region m corresponding to the target. c ≈1, in the background region m c ≈0; label y hn and mask parameter m c It is represented as a 19×19 two-dimensional matrix;

[0030] The reference frame label is obtained using formulas (2) and (3). and mask parameters And calculate the response graph s of the reference frame obtained in step ①.ref and reference frame labels The difference between them, r(s,c), is shown in equation (5):

[0031] r(s,c)=v c ·(m c s+(1-m c max(0,s)-y hn (5)

[0032] In the formula, s is the response graph of the reference frame. ref ;v c For spatial weights; m c Mask parameters for reference frame y hn Labels for reference frames

[0033] ③ Regularization is applied to the difference r(s,c) obtained in step ② to obtain L(f), as shown in equation (6), and this is used as the reference frame response diagram s. ref and reference frame labels The loss between the two is backpropagated to optimize the classification filter.

[0034]

[0035] Where, Let x be the training sample set, where x j c represents the training sample features extracted by the feature extraction network. j The coordinates of the target center in the sample are already labeled, i.e., the coordinates of the center point of the labeled bounding box; * represents convolution calculation; λ is the regularization factor; f is the optimized classification filter.

[0036] The optimized classification filter f is a backpropagation optimized classification filter, using the steepest gradient descent method as shown in equation (7):

[0037]

[0038] In the formula, f (i) Represents the classification filter after the i-th optimization; α represents the learning rate; Represents gradient calculation;

[0039] (1.3) Take the subsequent frames in the current sequence to be tracked as test frames and perform feature extraction. That is, use a feature extraction network to extract features from the test frames to obtain the features x of the test frames. test ;

[0040] (1.4) Perform convolution processing using the classification filter new_filter optimized in step (1.2) and the features of the test frame obtained in step (1.3), as shown in formula (8).

[0041] s test =x test *f new (8)

[0042] s test This is the test frame response graph; where x test Features representing the test frame; * represents convolution calculation; f new The optimized classification filter is new_filter;

[0043] (1.5) The response graph s of the test frame obtained in step (1.4) test The tracking status of the current test frame is determined, and the tracking status is divided into three categories: normal, uncertain, and undetectable. The normal state indicates that the test frame scene is simple, and the target and background can be easily identified through a classification task. The uncertain state indicates that the current test frame scene is complex, affected by interference and background, making it difficult to accurately identify the target and background. The undetectable state indicates that the current test frame scene is complex, obstructed, or the target and background cannot be identified. The specific implementation consists of the following steps:

[0044] (1.5.1) The response graph s of the test frame obtained in step (1.4) test Determine the tracking status of the current test frame. The tracking status is divided into normal status, uncertain status, and tracking not found status.

[0045] ①If the response graph s of the test frame test If only the target center has the highest response, it means that there is only the target to be tracked or there is a clear boundary between the target to be tracked and the background. That is, the tracking state of the current test frame is normal, which means that the scene of the test frame is simple and the target to be tracked and the background can be easily identified through the classification task.

[0046] ②If the response graph s of the test frame test The target region is the area selected around the position of the highest response score using the target bounding box size of the previous frame. At this time, the response scores in the target region are relatively messy, and the highest response score will drop suddenly. This indicates that the target to be tracked is in an uncertain state, which means that the target to be tracked is confused with the background or close to the interference.

[0047] ③If the response graph s of the test frame testIf the response score within the undetectable target area is more chaotic than that in the uncertain state, and the highest response score drops more significantly than in the uncertain state, it indicates that the target being tracked is severely occluded and belongs to the undetectable state.

[0048] (1.5.2) Calculate the response graph s of the test frame. test The variance of the response score in the target region and the highest response score in the test frame are used to determine the state of the target to be tracked in the current search region:

[0049] (1.5.2.1) Calculate the mean variance of the target region response score of the m frames preceding the test frame using a sliding method. The target region is the response map s of the test frame using the target bounding box size from the previous frame. tset The region around the highest response score position is selected, and the variance σ of the response score of the target region in the m frames prior to the test frame is recorded, as shown in Equation (9):

[0050]

[0051] In the formula, score i The score is assigned to each location in the target region of the response graph of the test frame. The response graph s of the test frame test The mean score of each location in the corresponding target area; w*h is the response graph s of the test frame. test The dimensions of the corresponding target area, including its width and height.

[0052] (1.5.2.2) Further calculate its average value according to formula (10), which is the mean variance of the response score of m frames.

[0053]

[0054] Where, σ j The variance of the target region response score for each frame within the target region of m frames;

[0055] (1.5.2.3) Meanwhile, the highest response score of the test frame is the response graph s of the test frame. test The maximum response score is recorded, and the highest response score max_score of the m frames preceding the test frame is recorded. The average of the highest response scores of the m frames is calculated using formula (11).

[0056]

[0057] In the formula max_score j Represents the highest response score per frame;

[0058] (1.5.2.4) Based on the normal state, uncertain state, and unfindable state described in step (1.5.1), combined with the mean variance of the target region response score of the m frames preceding the test frame obtained in step (1.5.2.2), And the average of the highest response scores of the m frames preceding the test frame obtained in step (1.5.2.3), to determine the tracking status of the test frame:

[0059] If equation (12) is satisfied, it indicates that the tracking state of the test frame is in an uncertain state:

[0060]

[0061] If equation (13) is satisfied, it means that the tracking overstate of the test frame is a state that cannot be found.

[0062]

[0063] Other situations are considered normal.

[0064] in, The mean variance of the target region response score in the m frames preceding the current test frame; The response graph s of the test frame m frames prior to the test frame. test The average score of the highest response; k1 and k2 are proportionality coefficients.

[0065] Step 2: Regression task during the tracking process:

[0066] (2.1) The response graph s of the test frame obtained from the classification task in step (1.4) test Predict the location of the target center;

[0067] (2.1.1) The response graph s of the test frame obtained in step (1.4) test The response graph of the test frame s test The position coordinates corresponding to the maximum response score are taken as the first response point. If the tracking state of the test frame is normal and there are no interfering objects, the first response point will be used as the predicted target center position.

[0068] (2.1.2) In the response graph s of the test frame test The target region, i.e., the response map s of the test frame using the target bounding box size from the previous frame. test The region around the location of the maximum response score in the test frame is shown in the response graph s. test The area outside the selected area is considered the target area, and the location corresponding to the highest response score outside the target area is regarded as the second response point.

[0069] (2.1.3) When the highest response score of the second response point is greater than 0.5 times the highest response score of the first response point, the second response point is considered to be a target similar to the background.

[0070] (2.1.4) Assume the current position of the first response point is c1[x1,y1], the position of the second response point is c2[x2,y2], and the target center point of the previous frame of the current test frame is in the response graph s of the test frame. test The position is c0[x0,y0], which is the center point of the target area in the response diagram. Then the position offsets of c1 and c2 relative to c0 are shown in Equations (14) and (15), respectively:

[0071]

[0072]

[0073] (2.1.5) Determine the true location of the current target:

[0074] The offsets calculated according to equations (14) and (15) are within different value ranges, and different response point positions are returned as the predicted target center positions, as shown in equations (16) and (17):

[0075] c1,(d1>Ω&d2<Ω)|(d1>Ω&d2>Ω)|(d1 <d2&d1<Ω&d2<Ω) (16)

[0076] c2,(d1<Ω&d2>Ω)|(d1>d2&d1<Ω&d2<Ω) (17)

[0077] In equations (16) and (17) above, Ω represents the threshold value within the range of values;

[0078] (2.2) Based on the inherent characteristics of the pedestrian target during the motion process, the target center position predicted in step (2.1) is corrected to obtain the final predicted target center position. That is, according to the inherent characteristics of the pedestrian target during the motion process, the target tracking is normal under normal conditions, and the pedestrian target does not undergo drastic scale changes during the motion process, and the pedestrian's motion is relatively smooth. Therefore, the offset of the target center under normal tracking conditions is recorded. The relative offset of the target center between two frames in frame v is recorded by sliding frames, and its average value is calculated as the predicted offset. The value range of v is 12-18. If the tracking state is in an uncertain state, the predicted target center position obtained in step (2.1) is corrected by the predicted offset to obtain the final predicted target center position.

[0079] (2.3) Based on the final predicted target center position obtained in step (2.2) and the target bounding box of the previous frame, if the frame is the second frame, the target bounding box marked in the first frame is used as the initial candidate bounding box; if it is a subsequent frame, the bounding box predicted in step (2.4) is used as the initial candidate bounding box; and combined with the target bounding box marked in the reference frame as the reference bounding box, a set of candidate bounding boxes is generated around the final predicted target center position obtained in step (2.2); the specific implementation process is as follows:

[0080] (2.3.1) According to step (2.2), the final predicted target center position is obtained. The target bounding box of the previous frame is used as the initial candidate bounding box. If the frame is the second frame, the target bounding box marked in the first frame is used as the initial candidate bounding box. If it is a subsequent frame, the bounding box predicted in step (2.4) is used as the initial candidate bounding box. Around the final predicted target center position, a candidate bounding boxes with different proportions are randomly generated to form the initial candidate bounding box set.

[0081] (2.3.2) Based on the inherent characteristics of pedestrian targets during movement, the scale changes of pedestrian targets are relatively gentle during movement. Combined with the target bounding box of the manually annotated reference frame in step (1.1), it is used as a reference candidate bounding box. Around the center position of the final predicted target, b reference candidate bounding boxes with different proportions are randomly generated to form a reference candidate bounding box set.

[0082] (2.3.3) The initial candidate bounding box set obtained in step (2.3.1) and the reference candidate bounding box set obtained in step (2.3.2) are fused to obtain a+b candidate bounding boxes as the candidate bounding box set;

[0083] In step (2.3.1), the value of 'a' ranges from 7 to 15; in step (2.3.2), the value of 'b' ranges from 3 to 7.

[0084] (2.4) The a+b optimizable candidate bounding boxes obtained in step (2.3) and the features of the test frame obtained in step (1.3) are fed into the bounding box prediction module to predict the target bounding box.

[0085] The target bounding box prediction process in step (2.4) includes the following steps:

[0086] (2.4.1) Since the specified tracking target is uncertain before annotation, the bounding box prediction process needs to combine the target information annotated in the reference frame. Therefore, it is necessary to extract the target information annotated in the reference frame. That is, the feature extraction network ResNet50 in step (1.1) is used to extract the features of the reference frame. The reference frame is fed into the ResNet50 feature extraction network and processed through four residual blocks in sequence. The output layer1 of Block3 and the output layer2 of Block4 of the reference frame are extracted as the features of the reference frame. These features are then fused after convolution and Pr Pooling and then passed through a fully connected layer to obtain the modulation vector, which can be used as the pedestrian target information annotated in the reference frame.

[0087] (2.4.2) For the test frame, such as Figure 4 As shown in the test frame branch, the features of the test frame obtained by ResNet50 in the feature extraction network in step (1.1) are used and passed through two convolutional layers. Then, PrPooling is performed on a+b bounding box regions to extract internal features. Combined with the modulation vector, i.e., the information of the pedestrian target marked in the reference frame, the Intersection over Union (IoU) of a+b bounding boxes is predicted through a fully connected layer. The gradient of the bounding box is calculated through the IoU. The a+b candidate bounding boxes are optimized to obtain the optimized candidate bounding boxes, which are used as a new set of candidate bounding boxes that can be optimized.

[0088] (2.4.3) Repeat step (2.4.2) for iterative optimization. After 5 iterations, take the average coordinates of the three candidate bounding boxes with the largest intersection-union ratio (IoU) as the coordinates of the predicted bounding box, which is the final predicted target bounding box.

[0089] (2.5) Update the classification filter new_filter based on the tracking state of the current test frame determined in step (1.5). During the movement of pedestrians, there will be scenes of occlusion, staggered movement between pedestrians, and blending into the background. By optimizing and updating the classification filter new_filter based on the tracking state of the current target, background or interference information can be prevented from being introduced. The optimization and update of the classification filter new_filter is performed based on the tracking state of the test frame as follows:

[0090] (2.5.1) When the tracking state of the current test frame is determined to be normal tracking state, the classification filter new_filter is optimized and updated every n frames or when an interference object is encountered; the value of n ranges from 15 to 25.

[0091] The optimization and update method for the classification filter new_filter is to use the test frame response map s2 and test frame labels obtained in step (1.4). In this implementation, the response graph s2 of the test frame is represented as a 19×19 two-dimensional matrix:

[0092] (i) Scale the bounding box predicted by step (2.4.3) onto a 19×19 two-dimensional matrix, and then the position c of the pedestrian target in the test frame in the 19×19 two-dimensional matrix can be obtained;

[0093] (ii) Calculate the target mask parameters m according to step ② in step (1.2.3). c and tag y hn Calculate the target mask parameters of the test frame using the following method and tags

[0094] (iii) Calculate the test frame response graph s using formula (2). test and test frame labels The residual between them is r(s,c); at this time, in equation (2), s is the response graph of the test frame. test ;v c For spatial weights; m c Target mask parameters for the test frame y hn Labels for test frames

[0095] (iv) Following the steps in step ③ of (1.2.3), obtain the test frame response diagram s. test and test frame labels The loss difference between the two is used to optimize and update the classification filter new_filter, resulting in a new optimized classification filter new_filter used for target foreground and background classification in subsequent test frames.

[0096] (2.5.2) When the tracking state of the current test frame is determined to be uncertain, the classification filter new_filter is not optimized or updated;

[0097] (2.5.3) When the tracking state of the current test frame is determined to be "unfinished", the classification filter new_filter will not be optimized or updated.

[0098] (2.6) By repeating steps (1.3) to (2.5), target recognition and detection are completed for each frame in the entire tracked video sequence, ultimately achieving single-target pedestrian tracking. The specific process is as follows:

[0099] (2.6.1) The information of the selected target pedestrian in the reference frame of the tracked video sequence is obtained through steps (1.1) to (1.2), and the resulting classification filter new_filter is used for target foreground and background classification in subsequent test frames;

[0100] (2.6.2) Each subsequent test frame repeats steps (1.3) to (2.5) until the last frame, thereby completing the detection of the specified pedestrian target in each frame of the entire video sequence, and finally realizing the tracking of a single specified pedestrian target in the reference frame in the entire video sequence.

[0101] Advantages of this invention: This invention designs a discriminative single-target pedestrian tracking method that integrates pedestrian characteristics and a novel online update strategy. It primarily studies the application of a discriminative model based on an online update strategy in single-target pedestrian tracking. Online training is one solution in discriminative models, requiring the prediction of a classification filter using information from the first frame. In subsequent tracking processes, if update conditions are met, the predicted results are used as new training samples to fill the training sample set. The classification filter is then optimized using this new sample set. However, during the update process, the current tracking status is uncertain. Updating every few frames may introduce excessive background or interference information, ultimately leading to inaccurate tracking or even drift. Furthermore, the predicted target center is not accurate enough under different states. To address these issues, this invention uses a response map obtained by convolving the classification filter and test frame features to determine the tracking status of the test frame, thus adjusting the update strategy accordingly. During pedestrian movement, there may be occlusion, staggered movement between pedestrians, or blending into the background, among other different states. Current techniques cannot determine the location of a pedestrian when the target is occluded, leading to drift in subsequent target tracking. Drift can also occur when the pedestrian moves away from other pedestrians. To address these occlusion and drift issues, this paper adjusts the predicted target center in the test frame based on pedestrian characteristics, resulting in more accurate tracking. Compared to other methods, the proposed method achieves higher accuracy for single-target pedestrian tracking and reaches real-time speed, combining both accuracy and speed, thus possessing practical value. [Attached Image Description]

[0102] Figure 1 This is a schematic diagram of the system framework of a single-target pedestrian tracking method that incorporates pedestrian characteristics and an online update strategy.

[0103] Figure 2 This is a structural diagram of the initializer module of an online update strategy pedestrian single-target tracking method that incorporates pedestrian characteristics, as described in this invention.

[0104] Figure 3This is a structural diagram of the optimizer module of the pedestrian single-target tracking method that integrates pedestrian characteristics and online update strategy.

[0105] Figure 4 This is a schematic diagram of the bounding box prediction module of a pedestrian single-target tracking method that incorporates pedestrian characteristics and an online update strategy.

Detailed Implementation Methods

[0106] like Figure 1 As shown, an online update strategy for single-target pedestrian tracking that integrates pedestrian characteristics is characterized by the following steps:

[0107] Step 1: Classification task during the tracking process:

[0108] (1.1) The features of the reference frame and each frame in the video sequence are extracted by the feature extraction network. That is, the first frame of the video sequence is selected as the reference frame, and the target to be tracked is selected by manually marking the target bounding box. In each subsequent frame of the video sequence, i.e. the test frame, the target to be tracked selected in the reference frame is identified and detected, so as to realize the tracking of the target in the video sequence, and thus realize the single target tracking process of pedestrians.

[0109] The feature extraction network in step (1.1) is a ResNet50, which consists of four residual blocks connected in series. The four residual blocks are named Block1, Block2, Block3, and Block4, and each residual block contains 50 convolutional operations; this is a well-known technique. Two convolutional layers are then connected after the ResNet50 to form the backbone network for feature extraction. Figure 1 As shown, this is used to extract features from the current reference frame or test frame image.

[0110] (1.2) Using the features of each frame extracted from the reference frame selected in step (1.1) and its manually annotated target bounding box, a classification filter is trained online by the model predictor to distinguish the foreground and background of the target to be tracked in subsequent frames and to predict the target position.

[0111] The model predictor in step (1.2), such as Figure 1 As shown, it consists of an initializer module and an optimizer module; the initializer module, as... Figure 2 As shown, the initial estimate of the classification filter can be effectively provided using only the appearance of the target to be tracked; the optimizer module, as... Figure 3As shown, the initial estimated classification filter init_filter is optimized to obtain the optimized classification filter new_filter. This filter is used to classify the target in front of and behind subsequent frames of the tracked video sequence and to roughly locate the center position of the target to be tracked.

[0112] The step (1.2) described above uses the information from the reference frame to train the classification filter new_filter online through the model predictor. This process consists of the following steps:

[0113] (1.2.1) Using the reference frame information of the currently tracked video sequence, including the reference frame image information and the manually specified target annotation bounding box, flip, mirror, blur and rotate data augmentation operations are performed on the reference frame respectively, and images with the respective operation effects are obtained. These images are combined into a set as the image set after data augmentation, and the corresponding annotation bounding box after data augmentation is obtained.

[0114] (1.2.2) The feature extraction network in step (1.1) is used to extract features from the image set processed in step (1.2.1) to obtain a set of training sample features, which are then fed into the model initialization module for initial estimation of the classification filter. The obtained initial estimated classification filter, init_filter, is as follows: Figure 2 As shown;

[0115] In step (1.2.2), the initial estimation of the classification filter by the model initialization module refers to extracting the features within the labeled bounding box by performing PrRoi Pooling on the features of the training samples. The resulting features are the target features, which are then used as the output, i.e., the initially estimated classification filter init_filter.

[0116] (1.2.3) The initial estimated classification filter init_filter obtained in step (1.2.2) and the feature extraction network in step (1.1) are used to extract features from the image set processed in step (1.2.1) to obtain a set of training sample features, which are sent to the optimizer module. The optimized classification filter new_filter is obtained through optimization and used for target foreground and background classification in subsequent test frames.

[0117] The specific implementation process of step (1.2.3) is as follows: Figure 3 As shown, it can be described using the following steps:

[0118] ① Calculate the reference frame response graph s ref :

[0119] The initial estimated classification filter init_filter from step (1.2.2) is convolved with the training sample features to obtain the response map s. ref The response diagram of the reference frame is shown in equation (1):

[0120] s ref =x ref *f init (1)

[0121] Where x ref The features representing the reference frame are the same as the features of the training samples; * represents convolution calculation; f init The initial estimated classification filter is init_filter;

[0122] ② Calculate the reference frame response graph s ref and reference frame labels The difference between them is r(s,c):

[0123] The response graph s of the reference frame calculated according to step (1.2.2) ref Its representation is a 19×19 two-dimensional matrix; by scaling the bounding box of the pedestrian target annotation in the reference frame onto the 19×19 two-dimensional matrix, the position c of the annotated pedestrian target in the 19×19 two-dimensional matrix is ​​obtained;

[0124]

[0125]

[0126] In formulas (2) and (3), t represents each position of a 19×19 two-dimensional matrix, and c represents the target position; The spatial coefficients, ρ, are obtained during training. k The distance calculation function is determined by formulas (4-1) and (4-2):

[0127]

[0128]

[0129] In this embodiment, N is 10. In a 19×19 two-dimensional matrix, the furthest distance between the positions of t and c is calculated as follows: Therefore, N is set to 10, i.e., k=9 represents all positions far from the target center, so the same processing can be performed; the value of τ is in the range of 0.45-0.55. Experiments show that when the value of τ is 0.5, the background and target corresponding areas can be better distinguished, and the experimental results are optimal.

[0130] Among them, y hn The true label represents the current frame response map; mc This represents the mask parameters, used to determine the region where the current target is located in a 19×19 two-dimensional matrix, within the region m corresponding to the target. c ≈1, in the background region m c ≈0; label y hn and mask parameter m c It is represented as a 19×19 two-dimensional matrix;

[0131] The reference frame label is obtained using formulas (2) and (3). and mask parameters And calculate the response graph s of the reference frame obtained in step ①. ref and reference frame labels The difference between them, r(s,c), is shown in equation (5):

[0132] r(s,c)=v c ·(m c s+(1-m c max(0,s)-y hn ) (5)

[0133] In the formula, s is the response graph of the reference frame. ref ;v c For spatial weights; m c Mask parameters for reference frame y hn Labels for reference frames

[0134] ③ Regularization is applied to the difference r(s,c) obtained in step ② to obtain L(f), as shown in equation (6), and this is used as the reference frame response diagram s. ref and reference frame labels The loss between the two is backpropagated to optimize the classification filter.

[0135]

[0136] In the formula, Let x be the training sample set, where x j c represents the training sample features extracted by the feature extraction network. j The coordinates of the target center in the sample are already labeled, i.e., the coordinates of the center point of the labeled bounding box; * represents convolution calculation; λ is the regularization factor; f is the optimized classification filter.

[0137] The optimized classification filter f in step ③ is a backpropagation optimized classification filter, which adopts the steepest gradient descent method as shown in equation (7):

[0138]

[0139] In the formula, f (i) Represents the classification filter after the i-th optimization; α represents the learning rate; Represents gradient calculation;

[0140] In the embodiment, the optimized classification filter f is the optimized classification filter new_filter obtained after five iterations of optimization, which is used for target foreground and background classification in subsequent test frames.

[0141] (1.3) Take the subsequent frames in the current sequence to be tracked as test frames and perform feature extraction. That is, use a feature extraction network to extract features from the test frames to obtain the features x of the test frames. test ;

[0142] (1.4) Perform convolution processing using the classification filter new_filter optimized in step (1.2) and the features of the test frame obtained in step (1.3), as shown in formula (8).

[0143] s test =x test *f new (8)

[0144] s test This is the test frame response graph; where x test Features representing the test frame; * represents convolution calculation; f new The optimized classification filter is new_filter;

[0145] (1.5) The response graph s of the test frame obtained in step (1.4) test The system determines the tracking status of the current test frame, which is divided into three states: normal, uncertain, and undetectable. The normal state indicates that the scene of the test frame is simple, and the target and background to be tracked can be easily identified through the classification task. The uncertain state indicates that the scene of the current test frame is complex, affected by interference and background, making it difficult to accurately identify the target and background to be tracked. The undetectable state indicates that the scene of the current test frame is complex, occluded, or the target and background to be tracked cannot be identified.

[0146] The specific implementation of step (1.5) consists of the following steps:

[0147] (1.5.1) The response graph s of the test frame obtained in step (1.4) test Determine the tracking status of the current test frame. The tracking status is divided into normal status, uncertain status, and tracking not found status.

[0148] ①If the response graph s of the test frame testIf only the target center has the highest response, it means that there is only the target to be tracked or there is a clear boundary between the target to be tracked and the background. That is, the tracking state of the current test frame is normal, which means that the scene of the test frame is simple and the target to be tracked and the background can be easily identified through the classification task.

[0149] ②If the response graph s of the test frame test The target region is the area selected around the position of the highest response score using the target bounding box size of the previous frame. At this time, the response scores in the target region are relatively messy, and the highest response score will drop suddenly. This indicates that the target to be tracked is in an uncertain state, which means that the target to be tracked is confused with the background or close to the interference.

[0150] ③If the response graph s of the test frame test If the response score within the undetectable target area is more chaotic than that in the uncertain state, and the highest response score drops more significantly than in the uncertain state, it indicates that the target being tracked is severely occluded and belongs to the undetectable state.

[0151] (1.5.2) Calculate the response graph s of the test frame. test The variance of the response score in the target region and the highest response score in the test frame are used to determine the state of the target to be tracked in the current search region:

[0152] (1.5.2.1) Calculate the mean variance of the target region response score of the m frames preceding the test frame using a sliding method. The target region is the response map s of the test frame using the target bounding box size from the previous frame. tset The region around the highest response score position is selected, and the variance σ of the response score of the target region in the m frames prior to the test frame is recorded, as shown in Equation (9):

[0153]

[0154] In the formula, score i The score is assigned to each location in the target region of the response graph of the test frame. The response graph s of the test frame test The mean score of each location in the corresponding target area; w*h is the response graph s of the test frame. test The dimensions of the corresponding target area, including its width and height.

[0155] (1.5.2.2) Further calculate its average value according to formula (10), which is the mean variance of the response score of m frames.

[0156]

[0157] Where, σ jThe variance of the target region response score for each frame within the target region of m frames;

[0158] (1.5.2.3) Meanwhile, the highest response score of the test frame is the response graph s of the test frame. test The maximum response score is recorded, and the highest response score max_score of the m frames preceding the test frame is recorded. The average of the highest response scores of the m frames is calculated using formula (11).

[0159]

[0160] In the formula max_score j Represents the highest response score per frame;

[0161] (1.5.2.4) Based on the normal state, uncertain state, and unfindable state described in step (1.5.1), combined with the mean variance of the target region response score of the m frames preceding the test frame obtained in step (1.5.2.2), And the average of the highest response scores of the m frames preceding the test frame obtained in step (1.5.2.3), to determine the tracking status of the test frame:

[0162] If equation (12) is satisfied, it indicates that the tracking state of the test frame is in an uncertain state:

[0163] If equation (12) is satisfied, it indicates that the tracking state of the test frame is in an uncertain state:

[0164]

[0165] If equation (13) is satisfied, it means that the tracking overstate of the test frame is a state that cannot be found.

[0166]

[0167] Other situations are considered normal.

[0168] in, The mean variance of the target region response score in the m frames preceding the current test frame; The response graph s of the test frame m frames prior to the test frame. test The average score of the highest response; k1 and k2 are proportionality coefficients.

[0169] In this embodiment, m = 25, k1 = 0.75, and k2 = 0.5. Let... The mean variance of the target region response score in the 25 frames preceding the test frame. The response graph s of the test frames preceding the test frame. testThe highest average response score is used. Furthermore, if the number of frames before the test frame is less than 25, the calculation process proceeds from the test frame backward until there are no more frames. The values ​​of 25 frames and the scaling factors k1 and k2 are optimal choices obtained through experimental results.

[0170] Step 2: Regression task during the tracking process:

[0171] (2.1) The response graph s of the test frame obtained from the classification task in step (1.4) test Predict the location of the target center;

[0172] Step (2.1) specifically refers to:

[0173] (2.1.1) The response graph s of the test frame obtained in step (1.4) test The response graph of the test frame s test The position coordinates corresponding to the maximum response score are taken as the first response point. If the tracking state of the test frame is normal and there are no interfering objects, the first response point will be used as the predicted target center position.

[0174] (2.1.2) In the response graph s of the test frame test The target region, i.e., the response map s of the test frame using the target bounding box size from the previous frame. test The region around the location of the maximum response score in the test frame is shown in the response graph s. test The area outside the selected area is considered the target area, and the location corresponding to the highest response score outside the target area is regarded as the second response point.

[0175] (2.1.3) When the highest response score of the second response point is greater than 0.5 times the highest response score of the first response point, the second response point is considered to be a target similar to the background.

[0176] (2.1.4) Assume the current position of the first response point is c1[x1,y1], the position of the second response point is c2[x2,y2], and the target center point of the previous frame of the current test frame is in the response graph s of the test frame. test The position is c0[x0,y0], which is the center point of the target area in the response diagram. Then the position offsets of c1 and c2 relative to c0 are shown in Equations (14) and (15), respectively:

[0177]

[0178]

[0179] (2.1.5) Determine the true location of the current target:

[0180] The offsets calculated according to equations (14) and (15) are within different value ranges, and different response point positions are returned as the predicted target center positions, as shown in equations (16) and (17):

[0181] c1,(d1>Ω&d2<Ω)|(d1>Ω&d2>Ω)|(d1 <d2&d1<Ω&d2<Ω) (16)

[0182] c2,(d1<Ω&d2>Ω)|(d1>d2&d1<Ω&d2<Ω) (17)

[0183] In equations (16) and (17) above, Ω represents the threshold value within the range of values;

[0184] (2.2) Combining the inherent characteristics of the pedestrian target during the movement process, the target center position predicted in step (2.1) is modified to obtain the final predicted target center position;

[0185] Step (2.2) specifically refers to: based on the inherent characteristics of the pedestrian target during the movement process, under normal conditions the target tracking is normal, and the pedestrian target does not undergo drastic scale changes during the movement process, and the pedestrian movement is relatively smooth. Therefore, the offset of the target center under normal tracking conditions is recorded, and the relative offset of the target center between two frames in v frames is recorded by sliding frames. The average value is calculated as the predicted offset, and the value of v ranges from 12 to 18. If the tracking state is in an uncertain state, the predicted target center position obtained in step (2.1) is corrected by the predicted offset, and finally the final predicted target center position is obtained.

[0186] In this embodiment, v = 16, which is the optimal choice obtained through experimental results.

[0187] (2.3) Based on the final predicted target center position obtained in step (2.2) and the target bounding box of the previous frame, if the frame is the second frame, the target bounding box marked in the first frame is used as the initial candidate bounding box; if it is a subsequent frame, the bounding box predicted in step (2.4) is used as the initial candidate bounding box; and combined with the target bounding box marked in the reference frame as the reference bounding box, a set of candidate bounding boxes is generated around the final predicted target center position obtained in step (2.2).

[0188] The specific implementation process of step (2.3) is as follows:

[0189] (2.3.1) Based on step (2.2), the final predicted target center position is obtained. The target bounding box of the previous frame is used as the initial candidate bounding box. If the current frame is the second frame, the target bounding box marked in the first frame is used as the initial candidate bounding box. If it is a subsequent frame, the bounding box predicted in step (2.4) is used as the initial candidate bounding box. Around the final predicted target center position, a candidate bounding boxes with different proportions are randomly generated to form the initial candidate bounding box set. In this embodiment, a is 10.

[0190] (2.3.2) Based on the inherent characteristics of pedestrian targets during movement, the scale changes of pedestrian targets are relatively gentle during movement. Combined with the target bounding box of the manually annotated reference frame in step (1.1), it is used as a reference candidate bounding box. Around the center position of the final predicted target, b reference candidate bounding boxes with different proportions are randomly generated to form a reference candidate bounding box set; in this embodiment, b is 4.

[0191] (2.3.3) The initial candidate bounding box set obtained in step (2.3.1) and the reference candidate bounding box set obtained in step (2.3.2) are fused to obtain a+b candidate bounding boxes as the candidate bounding box set;

[0192] In step (2.3.1), the value of 'a' ranges from 7 to 15; in step (2.3.2), the value of 'b' ranges from 3 to 7. In this embodiment, 'a' = 10 and 'b' = 4 are the optimal parameters obtained through experiments.

[0193] (2.4) The a+b optimizable candidate bounding boxes obtained in step (2.3) and the features of the test frame obtained in step (1.3) are fed into the bounding box prediction module to predict the target bounding box.

[0194] The target bounding box prediction process in step (2.4) includes the following steps:

[0195] (2.4.1) Since the specified tracking target is uncertain before annotation, the bounding box prediction process needs to combine the target information annotated in the reference frame. Therefore, it is necessary to extract the target information annotated in the reference frame, i.e.: Figure 4 As shown, the ResNet50 feature extraction network in step (1.1) is used to extract features from the reference frame. The reference frame is fed into the ResNet50 feature extraction network and processed through four residual blocks in sequence. The output layer1 of Block3 and the output layer2 of Block4 of the reference frame are extracted as features of the reference frame. These features are then fused after convolution and Pr Pooling and then passed through a fully connected layer to obtain the modulation vector, which can be used as the pedestrian target information labeled in the reference frame.

[0196] (2.4.2) For the test frame, such as Figure 4 As shown in the test frame branch, the features of the test frame obtained by ResNet50 in the feature extraction network in step (1.1) are used and passed through two convolutional layers. Then, PrPooling is performed on a+b (14 in the example) bounding box regions to extract internal features. Combined with the modulation vector, i.e. the information of the pedestrian target marked in the reference frame, the intersection-union ratio (IoU) of the a+b (14 in the example) bounding boxes is predicted through a fully connected layer. The gradient of the bounding box is calculated through the IoU. The a+b (14 in the example) candidate bounding boxes are optimized to obtain the optimized candidate bounding boxes, which are used as a new set of candidate bounding boxes that can be optimized.

[0197] (2.4.3) Repeat step (2.4.2) for iterative optimization. After 5 iterations, take the average coordinates of the three candidate bounding boxes with the largest intersection-union ratio (IoU) as the coordinates of the predicted bounding box, which is the final predicted target bounding box.

[0198] In this embodiment, after 5 iterations, the average coordinates of the three candidate bounding boxes with the largest IoU are taken as the coordinates of the predicted bounding box, which is the final predicted target bounding box.

[0199] (2.5) Update the classification filter new_filter according to the tracking status of the current test frame determined in step (1.5). During the movement of pedestrians, there will be scenes of occlusion, pedestrians moving in different directions and blending into the background. By optimizing and updating the classification filter new_filter according to the tracking status of the current target, the introduction of background or interference information can be prevented.

[0200] In step (2.5), the classification filter new_filter is optimized and updated based on the tracking state of the test frame as follows:

[0201] (2.5.1) When the tracking state of the current test frame is determined to be normal tracking state, the classification filter new_filter is optimized and updated every n frames or when an interference is encountered;

[0202] In step (2.5.1), the value of n ranges from 15 to 25. In this embodiment, n = 20. Since updating the classification filter is time-consuming, frequent updates can lead to a significant decrease in speed and introduce more background information, affecting the performance of the final overall model. Since the video is dynamically changing, if it is not updated for a long time, the classification effect of the classification filter will deteriorate, affecting the performance of the final overall model. The optimal choice of 20 frames was obtained through experimental results.

[0203] (2.5.2) When the tracking state of the current test frame is determined to be uncertain, the classification filter new_filter is not optimized or updated;

[0204] (2.5.3) When the tracking state of the current test frame is determined to be "unfinished", the classification filter new_filter will not be optimized or updated.

[0205] The optimization and update method for the classification filter new_filter in step (2.5.1) utilizes the test frame response map s2 and test frame labels obtained in step (1.4). In this implementation, the response graph s2 of the test frame is represented as a 19×19 two-dimensional matrix:

[0206] (i) Scale the bounding box predicted by step (2.4.3) onto a 19×19 two-dimensional matrix, and then the position c of the pedestrian target in the test frame in the 19×19 two-dimensional matrix can be obtained;

[0207] (ii) Calculate the target mask parameters m according to step ② in step (1.2.3). c and tag y hn Calculate the target mask parameters of the test frame using the following method and tags

[0208] (iii) Calculate the test frame response graph s using formula (2). test and test frame labels The residual between them is r(s,c); at this time, in equation (2), s is the response graph of the test frame. test ;v c For spatial weights; m c Target mask parameters for the test frame y hn Labels for test frames

[0209] (iv) Following the steps in step ③ of (1.2.3), obtain the test frame response diagram s. test and test frame labels The loss difference between the two is used to optimize and update the classification filter new_filter, resulting in a new optimized classification filter new_filter used for target foreground and background classification in subsequent test frames.

[0210] (2.6) By repeating steps (1.3) to (2.5), target recognition and detection are completed for each frame in the entire tracked video sequence, and single-target tracking of pedestrians is finally achieved.

[0211] The specific process of step (2.6) is as follows:

[0212] (2.6.1) The information of the selected target pedestrian in the reference frame of the tracked video sequence is obtained through steps (1.1) to (1.2), and the resulting classification filter new_filter is used for target foreground and background classification in subsequent test frames;

[0213] (2.6.2) Each subsequent test frame repeats steps (1.3) to (2.5) until the last frame, thereby completing the detection of the specified pedestrian target in each frame of the entire video sequence, and finally realizing the tracking of a single specified pedestrian target in the reference frame in the entire video sequence.

[0214] This embodiment utilizes the Python language and PyTorch framework to construct an online update strategy for single-target pedestrian tracking that incorporates pedestrian characteristics.

[0215] The main implementation operations involved are classification tasks and regression tasks. The classification task adopts a new update strategy, while the regression task corrects the target center position and adds candidate bounding boxes generated based on the first frame bounding box, which are our innovative points.

[0216] The video sequence is used as input, and the manually annotated pedestrian bounding boxes in the first frame serve as the pedestrian targets to be tracked in subsequent frames. The first frame image is augmented using data augmentation operations such as flipping, mirroring, and offsetting to obtain a training image set. The corresponding manually annotated pedestrian bounding boxes are then obtained according to the augmentation method, serving as the training sample set. Feature extraction from the training sample set is performed using a backbone network, which consists of a ResNet50 network pre-trained on the ImageNet dataset and a two-layer convolutional network trained offline. After processing, the training sample feature set is obtained. The training sample feature set is fed into the model predictor to predict the classification filter. First, the initializer module in the model predictor uses only the target appearance, i.e., the manually labeled bounding box, to effectively provide the initial estimate of the classification filter through Pr Pooling. Second, the initially estimated classification filter init_filter is extracted into the optimizer module. The optimizer module optimizes the initially estimated classification filter init_filter using the steepest gradient descent method. Experiments show that iterative optimization is performed 5 times with a learning rate of 0.6, i.e., setting i to 5 and α to 0.6 in equation (7) yields better results. Finally, the optimized classification filter new_filter is obtained. In subsequent frames of the current video sequence, the target object labeled in the reference frame is identified and detected, thereby achieving the tracking effect of a single specified pedestrian target in the current video sequence. During the tracking process of subsequent frames, the subsequent frames are used as test frames. The Siamese network concept is used to extract features by sharing the same backbone network with the reference frame. The extracted features are convolved with the optimized classification filter new_filter to obtain the response map. The predicted target center position is obtained according to step (2.1), and the predicted target center position is corrected according to the current tracking state to obtain the final predicted target center position.

[0217] The target bounding box size from the previous frame is predicted. If the current test frame is the second frame, the manually annotated bounding box size from the reference frame is used as the initial candidate bounding boxes. Then, 14 candidate bounding boxes are randomly generated at the final predicted target center position, combined with the manually annotated bounding box size. The IoU (Interchange of Value) of these 14 candidate bounding boxes is calculated through a regression branch. Gradient descent is used to directly optimize these 14 candidate bounding boxes. Finally, the average coordinates of the three candidate bounding boxes with the highest IoU are taken as the final predicted bounding box, completing the target recognition and detection for the current test frame. Then, recognition and detection are performed on each subsequent test frame of the current video sequence until the last frame, thus completing the recognition and detection of a specified pedestrian target in each frame of the entire video sequence, ultimately achieving tracking of a single specified pedestrian target throughout the entire video sequence.

[0218] Table 1 presents a comparison of the tracking performance of our method with other methods on pedestrian video sequence datasets. Our method achieves 30 FPS on a GTX1650, realizing real-time tracking. Our method combines accuracy and speed, and has certain practical value.

[0219] Ours ATOM DIMP SiamCAR SiamDW SiamRPN++ Ocean Success rate 0.740 0.696 0.681 0.655 0.650 0.633 0.617

[0220] Table 1.

Claims

1. A pedestrian single-target tracking method with an online update strategy that integrates pedestrian characteristics, characterized in that... It includes the following steps: Step 1: Classification task during the tracking process: (1.1) The features of the reference frame and each frame in the video sequence are extracted by the feature extraction network. That is, the first frame of the video sequence is selected as the reference frame, and the target to be tracked is selected by manually marking the target bounding box. In each subsequent frame of the video sequence, i.e. the test frame, the target to be tracked selected in the reference frame is identified and detected, so as to realize the tracking of the target in the video sequence, and thus realize the single target tracking process of pedestrians. (1.2) Using the reference frame information of the currently tracked video sequence determined in step (1.1), including the reference frame image information and the manually specified target annotation bounding box, flip, mirror, blur and rotate data augmentation operations are performed on the reference frame respectively, and images with the respective operation effects are obtained. These images are combined into a set as the data augmentation image set, and the corresponding annotation bounding box after data augmentation is obtained. Then, the classification filter is trained online through the model predictor to distinguish the foreground and background of the target to be tracked in subsequent frames and predict the target position. (1.3) Take the subsequent frames in the current sequence to be tracked as test frames and perform feature extraction. That is, use a feature extraction network to extract features from the test frames to obtain the features x of the test frames. test ; (1.4) Perform convolution processing using the classification filter new_filter optimized in step (1.2) and the features of the test frame obtained in step (1.3), as shown in formula (8): s test =x test *f new (8) Where s test This is the test frame response graph; where x test Features representing the test frame; * represents convolution calculation; f new The optimized classification filter is new_filter; (1.5) The response graph s of the test frame obtained in step (1.4) test The system determines the tracking status of the current test frame, which is divided into three states: normal, uncertain, and undetectable. The normal state indicates that the scene of the test frame is simple, and the target and background to be tracked can be accurately identified through the classification task. The uncertain state indicates that the scene of the current test frame is complex, affected by interference and background, making it difficult to accurately identify the target and background to be tracked. The undetectable state indicates that the scene of the current test frame is complex, occluded, or the target and background to be tracked cannot be identified. Step 2: Regression task during the tracking process: (2.1) The response graph s of the test frame obtained from the classification task in step (1.4) test Predict the location of the target center; (2.2) Combining the inherent characteristics of the pedestrian target during the movement process, the target center position predicted in step (2.1) is modified to obtain the final predicted target center position; (2.3) Based on the final predicted target center position obtained in step (2.2) and the target bounding box of the previous frame, if the current frame is the second frame, the target bounding box labeled in the first frame is used as the initial candidate bounding box; if it is a subsequent frame, the bounding box predicted in step (2.4) is used as the initial candidate bounding box. Then, based on the initial candidate bounding box, a candidate bounding boxes with different proportions are randomly generated around the final predicted target center position obtained in step (2.2), thus forming an initial candidate bounding box set. At the same time, based on the target bounding box of the manually labeled reference frame obtained in step (1.1), it is used as a reference candidate bounding box. Combining the inherent characteristics of the pedestrian target in the motion process, namely the relatively gentle scale change of the pedestrian target in the motion process, b reference candidate bounding boxes with different proportions are randomly generated around the final predicted target center position obtained in step (2.2), thus forming a reference candidate bounding box set. (2.4) The a+b optimizable candidate bounding boxes obtained in step (2.3) and the features of the test frame obtained in step (1.3) are fed into the bounding box prediction module to predict the target bounding box. (2.5) Update the classification filter new_filter based on the tracking status of the current test frame determined in step (1.5). During the movement of pedestrians, there will be scenes of occlusion, staggered movement between pedestrians, and blending into the background. By optimizing and updating the classification filter new_filter by judging the current tracking status of the target, the introduction of background or interference information can be prevented. (2.5.1) When the tracking state of the current test frame is determined to be normal tracking state, the classification filter new_filter is optimized and updated every n frames or when an interference is encountered: (i) Scale the bounding box predicted by step (2.4) onto a two-dimensional matrix, and then the target position c of the pedestrian target in the test frame in the two-dimensional matrix can be obtained; (ii) Calculate the target mask parameters of the test frame and tags In the formula, t represents each position of the two-dimensional matrix, and c represents the target position; The spatial coefficients, ρ, are obtained during training. k It is a distance calculation function: (iii) Calculate the test frame response graph s test and test frame labels The residual r(s,c) between them is: r(s,c)=v c ·(m c s+(1-m c )max(0,s)-y hn ) At this point, in equation (2), s is the response graph of the test frame. test ;v c For spatial weights; m c Target mask parameters for the test frame y hn Labels for test frames (iv) Following the steps in step ③ of (1.2.3), obtain the test frame response diagram s. test and test frame labels The loss difference between the two is used to optimize and update the classification filter new_filter, resulting in a new optimized classification filter new_filter for target foreground and background classification in subsequent test frames. By applying regularization to the difference r(s,c) obtained in step (iii), we can obtain L(f), that is: In the formula, Let x be the training sample set, where x j c represents the training sample features extracted by the feature extraction network. j The coordinates of the target center in the sample are already labeled, i.e., the coordinates of the center point of the labeled bounding box; * represents convolution calculation; λ is the regularization factor; f is the optimized classification filter; Use L(f) as the test frame response graph s test and test frame labels The loss between the two is backpropagated to optimize the classification filter. (2.5.2) When the tracking state of the current test frame is determined to be uncertain, the classification filter new_filter is not optimized or updated; (2.5.3) When the tracking state of the current test frame is determined to be "no state found", the classification filter new_filter will not be optimized or updated. (2.6) By repeating steps (1.3) to (2.5), target recognition and detection are completed for each frame in the entire tracked video sequence, and single-target tracking of pedestrians is finally achieved.

2. The pedestrian single-target tracking method based on an online update strategy incorporating pedestrian characteristics as described in claim 1, characterized in that... The feature extraction network in step (1.1) is a ResNet50 module structure; the ResNet50 module structure is composed of 4 residual blocks connected in series, and the names of the 4 residual blocks are Block1, Block2, Block3 and Block4 respectively; the ResNet50 module structure is followed by two convolutional layers, which are then combined to form a backbone network for feature extraction, used to extract features of the current reference frame or test frame image.

3. The pedestrian single-target tracking method based on an online update strategy incorporating pedestrian characteristics as described in claim 1, characterized in that... The model predictor in step (1.2) consists of an initializer module and an optimizer module. The initializer module can effectively provide an initial estimate of the classification filter using only the appearance of the target to be tracked. The optimizer module is used to optimize the initially estimated classification filter init_filter to obtain the optimized classification filter new_filter, which performs target foreground and background classification on subsequent frames of the tracked video sequence and predicts the center position of the target to be tracked for coarse localization.

4. The pedestrian single-target tracking method based on an online update strategy incorporating pedestrian characteristics as described in claim 3, characterized in that... The step (1.2) described above uses the information from the reference frame to train the classification filter new_filter online through the model predictor. This process consists of the following steps: (1.2.1) Using the reference frame information of the currently tracked video sequence, including the reference frame image information and the manually specified target annotation bounding box, flip, mirror, blur and rotate data augmentation operations are performed on the reference frame respectively, and images with the respective operation effects are obtained. These images are combined into a set as the image set after data augmentation, and the corresponding annotation bounding box after data augmentation is obtained. (1.2.2) Use the feature extraction network in step (1.1) to extract features from the image set processed in step (1.2.1) to obtain a set of training sample features, and send them to the model initialization module to perform initial estimation of the classification filter filter, and obtain the initial estimated classification filter init_filter; (1.2.3) The initial estimated classification filter init_filter obtained in step (1.2.2) and the feature extraction network in step (1.1) are used to extract features from the image set processed in step (1.2.1) to obtain a set of training sample features, which are sent to the optimizer module. The optimized classification filter new_filter is obtained through optimization and used for target foreground and background classification in subsequent test frames.

5. The pedestrian single-target tracking method based on an online update strategy incorporating pedestrian characteristics as described in claim 4, characterized in that... In step (1.2.2), the initial estimation of the classification filter by the model initialization module refers to extracting the features within the labeled bounding box by performing PrRoi Pooling on the features of the training samples. The resulting features are the target features, which are then used as the output, i.e., the initially estimated classification filter init_filter. The specific implementation process of step (1.2.3) can be described by the following steps: ① Calculate the reference frame response graph s ref : The initial estimated classification filter init_filter from step (1.2.2) is convolved with the training sample features to obtain the response map s. ref The response diagram of the reference frame is shown in equation (1): s ref =x ref *f init (1) In the formula, x ref The features representing the reference frame are the same as the features of the training samples; * represents convolution calculation; f init The initial estimated classification filter is init_filter; ② Calculate the reference frame response graph s ref and reference frame labels The difference between them is r(s,c): The response graph s of the reference frame calculated according to step (1.2.2) ref Its representation is a 19×19 two-dimensional matrix; by scaling the bounding box of the pedestrian target annotation in the reference frame onto the 19×19 two-dimensional matrix, the position c of the annotated pedestrian target in the 19×19 two-dimensional matrix is ​​obtained; In formulas (2) and (3), t represents each position of a 19×19 two-dimensional matrix, and c represents the target position; The spatial coefficients, ρ, are obtained during training. k The distance calculation function is determined by formulas (4-1) and (4-2): Among them, y hn The true label represents the current frame response map; m c This represents the mask parameters, used to determine the region where the current target is located in a 19×19 two-dimensional matrix, within the region m corresponding to the target. c ≈1, in the background region m c ≈0; label y hn and mask parameters m c It is represented as a 19×19 two-dimensional matrix; The reference frame label is obtained using formulas (2) and (3). and mask parameters And calculate the response graph s of the reference frame obtained in step ①. ref and reference frame labels The difference between them, r(s,c), is shown in equation (5): r(s,c)=v c ·(m c s+(1-m c )max(0,s)-y hn ) (5) In the formula, s is the response graph of the reference frame. ref ;v c For spatial weights; m c Mask parameters for reference frame y hn Labels for reference frames ③ Regularization is applied to the difference r(s,c) obtained in step ② to obtain L(f), as shown in equation (6), and this is used as the reference frame response diagram s. ref and reference frame labels The loss between the two is backpropagated to optimize the classification filter. In the formula, Let x be the training sample set, where x j c represents the training sample features extracted by the feature extraction network. j The coordinates of the target center in the sample are already labeled, i.e., the coordinates of the center point of the labeled bounding box; * represents convolution calculation; λ is the regularization factor; f is the optimized classification filter.

6. The pedestrian single-target tracking method based on an online update strategy incorporating pedestrian characteristics as described in claim 5, characterized in that... The optimized classification filter f in step ③ is a backpropagation optimized classification filter, which adopts the steepest gradient descent method as shown in equation (7): In the formula, f (i) Represents the classification filter after the i-th optimization; α represents the learning rate; This represents gradient calculation.

7. The pedestrian single-target tracking method based on an online update strategy incorporating pedestrian characteristics as described in claim 1, characterized in that... The specific implementation of step (1.5) consists of the following steps: (1.5.1) The response graph s of the test frame obtained in step (1.4) test Determine the tracking status of the current test frame. The tracking status is divided into normal status, uncertain status, and tracking not found status. ①If the response graph s of the test frame test If only the target center has the highest response, it means that there is only the target to be tracked or there is a clear boundary between the target to be tracked and the background. That is, the tracking state of the current test frame is normal, which means that the scene of the test frame is simple and the target to be tracked and the background can be easily identified through the classification task. ②If the response graph s of the test frame test The target region is the area selected around the position of the highest response score using the target bounding box size of the previous frame. At this time, the response scores in the target region are relatively messy, and the highest response score will drop suddenly. This indicates that the target to be tracked is in an uncertain state, which means that the target to be tracked is confused with the background or close to the interference. ③ If the test frame response graph s test The overall response score is more chaotic than the uncertain state corresponding to step ②, and the highest response score value shows a significant decrease compared to the highest response score value in the uncertain state corresponding to step ②, to the point that the highest response can no longer clearly point to the original target area, or its value has approached the average level of the background response. In this case, it is determined that the target to be tracked has encountered severe occlusion and its state is determined to be undetectable. (1.5.2) Calculate the response graph s of the test frame. test The variance of the response score in the target region and the highest response score in the test frame are used to determine the state of the target to be tracked in the current search region: (1.5.2.1) Calculate the mean variance of the target region response score of the m frames preceding the test frame using a sliding method. The target region is the response map s of the test frame using the target bounding box size from the previous frame. tset The region around the highest response score position is selected, and the variance σ of the response score of the target region in the m frames prior to the test frame is recorded, as shown in Equation (9): In the formula, score i The score is assigned to each location in the target region of the response graph of the test frame. The response graph s of the test frame test The mean score of each location in the corresponding target area; w*h is the response graph s of the test frame. test The dimensions of the corresponding target area, including its width and height. (1.5.2.2) Further calculate its average value according to formula (10), which is the mean variance of the response score of m frames. Where σ j The variance of the target region response score for each frame within the target region of m frames; (1.5.2.3) Meanwhile, the highest response score of the test frame is the response graph s of the test frame. test The maximum response score is recorded, and the highest response score max_score of the m frames preceding the test frame is recorded. The average of the highest response scores of the m frames is calculated using formula (11). In the formula, max_score j Represents the highest response score per frame; (1.5.2.4) Based on the normal state, uncertain state, and unfindable state described in step (1.5.1), combined with the mean variance of the target region response score of the m frames preceding the test frame obtained in step (1.5.2.2), And the average of the highest response scores of the m frames preceding the test frame obtained in step (1.5.2.3), to determine the tracking status of the test frame: If equation (12) is satisfied, it indicates that the tracking state of the test frame is in an uncertain state: If equation (13) is satisfied, it means that the tracking overstate of the test frame is a state that cannot be found. Other situations are considered normal. in, The mean variance of the target region response score in the m frames preceding the current test frame; The response graph s of the test frame m frames prior to the test frame. test The average score of the highest response; k1 and k2 are proportionality coefficients.

8. The pedestrian single-target tracking method based on an online update strategy incorporating pedestrian characteristics as described in claim 1, characterized in that... Step (2.1) specifically refers to: (2.1.1) The response graph s of the test frame obtained in step (1.4) test The response graph of the test frame s test The position coordinates corresponding to the maximum response score are taken as the first response point. If the tracking state of the test frame is normal and there are no interfering objects, the first response point will be used as the predicted target center position. (2.1.2) In the response graph s of the test frame test The target region, i.e., the response map s of the test frame using the target bounding box size from the previous frame. test The region around the location of the maximum response score in the test frame is shown in the response graph s. test The area outside the selected area is considered the target area, and the location corresponding to the highest response score outside the target area is regarded as the second response point. (2.1.3) When the highest response score of the second response point is greater than 0.5 times the highest response score of the first response point, the second response point is considered to be a target similar to the background. (2.1.4) Assume the current position of the first response point is c1[x1,y1], the position of the second response point is c2[x2,y2], and the target center point of the previous frame of the current test frame is in the response graph s of the test frame. test The position is c0[x0,y0], which is the center point of the target area in the response diagram. Then the position offsets of c1 and c2 relative to c0 are shown in Equations (14) and (15), respectively: (2.1.5) Determine the true location of the current target: The offsets calculated according to equations (14) and (15) are within different value ranges, and different response point positions are returned as the predicted target center positions, as shown in equations (16) and (17): c1,(d1>Ω&d2<Ω) / (d1>Ω&d2>Ω) / (d1 <d2&d1<Ω&d2<Ω)(16)c2, (d1<Ω&d2> Ω) / (d1>d2&d1<Ω&d2<Ω) (17) where Ω represents the threshold value in the range of values.

9. The pedestrian single-target tracking method based on an online update strategy incorporating pedestrian characteristics as described in claim 1, characterized in that... Step (2.2) specifically refers to: based on the inherent characteristics of the pedestrian target during the movement process, under normal conditions the target tracking is normal, and the pedestrian target does not undergo drastic scale changes during the movement process, and the pedestrian movement is relatively smooth, so the offset of the target center under normal tracking conditions is recorded, and the relative offset of the target center between two frames in v frames is recorded by sliding frames, and its mean is calculated as the predicted offset, with the value of v ranging from 12 to 18; if the tracking state is in an uncertain state, the predicted target center position obtained in step (2.1) is corrected by the predicted offset, and finally the final predicted target center position is obtained.

10. The pedestrian single-target tracking method based on an online update strategy incorporating pedestrian characteristics as described in claim 1, characterized in that... The specific implementation process of step (2.3) is as follows: (2.3.1) According to step (2.2), the final predicted target center position is obtained. The target bounding box of the previous frame is used as the initial candidate bounding box. If the current frame is the second frame, the target bounding box marked in the first frame is used as the initial candidate bounding box. If it is a subsequent frame, the bounding box predicted in step (2.4) is used as the initial candidate bounding box. Around the final predicted target center position, a candidate bounding boxes with different proportions are randomly generated to form the initial candidate bounding box set, where the value of a ranges from 7 to 15. (2.3.2) Based on the inherent characteristics of pedestrian targets during movement, the scale changes of pedestrian targets are relatively gentle during movement. Combined with the target bounding box of the manually annotated reference frame in step (1.1), it is used as a reference candidate bounding box. Around the center position of the final predicted target, b reference candidate bounding boxes with different proportions are randomly generated to form a reference candidate bounding box set; where the value of b is in the range of 3-7. (2.3.3) The initial candidate bounding box set obtained in step (2.3.1) and the reference candidate bounding box set obtained in step (2.3.2) are fused to obtain a+b candidate bounding boxes as the candidate bounding box set; The target bounding box prediction process in step (2.4) includes the following steps: (2.4.1) Since the specified tracking target is uncertain before annotation, the bounding box prediction process needs to combine the target information annotated in the reference frame. Therefore, it is necessary to extract the target information annotated in the reference frame. That is, the feature extraction network ResNet50 in step (1.1) is used to extract the features of the reference frame. The reference frame is fed into the ResNet50 feature extraction network and processed through 4 residual blocks in sequence. The output layer1 of Block3 and the output layer2 of Block4 of the reference frame are extracted as the features of the reference frame. These features are then fused after convolution and PrPooling, and after passing through a fully connected layer, the modulation vector is obtained, which can be used as the pedestrian target information annotated in the reference frame. (2.4.2) For the test frame, the features of the test frame obtained by ResNet50 in the feature extraction network in step (1.1) are processed through two convolutional layers, and then PrPooling is performed on a+b bounding box regions to extract internal features. Combined with the modulation vector, i.e. the information of the pedestrian target marked in the reference frame, the intersection-union ratio (IoU) of a+b bounding boxes is predicted through a fully connected layer. The gradient of the bounding box is calculated through the IoU, and the a+b candidate bounding boxes are optimized to obtain the optimized candidate bounding boxes, which are used as a new set of candidate bounding boxes that can be optimized. (2.4.3) Repeat step (2.4.2) for iterative optimization. After 5 iterations, take the average coordinates of the three candidate bounding boxes with the largest intersection-union ratio (IoU) as the coordinates of the predicted bounding box, which is the final predicted target bounding box. In step (2.5), the classification filter new_filter is optimized and updated based on the tracking state of the test frame as follows: (2.5.1) When the tracking state of the current test frame is determined to be normal tracking state, the classification filter new_filter is optimized and updated every n frames or when an interference object is encountered; where n ranges from 15 to 25. The optimized update method for the classification filter new_filter is: (i) Scale the bounding box predicted by step (2.4.3) onto a 19×19 two-dimensional matrix, and then the position c of the pedestrian target in the test frame in the 19×19 two-dimensional matrix can be obtained; (ii) Calculate the target mask parameters m according to step ② in step (1.2.3). c and tag y hn Calculate the target mask parameters of the test frame using the following method. and tags (iii) Calculate the test frame response graph s using formula (5). test and test frame labels The residual between them is r(s,c); at this time, in equation (5), s is the response graph of the test frame. test ;v c For spatial weights; m c Target mask parameters for the test frame y hn Labels for test frames (iv) Following the steps in step ③ of (1.2.3), obtain the test frame response diagram s. test and test frame labels The loss difference between the two is used to optimize and update the classification filter new_filter, resulting in a new optimized classification filter new_filter for target foreground and background classification in subsequent test frames. (2.5.2) When the tracking state of the current test frame is determined to be uncertain, the classification filter new_filter is not optimized or updated; (2.5.3) When the tracking state of the current test frame is determined to be "no state found", the classification filter new_filter will not be optimized or updated. The specific process of step (2.6) is as follows: (2.6.1) The information of the selected target pedestrian in the reference frame of the tracked video sequence is obtained through steps (1.1) to (1.2), and the resulting classification filter new_filter is used for target foreground and background classification in subsequent test frames; (2.6.2) Each subsequent test frame repeats steps (1.3) to (2.5) until the last frame, thereby completing the detection of the specified pedestrian target in each frame of the entire video sequence, and finally realizing the tracking of a single specified pedestrian target in the reference frame in the entire video sequence.

Citation Information

Patent Citations

  • Pedestrian tracking method and device and terminal

    CN110443210A

  • Online multi-target tracking method based on motion model and single-target clue

    CN111639570A