Pedestrian small target tracking method based on multilevel hierarchical data association
By employing a multi-level hierarchical data association method and utilizing depthwise separable convolution and nonlinear Kalman filtering algorithms, the problems of insufficient model complexity and accuracy in pedestrian small target tracking are solved, achieving efficient small target detection and tracking.
Patent Information
- Application Number
- CN202510964631.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-11-28
AI Technical Summary
Existing technologies for small pedestrian tracking suffer from problems such as complex model structure, large number of parameters, low small target detection accuracy, insufficient stability of prediction boxes, and insufficient matching accuracy between detection boxes and prediction boxes, resulting in poor performance of small pedestrian tracking on mobile devices.
A multi-level hierarchical data association method is adopted. The number of parameters is reduced by depthwise separable convolution and lightweight attention module, the receptive field is enhanced by dynamic weight allocation dilated convolution, and the Kalman gain is adjusted by nonlinear adaptive Kalman filtering algorithm. A lightweight re-identification network is introduced and a multi-level hierarchical data association method is adopted to perform data association in different ways for high and low confidence detection boxes.
It improves the accuracy of small target detection and trajectory prediction, reduces the number of target ID switching, and enhances the accuracy and real-time performance of target tracking.
Smart Images

Figure CN121033745A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of pedestrian target tracking, and in particular to a pedestrian small target tracking method based on multi-level hierarchical data association. BACKGROUND
[0002] Multi-target pedestrian tracking is widely used in monitoring security, autonomous driving, intelligent transportation and other fields, and therefore has high requirements for tracking accuracy and real-time performance. With the development of deep learning, multi-target tracking technology has developed rapidly. Methods based on TBD and JDE discard low bounding boxes with confidence less than a certain threshold after obtaining target detection results. These low bounding boxes may be caused by object occlusion, and direct discarding will result in a large number of missed detections and track interruptions, affecting the performance of target tracking. In view of this problem, Zhang et al. proposed the ByteTrack algorithm, which uses YOLOX as a detector to improve detection accuracy and uses low bounding boxes to improve data association methods without appearance feature matching, achieving fast tracking performance and achieving good results in the MOT17 dataset.
[0003] Although ByteTrack has superior performance, it has problems such as large parameter quantity, low small target detection accuracy, insufficient prediction box stability, and insufficient matching accuracy of detection boxes and prediction boxes, which result in poor performance in mobile small target pedestrian tracking. SUMMARY
[0004] In view of the above problems, the present application proposes a pedestrian small target tracking method based on multi-level hierarchical data association.
[0005] To achieve the above purpose, the present application adopts the following technical solutions:
[0006] A pedestrian small target tracking method based on multi-level hierarchical data association, comprising:
[0007] Step 1: Construct a pedestrian small target tracking method based on multi-level hierarchical data association; the method is composed of a feature extraction module DHU, a trajectory prediction module NSKF and a data association module MHDA; the DHU replaces ordinary convolution with depth separable convolution, introduces a lightweight attention module, reduces the parameter quantity while ensuring the accuracy of the model, uses dynamic weight allocation atrous convolution and improves the small target detection capability through a double fusion upsampling module; the NSKF adapts to the nonlinear system equation, adjusts the Kalman gain according to the target detection confidence and motion characteristics, and improves the trajectory prediction accuracy; the MHDA proposes a multi-level hierarchical data association method, uses cosine distance as the cost matrix for high confidence bounding boxes, and uses Mahalanobis distance as the cost matrix for low confidence bounding boxes for data association;
[0008] Step 2: track the small target of the pedestrian based on the constructed multi-level hierarchical data association.
[0009] Further, the deep separable convolution is used for feature extraction of each frame of the input video.
[0010] Further, after the deep separable convolution obtains the C5 feature layer, the dynamic weight distribution hole convolution is used to strengthen the small target receptive field and reduce the grid effect, and the inflation coefficients are 1, 2 and 3 respectively, so as to obtain a feature map with a size of 72x72x512.
[0011] Further, the CAE module is added after the 72x72x512 feature map; the CAE module performs global pooling, feature fusion and encoding, and attention weight generation on the 72x72x512 feature map, and then outputs D5 after weighting.
[0012] Further, the double fusion upsampling module includes deconvolution and sub-pixel convolution, and after the D5 feature map is fused through the deconvolution and sub-pixel convolution, a 288x288x64 high-resolution feature map P5 is obtained.
[0013] Further, a small target detection method based on anchor-free frame is used to detect the feature map P5, and the target detection confidence is obtained.
[0014] Further, the NSKF module includes a nonlinear Kalman filter system and a Kalman gain adjustment module.
[0015] Further, the linear Kalman filter system is transformed by UT, the prior distribution sampling points and their weights are obtained, the nonlinear function transmission is performed based on the state equation to obtain the state prediction mean and covariance, the UT transformation is used again, the predicted observation is obtained based on the nonlinear function transmission of the observation equation, and the observation prediction mean, variance and covariance are obtained by weighting, so as to obtain the nonlinear Kalman filter system.
[0016] The Kalman gain is adjusted according to the target detection confidence and the motion feature, and the effective confidence is calculated according to the following formula:
[0017]
[0018] Wherein C d is the target detection confidence, Δt is the time interval from the last high confidence, τ is the confidence decay constant, κ is the enhancement coefficient; the motion mutation index is calculated according to the following formula, and the motion feature is extracted:
[0019]
[0020] Wherein is the acceleration vector at time t, ωt is the angular velocity at time t, lambda is a rotation sensitive coefficient, and epsilon is a zero prevention constant;
[0021] The detection confidence and the motion feature are fused to construct a dual-mode decision mechanism, and Kalman gain is adjusted.
[0022] Further, the NSKF module is further provided with a re-identification module OSNet, and the OSNet uses a 3*3 convolution kernel.
[0023] Further, the MHDA module matches and updates the existing track with the detection frame.
[0024] Further, the MHDA module divides the detection result into high and low detection frames according to the DHU module, and the high and low detection frame threshold is set to 0.7.
[0025] The high-score detection frame is first associated with the existing track, the data association mode uses a cosine distance to construct an appearance similarity matrix for matching, the first data association is not successful in matching the track and the low-score detection frame for secondary association, and a Mahalanobis distance is used for position information matching.
[0026] Further, the cosine distance is calculated according to the following formula:
[0027]
[0028] Wherein, d(i,j) is the appearance matching degree between the i th prediction frame and the j th detection frame track, r j is a detection frame appearance feature vector, R i is a set of feature vectors of the nearest target multi-frame matching success, is the m th feature vector in the set R i
[0029] Further, the Mahalanobis distance is calculated according to the following formula:
[0030]
[0031] Wherein, d(i,j) is the motion matching degree between the i th prediction frame and the j th detection frame track, d j is the state vector of the j th detection frame, x i and s i are respectively the state vector and the covariance matrix of the i th prediction frame of the current frame predicted by the nonlinear adaptive Kalman filter.
[0032] Further, the existing track and the detection frame are matched and updated after the target tracking track.
[0033] Compared with the prior art, the present application has the beneficial effects:
[0034] (1) In order to solve the problems of complex model structure, excessive parameters and low detection accuracy of small targets, the DHU module is designed, the depth separable convolution is used and the lightweight attention module is introduced, so that the network parameters and calculation amount are reduced while the network accuracy is ensured; the complex structure of the FPN feature pyramid in YOLOX is abandoned, and the dynamic weight allocation dilated convolution is used to obtain more receptive fields that are beneficial to small targets; the sub-pixel convolution and the deconvolution are used to fuse the up-sampling and improve the resolution of the feature map, so that the small target detection capability is improved.
[0035] (2) In order to solve the problem that the noise matrix in the ordinary Kalman filter algorithm cannot adapt to the state change of the nonlinear system, the nonlinear adaptive Kalman filter (NSKF) algorithm is proposed, the lossless transformation is combined with the standard Kalman system to adapt to the nonlinear system equation, the Kalman gain is adjusted based on the target detection confidence and the motion feature, and the trajectory prediction accuracy is improved.
[0036] (3) In order to solve the problems of similar target error matching and low confidence shield target identity switching in ByteTrack, the lightweight re-identification network OSNet is introduced, and the multi-level hierarchical data association method (MHDA) is proposed, the appearance similarity cosine distance is added as the data association method for the high confidence detection frame matched for the first time; the motion position information is considered for the low confidence detection frame matched for the second time, the Mahalanobis distance is calculated, the target ID switching times are reduced, and the target tracking accuracy is improved. BRIEF DESCRIPTION OF DRAWINGS
[0037] Figure 1 The pedestrian small target tracking method based on multi-level hierarchical data association provided by the embodiment of the application is provided with an implementation flowchart;
[0038] Figure 2 The network structure diagram of the pedestrian small target tracking method based on multi-level hierarchical data association provided by the embodiment of the application is provided;
[0039] Figure 3 The dynamic weight allocation dilated convolution implementation flowchart provided by the embodiment of the application is provided;
[0040] Figure 4 The DHU module structure diagram provided by the embodiment of the application is provided;
[0041] Figure 5 The nonlinear Kalman filter method flowchart of the NSKF module provided by the embodiment of the application is provided;
[0042] Figure 6 The Kalman gain adaptive adjustment method flowchart provided by the embodiment of the application is provided;
[0043] Figure 7A flowchart of a MHDA module data association method provided by the embodiment of the present application is shown in the figure.
[0044] Figure 8 A result graph of each algorithm effect visualization provided by the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0045] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0046] Embodiment 1
[0047] In view of the problems of large parameter quantity of the current target tracking algorithm, difficulty in identifying small targets, and frequent trajectory interruption and identity switching caused by irregular motion, an embodiment of the present application provides a pedestrian small target tracking method based on multi-level hierarchical data association. First, a feature extraction (DHU) module is designed. Through a depth separable convolution lightweight model, a dynamic weight allocation atrous convolution is used to strengthen the small target receptive field, and an up-sampling double fusion module is constructed to improve the small target resolution. Second, a nonlinear adaptive Kalman filter (NSKF) algorithm is proposed to solve the problems that the ordinary Kalman filter (KF) algorithm cannot adapt to the nonlinear Gaussian system and the Kalman gain is fixed. Finally, a multi-level hierarchical data association method (MHDA) is proposed to design different cost matrices for data association for high and low confidence detection boxes. The overall flow of the pedestrian small target tracking method based on multi-level hierarchical data association is shown in the figure. Figure 1 The overall structure of the present application is shown in the figure. Figure 2
[0048] Embodiment 2
[0049] In the embodiment of the present application, the ordinary convolution in feature extraction is improved. The depth separable convolution is used to replace the original ordinary convolution, thereby forming a new lightweight extraction network. The generated C5 feature map is obtained through a dynamic weight allocation atrous convolution to obtain a receptive field more suitable for small target detection.
[0050] As shown in the figure, the dynamic weight allocation atrous convolution uses a 3-group atrous convolution parallel structure, a dynamic weight allocation module, and a feature weighting fusion module to extract target features. Figure 3 As an implementable manner, as shown in the figure, the specific implementation steps of the dynamic weight allocation atrous convolution are as follows:
[0051] Figure 3
[0052] The C5 feature map is processed by 3-branch hole convolution, the 3 branches are in parallel structure, and the expansion coefficients are 1, 2 and 3 respectively; a global average pooling operation is performed on the feature maps input by the 3 branches, the operation calculates the average value of all elements in each channel along the spatial dimension (height and width), and outputs a compressed feature vector; the compressed feature vector is input into a first full connection layer, the output dimension of the first full connection layer is set to one quarter of the input channel number, the output of the first full connection layer is processed by using a ReLU activation function, the activated features are input into a second full connection layer, the output dimension of the second full connection layer is fixed to 3, corresponding to the 3 branches, and the output of the second full connection layer is subjected to Softmax normalization processing, so that the three output values are converted into a probability distribution, ensuring that the sum of the three weight values is always equal to 1, through the above operation, the 3-branch hole convolution dynamically allocates weights, realizes small target feature extraction of the C5 feature map, and obtains 3 branch feature maps.
[0053] The 3 branch feature maps are weighted and fused, first, padding operation is performed to ensure that the feature maps of each branch have the same height and width, channel-level weighting is performed on each branch feature map, the feature map of branch 1 is multiplied by the corresponding weight value, the feature map of branch 2 is multiplied by the corresponding weight value, and the feature map of branch 3 is multiplied by the corresponding weight value, the three weighted feature maps are spliced in the channel dimension, the splicing method is that the three weighted feature maps are connected along the channel direction, and a fused feature map is output, the number of channels of the fused feature map is equal to the sum of the number of channels of the three branches, and finally a feature map with a size of 72*72*512 is obtained.
[0054] The embodiment of the application reduces the model parameter quantity by applying the depth separable convolution, lightens the model, discards the complex feature pyramid structure, uses the dynamic weight allocation hole convolution to obtain the small target receptive field, improves the small target detection precision, and further lightens the algorithm.
[0055] Embodiment 3
[0056] As shown in FIG. 2, on the basis of embodiment 2, the 72*72*512 feature map is input into a CAE module to obtain a feature map D5, so as to further improve the quality of small target feature representation and enhance the modeling ability of the model to context information. Figure 4 To avoid that the low-resolution feature map affects the small target detection effect in the detection stage, the D5 feature map is input into a double fusion upsampling module to improve the feature resolution, the double fusion upsampling module combines the deconvolution and sub-pixel convolution, and the working principle of the deconvolution is shown in formula 1:
[0057] out_size=stride*(in_size-1)+kernel-2*padding (1)
[0058] In this invention, stride=2 and kernel_size=4 are set. When the feature map is passed through the feature extraction module, the feature map size is 72×72×512. After deconvolution, a high-resolution feature map of 288×288×64 is obtained, which is more suitable for the detection of small targets.
[0059] The zero-padding region of deconvolution is invalid information and is not conducive to gradient optimization. The sub-pixel convolution can obtain a high-resolution feature map with less information loss.
[0060] This invention improves the model's attention to small targets through a CAE attention mechanism. The obtained D5 feature map is then processed by a dual fusion upsampling module to obtain a high-resolution feature map P5, which is beneficial for anchor-free detection in the next embodiment.
[0061] Example 4
[0062] Based on Example 3, the P5 feature map is subjected to the anchorless detection. Three convolutions are performed on the P5 feature layer: heat map prediction, the final result of which represents whether there is an object at each heat point and the type of object; center point prediction, the final result of which represents the offset of the center of each object from the heat point; and width and height prediction, the final result of which represents the predicted width and height of each object.
[0063] like Figure 2 As shown, the detection box confidence is obtained through this embodiment and used in the NSKF and MHDA modules described later.
[0064] The anchor-free detection strategy of this invention adaptively adjusts the anchor frame according to the target size, making it more suitable for small target detection.
[0065] Example 5
[0066] The embodiments of the present invention provide a detailed design for the NSKF module, such as... Figure 5 As shown, the linear system of the Kalman filter algorithm undergoes a Time-Lapse (UT) transformation to obtain the prior distribution sampling points and their weights. Based on the state equation, a nonlinear function transfer is performed to obtain the state prediction mean and covariance. The UT transformation is then applied again, and a nonlinear function transfer is performed based on the observation equation to obtain the predicted observations. The observation prediction mean, variance, and covariance are obtained through weighted summation. Finally, the Kalman filter gain is calculated to update the system's state estimate and estimated variance. The Kalman gain adaptive adjustment process is as follows: Figure 6 As shown.
[0067] Calculate the effective confidence level according to Formula 2:
[0068]
[0069] In the formula C dLet be the target detection confidence level, Δt be the time interval from the last high confidence level, τ be the confidence level decay constant, and κ be the enhancement coefficient.
[0070] The motion mutation index is calculated using the following formula to extract motion features:
[0071]
[0072] In the formula Let ω be the acceleration vector at time t. t Let ω be the angular velocity at time t, λ be the rotational sensitivity coefficient, and ε be the zero constant.
[0073] A dual-modal decision matrix is constructed by fusing the detection confidence and the motion features, and the Kalman gain is calculated according to Formula 4:
[0074] K = η·K b ⊙W s (4)
[0075] In the formula, η is the gain scaling factor, and K b W represents the current Kalman gain. s This is the state-related weight matrix.
[0076] This invention addresses the adaptability of Kalman filtering to nonlinear systems and, by combining the detection box confidence and target motion characteristics described in the previous embodiment, adaptively adjusts the Kalman gain to improve Kalman prediction accuracy.
[0077] Example 6
[0078] This invention embodiment details the MHDA module, which associates the original trajectory with the detection box to update the target tracking trajectory. The specific process is as follows:
[0079] First, high-confidence detection boxes are used for the first data association. Since the EIoU distance does not consider the target similarity features, it is easy to cause the problem of incorrect matching of similar or small targets. Therefore, cosine distance is used to construct an appearance similarity matrix to retain targets with similar appearance features, and then the matching accuracy is improved by EIoU distance.
[0080] Trajectories that failed to match in the first data association are re-associated with low-confidence detection boxes. Since the target appearance features of low-confidence detection boxes are not easily identifiable, Mahalanobis distance is used for position information matching. Target boxes whose position changes do not conform to the motion trend are discarded. Then, EIoU distance is used to improve the matching accuracy.
[0081] Unmatched tracks will be deleted, and a limit of 25 frames will be set for retaining unmatched tracks. The overall data association process is as follows: Figure 7 As shown.
[0082] In this embodiment of the invention, the target detection box is classified according to its confidence level, and an appropriate data association method is selected to improve data matching accuracy and target tracking performance.
[0083] To verify the effectiveness of the pedestrian small target tracking method based on multi-level hierarchical data association provided by this invention, the present invention also provides the following experiments, as detailed below:
[0084] 1. Dataset preparation: This invention is evaluated on three benchmark datasets to ensure comprehensive validation in different scenarios: MOT17 consists of 14 sequences and 1325 frames, representing moderate pedestrian density; MOT20 consists of 8 sequences and 13410 frames, capturing extreme crowd density and severe occlusion under low light conditions; VisDrone consists of 167 sequences and 263908 frames, all of which are small targets on drone platforms.
[0085] 2. Experimental environment and parameter settings: The experimental environment is shown in Table 1.
[0086] Table 1 Experimental Environment
[0087]
[0088] This invention uses Python 3.8 and CUDA 11.1. Hyperparameter settings: The model is trained using a momentum-based stochastic gradient descent algorithm with a momentum factor of 0.9, a weight decay coefficient of 0.0005, a batch size of 6, and a total training time of 50 epochs. The initial learning rate is 0.0001, the optimizer is SGD, and a cosine annealing strategy is used. The first epoch serves as a warm-up, the learning rate is reduced to 0.00001 after 30 epochs, and data augmentation is disabled for the last 10 epochs. The input image size is 800×1440, and the training cycle is 45 hours.
[0089] 3. The official evaluation metrics provided by MOT Challenge, namely FP, FN, IDS, FPS, MOTA, and IDF1, are adopted. To verify the effect of the model lightweight improvement, this invention adds Params and GFLOPs as evaluation metrics.
[0090] 4. Based on the above evaluation metrics, the method of this invention is compared with other mainstream target tracking algorithms. The tracking performance is shown in Table 2:
[0091] Table 2 Performance Comparison of Various Methods
[0092]
[0093]
[0094] As can be seen from Table 2, the method proposed in this invention has significantly improved performance on the three datasets MOT17, MOT20, and Vis-Drone.
[0095] In addition, combined Figure 8 It can also be seen that, through visualization, the pedestrian small target tracking method based on multi-level hierarchical data association provided by this invention has better tracking performance.
[0096] All user information and data involved in this application are authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data comply with the relevant laws, regulations and standards of the relevant regions. Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A pedestrian small target tracking method based on multi-level hierarchical data association, comprising a feature extraction module DHU, a trajectory prediction module NSKF, and a data association module MHDA, characterized in that, The DHU replaces ordinary convolution with depthwise separable convolution and introduces a lightweight attention module, reducing the number of parameters while maintaining model accuracy. It uses dynamically weighted dilated convolution to obtain the receptive field for small targets and improves the detection capability of small targets through a dual-fusion upsampling module. The NSKF adapts to nonlinear system equations and adjusts the Kalman gain according to the target detection confidence and motion features to improve trajectory prediction accuracy. The MHDA proposes a multi-level hierarchical data association method, using cosine distance as the cost matrix for high-confidence detection boxes and Mahalanobis distance calculation for low-confidence detection boxes.
2. The pedestrian small target tracking method based on multi-level hierarchical data association according to claim 1, characterized in that, After obtaining the C5 feature layer through the depthwise separable convolution, a dilated convolution with dynamic weight allocation is used. The dilated convolution has three branches with dilation coefficients of 1, 2, and 3. After weight allocation, the three branches obtain different weight coefficients. A feature map of size 72×72×512 is obtained through feature weighted fusion. The feature map is input into the CAE module, which performs global pooling, feature fusion and encoding, and attention weight generation before outputting a weighted D5.
3. The pedestrian small target tracking method based on multi-level hierarchical data association according to claim 1, characterized in that, The dual-fusion upsampling module includes deconvolution and subpixel convolution. After fusing the D5 feature map through the deconvolution and subpixel convolution, a high-resolution feature map P5 of 288×288×64 is obtained. The feature map P5 is detected using a small target detection method without anchor boxes, and the target detection confidence is obtained.
4. The pedestrian small target tracking method based on multi-level hierarchical data association according to claim 1, characterized in that, The NSKF module includes a nonlinear Kalman filter system and a Kalman gain adjustment module.
5. The pedestrian small target tracking method based on multi-level hierarchical data association according to claim 4, characterized in that, The linear Kalman filter system is transformed by UT to obtain the prior distribution sampling points and their weights. Based on the state equation, a nonlinear function transfer is performed to obtain the state prediction mean and covariance. The UT transformation is used again, and a nonlinear function transfer is performed based on the observation equation to obtain the predicted observations. The observation prediction mean, variance and covariance are obtained by weighting to obtain the nonlinear Kalman filter system. The Kalman gain is adjusted based on the target detection confidence level and motion characteristics, and the effective confidence level is calculated according to the following formula: Where C d Let Δt be the target detection confidence level, τ be the time interval from the last high confidence level, κ be the confidence attenuation constant, and κ be the enhancement coefficient. The motion mutation index is calculated using the following formula to extract motion features: in Let ω be the acceleration vector at time t. t Let ω be the angular velocity at time t, λ be the rotational sensitivity coefficient, and ε be the zero constant. A dual-modal decision-making mechanism is constructed by fusing the detection confidence and the motion features, and the Kalman gain is adjusted.
6. The pedestrian small target tracking method based on multi-level hierarchical data association according to claim 1, characterized in that, The MHDA module updates the existing trajectory by matching it with the detection box.
7. The pedestrian small target tracking method based on multi-level hierarchical data association according to claim 6, characterized in that, The MHDA module divides the detection results into high and low detection boxes according to the DHU module, and the threshold of the high and low detection boxes is set to 0.
7. The high-scoring detection box is first associated with data. The data association method uses cosine distance to construct an appearance similarity matrix for matching. The trajectory that fails to match the first data association is second associated with the low-scoring detection box, and Mahalanobis distance is used for location information matching. The cosine distance is calculated according to the following formula: Where d(i,j) is the appearance matching degree between the trajectory of the i-th predicted box and the j-th detected box, and r j R is the appearance feature vector of the detection box. i This is the set of feature vectors that were successfully matched from the nearest multiple frames to the target. For set R i The m-th eigenvector; The Mahalanobis distance is calculated according to the following formula: d(i,j)=min{(d j -x i ) T S i -1 (d j -x i )} Where d(i,j) is the motion matching degree between the trajectories of the i-th predicted box and the j-th detected box, d j Let x be the state vector of the j-th detection box. i and s i These are the state vector and covariance matrix of the i-th prediction box in the current frame, respectively, predicted by the nonlinear adaptive Kalman filter.